Ali Fakhar
Measurement

AI marketing ROI: measure value beyond time saved

Evaluate an AI workflow honestly: time accepted output, not generated output, count setup, review and maintenance, and keep released capacity, cash savings and commercial gains apart. A worked model with a sensitivity check.

ai-marketing-roi

The short version

  • Compare accepted output, including review, correction and rejected work, against a measured manual baseline.
  • Count setup, running, maintenance, failure and training costs over a stated period.
  • Released capacity, cash savings and incremental contribution are different; adding them together double-counts.
  • Test the assumptions. Review time and volume often decide whether a workflow pays back.
On this page

“The tool saves us ten hours a week” is a familiar way to report the value of AI in a marketing team, and it is usually incomplete. It counts how fast the first draft appears. It rarely counts the time spent checking it, fixing it, throwing some of it away, setting the workflow up and keeping it running.

A fair assessment compares accepted output, not generated output: work that met the quality bar, with review and correction included, against a measured manual baseline. It adds every cost over a stated period. Then it separates three kinds of value that are often blended: capacity released, cash actually saved and commercial gains that can be shown to come from the change. Many workflows produce useful capacity and no cash saving. That is a legitimate result, as long as you call it what it is.

Define success for this task first

Before you measure anything, say what the workflow is for. The answer changes what you measure:

  • Accepted output: more usable briefs, reports or drafts per month, at the same quality.
  • Timeliness: the Monday report ready before the Monday meeting.
  • Quality: fewer errors reaching leadership or customers.
  • A business outcome: more qualified opportunities, lower cost per opportunity.

A workflow can succeed on one and fail on another. Faster reports that leadership still does not trust have not solved the problem that justified the project. If you chose the pilot using a clear bottleneck, as described in AI marketing strategy: how to choose what to automate first, the success measure should follow from it.

Measure accepted output, not generated output

Time the task both ways on comparable work. For the AI-assisted version, record every step a person touches:

  1. preparing inputs and instructions
  2. generating the output
  3. reviewing it against the quality criteria
  4. correcting what is wrong
  5. discarding output that cannot be fixed, and doing that item another way

The last step is the one reports leave out. If one output in five is rejected and redone manually, that time belongs in the average. Divide total time by the number of accepted outputs, not the number generated.

Do the same for the manual baseline, including its own checking. Manual work has review costs too, and ignoring them flatters the AI version by comparison.

Quality needs the same care as time. Anthropic’s engineering guide to evaluating AI agents defines an evaluation simply: give an AI an input, then apply grading logic to its output to measure success. For a marketing workflow, the grading logic is your reviewer’s checklist, applied consistently to real examples. The checklist for content is in AI content quality: an editorial checklist before publishing.

Count all the costs

The subscription is usually the smallest and most visible cost. Over a stated period, list:

  • Setup: designing the workflow, writing instructions, building the test set, integration work. Often the largest single cost.
  • Running costs: subscriptions, usage-based API charges, any additional software.
  • Maintenance: updating instructions when the process changes, fixing breakages, re-testing after model or tool updates.
  • Failure handling: the time spent when something goes wrong, and the cost of any error that escapes review.
  • Training: getting the team to use it properly.

Decide how to treat people’s time before you start. If staff time is already paid for, saving it does not reduce cash costs; it frees capacity. Mixing the two, for example by subtracting salaried time as if it were an invoice, is an easy way for an ROI model to double-count.

Capacity, cash and contribution are different things

Keep three columns separate:

  • Released capacity: hours freed for other work. Valuable, but only real if the time is used for something that matters.
  • Realised cost reduction: money that actually stops being spent, such as a contractor no longer needed or an agency fee reduced.
  • Incremental contribution: additional profit from a commercial change, after the variable costs of delivering it, shown with an appropriate comparison.

Adding the three into one total usually counts the same benefit twice. The hours released to write more campaigns are not also a cash saving, and the campaigns’ revenue cannot be credited to AI without evidence that it would not have happened anyway.

A worked example: capacity without cash

This example is illustrative. The task, times, rates and costs are invented to show the method; none are results from my work or a client’s.

A team produces forty accepted outputs a month of a recurring task, such as campaign briefs checked against a template.

  • Manual: 60 minutes per accepted output.
  • AI-assisted: 15 minutes preparing and generating, 20 minutes reviewing and 10 minutes correcting: 45 minutes per accepted output, including redone rejects.
  • Time released: 15 minutes × 40 = 10 hours a month.
  • Internal valuation: £40 an hour, used only to compare capacity figures; nobody’s pay changes.
  • Maintenance: 2 hours a month of staff time.
  • Setup: 60 hours of staff time, spread over 12 months: 5 hours a month.
  • Tool subscription: £100 a month, a real cash cost.
Illustrative monthly value model
LineHoursValue at £40 an hourType
Time released on the task10£400Capacity
Maintenance−2−£80Capacity
Setup, spread over 12 months−5−£200Capacity
Net capacity released3£120Capacity
Tool subscription–−£100Cash
Modelled monthly value–£20Mixed

The model shows a small positive number, £20 a month. Look at what it is made of: £120 of capacity value, an estimate, minus £100 of real cash. The team spends £100 more a month than before and gains three hours of capacity. No cash profit has been established. The workflow is worth keeping only if those three hours go to work that matters more than £100, or if the capacity lets the team stop paying for something else.

That is an honest result, and it is still useful to know.

Sensitivity: what breaks the model

Small changes in the inputs move the answer a lot. Test the assumptions most likely to be wrong:

How the illustrative result changes with its assumptions
ScenarioNet capacity per monthModelled monthly value
Base case: 40 tasks, 45 minutes each3 hours£20
Review takes 10 minutes longer (55 minutes a task)−3 hours 40 minutesabout −£247
Volume halves to 20 tasks−2 hours−£180
Volume doubles to 80 tasks13 hours£420

Two lessons. Review time is often the deciding variable, so measure it rather than estimating it. And volume matters: fixed costs of setup and maintenance need enough repetitions to pay back. At these assumptions the break-even point is 38 tasks a month.

When you can call it ROI

Return on investment has a specific meaning. For a realised financial return, use:

ROI = (measured incremental financial benefit − incremental cost) ÷ incremental cost, over a stated period

State the period and exactly what counts as benefit. For commercial gains, use contribution after the relevant variable costs, not revenue. Pipeline value is a forecast, not realised revenue. And do not credit a sales increase to AI unless you have a comparison that shows what would have happened without it: a holdout group, a staggered rollout or a structured test (see marketing experiments). Connecting marketing activity to pipeline properly is covered in B2B marketing metrics.

In the worked example there is no measured financial benefit yet, so there is no ROI to report, only a capacity model. That distinction protects your credibility when finance reviews the numbers.

Stop, revise or scale

With the model built, the decision becomes clearer:

  • Stop if accepted-output time is not lower than the baseline, quality is worse, or the model only works under optimistic assumptions.
  • Revise if review or correction dominates. Better inputs, clearer instructions or a narrower scope often cut review time more than a new tool does.
  • Scale if the workflow holds up on real volume, the released capacity has a named use, and the sensitivity table does not flip negative under realistic assumptions.

If scaling means giving the system more autonomy, weigh the added checking and failure costs first: AI agents vs marketing automation.

Cost and value worksheet

Copy this into a spreadsheet. Keep the “Type” column: it stops capacity and cash from being added together by accident.

Blank cost and value worksheet
LineManual baselineAI-assistedTypeSource of the figure
Accepted outputs per month[n][n]Volume[Log or estimate]
Minutes per accepted output, including review and rework[min][min]Capacity[Timed sample]
Setup hours, spread over [n] months–[h]Capacity[Project log]
Maintenance hours per month[h][h]Capacity[Log]
Subscriptions and usage charges[£][£]Cash[Invoices]
Costs that actually stop–[£]Cash[Contract or invoice]
Measured incremental contribution–[£ or none]Contribution[Comparison method]
Decision–[Stop, revise or scale]–[Owner and date]

Sources and further reading

  1. Anthropic: Demystifying evals for AI agents

Questions worth asking

Does saving time count as ROI?

Not on its own. Time saved is released capacity; it becomes a financial return only if it reduces a real cost or is used for work that produces measurable additional contribution. Report capacity as capacity, and keep the ROI formula for measured financial benefit against incremental cost.

Should review time be included?

Yes, along with correction and any output that is rejected and redone. Divide total time by accepted outputs. In the worked example, ten extra minutes of review per task turns a small positive result into a clearly negative one.

Can we attribute more leads to AI?

Only with a comparison that shows what would have happened without it, such as a holdout group, a staggered rollout or a structured test. A rise in leads after adopting a tool may have other causes, and leads or pipeline are not realised revenue.

Ali Fakhar
About the author

Ali Fakhar

Ali Fakhar is a London-based marketer working across growth, paid media, content and practical AI.

More about Ali ↗
Analytics preferences

Current preference: off