Run cost is the resources and payments needed to operate a workflow at a specified demand, quality, and service level. Build cost is incurred once and is visible in the approval. Run cost scales with use, arrives after the decision, and lands on a budget the approval did not mention.

Four buckets

For a reporting period, total cost is dedicated fixed cost, plus metered usage multiplied by its dated unit price, plus recorded staff hours multiplied by a stated valuation rate, plus allocated shared cost.

The four have to be disjoint. Counting a contractor’s hours and the invoice covering those same hours produces a total that is wrong by the size of the larger item. A labour rate that already carries overhead must not receive overhead again.

The result is a mixed resource-cost view and has to be labelled as one. Actual payments, commitments, and accounting-period charges are reconciled separately, because allocated salary represents resource consumption without being incremental cash.

Two denominators

The unit is an eligible task or demand episode. One episode contains several model calls, retrieval, tool use, and a human completion step, and counting calls instead of episodes measures the implementation rather than the work.

Divide total cost by eligible episodes for cost per attempt. Divide by accepted completed episodes for cost per outcome. Both numerators include the cost of failed and incomplete episodes, because those consumed resources.

The gap between the two denominators is the whole subject. A workflow that costs little per attempt and completes half its attempts is expensive.

Why the model bill is the wrong thing to optimise

Take a month with 10,000 eligible episodes, 1,000 of fixed cost, 400 of metered charges, 40 review hours valued at 60 an hour, and 600 of allocated shared cost. Total is 4,400. If 8,000 episodes meet the completion standard, that is 0.44 per attempt and 0.55 per accepted outcome.

Now halve the metered charges to 200. Suppose the cheaper configuration needs 60 review hours instead of 40, and 7,000 episodes qualify instead of 8,000. Total rises to 5,400, and cost per accepted outcome rises to about 0.77.

The model bill fell by half and the economics got forty percent worse. Metered inference was nine percent of the original total; human review was fifty-five percent. Optimising the smallest bucket while degrading the largest is the standard outcome of measuring only what the provider itemises.

This is why the measurement guidance is to track total cost per defined use-case outcome, and to keep tracking it as the system changes. Every prompt, model, retrieval, or workflow change reopens the estimate.

What gets left out

Recurring evaluation, monitoring, support, idle capacity, and data refresh are running costs and are recorded separately from initial setup. Human review is recorded by activity rather than as a single pool, because the exception path and the routine check have different volumes and different people.

Where review effort is forecast rather than observed, expected hours are episodes multiplied by review share multiplied by minutes per review, divided by sixty — and that arithmetic is valid only when all three inputs describe the same population. Applying an exception rate measured on one cohort to a different cohort’s volume is a common way to produce a confident wrong number.

The rule

What stays fixed is that cost attaches to episodes and outcomes, not to components. What changes is the mix across the four buckets, and the mix moves whenever the configuration does, which means a run-cost model is dated and expires.

Where the model breaks

The denominators come from different cohorts. Combining this period’s spend with a different cohort’s outcomes produces a ratio with no referent. A zero denominator means unavailable, not free.

The owner is discovered late. The sponsor who approves the build, the budget holder who funds it, and the team that will carry the run cost after transfer are three roles, and the third is the one absent from the approval meeting. A prototype whose run cost has no named future owner has an unfunded liability with a launch date.

Not to be confused with

Cash. Allocated existing salary shows resource use without additional payroll outlay. Both views are legitimate and they answer different questions, so a single total labelled as spending is one of them mislabelled.

Build cost. The prototype-to-production multiple stops at operational acceptance. Everything here starts there.