STACKOPTIMA Editorial · Published · Reviewed

Begin with a completed unit of work

A monthly API bill is easy to count. The economic value of that bill is harder to interpret. A useful cost model starts with a completed unit of work: an accepted extraction, a resolved support case, a reviewed report, or another outcome the product can recognize. Calls and tokens are intermediate quantities.

This distinction matters when one candidate needs multiple attempts and another regularly succeeds first time. It also matters when a model generates more material than users need. Counting activity without counting accepted work can make an inefficient system look inexpensive simply because its billing unit is small.

Define acceptance before comparing costs. The definition might require a valid schema, correct values, source support, and delivery within a time limit. Human review can remain part of the workflow, but its cost should be visible rather than absorbed into an unexplained operating expense.

Make the arithmetic reproducible

For ordinary uncached text usage, estimate token spend as input tokens divided by one million times the input rate, plus output tokens divided by one million times the output rate. Rates and billing categories vary by provider; verify the exact model version and endpoint before relying on the result.

Consider an original illustrative scenario of 100,000 monthly calls, each using 1,000 input tokens and 300 output tokens. At hypothetical rates of $2 per million input tokens and $8 per million output tokens, the token estimate is $200 plus $240, or $440. These figures are teaching assumptions, not current provider prices.

Now suppose 10% of calls require one identical retry. Token spend rises to $484 under the same assumptions. If the product also pays for search, tool execution, or a reviewing model, those charges require separate rows. Do not hide them inside an unexplained multiplier that the reader cannot inspect.

Treat hosting as an architecture choice

A managed model API generally charges for inference through its own billing terms. An application using that API may still need web hosting, a database, queues, and storage. It does not automatically need a dedicated GPU. Adding a GPU estimate to every API recommendation can count an unnecessary deployment component.

Self-hosting changes the cost structure. Calculate actual billed accelerator hours, supporting resources, storage, and operations. A hypothetical resource at $1.20 per hour running continuously for a 30-day month costs $864 before extras: $1.20 times 24 times 30. Twenty-four hours is one day, not a monthly conversion.

Utilization makes the comparison sensitive. A resource that sits idle for most of the day still creates charges under a fixed hourly arrangement. Batching, scheduling, or scaling may improve the economics, but only if the workload tolerates the associated delays and the platform supports the intended behavior.

Model discounts as conditional

Caching is useful when reusable material and provider rules align. Anthropic's prompt-caching documentation distinguishes cache creation from subsequent reads and describes validity conditions. A budget should therefore separate cache writes, reads, and ordinary input instead of applying a universal discount to all tokens.

Create a base case, a low-reuse case, and a demand-spike case. Include output length growth, because verbose responses can increase both cost and completion time. Record contractual commitments separately from on-demand rates. Google's cost-optimization framework places spending decisions in the context of business value; a lower bill is not automatically a better product.

Compare the decision that can actually change

Add engineering time, incident response, review effort, and migration work where they materially differ between candidates. Avoid false precision: a range is more honest than a dollar-exact forecast built on guessed traffic. Keep one-time transition costs separate from steady monthly operations.

Finish with two numbers: expected monthly spend and cost per accepted task. Then identify the assumption most likely to reverse the choice. If utilization, retry rate, or review time dominates the result, run an experiment on that variable before negotiating a commitment. A useful estimate helps decide what to measure next.

Sources and further reading

Apply this to your workload →

All insights · Editorial policy and corrections