Zarif Automates
Enterprise AI10 min read

How to Forecast and Budget Token-Based AI Costs

ZarifZarif
||Updated August 12, 2026

The wrong way to budget AI is to take last month's token bill and add 20 percent.

That forecast has no explanation for adoption, seasonality, agent loops, model changes, longer context, new tools, quality failures, or a vendor replacing raw tokens with credits. It produces a number but not a decision model.

A defensible forecast starts with work: how many business events arrive, which workflows run, which steps call a model, what one successful run consumes, and how demand changes.

Definition

An AI workload forecast estimates future cost by multiplying expected business volume by the measured resource shape and success rate of each workflow, then applying model, tool, platform, and commercial rates.

TL;DR

  • Forecast by workflow and business event, not one blended token-growth percentage
  • Use distributions or percentiles because averages hide power users and long agent runs
  • Separate demand, technical efficiency, rate, quality, and commitment assumptions
  • Build low, base, and high cases with named triggers and management actions
  • Forecast cash spend and effective cost separately when credits or commitments expire
  • Reforecast monthly and when a model, prompt, workflow, rate card, or adoption policy changes materially

The Seven Inputs to an AI Budget

For each workflow, collect:

  1. Business volume: cases, documents, calls, messages, transactions, or users
  2. Eligibility rate: share of events that enter the automated path
  3. Model-call rate: calls made per eligible event
  4. Token shape: input, cached input, and output tokens per call
  5. Tool cost: searches, third-party APIs, embeddings, and other metered services
  6. Success shape: retries, failures, escalations, and accepted outcomes
  7. Commercial structure: rates, discounts, commitments, credits, rollover, and overage

Do not hide these in one “cost per request” assumption. Keeping them separate lets leadership understand why the forecast changed.

The Base Forecast Formula

For one workflow and model:

Monthly calls = business events × eligibility rate × calls per eligible event

Monthly token cost = monthly calls × average cost per model call

Monthly workflow cost = token cost + tools + orchestration + data + human review + allocated platform cost

Cost per accepted outcome = monthly workflow cost ÷ accepted outcomes

For a credit-priced SaaS product:

Monthly credits = sum of operation volume × credit weight per operation

Effective credit cost = committed credit spend ÷ credits actually used

The effective cost is higher than the contracted unit price when credits expire unused.

Step 1: Build a Workflow Inventory

Do not start with providers. Start with business use cases.

WorkflowDemand unitOwnerModel stepsOutcome unit
Support triageNew caseSupport Ops1Correctly routed case
Contract reviewContractLegal Ops3Accepted review
Sales-call analysisRecorded callRevOps2Used insight
Content draftingBriefMarketing2Approved draft

Record production, pilot, and proposed workflows separately. A pilot growing from 50 tests to 50,000 transactions is not ordinary month-over-month growth.

Step 2: Measure the Workload Shape

One average is not enough. Capture at least:

  • Median
  • 75th percentile
  • 95th percentile
  • Maximum with an explanation

Measure input, cached input, output, calls per workflow, tool usage, latency, retries, and human review.

Why percentiles? A document workflow may average 20,000 input tokens while a small group of long contracts consumes half the budget. An agent may usually call two tools but occasionally loop 30 times.

If direct token counts are unavailable because a SaaS vendor bills credits, use feature-level operations and credit burn. Require exports or reports that attribute usage by feature, team, and owner.

Step 3: Separate Five Forecast Drivers

Demand

Business events, active users, adoption, seasonality, and new rollout cohorts.

Technical efficiency

Context length, cache rate, calls per workflow, output length, routing, and retry behavior.

Quality

Evaluation pass rate, acceptance, escalation, and human-review minutes.

Rate

Model prices, tool prices, data-residency premiums, vendor credit weights, and volume discounts.

Commercial position

Committed amount, included allowance, expiration, overage, minimums, and price protection.

A variance report should attribute change to one of these drivers. “AI spend was over budget” is not actionable. “Eligible support volume grew 12 percent and retries rose from 4 to 11 percent after a prompt change” is.

Step 4: Build Low, Base, and High Cases

Do not use arbitrary plus-or-minus percentages. Define a story for each case.

Low case

  • Rollout reaches only planned users
  • Deterministic filters remove more known cases
  • Cache hit rate improves
  • Model mix shifts toward smaller models
  • No major new use case launches

Base case

  • Approved adoption plan occurs
  • Current workflow efficiency holds
  • Known price changes are included
  • Normal seasonal peaks occur

High case

  • Adoption accelerates
  • Transaction or document volume reaches the upper business forecast
  • Long-context share grows
  • Retries or tool calls reach the 95th percentile
  • One planned workflow launches early

Attach triggers to the scenarios. If actual calls are 15 percent above base for two consecutive weeks, activate the high-case review rather than waiting for month end.

A Worked Annual Budget

Assume a support-triage workflow. These figures are illustrative.

Base assumptions

  • 200,000 new cases per month
  • 70 percent eligible for automated triage
  • 1.2 model calls per eligible case
  • Direct model and tool cost of $0.012 per call
  • 4 percent retry overhead already excluded from the call count
  • 90 percent accepted routing decisions
  • $4,000 monthly platform and data allocation
  • Average human review cost of $0.03 per eligible case

Calculation

Monthly model calls:

200,000 × 70 percent × 1.2 = 168,000

Model and tool cost before retry:

168,000 × $0.012 = $2,016

Retry-adjusted model cost:

$2,016 × 1.04 = $2,096.64

Human review cost:

140,000 eligible cases × $0.03 = $4,200

Fully loaded monthly cost:

$2,096.64 + $4,000 + $4,200 = $10,296.64

Accepted outcomes:

140,000 × 90 percent = 126,000

Cost per accepted routing decision:

$10,296.64 ÷ 126,000 = about $0.082

Annual base budget before growth:

$10,296.64 × 12 = $123,559.68

The direct model charge is only about one-fifth of the fully loaded cost. Optimizing model tokens without examining human review would miss the larger lever.

Step 5: Add Growth and Seasonality

Model business volume separately from adoption.

For each month:

Forecast events = baseline events × business-volume factor × adoption factor × seasonal factor

Do not compound every driver blindly. A 20 percent growth plan and a holiday peak may affect different customer segments.

For new workflows, use ramp cohorts:

  • Pilot: limited events and mandatory review
  • Controlled production: one team or segment
  • Expansion: additional teams with measured quality
  • Mature: steady-state adoption and optimization

Associate a gate with each step. Expansion should depend on outcome quality and unit economics, not the calendar alone.

Step 6: Model Commitments and Credits

Usage discounts can lower unit price and raise total cost if the commitment is wrong.

Track four quantities:

  • Contracted commitment
  • Forecast eligible usage
  • Forecast effective usage after failures and exclusions
  • Buffer for approved growth

Calculate:

Commitment coverage = forecast billable usage ÷ committed usage

Coverage below 100 percent implies waste risk. Coverage far above 100 percent implies overage or throttling risk.

For expiring credits, show:

  • Expected exhaustion date
  • Expected unused balance at expiration
  • Effective unit rate after unused commitment
  • Cost of buying overage versus amending commitment

Do not create low-value AI work merely to consume sunk credits. Unused commitment is a procurement lesson, not a reason to generate waste.

Step 7: Forecast Cash and Economics Separately

The cash invoice and economic consumption may occur at different times.

  • Prepaid credits create cash spend before use
  • Annual platform fees may be recognized or allocated monthly
  • Overages may be billed in arrears
  • Internal labor cost may lag in reporting
  • Unused credits raise effective unit cost at expiration

Maintain two views:

  1. Cash forecast: invoices and payment timing
  2. Unit-economics forecast: fully loaded cost assigned to useful outcomes

Finance needs both for liquidity, accounting, sourcing, and investment decisions.

Step 8: Add Guardrails to the Forecast

Good controls act before cost, not only after it.

Informational alerts

Notify owners at 50, 75, and 90 percent of the monthly workflow budget.

Soft controls

Route low-priority work to batch processing, a smaller model, or delayed execution.

Hard controls

Pause noncritical workloads, cap agent iterations, limit document size, or require approval for overage.

Value controls

Stop expansion when acceptance, outcome, or cost-per-outcome thresholds deteriorate.

n8n can enforce workflow-level conditions, schedules, route selection, and notifications before an external model is called. This is valuable when the vendor API supplies usage data and the organization has defined safe fallback behavior.

Monthly Forecast Review

Review the top use cases by spend, growth, and variance.

  1. Actual versus forecast business volume
  2. Eligibility and model-call rate
  3. Token and tool shape
  4. Retry, failure, and acceptance rate
  5. Fully loaded cost per outcome
  6. Commitment use and expiration risk
  7. Rate-card or model changes
  8. Optimization actions and expected impact
  9. Next-month low, base, and high case
  10. Decisions: expand, optimize, renegotiate, cap, or stop

The AI FinOps operating model assigns ownership for this cadence.

Forecasting Mistakes to Avoid

Applying one growth rate to the entire AI bill

Different workflows have different demand, adoption, and cost shapes.

Forecasting from list rates

Use contracted rates and current model routing. Preserve list rates for savings analysis.

Ignoring retry and failure spend

Failed work consumes resources even when it creates no outcome.

Using a single average

Model percentiles and peak periods. Agent loops and long documents create heavy tails.

Assuming cheaper models mean a lower budget

Demand can grow faster than unit cost falls. Quality can also change review labor.

Treating unused credits as savings

A lower invoice caused by unused commitment can still be a poor effective unit rate.

After building the forecast, use how to reduce LLM token costs to change the technical drivers and the CFO token-economics guide to keep the outcome denominator honest.

Frequently Asked Questions

How do you forecast token usage?

Forecast business-event volume, multiply by the eligible share, model calls per event, and measured input, cached input, and output token distributions. Add tools, retries, quality failures, platform cost, and commercial terms.

How much buffer should an AI budget include?

Use a buffer tied to a documented high case rather than a universal percentage. The buffer should reflect approved adoption, workload variance, seasonality, model changes, and the organization's ability to throttle or defer work.

Should AI budgets be owned by IT or business units?

Use shared ownership. Business units own demand and value; Engineering owns workflow efficiency and quality; FinOps or Finance owns allocation and forecast; Procurement owns commercial terms. Central platform costs can be allocated through agreed drivers.

How often should an AI budget be reforecast?

Review monthly for production portfolios and more frequently during a major rollout. Reforecast when demand, model routing, prompt shape, rate cards, credit weights, success rates, or commercial commitments change materially.

Can n8n enforce an AI budget?

n8n can evaluate counters or cost data, route work, cap iterations, require approval, delay jobs, or send alerts. It needs reliable usage data and safe fallback behavior; it is an enforcement layer, not the source of the budget policy.


Sources and Further Reading

Zarif

Zarif

Zarif is an AI automation educator helping thousands of professionals and businesses leverage AI tools and workflows to save time, cut costs, and scale operations.