The shift from experimental AI use to production deployment has exposed a fundamental mismatch: most finance teams still think in line items, while AI vendors bill by the token. For mid-market firms, the gap between technical consumption and financial planning is now a source of budget overruns, procurement disputes, and stalled projects.
This article explains why per-token pricing is giving way to workload-level budgeting, how mid-market firms are implementing the change, and what risks remain.
Why per-token pricing is failing as an operational metric
Per-token pricing is precise but not useful for planning. A token is a unit of text processing, but it does not map cleanly to a business outcome. A customer-support chatbot may consume 10,000 tokens per query, while a document summarisation tool may use 50,000 tokens per document. Neither figure tells a finance team what the work is worth or what it should cost.
Mid-market firms are discovering that token-level tracking creates three problems:
- Cost attribution is fuzzy. When multiple teams share an API key, it is hard to know which department drove the spend.
- Forecasting is unreliable. Token consumption varies with input length, model version, and prompt complexity, making monthly budgets a guess.
- Optimisation is misdirected. Teams focus on reducing token count rather than improving business value, leading to prompt compression that degrades output quality.
As a result, finance leaders are asking a different question: not "how many tokens did we use?" but "what did this workload cost, and what did it deliver?"
The shift to workload-level budgeting
Workload-level budgeting treats an AI use case as a discrete cost centre. Instead of tracking tokens, firms define a workload—such as "customer support triage" or "contract review"—and assign a monthly budget based on expected volume and unit economics.
For example, a mid-market logistics firm might budget £8,000 per month for an AI-powered shipment tracking assistant. The budget is set by expected query volume and an acceptable cost per resolved query, not by token price. The team then monitors cost per workload, not cost per token.
This approach has several operational advantages:
- Clearer ownership. Each workload has a named owner who is accountable for cost and performance.
- Better forecasting. Workload volumes are easier to predict than token counts, especially when tied to business cycles.
- Simpler procurement. Negotiations with vendors shift from per-token rates to committed workload volumes, which can yield discounts and more predictable invoices.
Several mid-market firms are also using internal chargeback mechanisms, where business units pay for the workloads they consume. This creates a direct link between AI spend and departmental budgets, reducing the risk of uncontrolled usage.
How to implement workload-level budgeting
Transitioning from token-based to workload-based budgeting requires changes in three areas: measurement, governance, and tooling.
Measurement: define the workload unit
The first step is to define what constitutes a workload. This is not always obvious. A workload could be a single API call, a multi-step agentic task, or a batch process. The key is to choose a unit that is meaningful to the business and stable over time.
For example, a legal document review workload might be defined as "one contract reviewed." A customer support workload might be "one ticket resolved." The unit should be measurable, repeatable, and tied to a business outcome.
Once the unit is defined, firms need to track the cost per unit. This requires logging not just token usage but also the context window, model version, and any retries or fallback calls. Many AI platforms now offer usage analytics that can be exported to finance systems, but mid-market firms often need to build a lightweight internal dashboard to aggregate data across vendors.
Governance: assign ownership and approval thresholds
Workload-level budgeting works best when there is clear ownership. Each workload should have a named budget holder who is responsible for cost and performance. This person should have the authority to approve changes to the workload, such as switching to a cheaper model or reducing the frequency of retries.
Firms should also set approval thresholds for new workloads. If a team wants to launch a new AI use case, they must submit a business case that includes expected volume, cost per unit, and a payback period. This prevents the proliferation of low-value AI experiments that drain the budget.
Tooling: use cost-management platforms and internal dashboards
Mid-market firms are increasingly using third-party cost-management platforms that sit between the AI vendor and the internal team. These platforms provide real-time cost tracking, anomaly detection, and budget alerts. They also allow firms to set per-workload budgets and receive notifications when spend exceeds a threshold.
For firms that prefer to build in-house, a simple spreadsheet or BI dashboard can suffice, but it must be updated regularly and integrated with the vendor's usage API. The goal is to have a single source of truth for AI spend, rather than relying on monthly invoices that arrive after the fact.
Commercial impact: what changes for buyers and vendors
The move to workload-level budgeting has commercial implications for both sides of the market.
For buyers, the shift enables better negotiation. When a firm can articulate the cost per workload, it can compare vendors on a like-for-like basis. It can also push for volume discounts based on committed workloads, rather than accepting per-token rates that are opaque and volatile.
For vendors, the trend is a double-edged sword. On one hand, workload-based pricing can simplify sales and reduce churn, because customers are less likely to be surprised by a large bill. On the other hand, vendors may face pressure to lower prices as buyers become more sophisticated about cost per outcome.
Some vendors are already responding by offering workload-based pricing tiers, such as a flat monthly fee for a defined number of workloads. This is particularly common in vertical applications, such as legal or healthcare AI, where the workload is well-defined and the value is clear.
Risks and unknowns
Workload-level budgeting is not a silver bullet. There are several risks and unknowns that mid-market firms should consider.
- Workload definition drift. As AI models improve, the same workload may require fewer tokens or more complex reasoning. The cost per unit may change, making historical comparisons difficult.
- Hidden costs. Token usage is only one part of the cost. Firms must also account for data storage, network egress, and human review time. These costs are often excluded from per-token pricing but can be significant.
- Vendor lock-in. If a firm negotiates a workload-based contract with one vendor, it may be harder to switch to a cheaper alternative later, especially if the workload is tightly integrated with the vendor's API.
- Quality trade-offs. Reducing cost per workload may lead to lower-quality outputs if teams optimise for token count rather than outcome. Firms need to monitor quality metrics alongside cost.
FY Outlook
The shift to workload-level budgeting is likely to accelerate as AI becomes a larger line item in mid-market budgets. Finance teams will demand more transparency from vendors, and procurement will increasingly treat AI as a managed service rather than a utility.
We expect to see more vendors offer workload-based pricing plans, and more third-party tools emerge to help firms track cost per workload. However, the lack of standardisation in workload definitions will remain a challenge. Firms will need to develop their own internal standards and be prepared to adjust them as the market evolves.
Conclusion
Per-token pricing is a useful technical metric, but it is not a sound basis for business budgeting. Mid-market firms that move to workload-level budgeting gain clearer ownership, better forecasting, and stronger procurement leverage. The transition requires effort in measurement, governance, and tooling, but the payoff is a more predictable and accountable AI cost structure.
As AI becomes embedded in core operations, the firms that treat it as a managed workload rather than a metered utility will be better positioned to scale efficiently and avoid budget surprises.
Why It Matters
For mid-market firms, AI spend is no longer a small experimental line item. As usage scales, per-token pricing creates unpredictable costs and poor accountability. Workload-level budgeting aligns AI spend with business outcomes, enabling better forecasting, clearer ownership, and stronger vendor negotiation. Firms that fail to adopt this approach risk budget overruns and stalled AI initiatives.



