Skip to content
ARC Research
ReportCost controlSeptember 4, 2026· 10 min read

Metered Agents: When the Bill Depends on Which Model Answered

Coding-agent and platform pricing moved from flat seats to routed models and agent traffic, while survey evidence says roughly half of organizations are over their AI plan and almost none pull back. A brief on where the money actually goes and how to govern it by workflow.

Key findings
  • Coding-agent billing now follows the model that answered: Cursor's Auto moved to routed-model pricing on 24 August 2026, adding a per-million-token processing rate on third-party models for teams and enterprises, with the legacy enterprise flat rate expiring 7 September 2026.
  • Anthropic's own documentation puts Claude Code at roughly $13 per developer per active day and $150–250 per developer per month across enterprise deployments, with wide variance by model choice and usage pattern.
  • Enterprise platforms are metering agent traffic rather than seats, so the more work an agent does against your system of record, the more the meter runs.
  • Survey evidence says the governance layer has not caught up: a large share of organizations report AI spend over plan, few come in under, and only a small minority pause or scale back when it happens.

What changed in the billing model

For three years, AI tooling was priced like software: a seat, a month, a predictable line item. That is over. Coverage of Cursor's 24 August 2026 change describes Auto moving off a single flat token rate to the routed model's list price, plus a $0.25-per-million-token processing rate on eligible third-party model requests for Teams and Enterprise — applied to input, output, and cached tokens, including bring-your-own-key traffic. Legacy enterprise Auto keeps the old flat rate only until 7 September 2026.[1]

Anthropic documents the other half of the picture for Claude Code: cost is token consumption, averaging about $13 per developer per active day and $150–250 per developer per month across enterprise deployments, with 90% of users staying under $30 per active day. Spend limits and workspace-level rate limits exist, which is an admission that the default is unbounded.[2]

Platform vendors are following. When Salesforce exposed its platform as agent-callable APIs and MCP tools, it also tied revenue to agent activity: customers pay for headless API consumption priced by usage, and contract separately for model inference. Agent traffic, not seat count, drives the meter.[3]

The operational consequence is specific. A long agent session pays input rates on a transcript that keeps growing, and under routed pricing the same loop can cost several times more on a day the router favors a frontier model. Budgeting a month of agent work now requires assumptions about routing behavior that the buyer often cannot observe.[1]

What the survey evidence shows

Independent survey data says control is lagging demand. ETR's 2026 spending research reports IT budget growth around 3.8% for the year, with 47% of organizations reporting AI spend moderately or substantially over plan, just 6% under, and 10% with no formal AI budget to measure against at all. When overruns hit, only 17% pause or scale back; the rest seek supplemental budget, absorb it, or reallocate.[4]

Where does the reallocation come from? According to the same research, external contractors and consultants first (61%), then legacy infrastructure modernization (49%), then non-AI software licenses (36%). Security is the notable exception — few organizations raid the security budget to fund AI overruns.[4]

A separate survey of technical leaders at mid-market and enterprise companies found 43% already over their AI budgets, most with annual AI budgets between $100K and $1M, and a stated preference for guardrails over behavior mandates: faced with a hypothetical 20% price increase, the most common responses were consolidating vendors and auditing usage with unclear ties to impact.[5]

Analyst estimates put worldwide AI spending in the trillions for 2026, and FinOps practitioners rank controlling token cost and usage in SaaS-delivered software as a top concern — largely because invoices arrive after the fact and rarely explain which workflow, team, or customer caused the charge.[6]

Grounded outcomes for operators

1) Measure cost per workflow, not cost per tool. A seat price tells you nothing about whether a recurring business output got cheaper. Attribute every paid request to a team, a workflow, and where possible a customer.

2) Set caps and spike alerts before the next invoice, not after. Detection in minutes beats reconciliation in weeks, and every major vendor now exposes some form of limit or admin API.

3) Budget for routing variance. If you cannot observe which model answered, treat the top of the routed price range as your planning number.

4) Bound the loops. Retry budgets, context-growth limits, and step ceilings are cost controls as much as reliability controls — an agent that retries indefinitely is a billing event, not a feature.

5) Include the invisible costs. Review time, exception handling, integration work, and the manual process still running in parallel during a pilot are part of the total, and they are usually what makes a promising workflow uneconomic.

6) Charge it back. Agencies and internal platform teams that cannot bill agent spend to the client or cost center that caused it end up subsidizing the enthusiastic.

7) Expect scrutiny on external spend. Consulting is the first line cut to cover an AI overrun, which is why work should be scoped short, tied to a number, and defensible inside the same quarter it is bought.

Limitations and how to read this brief

Vendor pricing moves quickly; the figures cited here carry dates for that reason, and negotiated enterprise agreements can differ from public list terms. Survey percentages describe responding populations, not every company. Where a pricing change is documented through industry coverage rather than a primary vendor page, we label it as coverage. Confirm current terms with your vendor before planning against them, and confirm your own baseline before quoting anyone's per-developer average — including ours.

Sources & citations

Primary and secondary sources used in this brief. Open the original document to verify claims in context.

  1. [1] Industry coverage of Cursor's Auto pricing change (24 August 2026). Cursor Auto Pricing: What the August 24 Change Actually Costs. CellCog, 2026.
  2. [2] Anthropic. Claude Code — Manage costs effectively. Claude Code documentation, 2026.
  3. [3] Salesforce. Headless 360: platform capabilities as APIs, MCP tools, and CLI commands. Salesforce, 2026.
  4. [4] Enterprise Technology Research (ETR). AI Demand Is Strong. AI Budget Control Is Still Catching Up.. ETR Data Drop, 2026.
  5. [5] Retool. Managing AI costs in 2026: 101 tech leaders weigh in. Retool Blog, 2026.
  6. [6] CIO (citing Gartner and FinOps Foundation research). Nobody knows where their AI budget is going. CIO.com, 2026.

Want this applied to your stack?

Studio can score the paper against your environment: what to do first, what to ignore, who owns it.