All white paper

The Frontier Model Bill: Why AI Costs Behave Like Infrastructure, and How to Engineer Them Down

Why Claude and Frontier LLM Costs Keep Climbing, and How Enterprises Bring Them Down

L
Lyzr Team
Oct 7, 2026
23 min read

Frontier model spend is not a pricing problem. It is an allocation problem: enterprises send every task to the most capable model, resend full context on every call, and cannot see which agent spent what. 

The symptoms are now widespread. Uber used its full-year 2026 AI budget in four months, driven mainly by Claude Code. In McKinsey’s May 2026 Enterprise AI FinOps survey, 93% of the 75 qualified respondents had exceeded their AI budgets, and McKinsey’s experience is that 20% to 30% of AI spend is often unaccounted for. 

This paper argues for three shifts: 

Screenshot 2026 10 07 at 12.24.58 PM
The Frontier Model Bill: Why AI Costs Behave Like Infrastructure, and How to Engineer Them Down 34

The levers are large. In a worked coding example, prompt caching alone cut cost per task by 80% on Claude Opus 5.5. Batch processing halves the price of non-urgent work. Lyzr’s design target for moving a stable task to an owned small model is up to 90% lower inference cost on that task. 

The goal is to make every AI dollar traceable to a team, an agent and a result, and to pay frontier prices only where frontier intelligence is needed. 

1. From subscription to utility 

AI spend has stopped behaving like a software license. It now scales with usage, like cloud infrastructure, and usage is growing faster than budgets. 

Uber shows the pattern in miniature.

Uber exhausted its entire 2026 AI budget by April, four months into the year, after Claude Code spread across roughly 5,000 engineers faster than its finance models anticipated (Forbes). Uber rolled Claude Code out in December 2025; adoption climbed from 32% of engineers in February to 84% classified as agentic coding users by March (Forbes). 

Screenshot 2026 10 07 at 12.26.09 PM
The Frontier Model Bill: Why AI Costs Behave Like Infrastructure, and How to Engineer Them Down 35

Monthly cost per engineer averaged $150 to $250, with power users running $500 to $2,000; CTO Praveen Neppalli Naga reported spending $1,200 in a single two-hour session (Forbes). Internal leaderboards ranking engineers by Claude Code usage added an incentive to consume more. Uber responded with monthly spending limits on AI coding tools (Outlook Business). 

Screenshot 2026 10 07 at 12.26.23 PM
The Frontier Model Bill: Why AI Costs Behave Like Infrastructure, and How to Engineer Them Down 36

The market numbers tell the same story.

Enterprise generative AI spend rose from $1.7 billion in 2023 to $37 billion in 2025 (Menlo Ventures). Anthropic’s share of enterprise LLM API spend reached 40% in 2025, up from 12% in 2023 (Menlo Ventures). 

Anthropic reported a run-rate revenue of $14 billion in February 2026, and that Claude Code’s run-rate revenue had passed $2.5 billion (Anthropic). By the end of July 2026, the company-wide run rate had passed $65 billion, according to CNBC (CNBC). 

Screenshot 2026 10 07 at 12.26.49 PM
The Frontier Model Bill: Why AI Costs Behave Like Infrastructure, and How to Engineer Them Down 37

Claude now arrives inside software companies already use.

In August 2026, Salesforce and Anthropic announced Claudeforce. Claude became the reasoning model for the Atlas Reasoning Engine, the default in Agentforce Vibes and Agentforce Coworker, and the default model for Slack (Salesforce and Anthropic). Usage no longer flows only through contracts IT signed. 

The result is a budgeting problem most companies have not solved. In McKinsey’s May 2026 Enterprise AI FinOps survey, 93% of the 75 qualified respondents had exceeded their AI budgets (McKinsey). McKinsey also reports that, in its experience, 20% to 30% of AI spend is often unaccounted for, spread across vendors, tools and business units. 

What enterprises spend on seats alone.

Lyzr modeled internal Claude and ChatGPT seat spend for 756 enterprise and mid-market companies. The total is about $3 billion a year, concentrated in software, consulting, telecom and travel, and financial services. 

Screenshot 2026 10 07 at 12.27.22 PM
The Frontier Model Bill: Why AI Costs Behave Like Infrastructure, and How to Engineer Them Down 38

These are modeled estimates of seat spend for employee productivity, not reported contract values. The method is in Appendix B.

Treat this as the floor of the bill. Seats are the predictable part, priced per person per month. The API and product usage on top of them is metered per token, and that is where the overruns in this paper come from.

2. Anatomy of a frontier bill 

The price per token is rarely what drives the bill. What drives it is how many tokens each task consumes, and in agentic work that number is large, variable and mostly input. 

Start with list prices. Anthropic’s current models span a 10x range on input, and every model charges five times more for output than input. 

Screenshot 2026 10 07 at 12.28.39 PM
The Frontier Model Bill: Why AI Costs Behave Like Infrastructure, and How to Engineer Them Down 39

Those numbers look modest on a spreadsheet. Five mechanics multiply them. 

  1. One task becomes dozens of calls. A coding agent reads files, runs a tool, reads the result, tries again and runs tests. The user sees one task; the bill sees every call in the loop. 
  1. Context is resent on every step. Each call carries the system prompt, tool definitions, prior turns and tool outputs. Input grows with every step, so cost grows faster than the number of steps. 
  1. Output is priced at five times input. Verbose reasoning, long drafts and unrequested explanations are the most expensive tokens on the bill. 
  1. Failures are paid for twice. A retry, a loop that never converges or a wrong answer that a person corrects all consume tokens without producing a usable result. 
  1. Token counts can rise without any price change. Anthropic notes that Claude 4.7 and later models use a tokenizer that produces about 30% more tokens for the same text (Anthropic). A model upgrade can raise spend at an unchanged list price. 

The variance is as important as the level. Stanford research cited by McKinsey found that token usage can vary by up to 30 times for the same task, depending on how the agent executes it (McKinsey). Budgets built on average usage break on that spread. 

Screenshot 2026 10 07 at 12.30.10 PM
The Frontier Model Bill: Why AI Costs Behave Like Infrastructure, and How to Engineer Them Down 40

A worked example: one coding task 

Consider a coding agent working through a single fix in 30 turns. It starts with 20,000 tokens of context, adds 4,000 tokens per turn, and writes 1,500 tokens of output per turn. 

Screenshot 2026 10 07 at 12.32.14 PM
The Frontier Model Bill: Why AI Costs Behave Like Infrastructure, and How to Engineer Them Down 41

Across the task, that is 2.34 million input tokens and 45,000 output tokens. Input is 98% of the volume. 

Screenshot 2026 10 07 at 12.32.31 PM
The Frontier Model Bill: Why AI Costs Behave Like Infrastructure, and How to Engineer Them Down 42

Illustrative calculation by Lyzr at list prices. The caching column assumes the prior context is read from cache each turn and only new tokens are written. The task path is held constant across models; in practice a stronger model often finishes in fewer turns, which narrows the gap between rows.

Scale that up. At 500 engineers running four such tasks a working day, over 21 working days a month, the same work costs about $431,000 a month on uncached Opus 5.5 and about $85,000 with caching. The model, the engineers and the output are identical in both cases; only the plumbing changed. 

Screenshot 2026 10 07 at 12.36.01 PM
The Frontier Model Bill: Why AI Costs Behave Like Infrastructure, and How to Engineer Them Down 43

The lesson for both finance and platform teams is the same. The useful question is what a task costs end to end, and which parts of that cost are avoidable. 

3. The metric that matters: cost per successful outcome

A token bill is not an ROI metric. The number that matters is what it costs to produce one usable result: a resolved ticket, a processed claim, a merged pull request. 

McKinsey makes the same point: the unit of governance should be the completed business outcome (McKinsey). Across more than 100 agent implementations, Lyzr has seen monthly frontier-model spend range from about $50,000 at smaller companies to about $4 million at a large consumer goods company. In most of those cases, nobody could say what an individual completed task cost. 

Cost per successful outcome counts everything a result consumes, from the first model call to the last correction. 

Screenshot 2026 10 07 at 1.11.16 PM
The Frontier Model Bill: Why AI Costs Behave Like Infrastructure, and How to Engineer Them Down 44

A task counts as a successful outcome once it produces a usable result, first time or after a person corrects it; the cost of that correction sits in the top line. 

Why reliability is a cost line 

Language models are probabilistic, and errors compound across steps. A model that is right 99.9% of the time on each step completes a 1,000-step process correctly only 36.8% of the time. 

Shorter workflows are not immune. At 99% accuracy per step, a 50-step workflow succeeds end to end about 60% of the time. Every failure becomes a retry, an escalation or a person fixing the output. 

Screenshot 2026 10 07 at 1.11.45 PM
The Frontier Model Bill: Why AI Costs Behave Like Infrastructure, and How to Engineer Them Down 45

Why the cheapest model can be the most expensive 

Consider two options for the same task, where a human corrects each failure at a cost of $15.  

Screenshot 2026 10 07 at 1.12.25 PM
The Frontier Model Bill: Why AI Costs Behave Like Infrastructure, and How to Engineer Them Down 46

Illustrative figures. The cheaper model costs a fifth as much per call and 70% more per result.

This cuts both ways. It argues against routing purely on token price. It also argues against the reflex of buying a bigger model whenever reliability falls short, because, as Section 4 shows, architecture can often raise reliability more cheaply than model size can. 

4. The cost stack: seven layers of savings 

Cost optimization means paying for the right amount of intelligence on each step, and the levers work best in a fixed order. 

Each layer below either removes spend or makes the next layer more effective. Visibility comes first because no other lever can be measured without it. 

Screenshot 2026 10 07 at 1.15.41 PM
The Frontier Model Bill: Why AI Costs Behave Like Infrastructure, and How to Engineer Them Down 47

Sources: McKinsey, Anthropic pricing, Lyzr ShadowLM materials. Savings do not simply add up; each layer acts on what the previous one leaves.

Layer 0: Visibility and attribution 

You cannot optimize what you cannot see. Finance needs to know who spent each dollar, on which agent, model, feature and environment, and what the work achieved. 

An invoice from a model provider answers none of those questions. Neither does a dashboard that only reports spend after the fact. 

The control that matters is a budget that refuses. A spend alert is a notification; a ceiling that rejects calls once a team or agent has spent its allocation is a control. Uber’s experience shows the difference: the budget was only discovered to be gone after it was gone. 

Screenshot 2026 10 07 at 1.16.43 PM
The Frontier Model Bill: Why AI Costs Behave Like Infrastructure, and How to Engineer Them Down 48

Layer 1: Don’t call a model 

The cheapest inference is the one that never happens. Much of what agents do inside a sophisticated workflow is lookup, formatting or a fixed rule. 

Take a customer email asking for a refund. It looks like one task but is really five: 

Screenshot 2026 10 07 at 1.17.05 PM
The Frontier Model Bill: Why AI Costs Behave Like Infrastructure, and How to Engineer Them Down 49

Steps 2 and 3 are a database lookup and a fixed rule; they need no model at all. Steps 1 and 5 suit a small model. Only step 4 needs frontier-level judgment. 

Caching does the same for repeated work. Exact-match caching returns a stored answer for an identical request; semantic caching does so for requests that mean the same thing. Both suit support questions, classifications and recurring enrichment jobs. 

Hit rates depend on how repetitive the traffic is. A support queue built around a stable set of questions will hit the cache far more often than open-ended research work. Measure the hit rate on a sample of real traffic before sizing the saving, and evaluate semantic matches for quality, because two requests that look alike can need different answers. 

Layer 2: Call a smaller model 

Most tasks do not need the most capable model, but most teams send them there by default. McKinsey notes that users default to premium models because the trade-offs between quality, cost and latency are unclear to them. 

A routing layer classifies each request and sends it to the cheapest model that clears the quality bar. On Anthropic’s current list prices, Haiku 4.5 costs a quarter of Opus 5.5 and a tenth of Fable 5.1 per token. 

Routing is most powerful when the infrastructure is model-agnostic. If a workflow can move between Claude, GPT, Gemini and open-weight models without being rebuilt, each step can be placed on cost, quality and latency rather than habit. 

The arithmetic is simple. If 70% of the steps in a workflow can run on Haiku 4.5 and the remaining 30% stay on Opus 5.5, the blended price falls from $4.00 to $1.90 per million input tokens, a saving of about 52% on both input and output, provided the cheaper model clears the same quality bar on its steps. 

Screenshot 2026 10 07 at 1.17.28 PM
The Frontier Model Bill: Why AI Costs Behave Like Infrastructure, and How to Engineer Them Down 50

For platform teams: define a quality bar per task class, test cheaper models against it on real traffic, and route by class. Keep a fallback to the stronger model for task classes that fail the quality bar. 

Layer 3: Send less 

In agentic work, input dominates cost, so the biggest savings come from what is resent on every call. The worked example in Section 2 cut cost per task by 76% to 80% through prompt caching alone. 

Anthropic charges a tenth of the standard input price for cached reads on most models, 5% on Opus 5.5 and 2.5% on Fable 5.1 (Anthropic). Writing to the cache costs 1.25x, so it pays back after a single reuse. 

Screenshot 2026 10 07 at 1.17.51 PM
The Frontier Model Bill: Why AI Costs Behave Like Infrastructure, and How to Engineer Them Down 51

Beyond caching, the practical rules are simple: 

  • Summarize older turns instead of resending full transcripts. 
  • Retrieve only the document sections the step needs. 
  • Pass structured state between steps, not raw logs. 
  • Load only the tools a step can use; large tool definitions are resent on every call. 
  • Specify output formats and cap response length by use case. 

Layer 4: Pay less per call 

Not every workload needs an answer now. Overnight document processing, invoice classification, report generation and data enrichment can run asynchronously. 

Anthropic’s Batch API charges half the standard rate on both input and output, and the discount stacks with prompt caching (Anthropic). If the requirement is “ready by tomorrow morning,” paying for synchronous execution is waste. 

Screenshot 2026 10 07 at 1.18.08 PM
The Frontier Model Bill: Why AI Costs Behave Like Infrastructure, and How to Engineer Them Down 52

Layer 5: Architect for reliability 

When a workflow is not reliable enough, the usual answer is a bigger model. The Lyzr Six Sigma (6σ) Agent architecture takes the opposite approach: assume individual models will fail, and design the system so failures are caught. 

An Architect agent breaks the task into a dependency tree of atomic steps. For each step, several small micro-agents answer independently, and a voting coordinator accepts the majority. If the vote is contested, more agents are added until confidence is high. 

Screenshot 2026 10 07 at 1.18.24 PM
The Frontier Model Bill: Why AI Costs Behave Like Infrastructure, and How to Engineer Them Down 53

The math is binomial. If each micro-agent is wrong 5% of the time and errors are independent, five voters fail together on about 0.12% of steps. Thirteen voters reach Six Sigma territory, under 3.4 defects per million steps. 

Screenshot 2026 10 07 at 1.18.44 PM
The Frontier Model Bill: Why AI Costs Behave Like Infrastructure, and How to Engineer Them Down 54

Lyzr calculation. Assumes voter errors are independent. The single model is assumed to be stronger (1% error) than each voter (5% error), to show weaker voters outperforming a stronger model. 

That independence assumption matters. Five copies of the same model with the same prompt tend to make the same mistakes. In practice, independence comes from varying the models, prompts or retrieved context across voters, and from measuring agreement on real traffic. 

Does it save money? It depends on the price ratio. Lyzr’s published billing-dispute simulation, run on November 2025 pricing, compared one GPT-5.1 agent with five GPT-5 nano voters on a 50-step case: $0.88 against $0.18 per case. That rested on a 25:1 price gap between the two models. 

Screenshot 2026 10 07 at 1.26.29 PM
The Frontier Model Bill: Why AI Costs Behave Like Infrastructure, and How to Engineer Them Down 55

Within Anthropic’s current line-up, the gap is narrower. Running the same case (500,000 input and 25,000 output tokens) on list prices gives a different picture. 

Screenshot 2026 10 07 at 1.26.44 PM
The Frontier Model Bill: Why AI Costs Behave Like Infrastructure, and How to Engineer Them Down 56

Voting on every step with Haiku halves the cost of Fable 5.1 but costs more than a single Opus 5.5. The architecture pays off when it is used selectively: 

  • Vote only where confidence matters. Dynamic escalation starts with five voters and adds eight more when the vote is split. At a 5% voter error rate, escalating only 3–2 splits averages about 5.2 voters per step and leaves about 31 errors per million steps. Escalating every split vote reaches Six Sigma at about 6.8 voters per step. 
  • Take deterministic steps out of the model entirely, as in Layer 1. 
  • Share context across voters through caching, so the same document is not paid for five times. 
  • Use cheaper or owned voters. Open-weight models, or small models tuned through Layer 6, widen the price gap again. A tuned voter with a 1% error rate needs only seven agents for Six Sigma. 

The design principle is to spend extra compute only where the system needs more confidence. 

Layer 6: Own the model 

Some workloads stop needing a frontier model once they are understood. A compliance review, risk classification or support workflow becomes predictable after enough repetitions, and then paying frontier rates on every request no longer makes sense. 

ShadowLM is live today as part of Lyzr Nitro and as an open-source training SDK from Lyzr Research Labs, and runs inside Opencontroller. It moves one task at a time from a rented frontier model to a small open-weight model the enterprise owns. The agent keeps running throughout; only the model behind it changes. 

Screenshot 2026 10 07 at 1.27.48 PM
The Frontier Model Bill: Why AI Costs Behave Like Infrastructure, and How to Engineer Them Down 57

An evaluation gate sits between stages. A task advances only when quality holds against the frontier model on held-out real traffic and the savings exceed the cost. Every switch is reversible at the gateway. 

Lyzr’s design targets are up to 90% lower inference cost on migrated tasks, with up to 90% of day-to-day workloads eventually running on owned models. Data stays inside the enterprise’s environment, and the tuned weights belong to the enterprise. 

Two cautions keep the business case honest. The 90% figure covers inference only; training, hosting and the human review behind the training data are real costs. And training on captured prompts needs explicit consent, a retention period and a deletion path, because a tuned model cannot be un-trained. 

A migration pilot should report four numbers against the frontier baseline, measured on held-out real traffic: task accuracy, the share of outputs a reviewer corrects, p95 latency, and fully loaded cost per 1,000 requests including hosting. If any of the first three moves the wrong way, the task stays on the frontier model. 

5. Coding agents: the special case 

Coding agents are where frontier spend is rising fastest, and where blunt cost controls do the most damage. The goal is governed usage. 

They behave differently from chat. A request to fix one bug can trigger a long loop: search the repository, read dependencies, edit code, run tests, read failures and try again. That loop is exactly the pattern Section 2 showed: large, growing context resent on every turn. 

The spend also sits with individuals. Engineers who do not see the invoice reasonably treat the tool as free, as Uber’s COO observed (Quartz). And costs vary widely per person, from an average of $150 to $250 a month to as much as $2,000 for Uber’s power users. 

The early responses have been blunt. Uber introduced monthly limits per tool (Outlook Business). Microsoft moved engineers in its Experiences + Devices division from Claude Code to GitHub Copilot CLI by June 30, 2026, while keeping Claude models available inside Copilot CLI (The Verge, May 2026). Flat caps and tool swaps cut spend per tool or per person rather than per task, so they ration the valuable work along with the waste. 

What a coding-agent control layer needs 

Engineering leaders need the same layered controls as the rest of the estate, adapted to a fleet of developer machines: 

Screenshot 2026 10 07 at 1.29.19 PM
The Frontier Model Bill: Why AI Costs Behave Like Infrastructure, and How to Engineer Them Down 58

BaseCode is Lyzr’s control plane for coding agents. It manages Claude Code, Codex, Copilot, Gemini and opencode from one place, without replacing the tools engineers already use. 

Every request flows through the BaseCode gateway and is logged with tokens, latency and cost per person and project. Administrators can remap any agent’s Opus, Sonnet or Haiku model slots to another provider, hosted or self-hosted, behind scoped virtual keys. 

A Complexity Router sends simple edits to cheaper models and hard reasoning to frontier models. Budgets and token rate limits apply per group, model or employee, with hourly, daily or monthly resets, enforced at the gateway. 

The effect is that cost control stops being a policy document and becomes part of how the tools run. Engineers keep the agents they prefer; the organization decides which model each task deserves and what each team can spend. 

6. The spend you can’t see 

Part of every AI bill never reaches the AI budget. It sits on personal API keys, in business-unit purchases and inside agents nobody formally provisioned. 

Screenshot 2026 10 07 at 1.29.40 PM
The Frontier Model Bill: Why AI Costs Behave Like Infrastructure, and How to Engineer Them Down 59

Menlo Ventures estimates that product-led adoption drives 27% of enterprise AI spend, and that shadow AI pushes the real figure closer to 40% (Menlo Ventures). McKinsey warns that citizen developers can unintentionally create agents that consume millions of tokens a day (McKinsey). 

Screenshot 2026 10 07 at 1.29.51 PM
The Frontier Model Bill: Why AI Costs Behave Like Infrastructure, and How to Engineer Them Down 60

The typical pattern is unremarkable. An engineer connects an open-source agent framework to a personal API key, points it at internal systems and schedules it to run. It works, so nobody complains, and its spend lands on an expense report or an individual account rather than a cost center. 

Shadow agents are a security problem as well as a cost problem. In a Cloud Security Alliance survey of 418 IT and security professionals, 82% of organizations had discovered previously unknown AI agents in their environments in the past year (CSA). IBM’s 2025 Cost of a Data Breach report puts the cost premium for breaches at organizations with high levels of shadow AI at about $670,000 (IBM). 

Screenshot 2026 10 07 at 2.01.49 PM
The Frontier Model Bill: Why AI Costs Behave Like Infrastructure, and How to Engineer Them Down 61

Sources: Cloud Security Alliance, survey of 418 IT and security professionals; IBM Cost of a Data Breach 2025

Finding this spend requires continuous discovery rather than surveys. That means scanning cloud audit logs for model calls from unregistered workloads, inspecting clusters and developer machines for agent frameworks and stored provider keys, and flagging traffic that bypasses the approved gateway. 

One useful signal is spend with no registered owner: model consumption that cannot be traced to any known agent. It is usually the first place a discovery exercise finds money. 

7. Why vendor, observability and gateway tools are not enough 

Model providers’ own controls are improving, and enterprises should use them. But they govern one vendor’s traffic, and enterprise AI rarely lives inside one vendor. 

Anthropic’s Console tracks usage, and its Admin API exposes usage and cost data for organizations (Anthropic). That covers Claude traffic billed to that organization. It cannot see a coding agent on a separate key, a SaaS product with an embedded model, or a team’s agent running on another provider. 

It also cannot route. A vendor console can restrict which of its own models a group may use. It cannot send a step to a cheaper provider or to an enterprise-owned model, which is where Layers 2 and 6 find their savings. 

Observability tools have the opposite gap. They see across providers and report spend in detail, but most of them record a call after it has happened. A dashboard that shows an overrun is not the same as a budget that stopped it. 

AI gateways close part of that gap. They sit in the request path, so they can enforce budgets and route between providers. But they see only the traffic sent through them, so agents on personal keys stay invisible; they report cost per call rather than per outcome; and they route to models without producing the owned models that Layer 6 depends on. 

McKinsey’s recommendation is a centralized AI control plane: a management layer between users, applications, agents and models that provides visibility, enforces policy in real time and routes workloads by cost, quality and risk (McKinsey). It also advises keeping the ability to shift workloads between proprietary and open-weight models. 

Screenshot 2026 10 07 at 1.30.40 PM
The Frontier Model Bill: Why AI Costs Behave Like Infrastructure, and How to Engineer Them Down 62

8. A 90-day roadmap 

The first savings can start within two weeks, because the first move is a configuration change rather than a migration. 

Screenshot 2026 10 07 at 1.41.21 PM
The Frontier Model Bill: Why AI Costs Behave Like Infrastructure, and How to Engineer Them Down 63

Each phase funds the next. Attribution shows where caching and routing will save most, and those savings pay for the architectural work. Phase 3 puts the first task into shadow mode; each task that later moves to an owned model lowers the baseline for good. 

9. Reference architecture 

Lyzr maps onto the cost stack through two control layers: Opencontroller by Lyzr for agents and model calls across the estate, and BaseCode for coding agents on developer machines. 

Screenshot 2026 10 07 at 1.41.53 PM
The Frontier Model Bill: Why AI Costs Behave Like Infrastructure, and How to Engineer Them Down 64

Reference architecture · BaseCode and Opencontroller 

Opencontroller is additive rather than a migration. Model access is an OpenAI-compatible URL, so adoption means changing a base URL and a credential, not rewriting applications. It runs inside the customer’s own cloud, with about 11 ms of measured gateway overhead. 

Its budgets refuse calls once a scope has spent its allocation, rather than sending an alert. Its observability tracks cost and task success for every agent, the two inputs to cost per successful outcome. The ShadowLM module registers each tuned model in the gateway as an ordinary model, so migrating one agent is a reversible routing change. 

Two layers sit outside the gateways. Batch processing (Layer 4) is a provider discount that the gateway can route to. The Six Sigma architecture (Layer 5) is a way of building agents, which then run through the same gateway, budgets and traces as any other agent. 

The savings are already showing up in the market. In McKinsey’s May 2026 Enterprise AI FinOps survey, about a third of the organizations surveyed had already cut AI costs by 20% to 30% through active optimization (McKinsey). 

Where to start 

The usual first step is a two-week discovery. Read-only connectors scan an organization’s clouds, clusters and devices, and report how many agents it actually has, what they can reach and what they cost. For most organizations, that number makes the case for everything else in this paper. 

Appendix A: Claude pricing multipliers 

List prices per model are in Section 2. These multipliers change what a workload actually costs. 

Screenshot 2026 10 07 at 1.45.42 PM
The Frontier Model Bill: Why AI Costs Behave Like Infrastructure, and How to Engineer Them Down 65

Source: Anthropic pricing documentation, as of October 2026. 

Appendix B: Sector spend methodology 

The 756 companies were drawn from Lyzr’s list of enterprise and mid-market target accounts across nine sectors. The sector estimates in Section 1 are modeled, order-of-magnitude figures; companies do not disclose AI contract values, so these are not reported numbers. 

Each company’s estimate is built in five steps: 

  1. Representative revenue is taken from the midpoint of the company’s revenue band. 
  1. Employees are estimated as revenue divided by the sector’s typical revenue per employee. 
  1. AI-eligible seats are employees multiplied by the sector’s share of desk-based workers. 
  1. Active AI seats are eligible seats multiplied by the sector’s AI seat adoption rate, from 15% in manufacturing to 45% in software and consulting. 
  1. Spend is active seats multiplied by a blended $600 to $800 per seat per year, depending on deployment scale. 

The blended seat price reflects mid-2026 public benchmarks for Claude and ChatGPT enterprise plans. The model covers internal employee-productivity spend on Claude and ChatGPT only. It excludes API and product spend, which can be far larger for AI-native firms, and excludes other vendors such as Copilot and Gemini. 

Telecom, travel, hospitality and outsourced customer-support companies are reported as one sector. The source model split them by size, with large carriers, airlines, hotel groups and BPOs in one group and mid-market telecom and travel companies in another; the two groups are combined here so that the sector reflects one industry. 

The largest revenue band (above $50 billion) spans a very wide range, so figures for the biggest companies are the roughest. Sector averages also reflect the mix of companies in the sample. 

Appendix C: AI cost maturity scorecard 

Score your organization from 1 to 4 on each row, then add the scores. A total of 7 to 13 suggests reactive cost management; 14 to 20, visible but not controlled; 21 to 28, managed and optimizing. 

Screenshot 2026 10 07 at 1.46.33 PM
The Frontier Model Bill: Why AI Costs Behave Like Infrastructure, and How to Engineer Them Down 66

Sources 

  • Lyzr, Lyzr Six Sigma (6σ) Agent: Achieving Zero-Error AI for Enterprise Operations 
  • Lyzr, ShadowLM: The model layer, finally yours 
  • Lyzr, Opencontroller Capability Reference and BaseCode product brief 
Build with Lyzr

Try it in
Agent Studio
today.

From framework-agnostic design to production-grade agents, deployed in under 24 hours.