Every few months another team tells me the same story. They built an agent on a frontier model, it worked beautifully in the demo, and then it fell apart in production: too slow, too expensive, too unpredictable to ship.
The postmortem almost always lands on the model:
- “We need to wait for a faster frontier tier.”
- “We need a bigger context window.”
I think that diagnosis is wrong almost every time.
The model tier is a red herring. The real production-readiness lever is how aggressively you’ve decomposed the agentic workflow into deterministic code, leaving only narrow, well-scoped decision points for the LLM to actually reason about.
When a team fails with a cheap model, in my experience they didn’t fail because the model was too small. They failed because they handed it a job that was never a single, well-scoped decision in the first place. No model, however expensive, was going to make that job’s underlying ambiguity go away.
What Decomposition Changed in My Own Pipeline
I’ve hit this wall myself. I had a product that was flatly infeasible on a frontier-only architecture. Several sequential model calls per request pushed end-to-end latency past ten seconds. Cost scaled linearly with volume in a way the business couldn’t absorb. And roughly one in five outputs came back malformed or off-task, unpredictable enough that I couldn’t confidently ship it to a client.
The fix wasn’t a different model. It was ripping the pipeline apart, replacing most of what I’d been asking the LLM to “figure out” with plain coded logic, and reserving the model’s judgment for the handful of places where judgment was actually required.
Here is how the same request compared before and after that rework.
| Measure | Frontier-only architecture | After decomposition |
| End-to-end latency | Past ten seconds, driven by several sequential model calls | Roughly a third of the original |
| Token spend | Scaled linearly with volume, beyond what the business could absorb | A fraction of the original |
| Malformed or off-task outputs | Roughly one in five | A small single-digit percentage |
| Ready for a client | No, too unpredictable to ship | Yes, the fallback path caught what was left before it reached a client |

What follows is the framework I now use to decide where that line sits, and what has to be true around it for the savings to be real rather than deferred risk.
Cheap vs. Frontier: The Two Tiers This Framework Uses
Before the framework itself, it’s worth being specific about what “cheap” and “frontier” mean, because that split matters more than the framework.
This table shows the models I place in each tier.
| Tier | What it means | Examples |
| Cheap | The small, fast, low-cost tier every major lab now ships | Claude Haiku, GPT-4o mini or GPT-5 mini, Gemini Flash, and increasingly a locally hosted open-weight model for the highest-volume nodes |
| Frontier | The top reasoning tier | Claude Opus, or Claude Sonnet at its higher reasoning settings, GPT-5 or o-series reasoning models, Gemini’s Pro/Ultra tier |
You Choose a Model per Node, Not per App
This framework covers both kinds of jump. The size of the drop depends on the node, not on a blanket rule for the whole pipeline.
This table shows how far a node can step down, and why.
| Type of step-down | Example | When it fits |
| Three rungs | Opus-level model down to Haiku-level | The task never needed that much reasoning in the first place |
| One rung | Opus-class model down to Sonnet-class | The task still needs real reasoning, just not the most expensive version of it |
That’s the part people miss. You’re not choosing one tier for the app. You’re choosing a tier per node, and the gap between tiers can be one rung or three depending on what that specific node is doing.
Three Questions That Decide Each Node’s Model Tier

When I look at a pipeline node and decide whether it can run cheap or needs real reasoning power, I ask three questions, in order.
1. Is the Input Bounded?
If the node’s input space is genuinely open, it’s a frontier-reasoning job. That includes free-form user intent, ambiguous multi-domain requests, and anything where the “correct” interpretation depends on context the model has to infer.
If the input is schema-bound, single-purpose, and structurally similar every time, it’s a strong candidate for downgrading. Examples:
- Classify this ticket into one of eight categories.
- Extract these five fields from this document.
- Decide whether this response satisfies this rubric.
This matches what Anthropic’s own engineering team argues in its guidance on building effective agents: start with the simplest solution possible, and only keep a full agentic loop where the model genuinely must direct its own tool use.
2. Is the Decision Verifiable?
Can I write a check that tells me with confidence whether the output was correct, before it goes anywhere downstream? That check can be a schema, a regex, a business rule, or a second call.
- Nodes I can verify cheaply and deterministically are safe to downgrade. The verification layer catches what the smaller model gets wrong.
- Nodes where “correct” is a judgment call need the stronger model. No automated check can confirm the output, so there’s no safety net underneath it.
3. What Is the Blast Radius of a Wrong Answer?
Errors don’t all cost the same.
- A wrong classification that gets caught by validation and retried costs me a few hundred milliseconds.
- A wrong answer that silently propagates into a client-facing decision, a financial calculation, or an irreversible action costs a lot more.
I downgrade aggressively where errors are cheap and reversible. I keep frontier reasoning, or at minimum tighter human-in-the-loop review, wherever an error is expensive or hard to undo.
The Three Questions at a Glance
This table summarizes how each answer points a node toward a tier.
| Question | Downgrade toward the cheap tier when… | Keep frontier reasoning when… |
| Is the input bounded? | Input is schema-bound, single-purpose, and structurally similar every time | Input is open: free-form intent, ambiguous multi-domain requests, context the model must infer |
| Is the decision verifiable? | A schema, regex, business rule, or second call can confirm the output before it moves downstream | “Correct” is a judgment call that no automated check can confirm |
| What is the blast radius? | Errors are cheap and reversible, caught by validation and retried in a few hundred milliseconds | Errors propagate silently into client-facing decisions, financial calculations, or irreversible actions (or add tighter human-in-the-loop review) |
This lines up with independent framing I’ve seen elsewhere in the field: node classification should weigh whether a task is schema-bound and repetitive versus ambiguous and cross-domain, where errors compound silently rather than surfacing immediately.
I’d also add a fourth consideration specific to system design. The orchestrator, which decides how to decompose the task and holds the overall plan, is a different animal from the leaf nodes executing individual steps. I don’t downgrade that layer lightly.
Keep the Orchestrator Strong and Push the Workers Cheap

This is the objection I get most from engineers who know this space:
“Fine, but doesn’t the router or orchestrator layer itself need to be smart? Isn’t that just moving the problem?”
My answer is yes, and I don’t try to downgrade it. In my own architecture, I keep a clear split between two layers.
This table shows what each layer does and which tier it runs on.
| Layer | What it does | Model tier |
| Orchestrator | Decomposes the task, holds the shared plan and context, and decides how work gets distributed | Stays on a stronger model, because this is open-ended, cross-domain reasoning that needs frontier capability |
| Worker nodes | Each executes one narrow, bounded subtask | Pushed as hard as possible toward Haiku or GPT mini class models |
Why the Split Captures Most of the Savings
This isn’t just my own preference. It’s the architecture that shows up consistently once you look at where the money actually goes.
- Leaf nodes carry most of the spend. One cost breakdown I’ve seen puts execution and leaf-task tokens at roughly 70 to 85 percent of total spend in a typical agentic pipeline. That’s exactly why downgrading the leaf nodes captures most of the available savings, while the orchestrator can stay on the expensive model without meaningfully denting your cost curve.
- Heterogeneous systems perform better. NVIDIA’s research on small language models in agentic systems makes the same architectural case from a different angle. Systems that push routine subtasks to small models and reserve large models for genuinely open-ended reasoning reduce latency and cost while improving reliability, precisely because you’re not paying frontier prices for work that never needed frontier reasoning.
Map the Nodes First, Assign Tiers Last
The practical implication: don’t think of it as “choosing a model for the app.” Think of it as choosing a model per node, with the decomposition itself as the design decision that determines everything downstream.
A separate framework calls this the “minimum viable model” approach: structured per-node criteria around task narrowness, verifiability, and failure cost, rather than one model choice for the whole system.
That’s the mental model I’d hand to any founder or FDE starting this exercise:
- Map your pipeline into nodes.
- Classify each node against the three questions above.
- Only then start assigning model tiers.
Error Handling Is What Makes the Savings Real
“Proper error handling” is the part of this argument that gets hand-waved most often. It’s also the part that actually determines whether you’ve saved money or just deferred risk to production. Here’s what it concretely looks like in the systems I’ve built.
Wrap Every Model Call in a Sandwich
Every narrow AI touch point sits inside what I think of as a sandwich. This pattern has a name in the engineering community, the “sandwich architecture,” and it has become the standard shape of a production-safe agentic node for good reason.
This table shows the three layers of the sandwich.
| Layer | Type | What it does |
| Pre-step | Deterministic | Assembles exactly the context the model needs |
| Model call | Probabilistic | The model does its one narrow job |
| Post-step | Deterministic | Validates the output before anything downstream ever sees it |
Validate with Instructor and Reask on Failure
Concretely, that post-step is schema validation. I lean on a Python library called Instructor, which is built on top of Pydantic (Python’s standard data-validation library) and forces the model’s response into a strict, typed schema you define up front.
If the model’s output doesn’t match that schema, Instructor doesn’t just retry blindly. It feeds the specific validation error back into the prompt so the model can see exactly what it got wrong and self-correct. Engineers call this pattern “reasking,” and it’s a lot more effective than firing the same request at the model again and hoping.
If validation fails again after a couple of these reasking attempts, that’s not a signal to keep hammering the same model. It’s the trigger to escalate.
Escalate Up the Tier Instead of Retrying Forever
A layered escalation ladder I’ve found genuinely useful in production looks like this.
| Step | What happens |
| 1. Native schema enforcement | The first line of defense |
| 2. Instructor-style structured validation | Output is checked against the strict, typed schema |
| 3. Bounded healing loop | Two to three reasking attempts, each carrying the specific validation error |
| 4. Escalate up the tier | A frontier model with the same strict schema |
| 5. Relax the schema | Only if it comes to that |
| Throughout: circuit breaker | Stops calling a tier once it’s failing consistently, rather than burning tokens on a model that clearly can’t do the job right now |
Track Escalation Rate to Catch Silent Drift
The failure mode I watch for hardest isn’t the retry itself. It’s silent drift: a model quietly starts returning output in a slightly different shape than before.
If you’re not tracking escalation rate as its own metric, that drift can go unnoticed for a while, with more traffic falling back to the expensive tier than anyone budgeted for. So I treat escalation rate as an explicit operational metric, not an afterthought. If it creeps up, that’s a deploy-blocking signal, the same way a latency regression would be.
This is precisely the kind of production failure mode that doesn’t show up in a benchmark and only shows up once you’re running the thing for real.
Budget for the Cost of the Safety Net
One honest caveat worth naming: building this validation and escalation layer isn’t free.
This table shows where the cost of the safety net actually sits.
| Component | Cost profile |
| Rule-based routing | Adds under a millisecond |
| Embedding-based routing | Adds tens of milliseconds at most |
| Eval harness | Can cost more than the inference savings it was built to protect, if you’re not careful, because judge-model spend scales with every query, every candidate output, and every refresh cycle you run it on |
The fix isn’t to skip the eval layer. It’s to cascade the judge the same way you cascade the workers, and to budget engineering time for building and maintaining this infrastructure as a real cost of the savings, not a footnote.
The Model Tier Is a Detail, Not a Decision
None of this is an argument against frontier models. It’s an argument against defaulting to them out of caution instead of decomposing the problem first.
The teams I’ve seen get burned by “cheap” models weren’t burned by the model. They were burned by asking a narrow tool to do an open-ended job, with no validation layer to catch it when it inevitably guessed wrong.
The playbook, in short:
- Decompose the workflow into deterministic code and narrow decision points.
- Classify each node against reversibility, verifiability, and blast radius.
- Keep the orchestrator honest and the leaf nodes lean.
- Wrap every model call in error handling that actually escalates rather than just retries blindly.
Do that, and the model tier becomes what it should have been from the start: a detail, not a decision.
Book A Demo: Click Here
Join our Slack: Click Here
Link to our GitHub: Click Here
