All posts
AI Agents

Benchmarking Jev as an AI Router and Controller

Akshat Kumar
Akshat Kumar
Oct 6, 2026
16 min read
Benchmarking Jev as an AI Router and Controller

Can a model that never writes a single word be the brain of an AI agent? We made 21,314 calls to find out, and the whole experiment cost 72 cents.

The short version

If you read nothing else: Jev fits agentic RAG well, and it also did better at reranking than we expected.

  • Controller: Jev’s “is this context sufficient?” signal scored 0.90 AUROC on both single-passage and multi-hop tasks. Specialised local models drop to 0.69 on multi-hop.
  • Router: Jev was the best zero-shot router on all three tasks (0.80–0.91 accuracy) and the most consistently well calibrated. With labelled data, a small embedding classifier catches up: 10 examples per intent already beat Jev on Banking77 (0.875 vs 0.801).
  • Reranker: Jev raised NDCG@10 over retrieval alone by 6 points (SciFact) and 9 points (FiQA). bge-reranker-v2-m3 managed 0 and 3.
  • Cost: the full benchmark cost $0.72 of Jev credit. Decisions cost 0.02–0.07 per 1,000 calls, at about 350 ms median latency.

The problem: RAG is a chain of small decisions

We usually describe a RAG system as “retrieve, then generate”. In practice the answer quality is set by four small decisions made before the LLM writes anything:

  1. Route: does this message need retrieval at all, and if so, which kind? A poem request, a pasted document and a multi-step comparison all need different handling.
  2. Select: which knowledge base or index should be searched?
  3. Filter: of the 20 chunks retrieved, which ones belong in the prompt?
  4. Stop or continue: is what we have enough to answer, or should the agent search again?

Get any one of these wrong and the LLM is answering from the wrong evidence, confidently. Today, teams make these decisions with one of four tools, and each has a catch:

ToolHow it decidesThe catch
Hand-written rulesKeywords, regexes, if-statementsBrittle; breaks on phrasing it hasn’t seen
Embedding similarityNearest route or chunk by vector distanceMatches topic, not intent; can’t tell “related” from “answers it”
Small trained models (cross-encoders, classifiers)Trained for one task, e.g. MS MARCO relevanceStrong in their own domain, fragile outside it; need labelled data
An LLM callPrompt an LLM to decide, parse its textSlow (seconds), costly at volume, output must be parsed and can drift

A decision model like Jev pitches itself as a fifth option: an LLM’s flexibility, but with typed answers, probabilities and a price close to zero. That pitch is what we wanted to test.

Before you read: what you need to know

You don’t need to have run a benchmark before. If you know what RAG is, these two tables cover everything else in this article.

The concepts

TermWhat it means here
RAGRetrieval-augmented generation: fetch relevant text from a knowledge base, then give it to an LLM to answer from
Agentic RAGRAG where an agent loops: it can choose a strategy, search more than once, and decide when it has enough
ChunkA passage of a document, the unit that gets retrieved and put in the prompt
Dense retrievalFinding chunks by comparing embedding vectors of the query and the chunks
RerankerA second-stage model that re-sorts the retrieved chunks by true relevance
Cross-encoderThe usual reranker: reads the query and one chunk together and outputs a relevance score
RouterPicks which path a query takes: a tool, an index, or a strategy
ControllerDecides the agent’s next step, here: answer now or retrieve more
Zero-shotThe model gets no labelled examples, only a description of the task
Multi-hop questionNeeds facts from two or more documents combined, e.g. “Was the director of film X born before Y?”

The metrics

MetricUsed forHow to read it
NDCG@10Reranking0–1. How close the top 10 chunks are to the perfect order. +0.03 is a meaningful gain
MRR@10Reranking0–1. How high the first relevant chunk sits. 0.5 means it’s typically at position 2
Recall@5RerankingShare of relevant chunks that made the top 5, i.e. that would reach the LLM
AccuracyRoutingShare of queries sent to the right route
ECE (calibration error)Routing0 is perfect. How far a model’s stated confidence is from how often it’s actually right. 0.05 means “90% sure” really is about 85–95% right
Accuracy at 80% coverageRoutingAccuracy when the model handles only its most confident 80% and hands the rest to a fallback
AUROCController0.5 is a coin flip, 1.0 is perfect. How well the model’s score separates “enough evidence” from “not enough”, at any threshold

Meet Jev: a model that only decides

Jev is what TypeSafe calls a System One model, named after the fast, intuitive half of human thinking. You give it a state (any text or JSON) and some typed questions. It returns a probability for every possible answer and never generates prose. It went public in September 2026 at $0.042 per million input tokens, with output tokens free.

It answers three kinds of question:

TypeAsksReturnsOur use
noulA yes/no questionP(yes), 0–1“Is this chunk relevant?”, “Is this context sufficient?”
choicePick one of up to 255 named optionsThe choice + a probability per option“Which route?”, “Which index?”
scoreGrade against a 2–10 level rubricThe expected score + a probability per level“Rate relevance 0–3”

Here is a real call from our routing benchmark. We asked Jev how to handle a HotpotQA question:

image 32
Benchmarking Jev as an AI Router and Controller 4

And its answer, 494 input tokens later:

image 31
Benchmarking Jev as an AI Router and Controller 5

Where Jev could sit in a RAG pipeline

Here is the agentic loop we had in mind, with the four decisions from the problem section marked. Each highlighted step is a single Jev call, and each carries the score Jev got on that job in this benchmark. The rest of this article explains where those numbers come from.

The loop’s most important arrow is “no: search again”. Without a reliable step 4, an agent either stops too early and answers from half the evidence, or keeps searching and burns time and money.

Three questions we set out to answer

  1. Can Jev rerank? Jithin’s bet was no, since dedicated rerankers are trained on millions of relevance judgements and Jev was not. Akshat’s bet was that reranking is just scoring, and scoring is what Jev does.
  2. Can Jev route without training data? A router has to work from day one, before anyone has labelled a single query. Its confidence also has to be honest, or the fallback logic is useless.
  3. Can Jev tell an agent when to stop searching? This is the hardest one. Spotting that a relevant-looking passage doesn’t answer the question, or that a second piece of evidence is missing, is where RAG systems quietly fail.

Before you scroll on, make a guess. Which of the three do you think Jev did best at? Routing looks like the obvious answer. Hold that thought until Round 3.

What we tested

Jev (jev-1.13.0, TypeSafe AI) was tested zero-shot on six public tasks covering the three jobs in Jithin and Akshat’s thread. Every baseline is an open model run locally on an Apple M5 GPU. No LLM baseline was included.

JobTaskDataBaselines
RerankerRerank the top 20 dense-retrieved chunksBEIR SciFact and FiQA, 200 queries eachbge-base retrieval only; MiniLM-L6 cross-encoder; bge-reranker-v2-m3
RouterIntent routing, 77 classesBanking77, 770 queriesEmbedding zero-shot; DeBERTa NLI zero-shot; embedding classifier with 10 examples per class; logistic regression on the full training set
RouterRAG strategy: no retrieval, pasted context, single-hop, multi-hop, math tool600 queries from Dolly, NQ, HotpotQA, GSM8KSame as above
RouterKnowledge-base selector, 5 indexes360 BEIR queries (SciFact, FiQA, NFCorpus, TREC-COVID, NQ)Same as above
ControllerCan the agent answer from this passage?SQuAD v2, 800 (half unanswerable)roberta-base-squad2; MiniLM cross-encoder; DeBERTa NLI
ControllerAnswer now, or retrieve another hop?HotpotQA, 800 paired cases, one hop removed in halfSame as above

Jev reranking was tried three ways: one yes/no call per chunk, a 0–3 relevance score per chunk, and one call holding all 20 chunks with a yes/no question per chunk. Every Jev call is cached and logged, and the full suite cost under $1 of the $5 budget.

How we ran it

Everything ran on one laptop. The local models used its GPU, and Jev was called over its public REST API. The full suite took about 2 hours of wall-clock time, almost all of it local baselines. Each Jev pass took 1–3 minutes per task.

ComponentDetail
MachineApple M5 laptop, 10 CPU cores, 24 GB unified memory, macOS 27.0
Local inferencePyTorch 2.14.1 on the Apple GPU (MPS backend), fp32, no quantisation
SoftwarePython 3.12 (uv), sentence-transformers 6.1.0, transformers 5.18.0, datasets 5.0.1, scikit-learn 1.9.1
Jev endpointPOST api.typesafe.ai/v1/systemone, model jev-latest (served as jev-1.13.0)
Jev client16 parallel requests (8 for the 20-chunks-per-call variant), retry with back-off on 429/529, SQLite response cache, hard stop at $3.50 spend
Jev volume21,314 calls, 17.2M input tokens, 0.37 s mean latency, 3 failed requests, all succeeded on re-run

Steps, in order:

  1. Created a Python 3.12 project with uv and installed the libraries above.
  2. Downloaded every dataset from Hugging Face and drew the samples with a fixed seed (0), so every system saw identical inputs.
  3. Embedded the SciFact (5.2k docs) and FiQA (57.6k docs) corpora with bge-base-en-v1.5 and kept the top 20 chunks per query by cosine similarity.
  4. Ran the local baselines on the GPU: the two cross-encoders, the DeBERTa NLI classifier, roberta-squad2, and the embedding classifiers. Chunks were cut to 350 words, or 250 in the 20-chunks-per-call Jev variant.
  5. Sent the same inputs to Jev. Every response was cached, so re-running any script costs nothing.
  6. Scored everything in Python. Where a threshold was needed, it was tuned on one half of the data and tested on the other, then the halves were swapped and the two results averaged.
image 33
Benchmarking Jev as an AI Router and Controller 6

Two setup problems are worth knowing about. The local models competed for the one GPU when run in parallel, which slowed every baseline. Zero-shot NLI over 77 Banking77 labels also needed ~59k model passes and took about 20 minutes even batched.

Round 1: Jev as a reranker, and the result we didn’t expect

Jev beat both open rerankers on both datasets. It lifted NDCG@10 by 6 points over retrieval alone on SciFact and by 9 points on FiQA. The strongest open reranker, bge-reranker-v2-m3, gained 0 and 3 points.

SystemSciFact NDCG@10SciFact MRR@10FiQA NDCG@10FiQA MRR@10FiQA Recall@5
Jev, all 20 chunks in one call0.8360.8150.4760.5750.457
Jev, one yes/no call per chunk0.8250.7990.4780.5870.471
Jev, 0–3 score per chunk0.8210.7970.4730.5780.468
bge-base retrieval only0.7760.7440.3840.4720.374
bge-reranker-v2-m3 (568M)0.7740.7370.4180.5180.397
MiniLM-L6 cross-encoder (22M)0.7420.7020.3710.4470.356

The one-call version is the one to use. It matches per-chunk calls on quality at about half the tokens, because each Jev call carries roughly 300 tokens of fixed overhead.

Akshat’s “deterministic gate” also works: keep a chunk when Jev’s P(relevant) is 0.5 or more, and drop the rest.

  • SciFact: it keeps 1.7 chunks per query at 46% precision. bge-reranker keeps a similar number at 41%.
  • FiQA: it keeps 8–10 chunks and keeps 93–95% of the relevant ones. bge-reranker keeps 4 and loses 42% of the relevant ones.

Jev’s probabilities are rounded to two decimals, so near-ties are broken by retrieval order. That makes ranking within the top few chunks coarse.

Round 2: Jev as a router

Jev was the best zero-shot router on all three tasks by 11–51 accuracy points, with calibration error of 0.09 or less on every task, which no other system managed. A classifier trained on labelled examples still matched or beat it on accuracy.

SystemBanking77 accuracyRAG strategy accuracyKB selector accuracyCalibration error (ECE), range
Jev, zero-shot0.8010.8800.9140.04–0.09
Embedding zero-shot0.6880.3730.5750.06–0.25
DeBERTa NLI zero-shot0.5690.2850.6610.12–0.42
Embedding classifier, 10 labelled examples per class0.8750.7430.9030.12–0.45
Logistic regression, full labelled set0.9390.9030.9080.09–0.31

Calibration is the number that matters most for a router, because the confidence decides when to fall back to an LLM or a human. Routing only the 80% most confident queries lifts Jev’s accuracy to 0.87 on Banking77, 0.94 on RAG strategy and 0.97 on KB selection.

Jev’s main weakness is under-calling multi-hop. On the RAG strategy task it routed 52 of 120 HotpotQA multi-hop questions as single-hop, and 18 of 120 creative prompts as single-hop retrieval. Math, pasted context and single-hop were 359 of 360 correct. On KB selection, most errors were nutrition questions sent to the Wikipedia index (15 of 80).

Round 3: Jev as a controller, the decision that matters most

Jev is the strongest “answer now or retrieve more?” signal we tested. On multi-hop sufficiency it scores 0.90 AUROC, against 0.66–0.69 for every local model. Its default 0.5 cut-off is too eager to answer, though. Use about 0.8 instead.

SystemSQuAD v2 AUROCSQuAD v2 accuracy (tuned threshold)HotpotQA AUROCHotpotQA accuracy (tuned threshold)
Jev, yes/no: “is the context sufficient?”0.8980.8150.8970.822
Jev, choice: answer vs retrieve more0.8820.8190.8850.817
roberta-base-squad2 (trained on SQuAD v2)0.9320.8530.6870.649
MiniLM cross-encoder relevance0.6510.5990.6900.604
DeBERTa NLI zero-shot0.7360.6660.6600.630

The SQuAD-trained model wins only on its own dataset. It collapses on HotpotQA, where one paragraph can contain a plausible answer span while the second hop is missing. Jev holds the same 0.90 AUROC on both tasks. That is the property an agentic-RAG controller needs.

The threshold sets the trade-off between answering on missing evidence and spending a retrieval that wasn’t needed:

Jev P(sufficient) thresholdHotpotQA: answers with a hop missingHotpotQA: retrieves again needlesslySQuAD v2: answers unanswerableSQuAD v2: retrieves again needlessly
0.5 (default)31%7%39%6%
0.721%14%26%11%
0.8 (recommended)16%21%17%18%
0.99%38%9%34%

A plain yes/no question slightly beat (AUROC +0.01) framing the decision as an agent action (“answer” vs “retrieve_more”), so ask Jev about the evidence and keep the action logic in code.

Cost and latency

The whole benchmark used 17.2M Jev input tokens, which cost $0.72 of the $5 budget. Routing and controller decisions cost 0.02–0.07 per 1,000 calls, at about 340–380 ms median latency.

Jev useInput tokens per callCost per 1,000 callsp50 latencyp95 latency
Rerank 20 chunks, one call (SciFact)7,721$0.32480 ms595 ms
Rerank 20 chunks, one call (FiQA)4,644$0.20410 ms595 ms
Rerank 20 chunks, one call per chunk (SciFact)14,825$0.62432 ms693 ms
Router, 77 routes (Banking77)1,694$0.07342 ms594 ms
Router, 5 routes (RAG strategy)540$0.02366 ms536 ms
Controller, 1 passage (SQuAD v2)555$0.02381 ms725 ms
Controller, 4 paragraphs (HotpotQA)921$0.04336 ms418 ms

About 300 tokens of every call are fixed overhead, and each extra question adds only about 20 tokens. So the cheap pattern is one call that asks every question you need about the same state.

For comparison, bge-reranker-v2-m3 took 5.7–6.0 s per query on the laptop GPU while other jobs shared it. Per-chunk Jev latency was measured with 16 calls in parallel.

Observations and opinions

My read: Jev is best used as a judge of evidence, the controller in the loop, rather than as a drop-in model for any single stage. The table summarises what we saw and what to do about it. The three sections after it explain the points that need more context.

At a glance

ThemeWhat we sawWhat to do
Controller fitSame 0.90 AUROC on single-passage and multi-hop sufficiency. Every specialised model fell apart off its training distributionAdopt Jev as the controller first
RerankingWins by following instructions, not better relevance modelling. MiniLM made both SciFact and FiQA worse than retrieval aloneValidate on your own documents before shipping
Yes/no probabilitiesDecisive (42% at 0.10 or below, 11% at 0.90 or above) but lean towards yes: unanswerable questions were answered 31–39% of the time at 0.5Tune each threshold on about 100 labelled examples
Choice probabilitiesRouter calibration error stayed at 0.09 or below on every taskRoute if confident, else fall back to an LLM or human, with no tuning
Cold startWith 10 labelled examples per route, a 10 ms classifier was within 1 point of Jev on KB selection, 7 ahead on Banking77, 14 behind on RAG strategyLaunch on Jev, log its confident decisions as labels, move high-volume routes to a cheap classifier
Hidden structureMulti-hop questions read like single lookups; the second hop lives in the entitiesAsk a narrow question, e.g. “does answering need facts about two different entities?”
Call design~300 tokens of overhead plus ~20 per question, ~350 ms latency floorAsk every independent question about the same state in one call
PolicyAsking “is it sufficient?” worked as well as “answer or retrieve more?”Jev judges facts; code decides what happens next
Operations21,314 calls, 3 failures, all fine on re-run, under $1 totalCost is not a factor; latency and accuracy are

Why the controller is the best fit

The controller sees every kind of question, so it needs a signal that doesn’t depend on the domain. Jev’s sufficiency score held at 0.90 AUROC whether the task was a single passage or a multi-hop chain, while each specialised model was strong only on the data it was trained for. That consistency is what an agent loop needs.

Reranking: promising, but verify

Jev’s reranking win comes from following instructions. MS MARCO cross-encoders expect web-search queries, but SciFact queries are claims and FiQA queries are forum posts. Jev was told to look for evidence that supports or refutes the claim, and it did. I’d expect the gap to narrow on plain web-search queries.

Jithin’s instinct that a model not trained on relevance judgements would struggle was reasonable, but the data disagrees: reranking is a scoring problem, and Jev scores well. I still wouldn’t ship it as our reranker until it holds up on our own documents, since margins this large leave training-data overlap as a real possibility.

Using Jev well in a real system

Launch on Jev, graduate to a classifier. Jev is a cold-start router, not a forever router. Use it from day one, log its confident decisions as labels, and move high-volume routes to a cheap classifier later. Jev stays for low-confidence queries and new routes.

Batch your questions and keep the policy in code. Routing, reranking and sufficiency as separate sequential calls would add about a second per turn. One call asking all of them removes most of that. Keeping the thresholds and next-step logic in code also keeps the agent’s behaviour auditable and lets you change thresholds without re-prompting.

Caveats

The biggest open question is training-data overlap. Every dataset here is public, so Jev may have seen BEIR, SQuAD or HotpotQA during training, and its reranking margins are large enough to warrant that check.

  • No LLM baseline. An LLM-as-router or LLM-as-judge is the comparison readers will expect, and it isn’t here yet.
  • Small samples. Sample sizes are 200–800 per task, from a single run. Gaps under 2–3 points are within noise.
  • Shared GPU latencies. Local latencies are from a shared laptop GPU. A production GPU would run bge-reranker-v2-m3 several times faster.
  • No prompt tuning. Jev and the zero-shot baselines got the same untuned route descriptions. Banking77 used only its label names.
  • Dataset-derived routing labels. Routing labels come from each query’s source dataset. Some HotpotQA questions read as single-hop, which inflates Jev’s multi-hop error somewhat.

So, should you use Jev?

Remember the guess from the start? Routing was the obvious pick, but the controller was the real standout. It’s the hardest decision in the loop, and it’s where Jev beat every alternative, trained models included, by the widest margin: 0.90 AUROC on multi-hop evidence against 0.69 at best.

To Jithin’s original question: yes, Jev belongs in agentic RAG, as the controller first. And Akshat’s reranking idea deserved its test. It won that round too, pending a check on our own documents.

Use Jev whenReach for something else when
You need a decision from day one, with no labelled dataYou already have hundreds of labelled examples per route: a 10 ms classifier will match it
The decision is “is this evidence enough?”You need a ranking signal finer than two decimals
Inputs vary in domain or style (claims, forum posts, mixed queries)Your queries look exactly like a reranker’s training data (web search)
You need confidence you can act on: route, or fall backEvery millisecond counts: ~350 ms per call is a hard floor
You want the decision typed and auditable, with the policy kept in codeThe “decision” really needs reasoning steps, e.g. planning a multi-hop decomposition

The most interesting result is a design pattern, not a score. Let a decision model judge the evidence, keep the rules in code, and save the LLM for the one job only it can do: writing the answer.

Book A Demo: Click Here
Join our Slack: Click Here
Link to our GitHub: Click Here
Build with Lyzr

Try it in
Agent Studio

From framework-agnostic design to production-grade agents, deployed in under 24 hours.