Can a model that never writes a single word be the brain of an AI agent? We made 21,314 calls to find out, and the whole experiment cost 72 cents.
The short version
If you read nothing else: Jev fits agentic RAG well, and it also did better at reranking than we expected.
- Controller: Jev’s “is this context sufficient?” signal scored 0.90 AUROC on both single-passage and multi-hop tasks. Specialised local models drop to 0.69 on multi-hop.
- Router: Jev was the best zero-shot router on all three tasks (0.80–0.91 accuracy) and the most consistently well calibrated. With labelled data, a small embedding classifier catches up: 10 examples per intent already beat Jev on Banking77 (0.875 vs 0.801).
- Reranker: Jev raised NDCG@10 over retrieval alone by 6 points (SciFact) and 9 points (FiQA). bge-reranker-v2-m3 managed 0 and 3.
- Cost: the full benchmark cost $0.72 of Jev credit. Decisions cost 0.02–0.07 per 1,000 calls, at about 350 ms median latency.
The problem: RAG is a chain of small decisions
We usually describe a RAG system as “retrieve, then generate”. In practice the answer quality is set by four small decisions made before the LLM writes anything:
- Route: does this message need retrieval at all, and if so, which kind? A poem request, a pasted document and a multi-step comparison all need different handling.
- Select: which knowledge base or index should be searched?
- Filter: of the 20 chunks retrieved, which ones belong in the prompt?
- Stop or continue: is what we have enough to answer, or should the agent search again?
Get any one of these wrong and the LLM is answering from the wrong evidence, confidently. Today, teams make these decisions with one of four tools, and each has a catch:
| Tool | How it decides | The catch |
| Hand-written rules | Keywords, regexes, if-statements | Brittle; breaks on phrasing it hasn’t seen |
| Embedding similarity | Nearest route or chunk by vector distance | Matches topic, not intent; can’t tell “related” from “answers it” |
| Small trained models (cross-encoders, classifiers) | Trained for one task, e.g. MS MARCO relevance | Strong in their own domain, fragile outside it; need labelled data |
| An LLM call | Prompt an LLM to decide, parse its text | Slow (seconds), costly at volume, output must be parsed and can drift |
A decision model like Jev pitches itself as a fifth option: an LLM’s flexibility, but with typed answers, probabilities and a price close to zero. That pitch is what we wanted to test.
Before you read: what you need to know
You don’t need to have run a benchmark before. If you know what RAG is, these two tables cover everything else in this article.
The concepts
| Term | What it means here |
| RAG | Retrieval-augmented generation: fetch relevant text from a knowledge base, then give it to an LLM to answer from |
| Agentic RAG | RAG where an agent loops: it can choose a strategy, search more than once, and decide when it has enough |
| Chunk | A passage of a document, the unit that gets retrieved and put in the prompt |
| Dense retrieval | Finding chunks by comparing embedding vectors of the query and the chunks |
| Reranker | A second-stage model that re-sorts the retrieved chunks by true relevance |
| Cross-encoder | The usual reranker: reads the query and one chunk together and outputs a relevance score |
| Router | Picks which path a query takes: a tool, an index, or a strategy |
| Controller | Decides the agent’s next step, here: answer now or retrieve more |
| Zero-shot | The model gets no labelled examples, only a description of the task |
| Multi-hop question | Needs facts from two or more documents combined, e.g. “Was the director of film X born before Y?” |
The metrics
| Metric | Used for | How to read it |
| NDCG@10 | Reranking | 0–1. How close the top 10 chunks are to the perfect order. +0.03 is a meaningful gain |
| MRR@10 | Reranking | 0–1. How high the first relevant chunk sits. 0.5 means it’s typically at position 2 |
| Recall@5 | Reranking | Share of relevant chunks that made the top 5, i.e. that would reach the LLM |
| Accuracy | Routing | Share of queries sent to the right route |
| ECE (calibration error) | Routing | 0 is perfect. How far a model’s stated confidence is from how often it’s actually right. 0.05 means “90% sure” really is about 85–95% right |
| Accuracy at 80% coverage | Routing | Accuracy when the model handles only its most confident 80% and hands the rest to a fallback |
| AUROC | Controller | 0.5 is a coin flip, 1.0 is perfect. How well the model’s score separates “enough evidence” from “not enough”, at any threshold |
Meet Jev: a model that only decides
Jev is what TypeSafe calls a System One model, named after the fast, intuitive half of human thinking. You give it a state (any text or JSON) and some typed questions. It returns a probability for every possible answer and never generates prose. It went public in September 2026 at $0.042 per million input tokens, with output tokens free.
It answers three kinds of question:
| Type | Asks | Returns | Our use |
| noul | A yes/no question | P(yes), 0–1 | “Is this chunk relevant?”, “Is this context sufficient?” |
| choice | Pick one of up to 255 named options | The choice + a probability per option | “Which route?”, “Which index?” |
| score | Grade against a 2–10 level rubric | The expected score + a probability per level | “Rate relevance 0–3” |
Here is a real call from our routing benchmark. We asked Jev how to handle a HotpotQA question:

And its answer, 494 input tokens later:

Where Jev could sit in a RAG pipeline
Here is the agentic loop we had in mind, with the four decisions from the problem section marked. Each highlighted step is a single Jev call, and each carries the score Jev got on that job in this benchmark. The rest of this article explains where those numbers come from.
The loop’s most important arrow is “no: search again”. Without a reliable step 4, an agent either stops too early and answers from half the evidence, or keeps searching and burns time and money.
Three questions we set out to answer
- Can Jev rerank? Jithin’s bet was no, since dedicated rerankers are trained on millions of relevance judgements and Jev was not. Akshat’s bet was that reranking is just scoring, and scoring is what Jev does.
- Can Jev route without training data? A router has to work from day one, before anyone has labelled a single query. Its confidence also has to be honest, or the fallback logic is useless.
- Can Jev tell an agent when to stop searching? This is the hardest one. Spotting that a relevant-looking passage doesn’t answer the question, or that a second piece of evidence is missing, is where RAG systems quietly fail.
Before you scroll on, make a guess. Which of the three do you think Jev did best at? Routing looks like the obvious answer. Hold that thought until Round 3.
What we tested
Jev (jev-1.13.0, TypeSafe AI) was tested zero-shot on six public tasks covering the three jobs in Jithin and Akshat’s thread. Every baseline is an open model run locally on an Apple M5 GPU. No LLM baseline was included.
| Job | Task | Data | Baselines |
| Reranker | Rerank the top 20 dense-retrieved chunks | BEIR SciFact and FiQA, 200 queries each | bge-base retrieval only; MiniLM-L6 cross-encoder; bge-reranker-v2-m3 |
| Router | Intent routing, 77 classes | Banking77, 770 queries | Embedding zero-shot; DeBERTa NLI zero-shot; embedding classifier with 10 examples per class; logistic regression on the full training set |
| Router | RAG strategy: no retrieval, pasted context, single-hop, multi-hop, math tool | 600 queries from Dolly, NQ, HotpotQA, GSM8K | Same as above |
| Router | Knowledge-base selector, 5 indexes | 360 BEIR queries (SciFact, FiQA, NFCorpus, TREC-COVID, NQ) | Same as above |
| Controller | Can the agent answer from this passage? | SQuAD v2, 800 (half unanswerable) | roberta-base-squad2; MiniLM cross-encoder; DeBERTa NLI |
| Controller | Answer now, or retrieve another hop? | HotpotQA, 800 paired cases, one hop removed in half | Same as above |
Jev reranking was tried three ways: one yes/no call per chunk, a 0–3 relevance score per chunk, and one call holding all 20 chunks with a yes/no question per chunk. Every Jev call is cached and logged, and the full suite cost under $1 of the $5 budget.
How we ran it
Everything ran on one laptop. The local models used its GPU, and Jev was called over its public REST API. The full suite took about 2 hours of wall-clock time, almost all of it local baselines. Each Jev pass took 1–3 minutes per task.
| Component | Detail |
| Machine | Apple M5 laptop, 10 CPU cores, 24 GB unified memory, macOS 27.0 |
| Local inference | PyTorch 2.14.1 on the Apple GPU (MPS backend), fp32, no quantisation |
| Software | Python 3.12 (uv), sentence-transformers 6.1.0, transformers 5.18.0, datasets 5.0.1, scikit-learn 1.9.1 |
| Jev endpoint | POST api.typesafe.ai/v1/systemone, model jev-latest (served as jev-1.13.0) |
| Jev client | 16 parallel requests (8 for the 20-chunks-per-call variant), retry with back-off on 429/529, SQLite response cache, hard stop at $3.50 spend |
| Jev volume | 21,314 calls, 17.2M input tokens, 0.37 s mean latency, 3 failed requests, all succeeded on re-run |
Steps, in order:
- Created a Python 3.12 project with uv and installed the libraries above.
- Downloaded every dataset from Hugging Face and drew the samples with a fixed seed (0), so every system saw identical inputs.
- Embedded the SciFact (5.2k docs) and FiQA (57.6k docs) corpora with bge-base-en-v1.5 and kept the top 20 chunks per query by cosine similarity.
- Ran the local baselines on the GPU: the two cross-encoders, the DeBERTa NLI classifier, roberta-squad2, and the embedding classifiers. Chunks were cut to 350 words, or 250 in the 20-chunks-per-call Jev variant.
- Sent the same inputs to Jev. Every response was cached, so re-running any script costs nothing.
- Scored everything in Python. Where a threshold was needed, it was tuned on one half of the data and tested on the other, then the halves were swapped and the two results averaged.

Two setup problems are worth knowing about. The local models competed for the one GPU when run in parallel, which slowed every baseline. Zero-shot NLI over 77 Banking77 labels also needed ~59k model passes and took about 20 minutes even batched.
Round 1: Jev as a reranker, and the result we didn’t expect
Jev beat both open rerankers on both datasets. It lifted NDCG@10 by 6 points over retrieval alone on SciFact and by 9 points on FiQA. The strongest open reranker, bge-reranker-v2-m3, gained 0 and 3 points.
| System | SciFact NDCG@10 | SciFact MRR@10 | FiQA NDCG@10 | FiQA MRR@10 | FiQA Recall@5 |
| Jev, all 20 chunks in one call | 0.836 | 0.815 | 0.476 | 0.575 | 0.457 |
| Jev, one yes/no call per chunk | 0.825 | 0.799 | 0.478 | 0.587 | 0.471 |
| Jev, 0–3 score per chunk | 0.821 | 0.797 | 0.473 | 0.578 | 0.468 |
| bge-base retrieval only | 0.776 | 0.744 | 0.384 | 0.472 | 0.374 |
| bge-reranker-v2-m3 (568M) | 0.774 | 0.737 | 0.418 | 0.518 | 0.397 |
| MiniLM-L6 cross-encoder (22M) | 0.742 | 0.702 | 0.371 | 0.447 | 0.356 |
The one-call version is the one to use. It matches per-chunk calls on quality at about half the tokens, because each Jev call carries roughly 300 tokens of fixed overhead.
Akshat’s “deterministic gate” also works: keep a chunk when Jev’s P(relevant) is 0.5 or more, and drop the rest.
- SciFact: it keeps 1.7 chunks per query at 46% precision. bge-reranker keeps a similar number at 41%.
- FiQA: it keeps 8–10 chunks and keeps 93–95% of the relevant ones. bge-reranker keeps 4 and loses 42% of the relevant ones.
Jev’s probabilities are rounded to two decimals, so near-ties are broken by retrieval order. That makes ranking within the top few chunks coarse.
Round 2: Jev as a router
Jev was the best zero-shot router on all three tasks by 11–51 accuracy points, with calibration error of 0.09 or less on every task, which no other system managed. A classifier trained on labelled examples still matched or beat it on accuracy.
| System | Banking77 accuracy | RAG strategy accuracy | KB selector accuracy | Calibration error (ECE), range |
| Jev, zero-shot | 0.801 | 0.880 | 0.914 | 0.04–0.09 |
| Embedding zero-shot | 0.688 | 0.373 | 0.575 | 0.06–0.25 |
| DeBERTa NLI zero-shot | 0.569 | 0.285 | 0.661 | 0.12–0.42 |
| Embedding classifier, 10 labelled examples per class | 0.875 | 0.743 | 0.903 | 0.12–0.45 |
| Logistic regression, full labelled set | 0.939 | 0.903 | 0.908 | 0.09–0.31 |
Calibration is the number that matters most for a router, because the confidence decides when to fall back to an LLM or a human. Routing only the 80% most confident queries lifts Jev’s accuracy to 0.87 on Banking77, 0.94 on RAG strategy and 0.97 on KB selection.
Jev’s main weakness is under-calling multi-hop. On the RAG strategy task it routed 52 of 120 HotpotQA multi-hop questions as single-hop, and 18 of 120 creative prompts as single-hop retrieval. Math, pasted context and single-hop were 359 of 360 correct. On KB selection, most errors were nutrition questions sent to the Wikipedia index (15 of 80).
Round 3: Jev as a controller, the decision that matters most
Jev is the strongest “answer now or retrieve more?” signal we tested. On multi-hop sufficiency it scores 0.90 AUROC, against 0.66–0.69 for every local model. Its default 0.5 cut-off is too eager to answer, though. Use about 0.8 instead.
| System | SQuAD v2 AUROC | SQuAD v2 accuracy (tuned threshold) | HotpotQA AUROC | HotpotQA accuracy (tuned threshold) |
| Jev, yes/no: “is the context sufficient?” | 0.898 | 0.815 | 0.897 | 0.822 |
| Jev, choice: answer vs retrieve more | 0.882 | 0.819 | 0.885 | 0.817 |
| roberta-base-squad2 (trained on SQuAD v2) | 0.932 | 0.853 | 0.687 | 0.649 |
| MiniLM cross-encoder relevance | 0.651 | 0.599 | 0.690 | 0.604 |
| DeBERTa NLI zero-shot | 0.736 | 0.666 | 0.660 | 0.630 |
The SQuAD-trained model wins only on its own dataset. It collapses on HotpotQA, where one paragraph can contain a plausible answer span while the second hop is missing. Jev holds the same 0.90 AUROC on both tasks. That is the property an agentic-RAG controller needs.
The threshold sets the trade-off between answering on missing evidence and spending a retrieval that wasn’t needed:
| Jev P(sufficient) threshold | HotpotQA: answers with a hop missing | HotpotQA: retrieves again needlessly | SQuAD v2: answers unanswerable | SQuAD v2: retrieves again needlessly |
| 0.5 (default) | 31% | 7% | 39% | 6% |
| 0.7 | 21% | 14% | 26% | 11% |
| 0.8 (recommended) | 16% | 21% | 17% | 18% |
| 0.9 | 9% | 38% | 9% | 34% |
A plain yes/no question slightly beat (AUROC +0.01) framing the decision as an agent action (“answer” vs “retrieve_more”), so ask Jev about the evidence and keep the action logic in code.
Cost and latency
The whole benchmark used 17.2M Jev input tokens, which cost $0.72 of the $5 budget. Routing and controller decisions cost 0.02–0.07 per 1,000 calls, at about 340–380 ms median latency.
| Jev use | Input tokens per call | Cost per 1,000 calls | p50 latency | p95 latency |
| Rerank 20 chunks, one call (SciFact) | 7,721 | $0.32 | 480 ms | 595 ms |
| Rerank 20 chunks, one call (FiQA) | 4,644 | $0.20 | 410 ms | 595 ms |
| Rerank 20 chunks, one call per chunk (SciFact) | 14,825 | $0.62 | 432 ms | 693 ms |
| Router, 77 routes (Banking77) | 1,694 | $0.07 | 342 ms | 594 ms |
| Router, 5 routes (RAG strategy) | 540 | $0.02 | 366 ms | 536 ms |
| Controller, 1 passage (SQuAD v2) | 555 | $0.02 | 381 ms | 725 ms |
| Controller, 4 paragraphs (HotpotQA) | 921 | $0.04 | 336 ms | 418 ms |
About 300 tokens of every call are fixed overhead, and each extra question adds only about 20 tokens. So the cheap pattern is one call that asks every question you need about the same state.
For comparison, bge-reranker-v2-m3 took 5.7–6.0 s per query on the laptop GPU while other jobs shared it. Per-chunk Jev latency was measured with 16 calls in parallel.
Observations and opinions
My read: Jev is best used as a judge of evidence, the controller in the loop, rather than as a drop-in model for any single stage. The table summarises what we saw and what to do about it. The three sections after it explain the points that need more context.
At a glance
| Theme | What we saw | What to do |
| Controller fit | Same 0.90 AUROC on single-passage and multi-hop sufficiency. Every specialised model fell apart off its training distribution | Adopt Jev as the controller first |
| Reranking | Wins by following instructions, not better relevance modelling. MiniLM made both SciFact and FiQA worse than retrieval alone | Validate on your own documents before shipping |
| Yes/no probabilities | Decisive (42% at 0.10 or below, 11% at 0.90 or above) but lean towards yes: unanswerable questions were answered 31–39% of the time at 0.5 | Tune each threshold on about 100 labelled examples |
| Choice probabilities | Router calibration error stayed at 0.09 or below on every task | Route if confident, else fall back to an LLM or human, with no tuning |
| Cold start | With 10 labelled examples per route, a 10 ms classifier was within 1 point of Jev on KB selection, 7 ahead on Banking77, 14 behind on RAG strategy | Launch on Jev, log its confident decisions as labels, move high-volume routes to a cheap classifier |
| Hidden structure | Multi-hop questions read like single lookups; the second hop lives in the entities | Ask a narrow question, e.g. “does answering need facts about two different entities?” |
| Call design | ~300 tokens of overhead plus ~20 per question, ~350 ms latency floor | Ask every independent question about the same state in one call |
| Policy | Asking “is it sufficient?” worked as well as “answer or retrieve more?” | Jev judges facts; code decides what happens next |
| Operations | 21,314 calls, 3 failures, all fine on re-run, under $1 total | Cost is not a factor; latency and accuracy are |
Why the controller is the best fit
The controller sees every kind of question, so it needs a signal that doesn’t depend on the domain. Jev’s sufficiency score held at 0.90 AUROC whether the task was a single passage or a multi-hop chain, while each specialised model was strong only on the data it was trained for. That consistency is what an agent loop needs.
Reranking: promising, but verify
Jev’s reranking win comes from following instructions. MS MARCO cross-encoders expect web-search queries, but SciFact queries are claims and FiQA queries are forum posts. Jev was told to look for evidence that supports or refutes the claim, and it did. I’d expect the gap to narrow on plain web-search queries.
Jithin’s instinct that a model not trained on relevance judgements would struggle was reasonable, but the data disagrees: reranking is a scoring problem, and Jev scores well. I still wouldn’t ship it as our reranker until it holds up on our own documents, since margins this large leave training-data overlap as a real possibility.
Using Jev well in a real system
Launch on Jev, graduate to a classifier. Jev is a cold-start router, not a forever router. Use it from day one, log its confident decisions as labels, and move high-volume routes to a cheap classifier later. Jev stays for low-confidence queries and new routes.
Batch your questions and keep the policy in code. Routing, reranking and sufficiency as separate sequential calls would add about a second per turn. One call asking all of them removes most of that. Keeping the thresholds and next-step logic in code also keeps the agent’s behaviour auditable and lets you change thresholds without re-prompting.
Caveats
The biggest open question is training-data overlap. Every dataset here is public, so Jev may have seen BEIR, SQuAD or HotpotQA during training, and its reranking margins are large enough to warrant that check.
- No LLM baseline. An LLM-as-router or LLM-as-judge is the comparison readers will expect, and it isn’t here yet.
- Small samples. Sample sizes are 200–800 per task, from a single run. Gaps under 2–3 points are within noise.
- Shared GPU latencies. Local latencies are from a shared laptop GPU. A production GPU would run bge-reranker-v2-m3 several times faster.
- No prompt tuning. Jev and the zero-shot baselines got the same untuned route descriptions. Banking77 used only its label names.
- Dataset-derived routing labels. Routing labels come from each query’s source dataset. Some HotpotQA questions read as single-hop, which inflates Jev’s multi-hop error somewhat.
So, should you use Jev?
Remember the guess from the start? Routing was the obvious pick, but the controller was the real standout. It’s the hardest decision in the loop, and it’s where Jev beat every alternative, trained models included, by the widest margin: 0.90 AUROC on multi-hop evidence against 0.69 at best.
To Jithin’s original question: yes, Jev belongs in agentic RAG, as the controller first. And Akshat’s reranking idea deserved its test. It won that round too, pending a check on our own documents.
| Use Jev when | Reach for something else when |
| You need a decision from day one, with no labelled data | You already have hundreds of labelled examples per route: a 10 ms classifier will match it |
| The decision is “is this evidence enough?” | You need a ranking signal finer than two decimals |
| Inputs vary in domain or style (claims, forum posts, mixed queries) | Your queries look exactly like a reranker’s training data (web search) |
| You need confidence you can act on: route, or fall back | Every millisecond counts: ~350 ms per call is a hard floor |
| You want the decision typed and auditable, with the policy kept in code | The “decision” really needs reasoning steps, e.g. planning a multi-hop decomposition |
The most interesting result is a design pattern, not a score. Let a decision model judge the evidence, keep the rules in code, and save the LLM for the one job only it can do: writing the answer.
Book A Demo: Click Here
Join our Slack: Click Here
Link to our GitHub: Click Here
