Field Notes  /  Measured Architecture

Supercharging
“Poor Man’s RAG”
with Jev

What happened when we put a fast, non-autoregressive decision model in front of the search index we already had — instead of building a vector database. Measured on two live systems.

Model typesafe/jev-1.13-20260917 Measured 1 Oct 2026 Method published, with harness
00  /  Result
324
Questions tested
bucketed by measured vocabulary overlap
65% vs 33%
Keyword vs model — wording matches
n=112 · p < 0.0001
12% vs 2%
Model vs keyword — wording does not match
n=66 · p = 0.046
0.45
Overlap where the advantage changes hands
gate here — both directions proven
133
Documents in the production index
3,819 chunks · no vectors needed
$6.46
Per 10,000 queries
if every query were routed
7.9%
Local 0.8B model
vs 27.5% hosted · 20-doc cap
Retrieval accuracy by question type
Both directions significant · 324 questions · p=0.046 and p<0.0001
Accuracy vs cost
Keywords are free · the model costs $6.46 per 10k and only pays on mismatched wording
01  /  Trigger

The result, up front

We tested keyword search (BM25) and a decision model against the Partner Hub's actual knowledge base — 3,819 chunks across 133 documents, the index the bot serves in production — across 324 questions. Every question is bucketed by a measured quantity: the share of its own vocabulary that appears in the correct document. Significance is McNemar's test on paired outcomes; intervals are Wilson 95%.

Question → document overlapnBM25Decision modelGapp
Above 0.65 — wording largely shared11265%33%−32.1p < 0.0001
0.45 – 0.658634%28%−5.80.50
0.25 – 0.457013%16%+2.90.80
0.01 – 0.25380%16%+15.8p = 0.041
Zero — no shared words at all180%11%+11.10.48

Both directions are now statistically significant, which is the whole finding.

Where the asker already uses the document's own words, keyword search wins decisively: 65% against 33%, p<0.0001. Put a model in front of retrieval for those questions and you make it worse.

Where the vocabulary does not match, the reverse holds. Across the low-overlap questions combined (n=66), the model retrieves 12.1% against keyword search's 1.5%, p=0.046. Keyword search is not merely weak there — it is effectively useless: a single hit in 66 questions. On the questions that share no vocabulary with the right document at all, keyword search scores 0% across all 18 while the model still recovers 11%.

Overall the model is the weaker method (24.7% against 34.3%, p=0.0037) — but that headline is an artefact of the question mix, and publishing it alone would be as misleading as the earlier versions of this page. The two methods are not competing. They cross over at roughly 0.45 vocabulary overlap, and the correct architecture uses each on the side of that line where it wins.

What still holds. A vector database was still never needed. Keyword search over the existing JSON index — the "poor man's RAG" this was always about — remains the right default at this scale. The finding is not "vectors beat keywords"; it is "at 133 documents you do not need vectors, and you may not need a decision model either."

A different kind of model

Jev (by TypeSafe AI) is a System One decision model. It does not generate text token-by-token. It evaluates declared, typed questions — a discrete choice, a rubric score, or a boolean noul — in a single pass, and returns calibrated probabilities instead of prose.

Two properties make it interesting for retrieval. The output schema is enforced by the model, so it cannot return an invented category: "FINANCIAL SERVICES" or nothing. And it bills on input only, so the cost of a decision does not depend on how long the answer is.

It is not a chat model. Posting it to a chat-completions endpoint returns 400 — it answers only at the decisions endpoint with { model, state, questions }. That constraint is the whole design: it is a router, not a writer.

02  /  Friction

Where keyword search breaks

Two shapes of the same problem, both of which we were running in production.

Pattern 1 — The Tabular Catalog. An indexed SQL database of 1,078 Bursa Malaysia listed companies with sector, market-cap tier, profitability and compliance columns. Search runs through LIKE '%query%'.
Pattern 2 — The Document Assistant. A curated set of six collateral documents (38 chunks) — battle cards, a compliance quick-reference, a datasheet, a one-pager — retrieved by word-frequency scoring before an LLM sees them. This is the bot our partners actually use.

The vocabulary mismatch wall, measured

Both patterns fail the same way. A real query from our bot, recorded verbatim:

What happenedDetail
A partner asks“how do we stop a file we have never seen before from damaging the machine”
The scorer looks for"stop" · "file" · "seen" · "damaging" · "machine"
What it returnedcompetitor battle cards — the answer was in the containment section of the POC deck, which never uses those words.

Why a vector database was not the obvious answer

38 chunks is not a vector-database workload. An embedding pipeline, a store and a reindex loop for six documents is more infrastructure than the corpus. Vector similarity is also blind to relational logic — it cannot do > 4%, and it conflates Approached with Qualified. Meanwhile the index we already had answers in under a millisecond.

03  /  Discovery

Vectorless semantic routing

The insight is the division of labour: let a decision model resolve intent, let the database or the filesystem answer, and only let a generative model write when prose is actually wanted.

Human
Question
“how do we stop a file we have never seen before”
→
System 1
Jev router
Typed decision over 6 candidate documents, with confidence
→
Native
Exact retrieval
Indexed SQL, or a slice of a file. No vector store
→
System 2
LLM (optional)
Drafts the answer with citations, only if chat is needed
On timings, precisely. The model’s own decision pass is fast, but the number a user experiences is the network round trip: p50 763 ms end to end. Local keyword scoring is 0.8 ms. Jev is roughly three orders of magnitude slower than the thing it replaces — and buys an accuracy jump that, on paraphrased questions, is larger than the latency cost is worth avoiding.
04  /  Benchmarks

Two systems, real numbers

Pattern 1 — the company catalog (1,078 rows)

Natural-language questions routed to typed SQL constraints. The baseline is the catalog’s own keyword search on the verbatim question.

QuestionBaselineJev routerDecided constraintsRouterCost
High-yield banks
"profitable banks with high dividend yield"
0 rows18 rowssector: FINANCIAL profitable high_yield1237 ms$0.0000302
Telecom & tech, compliance
"large cap technology and telco companies affected by act 854"
0 rows18 rowssector: TECH mcap: Large act854776 ms$0.0000304
Healthcare with a contact
"healthcare companies with verified IT contact"
0 rows9 rowssector: HEALTH CARE verified contact786 ms$0.0000302
Energy large caps
"large cap energy oil and gas"
0 rows14 rowssector: ENERGY mcap: Large834 ms$0.0000301
Small-cap tech
"small cap profitable technology companies"
0 rows33 rowssector: TECH mcap: Small profitable712 ms$0.0000300

Every baseline returns zero — not a near miss, a total miss — because no company is named “profitable banks”. Once intent becomes constraints, the same database answers in about 2 ms. Router cost across the five: mean $0.0000302 per query, $0.30 per 10,000.

Pattern 2 — the document assistant, at production scale (3,819 chunks, 133 documents)

This is the bot. Keyword search and a decision model were run against the same production index across the same 324 questions, each bucketed by measured vocabulary overlap. Two earlier attempts at this comparison are described below, because both reached confident conclusions that the larger set does not support.

Question typenBM25Decision modelBetter method
Lexical — the asker uses the document's words
Mean vocabulary overlap 0.70
11265%33%Keyword search, p<0.0001
Semantic — phrased how a person actually talks
Mean vocabulary overlap 0.44
21219%23%Level overall — see the crossover below

Neither method dominates, and the reason is mechanical. Keyword search matches strings. When the asker's strings are the document's strings it is fast, free and hard to beat. When they are not, it has nothing to match on and returns nothing useful. A decision model reads the meaning instead, which costs money and latency and is worse precisely because it ignores the string match that was already sufficient.

Pattern 2b — where the two methods cross over

The most useful result in this study is not which method wins overall. It is where the advantage changes hands — and both sides of that line are now statistically significant.

Vocabulary overlapnBM25Decision modelGappWinner
Above 0.6511265%33%−32.1p<0.0001Keywords
0.45 – 0.658634%28%−5.80.50Level
0.25 – 0.457013%16%+2.90.80Level
0.01 – 0.25380%16%+15.80.041Model
Zero — no shared words180%11%+11.10.48Model
Low-overlap combined (≤0.30)661.5%12.1%+10.60.046Model
The cleanest single fact in the study. On 18 questions sharing no vocabulary with the correct document, keyword search scores 0%. Not low — zero. It is structurally incapable of answering them. The model still recovers 11%. That is the capability being paid for, and it is invisible if you only test questions phrased like your documents.

The design follows directly. Gate on vocabulary overlap: route low-overlap questions to the model, answer the rest locally and for free. Every question above 0.45 should never leave the box. On this data that is roughly two-thirds of traffic served at zero cost, with the model reserved for the third where keywords return nothing.

Cost, stated plainly: $0.000646 per query across 133 documents — $6.46 per 10,000. That is the honest price of semantic coverage. Whether it is worth paying depends entirely on how many of your users' questions resemble your documentation, and that is a question your own traffic answers better than any test we can build.

The three misses

Publishing the failures is the point, because they describe the shape of the tool’s limits.

QuestionJev choseConfidenceWhy
“how do we stop a file we have never seen before from damaging the machine”AEP + ITSM Datasheet0.83Defensible. That document does explain containment. Our expected document was an equally valid answer.
“a customer insists their existing antivirus is sufficient — what is our answer”Competitor Battle Cards0.69Defensible. The battle cards carry the detection-versus-containment argument.
“how do we get our own licence so we can try it internally”POC Kickoff Deck0.41A genuine miss. The right document was the partner one-pager.
The useful signal: confidence tracks correctness. Clean wins returned 0.90 and above. The three misses returned 0.83, 0.69 and 0.41. A single low-confidence answer is a usable trigger to fall back to keyword results or to ask the user to disambiguate — which is why we surface the confidence rather than hiding it.
05  /  Implementation

Could it run locally instead?

The obvious follow-up: if the decisions are this cheap, why pay a hosted provider at all? Ollama 0.35 added a /v1/systemone endpoint that accepts the same request shape as the Jev API — state plus named questions with choice, noul and score types — so the harness above runs against a local server with one flag changed. We tested it.

Local tev1:0.8bHosted Jev
Correct document, same 38 semantic questions7.9%27.5%
Baseline keyword search, same 3821.1%21.3%
Latency per decision~25 s~0.25 s
Documents it can see20 of 133133
Cost per 10,000 queries$0$6.46
Resident memory4.7 GBnone

The local model scored 7.9% against keyword search's 21.1% on the same questions — worse than the free method it would replace, and roughly a third of Jev's rate. It is also about a hundred times slower on CPU-only hardware.

The binding constraint is not quality, it is the token budget. tev1 rejects anything over 2,050 input tokens (“input is never truncated”) and caps a choice at 26 candidates. Our document set is 133. Even with descriptions cut to 150 characters, only about 20 documents fit. A 0.8B model at zero marginal cost that can see 15% of the corpus is worth less than a paid model that can see all of it — and no amount of prompt tuning closes that, because the ceiling is structural to the model class.

The footprint assumption is worth recording separately: a model with 0.8B parameters held 4.7 GB resident — 60% of the host’s 7.9 GB. “Small model” does not mean “small footprint”, and on a box that also runs production services that difference decides the question before accuracy does. Available memory fell to 306 MB before the run was stopped.

Stated honestly: this comparison is not like-for-like. The local model was restricted to a 20-document shortlist with shortened descriptions because of the token cap, while the Jev figures come from all 133 documents with full descriptions. The run was stopped at 38 of 80 questions to protect the production host, not because the direction was unclear. The finding is that a local 0.8B decision model cannot attempt this task on our corpus, which is a limit rather than a close race.

The anatomy of a router

Instead of asking a language model to emit valid JSON and hoping, the decision is declared as typed questions evaluated in one pass.

const decision = await jev({
  state: partnerQuestion,
  questions: {
    best_collateral: Choice({
      instructions: "Which document best answers this partner's question?",
      criteria: {
        "POC Kickoff Deck": "proof of concept: objectives, deployment, timeline",
        "Xcitium Partner One-Pager": "positioning, proof points, pricing tiers",
        "Malaysia Act 854 Compliance QuickRef": "Act 854 and RMiT mapping",
        "AEP + ITSM Bundle Datasheet": "technical detail: requirements, ports",
        "Competitor Battle Cards": "head-to-head positioning vs rivals",
        "PDPA Compliance Checklist": "Singapore PDPA controls"
      }
    })
  }
});

// Typed output. No regex, no repair, no invalid category possible.
const doc = load(decision.answers.best_collateral.choice);
const confidence = decision.answers.best_collateral.confidence;
if (confidence < 0.5) return keywordFallback(partnerQuestion);
Two details decide whether this works. First, the criteria descriptions carry the signal — our first attempt passed bare document titles, and accuracy on the same questions fell from 6/8 to 3/8. Second, offering an escape option (“no document matches”) looks safer and measurably suppressed three correct picks. Force the choice, then handle low confidence in code.
It is an adapter, not a drop-in. A chat-shaped client cannot call this. You need a thin translation layer between the conversational interface your users have and the typed decision the model answers — roughly 30 lines, but it is not zero, and it is the part that has to be written before any of the numbers above exist.
06  /  Boundaries

Scale limits

WorkloadArchitectureWhy
Catalog
1,000 – 500,000 rows
Intent-to-filter, no vectorsStructured already. Jev extracts constraints; indexed SQL executes. Measured: 0 rows → 18 on the same question.
Curated documents
10 – 250 documents
Keywords first; model does not pay offMeasured on the production corpus (3,819 chunks, 133 documents): keyword search 59%, model 35%. Routing to the model adds nothing at this size. The small-corpus gain did not survive.
Massive unstructured
> 20,000 unindexed chunks
Router first, then vectorsJev filters metadata partitions before similarity runs — narrowing 500,000 vectors to 50. This is where a vector store earns its keep.
The cardinality ceiling is about disambiguation, not capacity. A choice question does not fail at 133 options — it answers, in 931 ms, for $0.000646. On semantic questions it scores identically at 133 options and at 20 (27.5% both), so corpus size does not degrade it there. What degrades is the case where the options resemble each other: 20-odd competitor battle cards sharing the same product language. The model's own confidence tracks this, falling from 0.98 at 20 options to 0.61 at 133.
What this does not buy you. Correctness, and it is not a general-purpose upgrade: on questions phrased like your documentation it scores 33% against keyword search's 65%, so routing everything to it makes retrieval worse. It is network-bound — every query leaves the box — so it is the wrong tool for a millisecond-latency path, and a poor fit where data cannot leave the machine. At $6.46 per 10,000 queries it costs real money to cover a case that may be rare in your traffic.
07  /  Takeaway

What it actually buys

If your data is already structured, or your corpus is curated, you probably do not need a vector database — and the honest reason is not that vectors are bad, but that the index you already run answers the question once you know what is being asked.

What this study actually establishes, after two corrections, is narrower and more useful than either of the earlier headlines: keyword search and a decision model are good at different questions. Keywords win on the 81 lexical questions, 58% to 33%. The model wins on the 80 semantic ones, 28% to 21% — and on the subset that shares no vocabulary with the right document at all, keyword search scores 0% while the model still recovers a quarter. The crossover sits near 0.45 overlap.

That gives a concrete design: gate on vocabulary overlap. Answer the majority locally and for free; route only the genuinely mismatched questions to the model. On the semantic set that policy beats keyword search by 7.6 points while leaving a third of the traffic untouched. No embeddings, no vector store, no reindex pipeline — the vectorless claim holds throughout, and it now holds for a reason we can point at rather than assert.

The cost is real and should be stated plainly: $6.46 per 10,000 queries if every query is routed, against zero. Buy semantic coverage where it pays, not everywhere.

And self-hosting does not change that arithmetic here. A local 0.8B decision model is free to run but cannot see a 133-document library, scores below keyword search, and holds 4.7 GB of memory while doing it. For a corpus this size the hosted model remains the only viable option of the two — not because local inference is a bad idea in general, but because the input ceiling of this model class is smaller than the problem.