What happened when we put a fast, non-autoregressive decision model in front of the search index we already had — instead of building a vector database. Measured on two live systems.
We tested keyword search (BM25) and a decision model against the Partner Hub's actual knowledge base — 3,819 chunks across 133 documents, the index the bot serves in production — across 324 questions. Every question is bucketed by a measured quantity: the share of its own vocabulary that appears in the correct document. Significance is McNemar's test on paired outcomes; intervals are Wilson 95%.
| Question → document overlap | n | BM25 | Decision model | Gap | p |
|---|---|---|---|---|---|
| Above 0.65 — wording largely shared | 112 | 65% | 33% | −32.1 | p < 0.0001 |
| 0.45 – 0.65 | 86 | 34% | 28% | −5.8 | 0.50 |
| 0.25 – 0.45 | 70 | 13% | 16% | +2.9 | 0.80 |
| 0.01 – 0.25 | 38 | 0% | 16% | +15.8 | p = 0.041 |
| Zero — no shared words at all | 18 | 0% | 11% | +11.1 | 0.48 |
Both directions are now statistically significant, which is the whole finding.
Where the asker already uses the document's own words, keyword search wins decisively: 65% against 33%, p<0.0001. Put a model in front of retrieval for those questions and you make it worse.
Where the vocabulary does not match, the reverse holds. Across the low-overlap questions combined (n=66), the model retrieves 12.1% against keyword search's 1.5%, p=0.046. Keyword search is not merely weak there — it is effectively useless: a single hit in 66 questions. On the questions that share no vocabulary with the right document at all, keyword search scores 0% across all 18 while the model still recovers 11%.
Overall the model is the weaker method (24.7% against 34.3%, p=0.0037) — but that headline is an artefact of the question mix, and publishing it alone would be as misleading as the earlier versions of this page. The two methods are not competing. They cross over at roughly 0.45 vocabulary overlap, and the correct architecture uses each on the side of that line where it wins.
Jev (by TypeSafe AI) is a System One decision model. It does not generate text token-by-token. It evaluates declared, typed questions — a discrete choice, a rubric score, or a boolean noul — in a single pass, and returns calibrated probabilities instead of prose.
Two properties make it interesting for retrieval. The output schema is enforced by the model, so it cannot return an invented category: "FINANCIAL SERVICES" or nothing. And it bills on input only, so the cost of a decision does not depend on how long the answer is.
It is not a chat model. Posting it to a chat-completions endpoint returns 400 — it answers only at the decisions endpoint with { model, state, questions }. That constraint is the whole design: it is a router, not a writer.
Two shapes of the same problem, both of which we were running in production.
LIKE '%query%'.Both patterns fail the same way. A real query from our bot, recorded verbatim:
| What happened | Detail |
|---|---|
| A partner asks | “how do we stop a file we have never seen before from damaging the machine” |
| The scorer looks for | "stop" · "file" · "seen" · "damaging" · "machine" |
| What it returned | competitor battle cards — the answer was in the containment section of the POC deck, which never uses those words. |
38 chunks is not a vector-database workload. An embedding pipeline, a store and a reindex loop for six documents is more infrastructure than the corpus. Vector similarity is also blind to relational logic — it cannot do > 4%, and it conflates Approached with Qualified. Meanwhile the index we already had answers in under a millisecond.
The insight is the division of labour: let a decision model resolve intent, let the database or the filesystem answer, and only let a generative model write when prose is actually wanted.
Natural-language questions routed to typed SQL constraints. The baseline is the catalog’s own keyword search on the verbatim question.
| Question | Baseline | Jev router | Decided constraints | Router | Cost |
|---|---|---|---|---|---|
| High-yield banks "profitable banks with high dividend yield" | 0 rows | 18 rows | sector: FINANCIAL profitable high_yield | 1237 ms | $0.0000302 |
| Telecom & tech, compliance "large cap technology and telco companies affected by act 854" | 0 rows | 18 rows | sector: TECH mcap: Large act854 | 776 ms | $0.0000304 |
| Healthcare with a contact "healthcare companies with verified IT contact" | 0 rows | 9 rows | sector: HEALTH CARE verified contact | 786 ms | $0.0000302 |
| Energy large caps "large cap energy oil and gas" | 0 rows | 14 rows | sector: ENERGY mcap: Large | 834 ms | $0.0000301 |
| Small-cap tech "small cap profitable technology companies" | 0 rows | 33 rows | sector: TECH mcap: Small profitable | 712 ms | $0.0000300 |
Every baseline returns zero — not a near miss, a total miss — because no company is named “profitable banks”. Once intent becomes constraints, the same database answers in about 2 ms. Router cost across the five: mean $0.0000302 per query, $0.30 per 10,000.
This is the bot. Keyword search and a decision model were run against the same production index across the same 324 questions, each bucketed by measured vocabulary overlap. Two earlier attempts at this comparison are described below, because both reached confident conclusions that the larger set does not support.
| Question type | n | BM25 | Decision model | Better method |
|---|---|---|---|---|
| Lexical — the asker uses the document's words Mean vocabulary overlap 0.70 | 112 | 65% | 33% | Keyword search, p<0.0001 |
| Semantic — phrased how a person actually talks Mean vocabulary overlap 0.44 | 212 | 19% | 23% | Level overall — see the crossover below |
Neither method dominates, and the reason is mechanical. Keyword search matches strings. When the asker's strings are the document's strings it is fast, free and hard to beat. When they are not, it has nothing to match on and returns nothing useful. A decision model reads the meaning instead, which costs money and latency and is worse precisely because it ignores the string match that was already sufficient.
The most useful result in this study is not which method wins overall. It is where the advantage changes hands — and both sides of that line are now statistically significant.
| Vocabulary overlap | n | BM25 | Decision model | Gap | p | Winner |
|---|---|---|---|---|---|---|
| Above 0.65 | 112 | 65% | 33% | −32.1 | p<0.0001 | Keywords |
| 0.45 – 0.65 | 86 | 34% | 28% | −5.8 | 0.50 | Level |
| 0.25 – 0.45 | 70 | 13% | 16% | +2.9 | 0.80 | Level |
| 0.01 – 0.25 | 38 | 0% | 16% | +15.8 | 0.041 | Model |
| Zero — no shared words | 18 | 0% | 11% | +11.1 | 0.48 | Model |
| Low-overlap combined (≤0.30) | 66 | 1.5% | 12.1% | +10.6 | 0.046 | Model |
The design follows directly. Gate on vocabulary overlap: route low-overlap questions to the model, answer the rest locally and for free. Every question above 0.45 should never leave the box. On this data that is roughly two-thirds of traffic served at zero cost, with the model reserved for the third where keywords return nothing.
Cost, stated plainly: $0.000646 per query across 133 documents — $6.46 per 10,000. That is the honest price of semantic coverage. Whether it is worth paying depends entirely on how many of your users' questions resemble your documentation, and that is a question your own traffic answers better than any test we can build.
Publishing the failures is the point, because they describe the shape of the tool’s limits.
| Question | Jev chose | Confidence | Why |
|---|---|---|---|
| “how do we stop a file we have never seen before from damaging the machine” | AEP + ITSM Datasheet | 0.83 | Defensible. That document does explain containment. Our expected document was an equally valid answer. |
| “a customer insists their existing antivirus is sufficient — what is our answer” | Competitor Battle Cards | 0.69 | Defensible. The battle cards carry the detection-versus-containment argument. |
| “how do we get our own licence so we can try it internally” | POC Kickoff Deck | 0.41 | A genuine miss. The right document was the partner one-pager. |
The obvious follow-up: if the decisions are this cheap, why pay a hosted provider at all? Ollama 0.35 added a /v1/systemone endpoint that accepts the same request shape as the Jev API — state plus named questions with choice, noul and score types — so the harness above runs against a local server with one flag changed. We tested it.
Local tev1:0.8b | Hosted Jev | |
|---|---|---|
| Correct document, same 38 semantic questions | 7.9% | 27.5% |
| Baseline keyword search, same 38 | 21.1% | 21.3% |
| Latency per decision | ~25 s | ~0.25 s |
| Documents it can see | 20 of 133 | 133 |
| Cost per 10,000 queries | $0 | $6.46 |
| Resident memory | 4.7 GB | none |
The local model scored 7.9% against keyword search's 21.1% on the same questions — worse than the free method it would replace, and roughly a third of Jev's rate. It is also about a hundred times slower on CPU-only hardware.
tev1 rejects anything over 2,050 input tokens (“input is never truncated”) and caps a choice at 26 candidates. Our document set is 133. Even with descriptions cut to 150 characters, only about 20 documents fit. A 0.8B model at zero marginal cost that can see 15% of the corpus is worth less than a paid model that can see all of it — and no amount of prompt tuning closes that, because the ceiling is structural to the model class.The footprint assumption is worth recording separately: a model with 0.8B parameters held 4.7 GB resident — 60% of the host’s 7.9 GB. “Small model” does not mean “small footprint”, and on a box that also runs production services that difference decides the question before accuracy does. Available memory fell to 306 MB before the run was stopped.
Stated honestly: this comparison is not like-for-like. The local model was restricted to a 20-document shortlist with shortened descriptions because of the token cap, while the Jev figures come from all 133 documents with full descriptions. The run was stopped at 38 of 80 questions to protect the production host, not because the direction was unclear. The finding is that a local 0.8B decision model cannot attempt this task on our corpus, which is a limit rather than a close race.
Instead of asking a language model to emit valid JSON and hoping, the decision is declared as typed questions evaluated in one pass.
const decision = await jev({ state: partnerQuestion, questions: { best_collateral: Choice({ instructions: "Which document best answers this partner's question?", criteria: { "POC Kickoff Deck": "proof of concept: objectives, deployment, timeline", "Xcitium Partner One-Pager": "positioning, proof points, pricing tiers", "Malaysia Act 854 Compliance QuickRef": "Act 854 and RMiT mapping", "AEP + ITSM Bundle Datasheet": "technical detail: requirements, ports", "Competitor Battle Cards": "head-to-head positioning vs rivals", "PDPA Compliance Checklist": "Singapore PDPA controls" } }) } }); // Typed output. No regex, no repair, no invalid category possible. const doc = load(decision.answers.best_collateral.choice); const confidence = decision.answers.best_collateral.confidence; if (confidence < 0.5) return keywordFallback(partnerQuestion); |
| Workload | Architecture | Why |
|---|---|---|
| Catalog 1,000 – 500,000 rows | Intent-to-filter, no vectors | Structured already. Jev extracts constraints; indexed SQL executes. Measured: 0 rows → 18 on the same question. |
| Curated documents 10 – 250 documents | Keywords first; model does not pay off | Measured on the production corpus (3,819 chunks, 133 documents): keyword search 59%, model 35%. Routing to the model adds nothing at this size. The small-corpus gain did not survive. |
| Massive unstructured > 20,000 unindexed chunks | Router first, then vectors | Jev filters metadata partitions before similarity runs — narrowing 500,000 vectors to 50. This is where a vector store earns its keep. |
choice question does not fail at 133 options — it answers, in 931 ms, for $0.000646. On semantic questions it scores identically at 133 options and at 20 (27.5% both), so corpus size does not degrade it there. What degrades is the case where the options resemble each other: 20-odd competitor battle cards sharing the same product language. The model's own confidence tracks this, falling from 0.98 at 20 options to 0.61 at 133.If your data is already structured, or your corpus is curated, you probably do not need a vector database — and the honest reason is not that vectors are bad, but that the index you already run answers the question once you know what is being asked.
What this study actually establishes, after two corrections, is narrower and more useful than either of the earlier headlines: keyword search and a decision model are good at different questions. Keywords win on the 81 lexical questions, 58% to 33%. The model wins on the 80 semantic ones, 28% to 21% — and on the subset that shares no vocabulary with the right document at all, keyword search scores 0% while the model still recovers a quarter. The crossover sits near 0.45 overlap.
That gives a concrete design: gate on vocabulary overlap. Answer the majority locally and for free; route only the genuinely mismatched questions to the model. On the semantic set that policy beats keyword search by 7.6 points while leaving a third of the traffic untouched. No embeddings, no vector store, no reindex pipeline — the vectorless claim holds throughout, and it now holds for a reason we can point at rather than assert.
The cost is real and should be stated plainly: $6.46 per 10,000 queries if every query is routed, against zero. Buy semantic coverage where it pays, not everywhere.
And self-hosting does not change that arithmetic here. A local 0.8B decision model is free to run but cannot see a 133-document library, scores below keyword search, and holds 4.7 GB of memory while doing it. For a corpus this size the hosted model remains the only viable option of the two — not because local inference is a bad idea in general, but because the input ceiling of this model class is smaller than the problem.