If your corpus is small enough to fit in a prompt, you may not need a retrieval architecture at all. That is the sharpest finding from KyGround, a new benchmark posted on arXiv in October 2026 that directly compares tool-calling retrieval against vector RAG on a small Greek-English knowledge base, and it comes with a twist: the tool-calling design lost. The preprint reports that vector RAG answered 95% of canonical Greek questions correctly while tool retrieval managed 72%, and placing the whole knowledge base, about 26,000 tokens, directly in the prompt reached 99.3%.
One caveat belongs up front, because it shapes how much weight any of this should carry: KyGround is a single, small-scale, unreplicated preprint covering one language pair and one corpus. No independent replication of its numbers has appeared. Treat what follows as a well-measured result from one study, not a settled law, and treat the decision framework around it as the part that transfers to your own work.
Three ways to answer a question from a small closed corpus
For a closed corpus, docs, glossaries, terminology, product records, you have three realistic designs.
Whole-KB prompting. If the corpus fits in the model’s context window with room to spare, skip retrieval and put everything in the prompt. No embeddings, no vector store, no tool plumbing. The KyGround control run did exactly this with its ~26,000-token base and scored 99.3%, the best result in the study.
Vector RAG. The standard pipeline, documented in the MultiHop-RAG paper: “We partition the documents in the MultiHop-RAG knowledge base into chunks, each consisting of 256 tokens. We then convert the chunks using an embedding model and save the embeddings into a vector database.” At query time, embed the question, run cosine top-K over those chunks, and feed the winners to the model. This is the stack with real operational weight: an embedding model to choose and maintain, a chunking policy, an index, and a reindexing job every time the corpus changes.
Tool-calling lookup. Give the model lookup tools, a search endpoint, a get-record-by-id, a glossary query, and let it decide what to call. The appeal is architectural: a tool endpoint replaces the embedding pipeline and vector store, and the model composes lookups the way an agent would.
KyGround runs the comparison under tight controls: one router and answer model across all six conditions, a shared base prompt and knowledge base, and a control condition that gives vector retrieval the same router, prompt and function-calling mechanism as the tool agent. Its authors “know of no head-to-head comparison of tool-calling and vector retrieval in Greek.” The comparison did not go the tool agent’s way.
What KyGround actually measured
The benchmark, described in the KyGround preprint, consists of “198 questions drawn from the platform’s published records, with gold answers verified automatically against the records, each question posed in up to nine surface forms and scored by a deterministic, Greek-aware scorer.” The surface forms matter more than the headline accuracy: they are how the study tests robustness to the way people actually type Greek, with and without diacritics, capitalized or not, or transliterated into Latin characters (“Greeklish”).
The headline numbers:
| Condition | Vector RAG | Tool-calling retrieval |
|---|---|---|
| Canonical Greek questions | 95% | 72% |
| Unaccented / capitalised variants (accuracy lost) | at most 2 points | about 20 points |
| Greeklish variants (accuracy lost) | about 21–32 points | about 21–32 points |
| Tool agent with stemmed-token search (ablation) | — | 83.8% canonical |
| Whole KB in prompt (control) | 99.3% canonical Greek | — |
On canonical questions, vector RAG beat the tool agent by 23 points. On the orthographic variants, the gap widened in a specific and telling way: stripping accents or capitalizing input, forms the paper says Greek users often type, cost the tool agent roughly 20 accuracy points while vector RAG lost at most 2. Greeklish broke both designs roughly equally, costing each 21 to 32 points, so neither architecture survives every surface form a real user base will produce.
Why the tool agent lost
The instinct when a tool-calling agent underperforms is to blame the model’s tool use: malformed calls, wrong arguments, bad sequencing. The evidence points elsewhere. A multi-turn tool-calling diagnostic found that the strong closed anchor model, gpt-5.4, emitted the gold action on at least 99% of tool-call cases, with near-zero shifts under perturbation. That figure is Gold Action Recall, whether the model emits a tool call at all, and the paper warns that “a model that blindly calls tools scores the same,” so near-ceiling emission does not settle whether execution follows. KyGround’s own decomposition does the locating: wrong answers despite retrieved evidence were rare, and 19.3% of the tool agent’s answers were abstentions after its searches returned nothing, per the preprint. The tool agent’s losses sit in the lookup and matching step: the tool returning no record because the model’s arguments did not occur verbatim in one, or failing to match a user’s surface form to the record’s stored form.
That failure mode has a structural explanation. A vector index normalizes and embeds text into a space where accented and unaccented Greek forms land near each other, which is why RAG shrugged off the diacritic variants. A lookup tool backed by exact or lexical matching has no such forgiveness built in; if the endpoint expects canonical Greek and the user types without accents, somebody has to do the normalization, and in this study that somebody failed about a fifth of the time.
The paper’s ablation makes the point by fixing the tool rather than the model. Folding accents out of the interface’s free-text search removed the unaccented and capitals loss entirely. Matching stemmed tokens instead of literal phrases, the ablation reports, raised the tool agent’s canonical Greek accuracy from 71.6% to 83.8%, and on questions that do not quote a record’s title from 46.5% to 72.1%. The fix has limits: even with stemmed tokens the tool agent stayed 11.5 points below vector RAG on canonical Greek, and Greeklish accuracy stayed between 43.2% and 52.0% under both modes, because transliterated arguments match neither the Greek nor the English text of the records. Even so, about half the canonical gap was engineered away inside the search endpoint, which means the first move for a struggling tool agent is normalizing its search, not abandoning the design.
There is a second, more general reason to be careful with tool-calling retrieval: retrieving the right tool is itself a retrieval problem, and it is not solved. On the ToolRet benchmark, “even the best model (i.e., NV-embedd-v1) that demonstrates strong performance in conventional IR benchmarks, achieves an nDCG@10 of only 33.83 in our benchmark.” With a handful of tools this hardly matters. If your architecture assumes the model will pick correctly from dozens or hundreds of tools, you have reintroduced the retrieval problem one level up, with worse tooling for it.
The whole-knowledge-base control should change your default
The most consequential number in the study is not the RAG win. It is the control: the whole knowledge base, about 26,000 tokens, placed in the prompt, scored 99.3% in the KyGround runs.
Twenty-six thousand tokens was the whole of this study’s corpus: 103 published records. If your corpus is that size, fits comfortably in context alongside the conversation, and your query volume makes the per-request token cost acceptable, the evidence here says the retrieval layer is optional infrastructure. You would be paying for a chunking pipeline, an embedding job and a vector database to get accuracy you can have by pasting the corpus into the prompt.
The boundary conditions are real, though. The corpus has to fit. It has to change slowly enough that you are not rebuilding prompts constantly. And at high query volume, stuffing the whole base into every request costs more than retrieving two or three chunks: per question, whole-KB prompting averaged 26,294 input tokens against 1,612 for vector RAG and 6,517 for the tool agent, and the paper calls whole-KB prompting “about four times the input of the tool agent.” The study prices tokens, not dollars, so the economics are still your calculation to run. The point is narrower: for small corpora, “which retrieval architecture” is the second question. The first is whether you need retrieval at all.
A counter-argument worth keeping
Before anyone declares vector RAG the robust choice on orthographic messiness generally, one caveat cuts the other way. Research on dual-encoder dense retrieval under misspellings found that “robustness deteriorates when typos do not appear randomly. In detail, the most significant losses occur when typos appear on discriminative utterances.” Vector RAG’s two-point loss on KyGround’s variants is a Greek-diacritics result, where the variation is systematic and the embedding space has likely seen plenty of it. Typos landing on the exact tokens that distinguish one record from another are a different and harsher failure mode, and dense retrieval is not immune to it. Do not generalize the ≤2-point figure to your domain’s spelling chaos without testing it.
The same caution applies in reverse: if vector RAG underperforms on a morphologically complex language, switching architectures is not the only lever. ORPHEAS, a domain-specialized Greek, English embedding model, “outperforms state-of-the-art multilingual embedding models,” per its authors, showing that fine-tuning the embeddings can fix retrieval quality where a general-purpose model falls short. Embedding quality is a tunable component; architecture is a bigger commitment.
How much should you trust these numbers?
Three properties of the study limit how far the margins travel, and the paper’s own Limitations section concedes all three of them.
First, statistical power. KyGround’s own analysis, quoted in the preprint, states that “the study can detect differences of about 14 points with 80% power, so the smaller differences between the Greeklish gaps are reported as estimates.” The 23-point canonical gap between RAG and tool retrieval clears that floor. The 21-to-32-point Greeklish ranges do not resolve into precise per-architecture losses; they are estimates with wide error bars.
Second, literal-matching bias. The study reports that 105 of its 148 primary questions quote a record’s exact title. Questions that quote titles favor any retrieval method keyed on literal text, and they flatter both architectures relative to what users’ own paraphrased wording would produce. If your users ask “how do I export last quarter’s invoices” rather than quoting the “Quarterly Invoice Export” record title, expect real-world accuracy below the reported 95% and 72% (KyGround tables).
Third, the model behind every number is a stand-in. All six conditions used Claude Haiku 4.5 as router and answer model, and the tool agent is a reconstruction of the deployed one: seven of the nine tool names, all nine tool descriptions and the full system prompt were rebuilt for the study. The preprint notes that Claude Haiku 4.5 replaced the deployed GPT-5-family models, so the accuracies estimate what the designs achieve with this model, and that “another router may transliterate more or less often.” A different router could shift the tool-retrieval number in either direction. And none of this has been independently replicated.
A bake-off you can run on your own corpus
The transferable artifact from this study is not its verdict but its method. You can reproduce the shape of it in a week:
- Build a question set from real usage. Pull actual user queries if you have them; write them against your records if you do not. KyGround scored 148 answerable test items and still only detected ~14-point differences at 80% power, per the preprint, so treat 150 to 200 questions per condition as the floor for margins of that size to mean anything. Fewer questions means only huge gaps are real.
- Stratify by surface form. For each question, generate the variants your users actually type: missing diacritics, all caps, transliteration, abbreviation, the misspellings in your search logs. KyGround used up to nine forms per question. The variants are where the architectures diverged, so a canonical-only evaluation would have missed the study’s most useful finding.
- Verify answers deterministically where you can. KyGround checked gold answers against published records automatically with a Greek-aware scorer. If your answers can be checked against structured ground truth, do that; LLM-judged scoring adds noise exactly where your margins are smallest.
- Run all three arms. Whole-corpus-in-prompt (if it fits), vector RAG with your actual embedding model, and the tool agent with your actual endpoints. Report per-surface-form accuracy, not just the canonical average.
- Check the bias in your own questions. Count how many of your questions quote record titles verbatim. If the share is high, your numbers overstate real-world accuracy, and you should know that before you present them.
If you do this and the tool agent wins on your corpus, build it, the KyGround result is one study on one corpus and your measurement outranks it. The point of the study is that the platform behind it deployed the tool agent before anyone ran this comparison, and the deployed design turned out to be the loser.
Where tool-calling still earns its keep
The case for tool-calling retrieval was never really about small static corpora, and the evidence supports it where the other designs cannot follow.
Live and gated data. When records change constantly or sit behind per-user permissions, retrieval needs enforcement on the path, and a tool endpoint is the natural place for it. The cost is measurable: in a multitenant enterprise retrieval architecture, “the gated search path adds ∼19ms: auth server round-trip (∼14ms), ABAC policy evaluation (<1ms), and per-tenant store lookup (∼5ms).” Nineteen milliseconds is cheap insurance for correct access control, and you pay it regardless of which retrieval design sits behind the gate.
Large, messy, continuously updated corpora. At the far end of the corpus axis, retrieval pays for itself clearly. A supply-chain knowledge base system credits its RAG integration, built from unstructured communications, with “the resolution of about 50% of future tickets.” Whole-KB prompting is not an option there, and the retrieval layer demonstrably carried weight.
Many tools, not many chunks. When the application’s value is composing capabilities, checking inventory, then pricing, then a policy lookup, tool-calling is the architecture, with the ToolRet caveat that selecting among many tools is its own hard retrieval problem.
There is also a hybrid worth knowing about for small corpora: a preregistered comparison on a 24-paper corpus found an LLM-compiled wiki produced answers “1.9× longer in visible text (740 vs 385 words per answer) with 24.3 claims against 11.5” compared to single-round vector RAG, the study reports. Whether longer, denser answers are better depends on your application, and that study also used a single model, Claude Opus 4.7, for every generating step in both systems. It is another single-study result, but it shows the design space is wider than two endpoints.
Which architecture to build
Work through the axes in order.
Does the corpus fit in context? If it is on the order of tens of thousands of tokens and changes slowly, start with whole-KB prompting. KyGround’s 99.3% whole-KB result says you may be done. Spend the saved effort on evaluation instead of infrastructure.
If you need retrieval, how messy is the query surface? If your users type with dropped diacritics, inconsistent capitalization, or transliteration, and your corpus is small, the only direct evidence we have says vector RAG absorbs that mess far better than a lookup-tool agent (at most 2 points lost versus about 20). But try the cheap fix before the architecture swap: KyGround’s ablation removed the diacritic loss entirely and recovered about half the canonical gap just by folding accents and matching stemmed tokens inside the tool’s search. If permissions or live data keep you on tool endpoints anyway, that is the first change to make. And check the result on your domain, because the typo-robustness literature shows dense retrieval can fail hard when errors hit discriminative terms.
Does the data move or sit behind permissions? High update cadence or access control pushes you toward tool endpoints, and budget for the gating latency (roughly 19ms in the one measured design) either way.
Have you measured it on your corpus? This is the load-bearing question. The practical verdict the evidence supports is this: run a powered, surface-form-stratified bake-off on your own data before committing to either stack, and treat the embedding pipeline as optional until your own evaluation says otherwise. KyGround’s result, vector RAG at 95% against tool retrieval at 72% on canonical questions per the preprint, with the robustness gap wider still, shifts the burden of proof onto anyone claiming tool-calling is the obvious default for small corpora. It does not settle the question for yours. One preprint, one language pair, one corpus, title-heavy questions, a ~14-point sensitivity floor, and no independent replication is enough to change what you test first. It is not enough to skip the test.
Frequently Asked Questions
How much does whole-KB prompting cost compared to vector RAG per question?
per question, whole-KB prompting averaged 26,294 input tokens against 1,612 for vector RAG and 6,517 for the tool agent, and the paper calls whole-KB prompting “about four times the input of the tool agent.”
What is the minimum number of questions needed to detect accuracy differences in a bake-off?
KyGround scored 148 answerable test items and still only detected ~14-point differences at 80% power, per the preprint, so treat 150 to 200 questions per condition as the floor for margins of that size to mean anything.

Join the discussion
Share a useful perspective or ask a question about this article.