Benchmark
How we measure the effect of a curated context layer — the corpus, the comparison, the measured figures, and the honest caveats.
The claim on our site is that an agent reading a curated context layer does the same work as an agent searching the raw data, with fewer tokens, fewer searches, and fewer tool calls. That's measurable, so we measure it — on a corpus anyone can obtain, with every paid call receipted.
The corpus
A pinned commit of a large, real, publicly available body of company documentation: 4,161 Markdown files, roughly 9.2 million tokens of cross-functional operational prose. It's the GitLab handbook, licensed CC BY-SA 4.0.
Real matters. Synthetic corpora reward whoever wrote them, and a corpus small enough to fit in a context window makes the whole question moot. Every configuration reads byte-identical content, and the harness verifies that before anything runs.
The comparison
Not question-and-answer trivia. Each run is a work session: one ambiguous, job-shaped brief — prepare a new manager's first-month pack, brief the leadership meeting — that the agent must complete autonomously, deciding for itself what to look up. Five such sessions, run three times each, per configuration.
The published comparison is the cleanest one: the same agent, the same model, the same ingested data — with and without the curated layer on top. Only retrieval changes.
Deliverables are graded against a checklist of specific, source-traceable facts, and a fact only counts if the grader can locate a verbatim supporting quote in what the agent actually retrieved. A separate closed-book control (same model, no data access) confirms the results reflect retrieval rather than the model's training memory.
What we measure
The question is token efficiency: both configurations were free to search as much as they needed to finish the job well — the deliverables were graded to the same standard, and quality held level between them — so what separates the arms is what it costs to get there.
Tokens per session
28% fewermean input tokens over 15 work sessions per arm
Searches per session
46% fewermean search calls per session (152 vs 282 in total)
Tool calls per session
32% fewermean tool calls per session (445 vs 658 in total)
Quality holding level is the condition that makes these numbers a result at all: you can always spend fewer tokens by answering worse, so a token count only means something alongside the grade of the work it bought.
The honest caveats
- Tokens are not money. Providers discount cached reads heavily, so on a well-cached stack the cost gap is smaller than the token gap. That is why we publish tokens, searches and tool calls, all of which we measured, and no pound figure: pricing one needs a provider's list rate and an assumption about how much of a re-sent prefix actually hits the cache, and neither is a measurement. What we can give you is the split, per configuration, from the sessions' own receipts: how much input was first-time reading, billed at the cache-write rate, and how much was the conversation's own prefix re-sent on later turns, which is the cacheable part. Price that against your own provider and your own hit rate. Machine-paced agent sessions cache well; human-paced sessions with pauses cache badly, because provider caches expire in minutes.
- One corpus. This is a single, unusually well-organised body of documentation — closer to a best case for raw search than a typical company's scattered tools. That cuts against us, which is the right direction for a published claim.
Reproducing it
The corpus is public and pinned to a commit. The session briefs, grading checklists and per-configuration token accounting are retained with the results, and every paid call is receipted, so any figure can be traced back to the calls that produced it.
If you re-run it and get something different, we would like to know.