Benchmark

How we measure the effect of a curated context layer — the corpus, the comparison, the measured figures, and the honest caveats.

The claim on our site is that an agent reading a curated context layer does the same work as an agent searching the raw data, with fewer tokens, fewer searches, and fewer tool calls. That's measurable, so we measure it — on a corpus anyone can obtain, with every paid call receipted.

The corpus

A pinned commit of a large, real, publicly available body of company documentation: 4,161 Markdown files, roughly 9.2 million tokens of cross-functional operational prose. It's the GitLab handbook, licensed CC BY-SA 4.0.

Real matters. Synthetic corpora reward whoever wrote them, and a corpus small enough to fit in a context window makes the whole question moot. Every configuration reads byte-identical content, and the harness verifies that before anything runs.

The comparison

Not question-and-answer trivia. Each run is a work session: one ambiguous, job-shaped brief — prepare a new manager's first-month pack, brief the leadership meeting — that the agent must complete autonomously, deciding for itself what to look up. Five such sessions, run three times each, per configuration.

The published comparison is the cleanest one: the same agent, the same model, the same ingested data — with and without the curated layer on top. Only retrieval changes.

Deliverables are graded against a checklist of specific, source-traceable facts, and a fact only counts if the grader can locate a verbatim supporting quote in what the agent actually retrieved. A separate closed-book control (same model, no data access) confirms the results reflect retrieval rather than the model's training memory.

What we measure

The question is token efficiency: both configurations were free to search as much as they needed to finish the job well — the deliverables were graded to the same standard, and quality held level between them — so what separates the arms is what it costs to get there.

Tokens per session

28% fewer
With Pensieve
542k
Searching raw data
753k

mean input tokens over 15 work sessions per arm

Searches per session

46% fewer
With Pensieve
10
Searching raw data
19

mean search calls per session (152 vs 282 in total)

Tool calls per session

32% fewer
With Pensieve
30
Searching raw data
44

mean tool calls per session (445 vs 658 in total)

Quality holding level is the condition that makes these numbers a result at all: you can always spend fewer tokens by answering worse, so a token count only means something alongside the grade of the work it bought.

The honest caveats

  • Tokens are not money. Providers discount cached reads heavily, so on a well-cached stack the cost gap is smaller than the token gap. That is why we publish tokens, searches and tool calls, all of which we measured, and no pound figure: pricing one needs a provider's list rate and an assumption about how much of a re-sent prefix actually hits the cache, and neither is a measurement. What we can give you is the split, per configuration, from the sessions' own receipts: how much input was first-time reading, billed at the cache-write rate, and how much was the conversation's own prefix re-sent on later turns, which is the cacheable part. Price that against your own provider and your own hit rate. Machine-paced agent sessions cache well; human-paced sessions with pauses cache badly, because provider caches expire in minutes.
  • One corpus. This is a single, unusually well-organised body of documentation — closer to a best case for raw search than a typical company's scattered tools. That cuts against us, which is the right direction for a published claim.

Reproducing it

The corpus is public and pinned to a commit. The session briefs, grading checklists and per-configuration token accounting are retained with the results, and every paid call is receipted, so any figure can be traced back to the calls that produced it.

If you re-run it and get something different, we would like to know.