Models
How Pensieve assigns models to background work and Chat, and which models we currently recommend.
Pensieve does its thinking with AI models — reading your documents and writing the pages that make up your company's map. Pensieve manages both the provider credentials and one universal model policy for every Context.
This page is our running guide to the models Pensieve operates. It is based on something more useful than leaderboards: how each model actually performs at building a context layer. A model that tops the public benchmarks is not automatically the one that maps a company best, so we test that directly and share what we find here.
The three tiers
Pensieve splits background work across three tiers. They do very different jobs, so the best model for one is often the wrong choice for another. Interactive conversation uses a separate Chat slot, which defaults to GPT-5.6 Luna.
Each slot also carries a reasoning effort, which controls the amount of reasoning requested for that work.
Low — The first read
Takes the first pass over new material as it arrives and decides where the evidence belongs. High volume, and more judgement than the name suggests — which is why we run it at more reasoning effort than the tier's name implies, rather than picking the cheapest thing available.
Default reasoning effort: low — set per slot, and yours to change.
| Model | Input / M | What we found |
|---|---|---|
| Current pick OpenAI: GPT-6 Luna | $ · $0.10 | Not yet benched for context-building — we publish findings here once it has been through our tests. |
| Z.ai: GLM 5.3 Flash | $ · $0.15 | Not yet benched for context-building — we publish findings here once it has been through our tests. |
| OpenAI: GPT-5.6 Luna | $ · $0.20 | Reliable first reader. Runs the low seat at raised reasoning effort in the map we currently rate highest. Keep it out of the medium tier — it lost that seat in the same tests. |
| MiniMax: MiniMax M3 | $ · $0.30 | Solid & cheap. Good coverage at the lowest cost, with slightly looser connections between pages. In our latest bench it struggled with the medium tier's editing protocol — use it in the low tier, not the workhorse seat. |
| Qwen: Qwen3.8 Flash | $ · $0.15 | Not yet benched for context-building — we publish findings here once it has been through our tests. |
| Qwen: Qwen3.7 Plus | $ · $0.32 | Dependable all-rounder. Nearly as capable as MiniMax M3 at effectively the same price — the natural pick if you would rather keep everything in the Qwen family. |
Medium — The workhorse
Where most of the work, and most of the quality, lives: reading documents in depth and writing what they say into the pages of your map. It is also the busiest tier, so its cost adds up the fastest.
Default reasoning effort: high — set per slot, and yours to change.
| Model | Input / M | What we found |
|---|---|---|
| Current pick Z.ai: GLM 5.3 Flash | $ · $0.15 | Not yet benched for context-building — we publish findings here once it has been through our tests. |
| OpenAI: GPT-5.6 Luna | $ · $0.20 | Reliable first reader. Runs the low seat at raised reasoning effort in the map we currently rate highest. Keep it out of the medium tier — it lost that seat in the same tests. |
| OpenAI: GPT-5.6 Terra | $$ · $2 | Not yet benched for context-building — we publish findings here once it has been through our tests. |
| SpaceXAI: Grok 4.5 | $$ · $2 | Previous workhorse. Won the medium tier in the previous campaign where cheaper models failed its agentic editing protocol. It remains the first fallback while Terra is validated on Pensieve's workload. |
| MiniMax: MiniMax M3 | $ · $0.30 | Solid & cheap. Good coverage at the lowest cost, with slightly looser connections between pages. In our latest bench it struggled with the medium tier's editing protocol — use it in the low tier, not the workhorse seat. |
| Qwen: Qwen3.7 Plus | $ · $0.32 | Dependable all-rounder. Nearly as capable as MiniMax M3 at effectively the same price — the natural pick if you would rather keep everything in the Qwen family. |
| Qwen: Qwen3.8 Flash | $ · $0.15 | Not yet benched for context-building — we publish findings here once it has been through our tests. |
High — The architect
The lightest-touch but highest-leverage tier: it designs the shape of your tree, writes your company overview, and audits the map for claims that have drifted. It runs rarely, so you can afford your smartest model here.
Default reasoning effort: high — set per slot, and yours to change.
| Model | Input / M | What we found |
|---|---|---|
| Current pick OpenAI: GPT-6.1 Sol | $$ · $2 | Not yet benched for context-building — we publish findings here once it has been through our tests. |
| OpenAI: GPT-6 Sol | $$ · $2 | Not yet benched for context-building — we publish findings here once it has been through our tests. |
| OpenAI: GPT-5.6 Sol | $$ · $2 | Not yet benched for context-building — we publish findings here once it has been through our tests. |
| MoonshotAI: Kimi K3 | $$ · $1.19 | Previous architect. Anchored the high tier in our best-measured map — strongest structural consistency we've benched, at a small share of the total build cost. It remains the first fallback while Sol is validated on Pensieve's workload. |
| Anthropic: Claude Opus 5.5 | $$$ · $4 | Not yet benched for context-building — we publish findings here once it has been through our tests. |
Models and billing
Chat and background inference are included in your plan. Model calls are not a separate customer charge. Plans & billing explains knowledge units and your allowance; Settings → Usage shows operational activity.
Up next
MCP reference