
faq-general
Does tool-mediated memory actually beat stuffing the context window?
General FAQ · Evidence
That is a measurable question, so we measure it. ScoreCrux is CueCrux's open benchmark; the headline findings as of July 2026:
- Recall at scale. At 2M-token scale, stuffed-context recall collapses: 44% down to 28% on one model, 28% down to 8% on another. Tool-mediated memory through the daemon holds 80-100% on the same tasks.
- Answer quality per token. Near-identical quality at a fraction of the context tokens, tool-mediated memory against vendor-native context-stuffing on the open ScoreCrux benchmark.
- Budget discipline. Hard token budgets return far fewer tokens than naive top-K retrieval.
The benchmark is open precisely so you do not have to take our word for the numbers: run it against your own stack and corpus. Receipts over vibes applies to our marketing claims too.