Notes from the context layer.
Engineering deep-dives, benchmarks and customer stories on persistent code context, knowledge graphs, and the AI coding tools we work with daily.
1–10 of 112 · page 1 of 12
We Measured Grep on Real Codebases. Here's What Actually Breaks.
We measured how coding agents search with grep on Home Assistant and Saleor: about 150 model runs and 13 million tokens. Grep finds the files, then buries them in noise, and it can't follow inheritance or code that never says the name. A bigger model doesn't fix it.
Grep, Then Jev: A $0.04 Judge for Picking the Right Code Files
Grep finds 87% of the files a question needs, then buries them in noise. We put TypeSafe's Jev classifier behind grep on 5,075 candidate files from Saleor, Kubernetes, Istio, Argo CD, cert-manager and etcd. It found more of the right files than Opus 5.5 at about 1% of the cost.
Agent Memory Products Remember Your Conversation. Nobody Is Remembering Your Codebase
Every memory tool shipped this year stores what you said. Measured sessions of Claude Code auto compaction reduced 132,000 tokens of accumulated state to roughly 2,300, a 98% reduction, and what survives is the transcript rather than the understanding. Multiply that by 40 engineers separately paying a model to work out what the same billing service does, and the cost of not writing the answer down becomes visible.
The Security Pass Rate on AI Generated Code Has Not Moved in a Year, and Better Models Did Not Help
Veracode has now evaluated more than 150 models and the average security pass rate sits at about 56%, against 55% in their first report. Across that same period coding benchmarks improved steadily, larger models did not outperform smaller ones, and vendor claims about security aware training did not correspond to measured outcomes. A capability that stays flat while adjacent capabilities improve is telling you the bottleneck is somewhere other than the model.
A Team of Agents Built a System in a Week and Left €15M of Unmaintainable Code
The number that gets quoted from that case study is the token bill, somewhere between €10M and €15M for one week of autonomous agent work. The number that should get quoted is the second half of the sentence, which is that the resulting code was described as nearly unmaintainable. The bill was paid once. The code has to be lived with.
Waiting for a Bigger Context Window Is a Bet Against Four Orders of Magnitude
At a conservative 1,000 tokens per source file, 10 million files is roughly 10 billion tokens. A 1 million token window holds 0.01% of that, and a 10 million token window holds 0.1%. Closing that by scaling needs something like a 10,000x increase rather than a 10x one, and quality per token moves against you while you try.
Agent Skills Tell the Model How to Work. Nothing in That Loop Checks the Result
Skills are the best thing to happen to agent workflows in the last year, and we write them ourselves. This is not an argument against using them. It is an argument about what category of thing a skill is, because a large number of teams are now treating a code review skill as the answer to whether AI generated code is safe to merge, and a skill structurally cannot be that.
We Built claude-context's Architecture Ourselves in October 2025, and It Failed on Accuracy
claude-context is Zilliz's open source code search MCP server, around 11,800 stars, chunking along AST boundaries and serving hybrid BM25 and vector retrieval with Merkle tree incremental reindexing. The engineering is competent. We cannot credibly write a takedown of it, because we shipped the same architecture in October 2025, watched it fail on accuracy, and threw it away. This is what that cost us to learn.
Augment Code Measured Our Thesis for Us, Then Stopped One Step Short
When Augment exposed its context engine over MCP in February 2026, Claude Code on Opus 4.5 gained around 80% on quality and Cursor gained around 71%, with code completeness up 60% and correctness improved fivefold. Somebody else measured that, on somebody else's product, and it is the strongest evidence anyone has produced that context architecture matters more than model selection. Here is where our architecture diverges from theirs.