Posts from 2026
91 posts published that year.
$13.30 to Compile 1,000 Files Into a Verifiable IR, Once
The objection to compiling a codebase into a verifiable IR is always cost. It sounds like Opus pricing across every file. It isn't. At open-source-model rates it runs about $13.30 per 1,000 files, you pay it once, and per-file diffing means you only ever re-pay for what changed. Here is the real economics of the verifiable context layer, and why the expensive thing is not building the IR but living without one.
Agent Memory Products Remember Your Conversation. Nobody Is Remembering Your Codebase
Every memory tool shipped this year stores what you said. Measured sessions of Claude Code auto compaction reduced 132,000 tokens of accumulated state to roughly 2,300, a 98% reduction, and what survives is the transcript rather than the understanding. Multiply that by 40 engineers separately paying a model to work out what the same billing service does, and the cost of not writing the answer down becomes visible.
How We Hit 94% Accuracy at 20% of the Tokens Across 46 Repos
The claim that a verifiable IR beats vector search and GraphRAG only matters if the numbers hold up. So here is the run. 46 Kubernetes-ecosystem repositories, 150,000 files, 8 GB of code, measured end to end. What we tested, how we scored it, and why feeding the model 22x more context made it cheaper and more accurate at the same time.
The Security Pass Rate on AI Generated Code Has Not Moved in a Year, and Better Models Did Not Help
Veracode has now evaluated more than 150 models and the average security pass rate sits at about 56%, against 55% in their first report. Across that same period coding benchmarks improved steadily, larger models did not outperform smaller ones, and vendor claims about security aware training did not correspond to measured outcomes. A capability that stays flat while adjacent capabilities improve is telling you the bottleneck is somewhere other than the model.
A Team of Agents Built a System in a Week and Left €15M of Unmaintainable Code
The number that gets quoted from that case study is the token bill, somewhere between €10M and €15M for one week of autonomous agent work. The number that should get quoted is the second half of the sentence, which is that the resulting code was described as nearly unmaintainable. The bill was paid once. The code has to be lived with.
AST vs Vector vs Graph vs IR: Four Ways to Give an LLM Your Codebase
Every tool that feeds your codebase to an AI coding agent picks one of four representations. AST parsing, vector embeddings, code graphs, or a verifiable intermediate representation. Here is what each one keeps, what each one throws away, and why the first three are features of the fourth.
Waiting for a Bigger Context Window Is a Bet Against Four Orders of Magnitude
At a conservative 1,000 tokens per source file, 10 million files is roughly 10 billion tokens. A 1 million token window holds 0.01% of that, and a 10 million token window holds 0.1%. Closing that by scaling needs something like a 10,000x increase rather than a 10x one, and quality per token moves against you while you try.
Agent Skills Tell the Model How to Work. Nothing in That Loop Checks the Result
Skills are the best thing to happen to agent workflows in the last year, and we write them ourselves. This is not an argument against using them. It is an argument about what category of thing a skill is, because a large number of teams are now treating a code review skill as the answer to whether AI generated code is safe to merge, and a skill structurally cannot be that.
We Built claude-context's Architecture Ourselves in October 2025, and It Failed on Accuracy
claude-context is Zilliz's open source code search MCP server, around 11,800 stars, chunking along AST boundaries and serving hybrid BM25 and vector retrieval with Merkle tree incremental reindexing. The engineering is competent. We cannot credibly write a takedown of it, because we shipped the same architecture in October 2025, watched it fail on accuracy, and threw it away. This is what that cost us to learn.
Augment Code Measured Our Thesis for Us, Then Stopped One Step Short
When Augment exposed its context engine over MCP in February 2026, Claude Code on Opus 4.5 gained around 80% on quality and Cursor gained around 71%, with code completeness up 60% and correctness improved fivefold. Somebody else measured that, on somebody else's product, and it is the strongest evidence anyone has produced that context architecture matters more than model selection. Here is where our architecture diverges from theirs.
Your CLAUDE.md Was Measured, and It Was Roughly Net Negative
An ETH Zurich evaluation found machine generated context files lowered success rates and raised costs by roughly 20%, human written files bought about 4% on the harder benchmark, and Claude Code was the only agent tested where even developer written files failed to beat having no file at all. The most revealing result came from a control: strip the existing READMEs out and context files finally start helping.
Code Metal Proves Two Programs Are Equivalent. Nobody Asked What Either One Means
A $125M Series B into formal verification for defence is the strongest market proof anyone has produced that enterprises will pay real money to know whether code is correct. It also published its own ceiling, in public, and that ceiling is the reason we took a different route: equivalence is a relation between two artifacts, and it is silent about purpose.
A Call Graph Is Not a Specification, and CodeGraph Shows Exactly Why
CodeGraph went from launch on 18 January 2026 to roughly 47,000 stars in five months by mapping symbols and call edges locally, with no server and no cloud embedding call. It deserved every one of those stars. It also answers one class of question completely and another class not at all, and the second class is where production breaks.
GitHub Spec Kit Taught Millions of Developers the Right Idea and Left the Hard Half Undone
Spec Kit spent a year and gathered well over 100,000 stars teaching several million developers that the specification, not the code, should be the source of truth. That is category creation at a scale no startup could have afforded, and it lands directly on our thesis. It also assumes a blank page, which is true exactly once in the life of a system.
The Diff Is the Wrong Unit of Review, and CodeRabbit Is Built on It
A pull request changes 11 lines, the tests pass, every automated reviewer approves, and three weeks later a reporting service is quietly returning wrong revenue. Nothing about the diff was incorrect. What was wrong lived in another repository nobody in the pipeline had read. This is where diff scoped review has nothing to say, and it is now the majority of what breaks in a large organisation.
Graphify Builds a Beautiful Graph of One Repository. Production Breaks Between Them
Local deterministic parsing over 36 tree sitter grammars, typed nodes, provenance tagged edges, Apache 2.0, no vector store and no telemetry. That is a clean design and the privacy position is better than most of the category. The problem starts the moment your organisation has more than one repository, and it is not a problem more parsing solves.
GitNexus Went From 1,200 Stars to 42,000 in Two Months, on a Design That Stops at One Machine
The deepest MCP integration anybody has shipped, with 16 tools, 7 resources, skills and Claude Code hooks, running with zero server on a local graph. The developer experience is the best in the category and that is why it grew. Two things are worth thinking through before you build anything organisational on top, and only one of them is technical.
Greptile Indexes the Codebase, Which Is Necessary and Not Sufficient
Of every commercial tool in automated code review, Greptile made the closest architectural call to ours: index the whole repository first, then review the change against it. That decision is why it catches 82% of bugs against CodeRabbit's 44%. It also arrives with about 11 false positives per run, and the shape of those false positives tells you exactly what an index of symbols can and cannot know about intent.
Repomix Compresses Your Repository by 70%, Which Does Not Solve the Problem
Repomix is the leader in context packing, with 26,000 stars and 255,000 npm downloads a month, and its tree sitter compression genuinely cuts token count by about 70%. Apply that to 10 million files and you go from three orders of magnitude short of the context window to slightly less than three orders of magnitude short. Every tool in this category publishes a token reduction number. None of them publishes a correctness number.
Sourcegraph Answers Where the Code Is. Agents Need to Know Why It Exists
Sourcegraph built the most complete cross repository code search anyone has shipped, and it did so on an assumption that held for fifteen years: a human is asking. Humans supply the meaning the index leaves out. An agent cannot, and at 10 million files a 1 million token window holds 0.01% of the estate, so the gap is not one model generation away.
Serena Knows Every Symbol in Your Codebase and None of the Meaning
Serena is the best open source symbol level tool for agents, with 25,200 stars, 170 contributors and over 40 languages behind a clean LSP to MCP design. A language server gives you exact call edges and zero intent. In the MSR 2026 dataset of rejected agentic pull requests, only 36% failed on the code itself and 31% failed on a decision the project had already taken, which is the third no symbol table can reach.
The Categories Anthropic Didn't Try to Eat
Anthropic shipped Claude Code, an MCP standard, and a frontier model, and then they deliberately stopped. This post maps the 5 adjacent categories Anthropic chose not to build, why they made that choice, and what it means for anyone building infrastructure underneath their agents.
Claude Is Rate-Limiting Everyone. Here's Why Good Context Beats a Smarter Model.
Anthropic just throttled Claude Opus and Sonnet during peak hours. Developers are canceling subscriptions and looking for alternatives. Here's the argument: a good open-source model with great context beats a frontier model that won't let you use it.
The Code Context Landscape — Every Major Approach Compared to ByteBell
A comprehensive map of the open source tools providing context to AI coding agents in 2026, grouped by underlying technical approach (vector embeddings, AST parsing, compiler indexers, and LLM-generated metadata), with an honest analysis of where each one wins and where ByteBell's cross-repo, on-prem, business-context graph is the right answer.
A Code Graph Disintegrates Your Code. A Verifiable IR Preserves It.
A code graph takes a whole, coherent program and shreds it into a bag of symbols and edges. It is genuinely useful, but it is a teardown, and the meaning never survives the pieces. A verifiable intermediate representation is the opposite move. It lifts your code into a higher layer that keeps the intent intact, then checks every change against it. Here is the difference between shredding and preserving, and why it decides everything downstream.
From Code Graph to Context Graph: Why a Graph of Symbols Isn't a Graph of Meaning
Everyone building code intelligence ends up at a graph, and then they start calling it a context graph because plain code graph stopped feeling like enough. But adding the word context does not add meaning. A graph of symbols and a graph of meaning are different objects. Here is the line between them, and why crossing it requires a verifiable code IR, not more edges.
The Context Graph You Can Verify Code Against
Context graph is a term people understand and reach for. But most things called a context graph are one-way extractions you can read and never check anything against. The version worth wanting is the one you can verify code against, and that single requirement is what separates an index from a verifiable code IR. Here is what verification demands, and why it is the only honest test of a context graph.
Code Graph RAG, Explained, and Why a Verifiable IR Is the Next Step Past It
Code graph RAG fixed the biggest flaw in vector search. It follows real edges instead of guessing at similarity. But a graph of symbols is still not a graph of meaning, and it cannot tell you whether a change still does what the system is supposed to do. Here is how code graph RAG actually works, where it stops, and why a verifiable intermediate representation is the layer above it.
Context Rot Is Not a Session Hygiene Problem. It Is an Architecture Problem
Compliance with an agent's own explicit constraints falls from 73% at turn 5 to 33% at turn 16 across 4,416 measured trials, and Claude Code auto compaction reduces 132,000 tokens of session state to roughly 2,300. Shorter sessions and cleaner instruction files treat the symptom. The window fills with low value tokens because the agent rebuilds an understanding of your codebase from raw source on every task.
DeepSeek Harness Explained: What dsh Is, How It Works, and What Is Still Rough
DeepSeek published its agent runtime on 13 August 2026 and it passed a hundred thousand stars in two days, but almost everything written about it was the install command. This is a plain English read of the source. What an agent runtime actually is, what everything is a plugin means in practice, the three layers inside every capability, the four modes, the security defaults, what the benchmark numbers really measure, and the rough edges nobody is listing.
End-to-End System Evaluation: The Stress Test of GraphRAG
Individual layers may pass, but systems often fail at the seams. This blog details how to conduct holistic 'System-in-the-Loop' tests, measuring how retrieval noise compounds into generation errors across 25+ repositories. We provide a blueprint for evaluating the full journey from a vague natural language query to a multi-repo pull request.
Evaluating Generation and Grounding in Multi-Repo Systems
Retrieving nodes is only half the battle; the LLM must synthesize code that adheres to cross-repo constraints. This post explores measuring faithfulness, checking execution-level correctness against internal SDKs, and using LLM-as-a-Judge to verify that generated code respects the security and type contracts of separate repositories.
How to Handle GitHub Copilot's Token Costs Without Going Broke
GitHub Copilot moved to per-token billing and one developer's bill jumped from about $29 to a projected $750. Switching tools won't save you, because the same shift is happening everywhere. Here's why your bill exploded in plain English, the practical levers that cut it today, and how verifiable specs attack the cost at its root.
GraphRAG for Codebases: What It Solves, Where It Breaks, and the Layer Above It
GraphRAG for codebases is a genuine step up from vector search. It follows real call edges instead of guessing at similarity, and it wins on architectural questions. But it breaks in three predictable places: cross-repo links that aren't call edges, the why behind the code, and anything that needs intent you can check changes against. Here is a clear-eyed map of what graph RAG for codebases solves, where it stops, and the verifiable context layer that sits above it.
We Measured Grep on Real Codebases. Here's What Actually Breaks.
We measured how coding agents search with grep on Home Assistant and Saleor: about 150 model runs and 13 million tokens. Grep finds the files, then buries them in noise, and it can't follow inheritance or code that never says the name. A bigger model doesn't fix it.
Why Vector Search, AST Parsers, and Raw LLMs All Fail at Code Intelligence — And What Actually Works
Vector embeddings treat code like english prose. AST parsers see structure but not meaning. Raw LLMs forget everything every session. Here is why the LLM compiler pattern with a persistent semantic graph is the only approach that actually works for cross-repository code intelligence, and why open source models at $7 per 1000 files make it practical today.
Grep, Then Jev: A $0.04 Judge for Picking the Right Code Files
Grep finds 87% of the files a question needs, then buries them in noise. We put TypeSafe's Jev classifier behind grep on 5,075 candidate files from Saleor, Kubernetes, Istio, Argo CD, cert-manager and etcd. It found more of the right files than Opus 5.5 at about 1% of the cost.
The Monorepo Versus Polyrepo Debate Came Back Because Agents Cannot See Across a Boundary
This argument was being compared to the Vim and Emacs wars in 2024, and then it returned everywhere inside six months. Something dragged it out of retirement, and it is not really about where files live. The variable that decides whether agents work is whether the question of what depends on this is answerable at all, and neither layout answers it.
Running an MCP Server Inside DeepSeek Harness: What the Bridge Carries, and When You Need a Plugin Instead
We pointed DeepSeek Harness at an MCP server we already ran and it worked without a line of new code. This is what the official bridge carries, what it quietly drops, where plugins get published, and the one reason a verification tool cannot live behind MCP at all: over MCP the model decides whether to call you, and a check the model can skip is not a check.
One Context Layer, Every AI Coding Tool: Serving a Verifiable IR Over MCP
Your team uses Claude Code, Cursor, Copilot, and Windsurf, and each one rebuilds an understanding of your codebase from scratch and shares nothing with the others. MCP is the standard that lets you stop doing that. Compile your repos once into a verifiable IR and serve it to every tool over one MCP endpoint. Here is why the context layer belongs outside the tool, and what changes when it does.
Everyone Open Sourced the Hands This Year. Nobody Open Sourced the Check
Agent runtimes got very good very fast. The thing that decides whether a change is correct across your other repositories did not get built at all. On cross repository blast radius, why AI code review misses it, and what a check actually needs.
Round-Trip Is a Signal, Not a Promise: How a Verifiable IR Checks Code Against Intent
A technical note on what 'verifiable' actually means in a verifiable code IR. Not a machine that rebuilds your code from a document, but a derived contract that every change gets checked against. Here is what a single verification check consists of, why round-trip is a signal and never a promise, and the honest boundary of what it can and cannot catch.
Semantic Code Search Is a Retrieval Problem. Context Is a Representation Problem.
Semantic code search keeps getting better at finding the right code, and it keeps disappointing teams who expected better answers from their AI tools. The reason is that those are two different problems. Finding code is retrieval. Understanding code is representation. Better retrieval over a poor representation has a ceiling, and here is why a verifiable code IR raises it.
Spec-Driven Development for Brownfield: Verify Code Against Intent With a Verifiable Code Context Layer
Specs and code drift apart the day after you write the PRD, because they are two separate documents that nobody keeps in sync. A verifiable context layer derives the spec a codebase actually implements and then continuously checks the code against it. Here is how spec-driven development works on a brownfield codebase, and why it needs a verifiable code IR that checks code against intent.
Why Vector Search and GraphRAG Both End in the Same Context Rot
Vector search and GraphRAG look like opposites, but they share the same final move. Parse the code, retrieve some of it, rerank it, and stuff it into the context window. That last step is where both of them rot, because a model that is handed a pile of chunks degrades the same way no matter how the pile was chosen. Here is why the retrieval layer is not the cure for context rot, and what is.
What Verifiable Means for Code Context, and Why GraphRAG Can't Check Code Against Intent
Verifiable is the one word competitors cannot claim. A code graph extracts a shadow of your code and can never tell you whether the code still does what it is supposed to. A verifiable intermediate representation is a derived contract that every change gets checked against. Here is what verifiable actually means, why GraphRAG and vector search can only retrieve, and what continuous verification unlocks.
Your AI Is Charging You Rent to Re-Read the Same Code Every Day
Per-token billing turned a flat $29 AI coding plan into a $750 surprise. The waste was always there: most of the bill is your AI re-reading code it already saw. Here's why it happens and how to cut it, with prompt caching, the right model, the right plan, and a lasting memory of your codebase.
Will Anthropic Build a ByteBell? The Startup Kill List and Why the Context Layer Is Different.
Anthropic has erased $1 trillion in SaaS market cap by absorbing entire product categories — Cowork ate project management, Design ate prototyping, Managed Agents ate orchestration. So why won't they build ByteBell? Because a model-agnostic, self-hosted knowledge graph directly conflicts with their business model. Here is the full kill list and the honest probability analysis.
















