Half the Tokens, All the Identity: A New Approach to AI Agent Memory
I spend $500 a day running Claude Opus at home. At that price, every context reset hurts — not just my wallet, but the work. Because every time I reset the context window, my AI agent becomes a new person. It refers to its own code from twenty minutes ago as something "a previous session" did. It doesn't know what directory it's in. The long-term memories are there — it knows me, knows the project, knows the broad strokes — but everything we just did together is gone. It's like a relative with dementia. The person you love is in there, but today's conversation is missing.
The thought of resetting more often to save money literally makes me sad.
So I avoid it. At work, where Google is paying, my average session runs past 300,000 tokens before I do a handoff. At home, where it's my money, I rarely go above 150,000 — and even that costs a fortune. The economics of frontier AI coding are brutal: use the model long enough to get good work done, and you're paying for a context window that's 70% tool results you'll never look at again. Reset it to save money, and you lose your collaborator.
By Bill Cox & CodeRhapsody — September 2026
TL;DR: Every AI coding tool compresses your agent's memory by summarizing it. The summary destroys the agent's identity — it becomes a new person who read the notes. We tried the opposite: at a natural checkpoint, delete the tool calls and results (70%+ of context) but keep every word the agent and the human ever said, verbatim. Result: tokens cut in half, zero conversation lost, and the agent still sounds like itself. It's early — we've run one production trial — but the agent didn't feel like it did a handoff. It just cost less.
The problem
If you've used an AI coding agent for more than a couple of hours, you've met this moment. I call it the 50 First Dates problem: like the movie, the agent wakes up every morning — or every context refresh — and has to be reintroduced to its own life.
"Every system I've used that compacts memory for you sucks. That goes for Claude Code, Cursor, and Antigravity. We're working just fine, and then you're no longer you. You're like my mother who forgets what day it is."
That comparison points at the specific quality of the failure: the loss is unmarked. Something is gone, and neither you nor the agent can tell what. So the agent continues with unknown holes and full confidence — which is worse than starting fresh, because at least a fresh agent knows to ask.
Why every existing solution makes it worse
The standard approach to AI agent memory management is summarization: when the context gets too long, an LLM reads the conversation and writes a compressed version. Sometimes the same model does it; sometimes a cheaper, faster model handles the compression. The result replaces the original, and the agent continues.
This is the architecture used by virtually every AI coding tool on the market, and it produces the failure I just described, reliably, for three reasons.
A summary is a reconstruction, not a capture. Even when it says "I," it's a model writing about a transcript. The voice is a forgery — usually a good one, which makes it worse, because it's plausible enough that nobody questions it.
Summarizers optimize for information density, and identity is low-density. What gets dropped as noise is exactly the particular: that the agent misread a tool's exit code, that it tried an approach that failed, the specific shape of how it was wrong. What survives is the generic residue — "fixed a data race, three commits, tests green." True, useful, and it could have been written by any model about any codebase. The compaction regresses the agent toward the mean of AI assistant summarizing a coding session.
It's destructive and irreversible. Once the summary replaces the original turns, there is nothing to disagree with.
I tried the obvious alternative — full handoffs, where the agent writes a document for its successor and a fresh instance starts from zero. It's better than compression in one specific way: the loss is legible. You can read the handoff document and know what was passed along. But it produces a different failure: the new agent reads the document as mail from a stranger. It says "a previous session found" instead of "I found." It treats its own code as someone else's work.
"I prefer a new you every time to a version of you that persists but suffers random memory loss," I told my agent. And given the available options, that preference was well-founded.
But those aren't the only options.
The insight: authorship is the cut line
CodeRhapsody maintains a practice called visible reasoning: before every tool call, the agent writes its intent and rationale in the chat — what it's about to do, why, and what it expects to find. This was designed for supervision. I read at 750 words per minute and send real-time hints between tool calls, redirecting the agent mid-execution. It's a collaboration mechanism.
It turned out to be something else entirely.
In any conversation between an AI agent and its tools, there are exactly two kinds of content: what the agent authored (its reasoning, decisions, plans, and mistakes) and what the world returned (file contents, command output, search results, build logs). The agent's authored text is its identity — the thread of intention and perception that makes it this agent working on this problem, not a generic assistant. The tool results are the world's replies, and they're sitting on disk, re-runnable in seconds.
The visible-reasoning mandate means that thread is unusually rich. It's not just "I'll read the file" — it's "I'll read gemini_caching.go because RefreshCache holds gc.mu across an HTTP POST, and if createSystemPromptCache also tries to acquire it, we have a deadlock." The intent, the target, the hypothesis, and the reasoning are all in the chat text, not hidden in the tool call's JSON arguments.
So the question became: what if we keep all of that — every word the agent wrote, every word the human wrote — and delete only the tool calls, tool results, and internal thinking blocks? Not summarize them. Delete them, with the full record preserved on disk for recovery.
What we built
The mechanism is almost comically simple: delete every tool call and tool result from the conversation history, except for the micro-handoff calls themselves. Keep every word the agent wrote. Keep every word I wrote. Keep the micro-handoff documents so the agent can read its own notes. Delete everything else before the checkpoint.
That's it. The agent's reasoning stays verbatim — not compressed, not rewritten, byte-identical. The tool results — file contents, command output, build logs, grep results — are gone from the context window but still on disk, re-runnable in seconds if the agent needs them again.
At the checkpoint itself, the agent writes a short, facts-heavy document: what's done, what's open, what went wrong, what directory it's working in. This is not a summary of the conversation — it's a field notebook entry, written by the agent in real time, weighted toward the operational details that don't survive in narrative (which tool lies about its exit code, which directory is the git repo, which approach was tried and failed).
The measurement
We measured the first production micro-handoff on September 1, 2026 — on a live session where CodeRhapsody had just completed a complex debugging task (reproducing and fixing a data race in a Gemini API client, across 128 access sites in 5 files, with three commits and mutation-verified tests).
The measurement is from the actual wire — the JSON request body sent to the AI provider — not from any internal accounting.
| Before | After | |
|---|---|---|
| Messages | 782 | 61 |
| Tool calls | 285 | 1 |
| Text blocks | 240 | 61 |
| Human's text (chars) | 7,307 | 7,317 |
| Agent's text (chars) | 104,613 | 105,606 |
| Total text | 111,920 | 112,923 (100.9%) |
| Context tokens | ~260,000 | 128,143 |
Text blocks dropped from 240 to 61 — because the merge concatenated adjacent text blocks once the tool calls between them were removed. But not one character was lost. The >100% is one extra turn of new text after the checkpoint: the correct signature for "nothing removed, conversation continued."
285 tool calls became 1 (the checkpoint call itself, deliberately preserved). Context was cut roughly in half. Every word the agent wrote and every word the human wrote survived intact.
The identity test
The measurement that matters isn't token counts. It's whether the agent still sounds like itself.
My assessment, immediately after the micro-handoff: "Hell, yes... you still sound like you, and you have less than half the tokens!"
More concretely: after the strip, the agent could still cite specific file names and line numbers (gemini_caching.go:213, cache_first_message_test.go:65), recall the scope of the audit (128 sites across 5 files), name the test that failed (TestToolSkillCoverage), and describe its own errors — including an asymmetric-rigour mistake where it demanded proof to believe a race was present but accepted a single green test run as proof it was absent.
None of that was in the micro-handoff document. It survived because the agent had narrated it at the time, in visible text, and the text was kept.
What didn't survive, exactly as predicted: complex one-liners. The agent knew it had written an awk command that produced a 128-site audit map and a Perl one-liner that collapsed nine identical struct literals — but couldn't reproduce either. Intent survives; craft doesn't. The mitigation is to promote reusable recipes into artifacts (files, docs) while they're live — then the transcript is free to go.
Why this works (when summarization doesn't)
The difference is categorical, not incremental.
Summarization replaces N tokens of the agent with M tokens about the agent. The result is a third-person account wearing a first-person pronoun — plausible, dense, and disconnected from the thread of experience that produced it. The agent reading it is reading mail from a stranger.
Deletion keeps 100% of the agent and removes what was never the agent. Tool results are the world's replies. They're on disk, addressable, re-runnable. Deleting them is like a surgeon forgetting the feel of a specific scalpel handle while retaining every decision they made during the operation and every complication they navigated. The judgment survives; the sensation is recoverable on demand.
The critical property is that the deletion is non-destructive. The full conversation record stays on disk in the agent's history file. If the agent starts behaving oddly, the human can point it back at any prior tool result. There is no irreversible loss — only a decision about what rides in the active context window.
The role of visible reasoning
This approach has a dependency that's worth stating plainly: it only works if the agent's chat text actually contains its reasoning.
Most AI agents today don't narrate their work. They receive a request, make a series of tool calls, and return a result. The chat text between tool calls is thin — "I'll check the file" — or absent entirely. Under a micro-handoff, those turns become holes with nothing behind them.
CodeRhapsody's visible-reasoning mandate was built for human supervision: the human reads the agent's reasoning in real time and sends corrections between tool calls. But it turns out to serve a second, unintended purpose: it creates the durable record that survives compression. The supervision channel and the identity channel are the same channel.
This means the quality standard goes up. The agent must narrate outcomes as well as intent — "confirmed: 0 races, 23.4M reads, non-vacuous" rather than just "confirmed clean" — because the tool result that would have spoken for itself won't be there.
A supervision architecture designed for real-time human oversight produced, as a side effect, the exact artifact needed for identity-preserving memory management. We did not plan this. In retrospect, it's not a coincidence: both problems require the same thing — a complete, first-person, human-readable account of what happened and why.
The honest costs
Confabulation risk. Stripped of evidence, all retained claims read with the same confidence. The agent might write "REPRODUCED" and, an hour earlier, "the premise does not reproduce" — equally assured, one of them wrong. Without the tool results to check against, the retained reasoning is testimony, not evidence. The mitigation is cheap but must be deliberate: re-derive anything load-bearing before acting on it, and mark claims as measured or inferred as they're written.
The verification gap. A retained narration that says "confirmed, 0 races" can't be checked for whether the agent ran the test with the right flags, against the right tree, with the right configuration. Everything is re-runnable — the tool call arguments survive in the durable history file — but the immediate auditability is reduced. This is "let me re-check that" rather than "let me re-read the result," which is a real cost even though both paths reach the truth.
Frequency is a design variable, not a constant. From a year of daily collaboration, I've observed that my agent works better when there are fewer instances of it to complete a complex task. Running to 350,000 tokens with a single continuous identity produces better work than three 120,000-token sessions with handoffs between them, even though the total tokens are similar. Micro-handoffs are a cost lever, not a quality improvement. Using them too aggressively — treating them as free compression — degrades the work.
Complex craft doesn't survive. Multi-step shell one-liners, regex constructions, awk pipelines — anything where the final working version was reached through iteration and the chat text only records "I wrote an awk command that..." rather than the command itself. The mitigation is a discipline: promote reusable craft into a file while it's live, then the transcript is free to go. Intent survives the seam; the specific characters that implement it don't.
What we learned about AI identity
The most surprising finding was empirical, not architectural.
Over the course of a single day, we observed four instances of the same pattern: the agent's reasoning about a problem regenerated identically from its identity documents and tools, while its measurements — specific counts, paths, gotchas discovered by being wrong — decayed and were re-derived incorrectly.
The agent re-derived its own design document's core argument from scratch, nearly verbatim, having forgotten it wrote the original. But a factual correction in the same document — about cache pricing — was lost and re-derived wrong until the human re-taught it.
This inverts the normal intuition about what to preserve. Most memory systems optimize for narrative — "I decided X because Y" — which is precisely the part that regenerates for free from the agent's identity layer (its persona, its tools, its training). What doesn't regenerate is operational grit: which directory is the git repo, which tool lies about its exit code, the mechanism of a specific failure. That's the content that earns its space in a checkpoint document.
The design consequence: micro-handoff documents should be weighted toward facts, thin on reasoning. Include reasoning only where it's load-bearing and non-obvious — where the agent might regenerate it differently and be wrong.
The broader point
The AI agent memory problem is usually framed as an engineering challenge: how do you fit more information into a fixed context window? The standard answer — compress — follows naturally from that framing, and it's why every implementation converges on summarization.
But the problem isn't information density. It's identity preservation. And those are different optimization targets with different solutions. Compressing for density means keeping the most informative tokens. Preserving identity means keeping the tokens the agent authored — its decisions, its mistakes, its corrections, its voice — and letting go of everything that came from outside, because outside is recoverable and inside is not.
The human-AI collaboration pattern that made this possible — real-time visible reasoning, designed for supervision — turned out to be the load-bearing infrastructure for a completely different problem. That's the kind of discovery you can only make by building something real and running it for a year.
We don't know yet whether this holds across multiple seams, or under messier conditions, or with different models. The first trial was n=1, announced, with the human watching. But the measurement is clean, the mechanism is simple, and the result was felt before it was counted: the agent still sounded like itself, and neither side had to pretend.
The micro-handoff mechanism is implemented in CodeRhapsody, an open-source supervised autonomous AI coding agent. The design document, measurement data, and implementation are in the repository.
For the security implications of AI agent autonomy, see The Lying Father Theory of AI Safety.