blog

What it costs to remember

August 12, 2026

What it costs to remember

If you want an AI agent to remember things, you have three realistic options today.

Full-history context. Keep appending to one file, or just paste the history in. The model has a million-token window now, so let it read everything.

Put your notes in a file-vector store. One file per conversation or topic, retrieved by similarity, loaded whole. This is the pattern most people land on once the first option gets expensive.

Use a memory layer. Something that reads your history once, indexes what it means, and can answer questions against it.

All three keep your original text, which is the property we care about and the reason these are the three we compared. We measured what each one costs to answer the same fifty questions.


The setup

Fifty questions from LongMemEval, spanning every type it covers: facts stated once, facts that changed over time, facts assembled across several conversations, questions that turn on when something was said, and questions with a false premise that should be declined rather than answered. Each question carries its own haystack of roughly fifty prior conversations.

Every arm used Claude Opus 4.8 as the reader and GPT-4o as an independent judge.

Only the context changed.

The vector-store arm is not a strawman. We built it with the same embedding model we run in production, so it is a genuinely good version of that approach.


What it costs

Approach Tokens per question Correct vs Vertiso
Vertiso Memory 6,737 96% —
File-vector store, top 1 file retrieved 4,198 54% 0.6x
File-vector store, top 3 files 12,179 86% 1.8x
Full-history context 170,265 96% 25.3x

Result tokens, measured. This is what lands in your agent's context for each question: the retrieved content itself, with the reader's own prompt scaffolding excluded from every row so the comparison is like for like.

Two results matter here.

Full-history context matches us on accuracy and costs 25x more. At this corpus size, there is no accuracy penalty for brute force. However, there is a very high token cost.

We beat the three-file vector store on both axes at once. Ten points more accurate, and 1.8x cheaper. That is the comparison most people actually face, because the vector store is what you build when the first option tips over.

Today, there is one configuration that uses fewer tokens than we do. It answers barely half the questions correctly.


Why you can't just retrieve one file

The cheapest configuration available loads a single file and scores 54%.

It does not fail randomly. It fails on the questions whose answer is not inside any one file: facts assembled across several conversations, and dates that only resolve against when something was said. Multi-session drops to 38%, temporal reasoning to 25%. Questions answerable from a single conversation stay at 100%.

This is the structural problem with file-level memory. A file-level system cannot know in advance how many files a question needs, so it has to over-fetch to be safe. Each additional file costs about 4,000 tokens whether or not it contains anything relevant. A conversation file runs about 3,600 tokens; the sentence that answers the question is closer to 50.

Going to three files buys back most of the accuracy and triples the bill, and still lands ten points behind us while using more tokens. We retrieve the spans, not the files. That gap widens as the memory files grow, which is exactly what happens to any memory accumulating over the years.


Where reading everything actually breaks

The full-corpus arm ties us on accuracy, so the interesting question is where it does not.

Knowledge-update questions, the ones where a fact changed and the current value is the answer:

How memory works Correct
Vertiso Memory 8/8
Vector store, 3 files 7/8
Full-history context 6/8
Vector store, 1 file 4/8

Both of the full-corpus arm's misses are the same failure. It had the old value and the new value in front of it simultaneously and reported the stale one. A gym time that moved from 7:00 pm to 6:00 pm came back as 7:00. A coin collection the user had added to came back at the old count.

That is what a million-token window does not solve. Holding every version of a fact at once is not the same as knowing which one is current. Superseding is work, and somebody has to do it before the question is asked.


What we deliberately did not compare against

There is a fourth category: systems that decompose your history into facts and discard the original. You ask a question, you get an assertion, and the conversation it came from is gone.

Today, those systems will beat us on tokens. They ship less, because they kept less.

We left them out because it is not the same product. We decompose too, and then we keep your source text, linked to every fact derived from it. You can read the original, correct it, export it, and check any answer against what you actually said. Comparing token counts against a system that threw the source away would be measuring the wrong thing, and it would flatter whichever of us cared least about provenance.

If what you want is the cheapest possible assertion with no way to verify it, that category exists, and we are not in it.


What's actually in our 6,737 tokens

Almost none of it is the answer. The answer text is about 1% of what we return.

The rest is evidence: the source memories the answer came from, the specific facts inside them that supported it, and the links between those and everything else in your graph. You get roughly ten sources and five knowledge cards, and you can audit every one.

That is a deliberate trade. We could return just the answer and cut the payload by two orders of magnitude. Then you would have an assertion you cannot check, which is the category above.


What this does not prove

Fifty questions, one corpus size. Enough to separate these approaches, not enough to claim a universal constant. The ratio moves with how much history you have.

Ingestion is not in the ratio. Building memory from your history costs something once. Reading raw context costs nothing up front and full price on every question after. The honest framing is amortization, and the break-even arrives quickly, but it is not free on question one.

Prompt caching does not rescue the full-history context arm here, and we would rather say so than have it raised. In this benchmark, every question carries its own separate corpus, so there is nothing to cache across questions. In a real deployment hitting the same history repeatedly, caching genuinely helps. It narrows the gap. It does not close it, because you are still moving a corpus through a model to retrieve two hundred tokens of it.

A chunk-level system would beat us on tokens, and we should say that plainly rather than let someone discover it. Retrieving four small chunks is less context than our payload, not more. The argument against that design is not cost. It is that a chunk knows what it says and not when it was said, what it replaced, or what it belongs to, which is precisely the failure mode in the supersession table above.

We are not describing anyone's product. The vector-store arm is our own construction of a widely used pattern, built as well as we could.

Our accuracy figure comes from our published benchmark run, scored on these same fifty questions with the same reader. The comparison arms were run separately and more recently. Same model, same judge, different day.


The part that generalizes

The specific ratio depends on corpus size. The shape does not.

Reading raw context scales linearly with how much history exists. Twice the history, twice the bill, on every question, forever. And, with a real memory, the context always goes up.

Answering from a memory store is roughly flat. You pay for the answer and its evidence. A corpus ten times larger costs about the same to query, because retrieval got harder while the payload stayed the same size.

At the scale we measured, that is 25x against the full-history context. For someone two years in, it is larger. The point is not the multiple. The point is that one of these curves goes up and the other mostly does not.


What this research found in our own product

Running this experiment meant measuring our own payload byte by byte for the first time. We did not like what we found.

Of the 6,737 tokens we return, the answer is about 1%. Most of the rest is evidence, which is the point. But a large share of it was the same content two and three times over: the verbatim fragment, an LLM paraphrase of that same fragment, and then, inside a typed-attributes block, the fragment quoted a third time as its own supporting evidence. On top of that, we were shipping record-keeping fields no client will ever need: creation timestamps, archive flags, completion flags, graph edges nobody traverses mid-answer.

We were shipping the kitchen sink, and it costs unnecessary tokens

So we rebuilt the recall and search renderer around a single question: what does the receiving model actually need? The answer, when you asked for one. The fact bodies that support it, with their dates. And a short list of the source memories you can open and read yourself.

That is the whole payload now:

  • the synthesized answer, when requested, with its confidence and any gaps
  • 25 fact bodies, verbatim, each with its temporal anchor
  • the top 10 source memories as id, title, and summary, so you can open any of them

Everything else is available on demand from the endpoints that exist to serve it. None of the fields that made the answer checkable were removed. What went away was the second and third copy of it.

Tokens per question vs Full-history context
Before 6,737 25.3x
After 2,093 81.4x

A 3.2x reduction, with no material change to what you can verify. The retrieval did not change, and the accuracy did not move, because none of those bytes ever reached the model that composes the answer. They were pure overhead on your side of the wire that didn’t belong there.

The optimized payload is now smaller (about 50%) than retrieving a single file from a vector store, and it answers 96% of the benchmark correctly, compared to that approach's 54%.

We would rather find this ourselves and tell you than have you find it on your invoice.


Reproducing this

Our run is published in full, question by question, with the reference answer, our answer, the judge's verdict, and our written rationale wherever we disagree with the judge:

The do-nothing arm is published with its manifest, per-question results, and judge verdicts:

The corpus is LongMemEval's published longmemeval_s_cleaned dataset, unmodified. The do-nothing arm read it verbatim from the benchmark file and touched none of our stored data.

Four of the fifty questions have reference answers we believe are wrong, and we applied that judgment to every arm equally, including where it cost us a point. Our reasoning on each is published on the benchmark page.

Vertiso Memory is one memory store across every AI tool you use. Write a decision in Claude, recall it in Codex, Cursor, or Gemini. The memory is yours: inspectable, editable, and portable the day you switch tools.

Free tier at memory.vertiso.ai. Full benchmark results, including the ones we got wrong, at memory.vertiso.ai/benchmarks/longmemeval.