← Back to Blog

Serving Memory vs. Engineering Memory

Mem0, Supermemory, Zep and Letta are building something real. It just isn't the thing your engineering team keeps losing.

Two products, one word

Every AI memory pitch sounds the same. "Your agent never forgets." "Persistent context." "Memory for AI."

So engineering leaders line the vendors up side by side and ask which one is best. Wrong question. They are not the same product. They are not even solving the same problem.

Two very different things are being sold under one word. One remembers the person your product talks to. The other remembers why your software is the way it is. I call them serving memory and engineering memory. Mix them up and you end up paying per API call to build a user profile of your own senior engineer.

I started building Recallium in 2025 because I was tired of watching agents relearn my codebase from markdown files that went stale the day they were written. This is the distinction I wish someone had drawn for me.

Serving memory remembers users. Engineering memory remembers decisions.

Serving memory lives in your product's runtime. A support bot remembers a customer's last three orders. A tutoring app remembers a student struggles with fractions. A shopping assistant remembers someone hates wool. The unit of memory is a person. The job is personalization. The scale is millions of end users, so the bill is metered by requests and tokens, because that is how traffic grows.

Engineering memory lives in your team's build process. The unit is a decision: what we chose, what we rejected, the constraint that forced it, and what replaced it later. The readers are the engineers and coding agents building the product. The scale is a team, so it is priced per engineer, because that is who it serves.

Serving memoryEngineering memory
RemembersThe person on the other side of the chatWhy the software is the way it is
Unit of memoryA fact about a userA decision, a fix, a constraint, with its reasoning
Who reads itYour product, on behalf of each end userEvery engineer and every coding agent on the team
What "current" meansThe latest preference winsThe latest decision wins, and the old one stays on record
ScaleMillions of end usersOne team, many repositories, many tools
Billed byRequests, tokens, queriesEngineer
Cost of a stale memoryA slightly worse replyA bug your security review already killed, shipped again

That last row is the whole argument. In serving memory, a stale fact is an annoyance. In engineering memory, a stale fact is an incident.

A stale memory is more dangerous than a missing one.

Where Mem0, Supermemory, Zep and Letta actually sit

These are good products. Let me place them fairly, using their own words and their own price pages, checked October 2026.

Mem0 calls itself drop-in memory infrastructure for AI agents and apps. Every tier on its pricing page lists unlimited end users and meters add requests and retrieval requests. Pro is $249 a month. That is serving memory, priced like serving memory.

Supermemory builds an auto-maintained profile for every user: stable facts plus recent activity. Its pricing draws credits down by tokens and queries. Pro is $19 with 3 team seats and unlimited end users; per-tag access arrives on Scale at $399. Serving memory again.

Zep describes itself as a unified context layer, built on a temporal context graph that tracks how facts change over time. Its pricing charges credits per episode, a chat message or a JSON payload, with users unmetered. Serving memory, with real temporal sophistication.

Letta is the interesting one. It comes from the MemGPT authors and builds stateful agents: the agent edits its own memory blocks, and periodic "dream" subagents reflect on recent conversations to consolidate them. Letta Code is a memory-first coding harness. This is not serving memory. It is agent memory: one long-lived agent's sense of self, living inside Letta's runtime.

"But they have Claude Code plugins"

They do. Mem0 and Supermemory ship real Claude Code plugins, and both offer shared memory scoped to a repository. Letta lets teams share agents. None of this is vaporware, and I won't pretend otherwise.

The difference is not whether they can store a coding fact. It is the shape of what they store, and the shape of the bill.

Read Supermemory's own Claude Code launch post. The idea, in their framing, is that Claude Code should know you, so the plugin injects your user profile at session start. That is serving memory pointed at a developer. The developer becomes the end user. Your preferences get remembered. Your team's decisions are still nobody's job.

Letta has the opposite trade-off. Its memory is rich, but it belongs to an agent. Your team does not run one agent. It runs Claude Code on one laptop, Codex on another and Cursor on a third, often in the same afternoon.

What engineering memory has to get right

Engineering memory is not serving memory with a coding plugin bolted on. It has five jobs serving memory was never designed for. This is what Recallium is built around.

1. Keep the why, not just the what

Git already keeps the what. commit 7d8f1c3: rotate auth tokens, drop the 30-day JWT expiry. Fine. Now ask your agent why.

Recallium keeps the part the commit can't. A security review found a replay risk in leaked 30-day tokens. Shorter-lived JWTs alone were rejected, because mobile sessions would drop every hour. Mobile clients must survive 24 hours offline. And this replaced the March decision, which is still on record.

That paragraph is two hours of an engineer and an agent working a production issue. Without engineering memory, it leaves with the session. Every. Single. Time.

2. Know which decision is current

A fact extractor that stores "uses 30-day JWTs" in March and "uses rotating refresh tokens" in September has handed your agent a coin flip.

Recallium keeps the full history and makes the current decision explicit. An agent that finds the March call also sees what replaced it, and why. Your agents stop suggesting the approach the team already walked away from.

3. Belong to the team, not the tool

Recallium is not Claude's memory, Codex's memory or Cursor's memory. It is your team's, and every tool is a client of it. Connect once over MCP and it works in Claude Code, Codex, Cursor, VS Code and 60+ other clients. Use Claude Code today, Codex tomorrow, a private model where you must. Models come and go. Institutional memory stays.

Memory stays personal until you share it, then you promote a memory or a whole project to the team. On Recallium Cloud, teams and role-based access are built in: the server decides who reads a project's memory, not a naming convention on a shared API key. Ask "what did Sam decide about auth?" and get Sam's reasoning, not a Slack archaeology dig. When Sam moves on, the reasoning stays.

4. Be tied to the code

A commit carries a trailer that points to the decision and the fix behind it. A memory carries the files it touched. Six months later, "why is this here?" still has an answer. Half-finished work lives in a workstream, so on Monday your agent, or a teammate's, resumes instead of re-reading the diff.

5. Happen on its own

No rituals. No "remember to save that." Recallium searches before your agent touches a file, stores what you worked out with the why attached, and ties it to the commit that shipped it. Your team keeps working the way it already works.

Notice what is missing from that list: anything about end users. That's the point.

The numbers: we won their benchmark anyway

Here is the irony. LongMemEval is a serving-memory benchmark: 500 questions about a user's past chat sessions. It is the home turf of every product in the previous section. It is not our job.

We ran it anyway, all 500 questions, abstentions included, one pass, no best-of-N. On the 30 September 2026 build, Recallium put the right past session on the page for 499 of 500 questions (99.8% hit@10). The first result was already right 95.4% of the time. And it served 3.8 results per question on average.

Let that sink in. Not a padded top ten. Not 200. Three or four.

LongMemEval-S, 500 questions. Vendors ran their tests separately and reader models differ; not a controlled head-to-head ranking. Sources: recallium.ai/benchmarks, github.com/mem0ai/memory-benchmarks, getzep.com/research, arXiv 2501.13956.
SetupQA accuracyReader modelResults read per question
Recallium98.4%Claude Opus 5.53.8
Mem0 (committed result file)93.4%GPT-5200
Recallium, official GPT-4o reader90.2%GPT-4o3.8
Zep, current research page90.2%GPT-5.450 graph items
Zep, peer-reviewed paper71.2%GPT-4oNot stated
No memory system, full history60.2%GPT-4oEvery past session

Two caveats. The vendors ran their tests separately, and the reader models differ, so this is not a controlled head-to-head ranking. Read the Mem0 row as a difference in how much context each system needs, not a scoreboard. The cleanest comparison is the GPT-4o one: same reader, same judge, 90.2% against the 71.2% in Zep's own paper. Supermemory publishes recall rather than answer accuracy, so it isn't in the table.

The column that matters to your bill is the last one. Every result an agent reads is tokens you pay for, every call, for every engineer. Against pasting the full history, Recallium gives the agent about 90% less to read. One agent reads the code and records what it learned; every teammate's agent gets it back in one search, so the team pays for those tokens once.

We also publish where we miss: 8 of 500 with the Opus 5.5 reader, five of them on temporal reasoning, which is the next target. Reader, judge, depth, date and every miss are on the benchmarks page. Check our work.

Side by side

Four of these five are built to remember someone. One is built to remember why.

Five memory products, side by side. As published on each vendor's own pricing and docs pages, checked 8 October 2026.
ProductKind of memoryRemembersMemory belongs toBilled by
RecalliumEngineeringDecisions, fixes and constraints behind the code, with what replaced whatThe team, across every MCP clientPer engineer
Mem0ServingFacts about each end userYour app's usersAdd and retrieval requests
SupermemoryServingAuto-maintained user profilesYour app's usersCredits for tokens and queries
ZepServingHow facts about users and data change over timeYour app's users and dataCredits per episode ingested
LettaAgentOne agent's own memory blocksA long-lived Letta agentAgents and model usage

Pick the memory that matches the job

Building a product that should remember its customers? Use serving memory. Mem0, Supermemory and Zep are built for exactly that, and they are good at it.

Want one long-lived agent that rewrites its own memory and lives in one harness? Look at Letta.

But if your team keeps relearning why the code is the way it is, if your agents keep proposing the fix you already rejected, if the reasoning walks out the door when someone changes teams, that is not a serving problem. That is an engineering memory problem.

Your code remembers what changed. Recallium remembers why.

Recallium Cloud is in a closed pilot. Start free on your own with 500 memories a month through December 2026, and bring the team onto Pro when you are ready. One command connects the coding agents on your machine:

npx -y recallium install

Join the waitlist or read why Recallium.

Sources

Also published on Medium →

Related Resources: