One memory across AI agents: what it actually takes
The memory built into a coding agent lives inside that vendor's tools. What I tell one agent stays with that agent, and the next one I open in the same project starts from nothing. A checked-in instruction file gets part of the way, since any agent can read the repository, but it belongs to the repository: it can't hold what spans projects, or what I wouldn't commit.
So I built a memory that sits outside the agents: a Postgres store behind six MCP tools,
remember_fact, remember_note, recall_facts,
search_notes, revise_fact and forget. It holds
facts, which record both when something was true and when I believed it, and notes, which
are searchable text. Today it runs locally, for one user, as a Claude Code plugin.
The tools were the easy part. The work was in four requirements around them, and the first three fail silently when you get them wrong.
I'm writing them down because I'm deciding whether to build a hosted version, lontar.dev, and because they hold no matter who builds it.
Each claim below says how I know it: built (in the code today), measured (by me, with the Claude Code version named where it matters), found (in Claude Code's shipped binary), documented (in a vendor's documentation, untested unless also marked measured) or designed (written down, not built).
The memory must know where you are, and the model must not be the one to say
Most of what I want an agent to remember belongs to one context. A client's deployment quirks belong with that client's repositories; how I like commit messages written belongs everywhere. So the store is divided into spaces, and the first design question was who picks the space for each read and write.
The obvious answer is a space parameter on every tool, and I didn't add one.
A parameter the model fills in is a parameter any text in its context can argue for: a
README, a web page, another tool's output. lontar's tools can't name a space, so a model
can't ask for another one. Built.
Instead the server decides once, at startup, from where it was started. Claude Code starts each session's MCP server in that session's working directory (measured on 2.1.286, and on 2.1.263 before it), and a config file maps directories to spaces:
default_space = "personal"
[[space]]
id = "personal"
paths = ["~/dotfiles"]
[[space]]
id = "work"
paths = ["~/codes/acme-*"]
reads_from = ["personal"]
Resolution runs in a fixed order: the LONTAR_SPACE environment variable if it
is set, then the longest path in the file that matches the working directory, then
default_space, which must be named explicitly. A path matches the directory
itself or anything beneath it, and a trailing * matches by plain prefix, so
~/codes/acme-* covers ~/codes/acme-api and
~/codes/acme-web. reads_from widens reads in one direction only:
here, work sessions see personal facts and personal sessions never see work ones. Every
write, revision and deletion goes to the session's own space. Built.
Two spaces claiming the same directory at the same depth is an error, not a guess:
matches "work" and "personal" equally; make one path more
specific. Picking one silently would hide the loser's memories and send every write to the
winner. Built.
The failure that stays silent is the boring one. Rename or move a directory and its entry
stops matching. Nothing errors; the session falls through to default_space.
If the default happens to be the space that directory belonged to, everything looks fine
for a long time while the config names a path that no longer exists. If it isn't, the
agent quietly reads and writes the wrong space. lontar doesn't check configured paths
against the disk yet, so the fix today is discipline: update the paths when you rename.
Built.
The memory must be there before the first prompt
recall_facts is a tool, and a tool is something the model decides to call. At
the start of a session it has the least reason to, because it doesn't yet know what it's
missing. So in Claude Code the current facts arrive before the first prompt, through a
SessionStart hook: no tool call, and nothing for the model to decide.
Built.
The hook runs lontar recall in the directory named in the hook's input, not
in its own working directory, and that command resolves the space through the same code as
the server, so the boundary between spaces lives in one place. It fires on
startup and clear only; after compact, a wall of
facts mid-session is interruption, not context. It also blocks session start, so it has a
budget of 800 ms and a hard timeout of 5 seconds. On my machine it takes 656 to 799 ms
over five runs (measured).
Every failure exits 0 and prints nothing: Postgres down, node or
jq missing, no config, timeout. The hook runs at the start of every session
in every space. A memory that sometimes forgets is an annoyance; one that blocks unrelated
work gets uninstalled. The price is in the hook's own comment: when it is broken, a
session simply has no memory, which looks exactly like an empty space.
Built.
The output shape is where that price comes due. Claude Code reads the context from a nested field:
{"hookSpecificOutput": {"hookEventName": "SessionStart", "additionalContext": "…"}}
I tested three shapes on Claude Code 2.1.286, each in a fresh headless session with a
marker string (measured). The nested shape arrived, labeled as hook
context. Plain text on stdout also arrived, as a hook success message. A flat
{"additionalContext": "…"} is valid JSON, the hook exits
0, and the model saw nothing. Pair that with a hook that is silent by design, and you have
a memory that never worked and never once complained.
The plugin as a whole can fail the same way. Some time after I moved the repository, the plugin stopped loading, because its marketplace entry still pointed at the old path, and sessions started with no memory and no memory tools. Nothing said so. I found it while checking something else.
It has to fit
A hook can't inject unlimited context. Claude Code's hooks documentation states the cap: a
hook's additionalContext, and its plain stdout, are each limited to 10,000
characters. I measured it before I found that sentence, with a hook that injects exactly N
characters of numbered 50-character lines, one fresh headless session per size, and the
model asked for the first and last line it could see:
| Injected characters | What the model saw | Claude Code |
|---|---|---|
| 10,000 | all 200 lines | 2.1.268, 2.1.285, 2.1.286 |
| 10,001 | the first 40 lines, then "Output too large (9.8KB). Full output saved to: …" | 2.1.285, 2.1.286 |
| 10,050 to 19,999 (six sizes) | the same 40-line preview | 2.1.285 |
The documentation and the measurement agree to the character (documented, measured).
What happens above the cap is the part worth reading twice. The output goes to a file, and the model gets the path and a preview of the first 2,000 characters. In the documentation's words, "Claude Code doesn't ask Claude to read the file." When I checked on 2.1.266 to 2.1.268, injected context didn't show in the transcript view either, so the person at the keyboard saw nothing. Past 10,000 characters, whether the rest of the memory reaches the model is the model's decision again, which is exactly what the hook was there to take away.
A memory store reaches the cap sooner than its own limits suggest.
recall_facts stops at 200 facts or about 10,000 tokens, but the token count
is estimated from the statements alone, and each fact renders with about 97 more
characters of metadata: a subject, an id, a date and a source. At 150-character
statements, 40 facts fit in 10,000 characters, a full recall is about 49,000, and the
preview keeps the newest eight. The line saying the recall was capped comes last, so it is
the first thing cut. That is arithmetic on the renderer, not something I have hit yet.
There is a third number. The Claude Code binary carries a table of limits,
reason:2000, stopReason:2000, systemMessage:4000, additionalContext:8000,
permissionDecisionReason:2000, and the additionalContext:8000 entry is there in 2.1.268 and still in
2.1.286 (found). A 10,000-character injection arrived whole, so it isn't
the SessionStart cutoff, and I don't know what it governs.
The hosted design caps the injection at 8,000 characters, under both numbers, rendered on the server. Facts about the current workspace come first, then facts with no subject, then the rest newest first, and when anything is cut the injection ends with a line inside the budget: "N more facts in lontar; call recall_facts to see them" (designed). The local hook has no cap of its own (built). A space that outgrows 10,000 characters keeps working, quietly, as a 2,000-character preview of its newest facts.
What another agent must supply
The six tools are plain MCP. Their schemas and behavior don't depend on the client, with
one exception: the server tags every write source=claude-code, a constant a
second client would have to replace with its own name (built). Everything
else that changes from one agent to the next comes down to two questions: how does the
agent tell the memory where you are, and can it load memory before the first prompt?
Cursor is my example of a second agent. Everything below comes from its documentation as I read it on October 1, 2026, and none of it is tested (documented).
The workspace signal exists in two forms. Cursor resolves
${workspaceFolder} inside an MCP server's command,
args, env, url and headers, so a
config can hand the server its workspace explicitly. And it lists MCP roots,
the protocol's own way for a client to report its workspace folders, as supported. What
the docs don't say is which directory a stdio server starts in. The local build resolves
its space from exactly that, so if it isn't the project, every session lands in
default_space, the same silent failure as a renamed directory, arriving from
a different direction.
Session-start loading exists too, with a caveat. Cursor's sessionStart hook
can return an additional_context field, documented as "Additional
context to add to the conversation's initial system context." The same page says the
hook "runs as fire-and-forget; the agent loop does not wait for or enforce a blocking
response." Whether a recall that takes most of a second lands before the first prompt
is the first thing I would measure.
Built-in memory needs none of this, because it never leaves the product it is built into. The vendor decides where it is stored, what it is scoped to and when it is written, and the next tool you open doesn't see it. Support for agents beyond Claude Code is planned, with tested setup guides at launch.
Where this stands
Everything marked built above runs today, for one user: me. It is a stdio MCP server over a Postgres I run myself, installed as a Claude Code plugin, and the repository is private for now. It will be open source (AGPL-3.0), whether or not the hosted version gets built.
The hosted version does not exist yet. It would be the same six tools over HTTP with a sign-in, and the 8,000-character session-start design above. Whether it gets built depends on a waitlist: if 280 people confirm a signup at lontar.dev before the window closes, I build it, and if not, lontar stays a personal tool and the list is deleted. The page has the date and the details.
Whichever way that goes, the requirements stand. A memory shared across agents has to know where you are without the model choosing, arrive before the first prompt, fit in what the agent will actually show the model, and get both of those signals from every agent it supports. If you are building any of this, I would like to compare notes.