Two soft glowing lights joined by a thin line on a dark monitor in a moonlit room, like two AI agents resting overnight
Back to blog
EngineeringAI Automation

Why Our AI Agents Sleep: Building Memory That Updates Itself

Plenty of memory tools exist for AI agents. We built our own anyway, one that corrects itself while the agents are idle. Here is why, how it works, and what we measured.

S
Sno AI Team
September 30, 2026
|
1 min read
|
116 views
Share:

In March you give an agent the staging server's address. In June the server moves, and you give it the new one. A memory that only collects things now holds both. The first one isn't wrong, exactly. It was true at the time. But every so often the agent will pick it, with complete confidence, like a man giving directions to a restaurant that closed years ago.

That's the problem with most memory for AI agents. Getting one to remember is easy. The trouble starts when what it remembers goes out of date.

We ended up building our own memory into Sno Station. I didn't especially want to. There are good memory tools out there already, some with tens of thousands of stars on GitHub, and the world doesn't need a tenth memory vendor. We're not trying to be one. But we needed three things, and we couldn't find them together.

One memory, every agent

The first thing is simple to say. We run more than one agent. Claude Code, Codex, sometimes OpenClaw and Hermes. Each one comes with its own idea of memory, and each one's memory stops at its own front door.

So you tell Claude Code that the team deploys on Thursdays, and on Friday Codex doesn't know. You're the one carrying the information between them. You're the cable.

In Sno Station there is one memory store per person, on your own machine, encrypted. Every agent reads from it and writes to it through a thin plugin. The plugin quietly captures what's worth keeping as you work and recalls what's relevant when a session starts. There's exactly one writer underneath, so two agents can't scribble over each other.

It's also what makes handoff work. Say one agent runs out of quota halfway through a job and the other takes over. The second one isn't starting cold. It knows what the first one knew. Without shared memory, handoff is just a second agent asking you to explain everything again.

And it means the memory belongs to you and not to a harness. Swap the agent out for a different one next year, and what it learned about you and your work stays put.

Replacing, not piling up

The second thing is the staging server.

A lot of memory systems only append. Every new fact goes on the pile. Nothing comes off. The demo goes fine. A few months later you've changed jobs, the API has changed, and the agent still remembers your old boss. An append-only memory holds every version and lets the model sort it out at the worst possible moment.

Ours works differently, and the rule is strict. When a new fact conflicts with an old one, a model looks at the pair and gives one of three answers: replace, keep, or not sure. That's all it's allowed to say. The model judges. Ordinary code does the rest. The old memory gets marked retired, the new one takes its place, and only current memories are ever put in front of an agent.

The model has no delete button. I'd like to stress that. It can't remove anything on its own, and a retired memory is still there if you ever need to look at what used to be true. It's sort of like the difference between striking a line through a ledger entry and tearing out the page.

How well the judging works

The judging is the risky part. If the model replaces a fact it should have kept, you've lost something true. That's worse than clutter.

So we trained a small model of our own for that one decision, and tested it against a frontier model from one of the big labs. Same exam for both: 223 pairs of facts, an older one and a newer one, each with a known right answer.

Our model chose the right action 97.5 percent of the time. The frontier model, 96.8.

The number I care about more is the wrong replacements, the cases where a true fact would have been thrown out. Ours did that 3.2 percent of the time. The frontier model, 4.1.

I don't want to oversell this. It's a small exam, 223 questions, from June, and both models are close to the ceiling on it. We've since written a harder one. It isn't a comparison with other memory tools either. We haven't run that yet, and I won't quote someone else's numbers as if we had.

What it does tell us is that a small model, running on a single graphics card, can make this one call about as well as the biggest ones. That matters because the judging happens constantly, and it happens on things you'd rather not send to anybody.

Why sleep

The third thing is when the work gets done.

You can't do all this while the agent is working. Each new fact means going back through everything the agent already remembers. It's slow, and the agent has a job to finish. So the heavy part happens when the agents are idle.

Inside the company we call that pass REM, after the stage of sleep where people seem to file away the day. I resisted the name at first. It felt cute. But it's accurate. The pass works through the day's sessions. Old facts get retired and contradictions get settled. There's less to remember afterwards, and what's left sits closer to how you work now. The memory is supposed to get smaller overnight, which is the opposite of what a pile does.

You choose how much of this leaves your machine. In the most private setting, none of it does. Your own agent's model does the consolidating and everything stays local. In the default setting your agents' own subscriptions do the thinking. And if you opt in, our models take the hardest calls, like the replace-or-keep decision above.

I should say where this stands. The shared store and the replacement rule are in daily use here. The sleep pass is the newest piece, and on our own machines it doesn't run every night yet. It had been skipping itself over a configuration mistake, and that's being fixed. I'd rather tell you than have you find out.

Why it's underneath everything

Here's the thing I didn't understand when we started. I thought memory was a feature. It's closer to a floor.

The nightly review, where agents read their own sessions and write down lessons, needs somewhere for the lessons to live. Handoff between agents needs it. A record of what an agent has done and how often it got things right the first time needs it. Later, anything we build in the cloud will be distilled from what this layer holds.

None of those would be worth building on a pile that only grows. They need a memory that can be wrong on Monday and corrected by Tuesday, without a person going in to fix it.

When the memory is wrong, the next agent inherits the mistake and you explain it again.

Until the server moves again.

Sno Station is open source. It's at sno.ai.

S

Written by Sno AI Team

Contributing writer at Sno.ai, sharing insights about AI, productivity, and knowledge management.

Related Articles

Comments

Comments coming soon. Configure Giscus at giscus.app

Why Our AI Agents Sleep: Building Memory That Updates Itself