
290 Sessions, 13 Lessons: What AI Agents Taught Themselves
Every night our AI agents read their own work and write down what they learned. After 290 sessions, 13 lessons survived and we kept 6. Here is what they found, and what it was worth.
Most mornings a lesson shows up at the top of my terminal. It says: execute the tests you wrote, and report pass or fail before you say you're done.
Nobody on our team wrote that sentence. An agent did, at night, after reading through a day of its own work and noticing that it kept writing tests and then not running them.
People call this recursive self-improvement, RSI for short. It sounds like something from a paper about the end of the world. In practice it's closer to a musician listening back to the recording after the concert. You hear the passage where you rushed. You make a note. Tomorrow you rush a little less.
We've had this loop running on our own machines since the middle of September. Most of what I read about RSI for AI agents is about what it might do someday. This is about what ours has done so far.
How the loop works
Sno Station runs it once a day. The steps are plain.
It collects the day's sessions from Claude Code and Codex. Each session gets a quick label: worth keeping or not, succeeded or failed. The kept ones go to a judge model, which looks for a moment where something was learned. Usually that's a human correcting the agent, though sometimes the agent catches itself, and sometimes the environment just bites it.
For each moment the judge writes a lesson with four parts. The situation that triggers it. The advice. The reason. And the evidence, quoted word for word from the session, with the line it came from.
Then the lesson has to get past two gates.
The first is mechanical. Every quote the lesson cites must appear, character for character, in the session it claims to come from. A lesson built on a quote that isn't there gets dropped. The second gate is a different model whose only job is to doubt. It scores each sentence of the lesson against the evidence, and a lesson that claims more than its evidence supports doesn't pass.
After that, a person. I read what's left and say keep or don't.
Lessons that survive get shown to agents at the start of later sessions, one line each, the ones that have helped most at the top. If a lesson looks relevant the agent can open the whole thing.
What came out
Here's the count, from the loop's own records on our build machine.
It collected 1,879 sessions. Of those, 290 were sent to the judge. 13 lessons made it through both gates. I kept 6. The other 7 are still sitting there, waiting for me to decide.
290 sessions, 13 lessons. That ratio surprised me at first. Then I thought about it. How many days of your own work contain something you'd write down and tape to the wall? Most days are just days.
And the lessons themselves are small. I expected something grander, if you want to know the truth. Here are three, quoted as the agents wrote them.
After a git mv package or directory rename in an npm workspace, before running the verification proof, explicitly check each package's gitignored build output directories (dist/, build/, out/) for stale artifacts. Move or delete them.
Before constructing each hunk of a patch, read the exact line range at that hunk's target to capture verbatim context. Never reuse identifiers or syntax from an earlier truncated read of the same file.
Label the capability "not documented in inspected sources" until an explicit contract or scoped test establishes its absence.
The first one cost us an afternoon the day it happened. A package got renamed, everything in the source was correct, and the install kept failing because an old compiled file was still lying in a folder git doesn't look at. The second is the kind of thing an agent does forty times a week: it remembers roughly what a file said and writes a patch against its memory instead of the file.
The third one I like most. An agent was comparing products, didn't find a feature in a competitor's docs, and reported that the competitor didn't have it. Silence isn't absence. That's a lesson I've had to teach humans.
It's the stuff a senior engineer carries around without knowing it. Scar tissue. The difference is that an agent starts every session with no scars at all.
What it's worth
I can't tell yet how much any of it helps.
The lessons were shown at the start of 5,595 sessions. An agent opened one in full 135 times. So about 2 percent of the time. The one-line version does most of the work, or nothing does. I can't always tell which.
The loop also checks its own results. After a lesson is shown, it asks whether the agent followed it and whether that helped. So far it has two clear answers of "followed, and it helped." Both for the same lesson. Seven more came back unclear.
And the lesson from the top of my terminal, the one about running your tests? After it went live, agents still skipped the tests sometimes. 4 failures in 12 sessions. Before the lesson it was 22 in 119. The sample is too small to say it got worse. It's certainly too small to say it got better.
So I don't know yet. I mean that plainly. If you forced me to put a value on it today, I'd say each kept lesson saves a mistake that costs somewhere between ten minutes and an afternoon, and that the same six mistakes come around often enough for that to matter. One night of judging 193 sessions cost about fifty cents. The price of being wrong about the value is low.
What I do trust is smaller than a measurement. I recognize the lessons. I read them and think, yes, that happened, and yes, that's what I would have told it.
What went wrong
The loop itself needed the same treatment it gives the agents.
Early on, a defect let draft lessons slip past the second gate without being checked at all. We threw every one of those out. The count above only includes what went through the whole process.
We had also built the loop too carefully. A canary, a hash check, crash recovery. One night we cut 724 lines of that out of it, and it ran better. A system for learning from mistakes had been wrapped in protection against mistakes it never made.
Then at the end of September the loop went quiet. A settings file had gone missing, and each night it woke up, found nothing it was allowed to send, and went back to sleep. It did this for a week before anyone noticed. A thing that fails silently at three in the morning is hard to love.
That one is on the list now. Probably it should become a lesson.
Where this goes
Thirteen lessons is not an agent rewriting itself. It's a notebook. A month ago our agents made the same mistake on Tuesday that they made on Monday, with no memory of either.
Some mornings the reminder is there and the tests still don't get run.
Sno Station is open source, and the nightly review is part of it. It's at sno.ai.
Written by Sno AI Team
Contributing writer at Sno.ai, sharing insights about AI, productivity, and knowledge management.
Related Articles
Comments
Comments coming soon. Configure Giscus at giscus.app


