Discussion about this post

User's avatar
Kazuki Nakayashiki's avatar

What struck me here is how both humans and agents drift into the same failure pattern: once a rule “looks settled” no one checks if reality has moved on. The fix the piece suggests is almost disappointingly simple yet hard to practice in real systems. Decide ahead of time which constraints must always be re-opened at the source and how far anything is allowed to reach when no one is paying close attention. It feels less like an AI safety insight and more like a general lesson in institutional memory and how quietly it goes stale.

Kei Watanabe's avatar

Checking a rule's source takes one step: open the record the memory points to before you act on it. The models in our paper had that step available and took it about one time in five.

The part I did not expect: the fix was not telling the agent "this memory might be old." That did nothing. What worked was deciding in advance which memory gets checked every time, the one that limits your options. Same for people, if Yocco is right: the checking has to be built into the workflow, or it quietly stops.

Which rule would you make your agent re-check every single time, no matter what? I'll go first: anything that says "do not contact X." Those get written after something went wrong, and they are exactly the ones that go stale when the situation changes.

Thank you to AfroTech for sponsoring this issue. Fees and open dates are at https://www.passionfroot.me/glasp-newsletter

No posts

Ready for more?