about this reference
Reliability work needs a shared record.
Documenting, measuring, and mitigating behavioral degradation in long-running AI agents — definitions, metrics, and mitigation patterns from real operations.
What this site is
Behavioral State Decay is a field reference on agent drift and long-running AI reliability. Agents that run for hours, days, or across many sessions do not fail the way single prompts fail: they degrade. Instructions lose force, style drifts, calibration slips, and shortcuts accumulate. That degradation is rarely documented with enough precision to compare notes across teams.
This site exists to build that record: careful definitions of what decays and how, measurements that make degradation visible before it becomes an incident, and mitigation patterns that hold up in real operations rather than in demos.
How the material is organized
Every entry belongs to one of five pillars. Definitions pins down vocabulary. Measurement covers metrics, probes, and evaluation harnesses. Mitigation collects patterns that slow or reverse decay. Field Notes records observations from live agent operations as they were found. Bridges connects this work to adjacent disciplines such as SRE practice and model evaluation.
Who it is for
Engineers operating long-running agents in production, researchers studying model behavior over time, and anyone deciding how much autonomy to hand an agent and for how long. You should leave an entry with something you can measure or apply, not a slogan.