Bridges
What Agent Reliability Can Borrow from Customer-Base Models
Marketing science spent decades inferring silent customer dropout from behavior alone. A guide to which of its habits, cohort curves, survival analysis, hazard thinking, transfer usefully to measuring agent drift, and which assumptions do not.

Marketing analytics and agent reliability engineering face structurally similar problems: an entity that once behaved as intended may be silently ceasing to, no event announces the transition, and the operator must infer the change from behavioral traces alone. A customer in a noncontractual setting does not file a notice of disengagement; they simply stop buying. A long-running agent does not report its own drift; it simply produces work that tracks the objective a little less each week. The two problems are not the same, and the models built for one do not apply unchanged to the other. But marketing science has a several-decade head start on inferring a hidden condition from silence, and some of its thinking is worth borrowing.
The structural resemblance
The classic formalization on the customer side is the "counting your customers" problem: given only a customer's repeat-purchase history in a noncontractual setting, infer whether they are still an active buyer or have silently churned. The Pareto/NBD model posed it in 1987, and the BG/NBD model made that modeling approach substantially easier to implement (Fader, Hardie, and Lee, 2005). The scope is specific, repeat purchases rather than engagement in general, and the key move is that "churned" is a hidden state, never directly observed, only inferred from the shape of behavior over time.
Agent drift shares that surface structure. "Drifted" is likewise a condition no log line records: nothing announces the moment an agent's operative goal diverged, per behavioral state decay. But the resemblance should not be overstated. BG/NBD assumes a Poisson repeat-purchase process while the customer is active, heterogeneity across customers, and an absorbing dropout state tied to transaction opportunities. None of that automatically describes agent drift, which is multidimensional, driven by dependent mechanisms rather than independent events, and at least partly reversible through intervention. The models do not port; the habits of thought do.
What transfers
- Longitudinal analysis. Both fields need behavior plotted against entity age (customer tenure, session interaction count), not just cross-sectional snapshots of quality.
- Censoring. A customer who has not bought recently may still be active; a session with no visible failures may still be drifting. Both fields must reason about observation windows that close before the outcome is known.
- Hazard thinking. Survival analysis and time-to-churn distributions have a natural agent analogue in time-to-drift distributions: the median lifetime and the shape of the hazard matter more than any single average.
- Cohort curves. Percent of a signup cohort still active at week N maps to percent of sessions still performing at baseline after N interactions.
- Heterogeneity. Customer models fit distributions over individual churn propensities rather than one global rate; agent systems should likewise expect time-to-drift to vary by workload, session hygiene, and task mix rather than clustering around one number.
- Triage habits. RFM-style scoring (recency, frequency, value) suggests drift-risk triage: which long-running sessions deserve inspection first.
- Intervention evaluation. Win-back campaigns judged against holdout controls correspond to re-anchoring patterns judged by whether a replayed suite recovers, per drift regression tests.
What does not transfer yet
The inferential machinery. BG/NBD's probability-alive estimate is a genuine posterior from an explicitly specified probabilistic model. The closest agent-side artifact today, a stability score such as the Agent Stability Index, is not that: it is a deterministic weighted composite with an alarm threshold, an observable score rather than a posterior over a modeled latent state. A true agent analogue of probability-alive would require an explicitly specified Bayesian, hidden Markov, survival, or state-space model over agent telemetry, which the field does not yet have. The drift study behind the ASI projects, in its simulations and under its assumptions, a median of 73 interactions before drift becomes detectable (Rath, arXiv:2601.04170); that is a property of the simulation, not a field constant to build cadences on.
Where the agent side is better off
Agent systems can run controlled replays, the drift regression suite, which customer analysts, who cannot re-run last quarter on a sandboxed copy of a customer, mostly cannot. Agent operators can also intervene on the entity's actual state, editing context and re-injecting goals, where marketers can only send stimuli and hope. Instrumentation such as state externalization narrows the inference gap, but does not close it: externalized files record the agent's declared state, not necessarily the state governing the model's next action, so instrumentation reduces uncertainty and improves diagnosis while the latent condition still has to be inferred and validated.
The practical takeaway: if your organization already has churn modeling competence, you have a real head start on measuring behavioral state decay in agents. The curve-reading habits, the censoring discipline, and the intervention-evaluation rigor transfer well; the models themselves have to be rebuilt for a different process, and pretending otherwise would repeat on agents the mistake of applying a purchase model to behavior it was never specified for.