Essay 03 · The library
Where Regime Detection Fails
False alarms, unstable labels, hindsight-flattered backtests and reflexivity — an honest field guide to the ways regime-detection AI misleads, and the questions that catch it out.
An honest series about detection has to end with its failure modes — not as a disclaimer bolted to the finish, but as the part that makes the rest usable. Detection failures come in families, and each family looks obvious once named.
The false economy of fast alarms
Latency and false alarms trade against each other; there is no free corner. A detector sensitive enough to catch a break on day one will also misfire several times a year on stretches that later prove to be ordinary turbulence. Every false alarm a strategy acts on costs turnover, spread and slippage; missed alarms cost drawdowns. Which is worse depends on who is trading and what a whipsaw does to their risk limits — which is why “the best performing detector” in a paper and the one worth running with real money are often different instruments entirely. A backtest that omits transaction costs is a rehearsal without gravity.
Labels are agreements, not discoveries
A two-state model of the same market finds different regimes than a five-state model. States, clusters, breakpoints — these are shapes the method is built to see, applied to data, not veins of truth running through it. “Crisis,” “calm,” “risk-on” — every label is a human agreement imposed after the clustering, and different honest analysts will draw the boundaries differently. Problems surface the moment a label is treated as a fact the market contains. Any write-up of regime work that cannot say in advance what is being labelled and when a label applies is describing a mood, not a market.
Hindsight flatters every chart
Looking at a long price series, regime boundaries seem obvious — the eye retroactively paints the calm years beige and the violent year red. Live, no one holds that chart; they hold the noisy left edge, and boundaries reveal themselves only in retrospect. The most seductive errors in this field leak that retrospective clarity forward:
- Label leakage. Evaluating a detector against regime labels that were themselves assigned using future data — the classic version of knowing the answer while pretending to compute it.
- Overnight switches. Backtests that quietly let a strategy know the regime changed the moment it changed, while the detector that must run in the real world needs days of confirming observations.
- Specification fishing. Trying twenty detector configurations, reporting the one that would have caught 2020, and never mentioning the nineteen.
The honest test is monotone: at each point in time, the detector may use only the past; every alarm it raises must be counted; every cost it triggers must be subtracted.
Reflexivity: detectors change what they detect
The quietest failure mode has no analogue in physics or medical imaging, because detections by many hands alter the market being detected. If a widely-imitated signal says a regime has broken, the resulting rush of repricing becomes part of the regime it claimed to find — sometimes amplifying the break, sometimes causing one. The February 2018 episode reads both ways: strategies premised on persistent calm had grown so crowded that their own unwinding helped make the break violent. A detector in a social system is both a rain gauge and a sprinkler. Any treatment of AI in investment that ignores this loop — that models market structure as weather rather than crowd behaviour — is describing a calmer field than the one that exists.
A skeptic’s checklist
Put together, the previous essays leave a portable test. When a paper, a vendor, or a headline claims a model “detects regime changes,” ask:
- Were regimes — and the labels applied to them — defined before the data was examined?
- Does the detector at each date use strictly the past? Or does any part of the evaluation borrow future knowledge?
- What is its documented false-alarm rate on calm stretches that never became a new regime?
- Are transaction costs and latency subtracted, priced in money rather than adjectives?
- Does the claimed edge survive small, boring changes to the specification — state count, hazard, window length?
- Could wider adoption of the same signal plausibly change its own accuracy — and does the writer acknowledge that?
Questions at this level are not hostility to machine learning; they are the courtesy this craft owes its claims. The role that remains for AI in investment is real but narrower than marketing suggests: not a prophet, but an instrument panel — fast probability updates, structure made visible, staleness made measurable. The panel earns trust the same way this site hopes to: by publishing what it cannot do alongside what it can. If an essay on this site ever fails its own checklist, that is precisely the kind of correspondence the inquiry desk exists to receive.