Research
Research notes and methodology.
Regression Testing for AI: Gates That Actually Block a Release
A CI gate that reports a score and lets the pipeline through is not a gate. It is a logging statement with a dashboard. Here is how to build AI regression tests that fail a build — and what has to be true before anyone will let them.
Detecting Drift in Production AI: Model, Data, and Policy
An AI system can degrade without anyone changing a line of code. Three distinct kinds of drift — model, data, and policy — produce different signals and demand different responses. Here is how to detect each one and when to re-evaluate.
Building Golden Datasets That Don't Rot
Every evaluation rests on a set of examples with known-correct answers. Those examples decay — through leakage, staleness, and quiet contamination — and a rotten golden set produces confident scores that mean nothing. Here is how to build one that survives.
How Many AI Decisions Must You Actually Review?
Reviewing everything is unaffordable. Reviewing only the failures the system already flagged is worse than useless — it guarantees you never find the failures it missed. Here is how to size a risk-based review sample.
Inter-Rater Reliability: Why AI Evaluation Needs Cohen's Kappa
If two qualified people score the same AI output differently, your evaluation is measuring your people as much as your model. Raw agreement hides this. Kappa exposes it — and disagreement, handled properly, is the most useful signal an evaluation produces.
Model Cards and System Cards as Evidence
Most model cards describe. Almost none prove. The difference between a document that satisfies a reader and one that survives an auditor is a small number of properties that are cheap to add and rarely present.
Methodology
SingleAxis AI Safety Framework.
SASF provides a structure for evaluation layers, evidence, failure classification, and review. It supports the evaluation method rather than serving as a standalone product.
Read the methodologyResearch collaboration
Working on healthcare AI evaluation or monitoring?
We welcome conversations with medical and technical researchers studying oversight, monitoring, and real-world reliability.
Contact SingleAxis Research