Research

Research notes and methodology.

Browse all briefings
01
CI/CD

Regression Testing for AI: Gates That Actually Block a Release

A CI gate that reports a score and lets the pipeline through is not a gate. It is a logging statement with a dashboard. Here is how to build AI regression tests that fail a build — and what has to be true before anyone will let them.

02
Monitoring

Detecting Drift in Production AI: Model, Data, and Policy

An AI system can degrade without anyone changing a line of code. Three distinct kinds of drift — model, data, and policy — produce different signals and demand different responses. Here is how to detect each one and when to re-evaluate.

03
Gold Tasks

Building Golden Datasets That Don't Rot

Every evaluation rests on a set of examples with known-correct answers. Those examples decay — through leakage, staleness, and quiet contamination — and a rotten golden set produces confident scores that mean nothing. Here is how to build one that survives.

04
Sampling

How Many AI Decisions Must You Actually Review?

Reviewing everything is unaffordable. Reviewing only the failures the system already flagged is worse than useless — it guarantees you never find the failures it missed. Here is how to size a risk-based review sample.

05
Methodology

Inter-Rater Reliability: Why AI Evaluation Needs Cohen's Kappa

If two qualified people score the same AI output differently, your evaluation is measuring your people as much as your model. Raw agreement hides this. Kappa exposes it — and disagreement, handled properly, is the most useful signal an evaluation produces.

06
Documentation

Model Cards and System Cards as Evidence

Most model cards describe. Almost none prove. The difference between a document that satisfies a reader and one that survives an auditor is a small number of properties that are cheap to add and rarely present.

View the complete briefing archive

Methodology

SingleAxis AI Safety Framework.

SASF provides a structure for evaluation layers, evidence, failure classification, and review. It supports the evaluation method rather than serving as a standalone product.

Read the methodology

Research collaboration

Working on healthcare AI evaluation or monitoring?

We welcome conversations with medical and technical researchers studying oversight, monitoring, and real-world reliability.

Contact SingleAxis Research
Healthcare AI Evaluation Research Notes — SingleAxis