Subpopulations
Compare performance across relevant clinical profiles, demographics, equipment or sites.
Robustness and failure points
A strong average result does not describe the conditions in which a model fails. QSTOM-IT puts the system under targeted stress to identify the populations, data and scenarios where performance degrades.
Testing
Compare performance across relevant clinical profiles, demographics, equipment or sites.
Test sensitivity to missing data, noise, format changes and quality variation.
Observe stability over time and the effect of changes in population or practice.
Study false negatives, false positives, silent errors and high-consequence scenarios.
Assess what happens when the use context differs from development data.
Examine hallucinations, source fidelity, variability, refusal behaviour and out-of-context use.
Access
Documentation, results, protocols and logs, without access to individual-level data.
Tests run inside infrastructure controlled by the organisation with restricted access.
Where necessary, work with infrastructure or partners appropriate to health data constraints.
Reporting
We document failure points, evidence level, reproduction conditions and potential consequences. A critical risk is not averaged away by strong results on other criteria.
FAQ
Duration depends on access level, number of scenarios and model complexity. A short scoping phase defines priority tests before work begins.
Yes, when the system has a clearly defined use case. The protocol should include sources, error scenarios, response variability and suitable evaluation criteria.
No. It documents failure conditions and observed results within a defined scope. It cannot prove the absence of every possible risk.
QSTOM-IT
We can design a stress-test protocol adapted to the model, data and decision at stake.