Use targeted rare-case testing: Which evaluation response best addresses both rarity and consequence?
Rare severe errors need targeted evidence; aggregate accuracy and random expansion may leave the critical uncertainty unresolved.
The question
A public-benefit screening model shows 98% overall accuracy, but only a few severe eligibility errors appear in historical data. Those errors can wrongly deny essential support, and the general test set contains few such cases. Which evaluation response best addresses both rarity and consequence?
Preparing for AIGP? Take the free 5-min readiness quiz →
- Report overall accuracy with a warning about limited severe-error evidenceA warning acknowledges uncertainty but does not generate evidence about rare, consequential errors or their operational impact.
- Lower the acceptance threshold to reduce potential denialsChanging the threshold without measuring its effect on severe false negatives and workload is not evidence-based.
- Collect a much larger random test set and rely on its aggregate accuracyRandom expansion may still contain too few severe cases and preserve the masking effect of aggregate accuracy.
- Use targeted rare-case testing, measure severe-error rates, and document uncertainty ✓Targeted testing increases evidence on rare cases, while severe-error measurement and uncertainty documentation address consequence and limited data.
The trap
Rare and severe does not call for aggregate accuracy alone; target the cases and report uncertainty around their error rates. How to remember it
Rare severe errors need targeted evidence; aggregate accuracy and random expansion may leave the critical uncertainty unresolved.
How many of these would you get right?
One of 1581 AIGP questions on Certsqill. Take a free five-minute check and see your score per domain — not one number, but which section to open tonight.
Test your AIGP readiness — freeMore Understanding How to Govern AI Development questions
- Link model, dataset, transformation, and test versions: What control most directly resolves the uncertainty? →
- Document data lineage for training and testing artifacts: What evidence should be required before approval? →
- Compare training distributions with deployment conditions: Which evidence most directly resolves that →
- All 426 Understanding How to Govern AI Development questions →
Part of the Certsqill AIGP question bank · Understanding How to Govern AI Development ·
Every answer, right and wrong, comes with its own explanation.