Build a targeted defect test set and require: Which evaluation change is most defensible?
Rare severe conditions need targeted representative evidence and a metric tied to detecting the consequential failure, not aggregate accuracy.
The question
A manufacturing inspection model detects a defect occurring in 0.5% of units; failures can cause serious equipment damage. A random test set produces 99.5% accuracy but contains few defective units. Which evaluation change is most defensible?
Preparing for AIGP? Take the free 5-min readiness quiz →
- Increase the training set while retaining accuracy as the primary metric.More training data may help learning, but accuracy still obscures whether rare dangerous defects are detected reliably.
- Replace real defects with synthetic defects and evaluate aggregate performance.Synthetic defects may support exploration but cannot alone establish real-world detection performance or realistic operating conditions.
- Use random testing and report accuracy with confidence intervals.Random testing may contain too few failures to estimate detection performance, while accuracy remains dominated by nondefective units.
- Build a targeted defect test set and require a detection-rate criterion. ✓Targeted testing supplies evidence for rare conditions, and a detection criterion reflects the severe consequence of missed defects.
The trap
Ask whether the test contains enough consequential cases to estimate the risk-relevant error. How to remember it
Rare severe conditions need targeted representative evidence and a metric tied to detecting the consequential failure, not aggregate accuracy.
How many of these would you get right?
One of 1581 AIGP questions on Certsqill. Take a free five-minute check and see your score per domain — not one number, but which section to open tonight.
Test your AIGP readiness — freeMore Understanding How to Govern AI Development questions
- Version the model: Which documentation choice is most useful? →
- Provenance documentation alone establishes suitability: Considering provenance documentation specifically, →
- The validation population represents the new users: Which prior assumption requires reassessment first? →
- All 426 Understanding How to Govern AI Development questions →
Part of the Certsqill AIGP question bank · Understanding How to Govern AI Development ·
Every answer, right and wrong, comes with its own explanation.