Treat the gap as possible overfitting and investigate: Before approval, what does this contrast most directly
The training-to-held-out decline supports investigating overfitting and generalization before approval.
The question
A housing application model scores 98% on training records and 79% on a separately collected held-out set. Both sets represent the intended applicants, and a data-quality review found no material quality difference between them. No subgroup breakdown is available. Before approval, what does this contrast most directly support?
Preparing for AIGP? Take the free 5-min readiness quiz →
- Avoid claiming fairness from aggregate performance alone.The missing subgroup breakdown prevents a fairness conclusion, but this option does not identify the principal implication of the observed performance gap.
- Check for postdeployment population drift.Drift is a monitoring concern after deployment; it does not explain the predeployment gap between training and held-out results.
- Treat the gap as possible overfitting and investigate generalization before approval. ✓The large decline on held-out applicants is evidence that performance may not generalize, so approval should await investigation.
- Approve the model conditionally while relying on future monitoring to determine whether the training score generalizes.Future monitoring cannot replace preapproval evidence about the already observed held-out performance gap, especially when approval depends on adequate generalization.
The trap
A high training score is not deployment evidence; investigate a substantial held-out decline. How to remember it
The training-to-held-out decline supports investigating overfitting and generalization before approval.
How many of these would you get right?
One of 1581 AIGP questions on Certsqill. Take a free five-minute check and see your score per domain — not one number, but which section to open tonight.
Test your AIGP readiness — freeMore Understanding How to Govern AI Development questions
- Postal-sector codes warrant proxy-bias analysis: What does this evidence most directly suggest? →
- The test estimate is likely inflated by record duplication: What interpretation is best supported? →
- Compare line-specific errors: What conclusion is most justified? →
- All 426 Understanding How to Govern AI Development questions →
Part of the Certsqill AIGP question bank · Understanding How to Govern AI Development ·
Every answer, right and wrong, comes with its own explanation.