Treat the labeled sample as representative of every: Which validation conclusion is not justified?
Labels limited to completed tickets and larger customers do not establish performance for all submitted requests and customer groups.
The question
A software procurement team reviews a vendor’s 98% accuracy claim. The vendor labeled only completed support tickets; abandoned requests and tickets from smaller customers were excluded. The tool will serve all submitted requests. Which validation conclusion is not justified?
Preparing for AIGP? Take the free 5-min readiness quiz →
- Accept the claim after confirming larger customers dominate the labeled sample.Dominance by larger customers reinforces selection bias and cannot justify performance for all submitted requests and customer groups.
- Treat the labeled sample as representative of every submitted request and customer group. ✓Excluding abandoned requests and smaller customers prevents the reported accuracy from being assumed representative of the deployment population.
- Review accuracy separately for completed and abandoned requests.This is an appropriate follow-up analysis, but it is not itself the invalid conclusion in the vendor’s claim.
- Confirm that the records have labels.The vendor does have labeled records; the problem is their coverage, not the existence of labels.
The trap
Ask which deployment cases lack labels, then test whether their absence could change the reported metric. How to remember it
Labels limited to completed tickets and larger customers do not establish performance for all submitted requests and customer groups.
How many of these would you get right?
One of 1581 AIGP questions on Certsqill. Take a free five-minute check and see your score per domain — not one number, but which section to open tonight.
Test your AIGP readiness — freeMore Understanding How to Govern AI Development questions
- The validation population represents the new users: Which prior assumption requires reassessment first? →
- Evaluate it on a held-out set excluded from training: Which evaluation design best distinguishes learning from →
- Group duplicate shipment records before creating training: What control directly addresses this evaluation →
- All 426 Understanding How to Govern AI Development questions →
Part of the Certsqill AIGP question bank · Understanding How to Govern AI Development ·
Every answer, right and wrong, comes with its own explanation.