Clarify criteria: What evidence should it prioritize before interpreting model performance?
Low agreement undermines the evaluation target; clarify criteria and adjudicate ambiguity before drawing conclusions from model scores.
The question
An employment-screening team has three reviewers label applications for a defined job-related criterion. Agreement is low, and the model performs inconsistently against their labels. The team has not documented ambiguous cases or an adjudication rule. What evidence should it prioritize before interpreting model performance?
Preparing for AIGP? Take the free 5-min readiness quiz →
- Increase the number of labeled applications immediatelyMore labels may add volume without improving validity if reviewers continue applying an undefined or inconsistent criterion.
- Use majority labels as the definitive outcome without reviewing disagreementsMajority voting can conceal systematic ambiguity and does not explain whether disagreement reflects unclear criteria or reviewer error.
- Train the model longer against the current labelsAdditional optimization may reproduce inconsistent judgments rather than establish a reliable target for evaluation.
- Clarify criteria, measure agreement, and adjudicate ambiguous cases ✓Documented criteria, agreement measurement, and adjudication address whether labels provide a dependable evaluation target.
The trap
When reviewers disagree, validate the target labels before tuning or judging the model. How to remember it
Low agreement undermines the evaluation target; clarify criteria and adjudicate ambiguity before drawing conclusions from model scores.
How many of these would you get right?
One of 1581 AIGP questions on Certsqill. Take a free five-minute check and see your score per domain — not one number, but which section to open tonight.
Test your AIGP readiness — freeMore Understanding How to Govern AI Development questions
- Test re-identification and residual disclosure risks: What is the most defensible response? →
- Link model, dataset, transformation, and test versions: What control most directly resolves the uncertainty? →
- Use targeted rare-case testing: Which evaluation response best addresses both rarity and consequence? →
- All 426 Understanding How to Govern AI Development questions →
Part of the Certsqill AIGP question bank · Understanding How to Govern AI Development ·
Every answer, right and wrong, comes with its own explanation.