Track live prediction accuracy against realized outcomes: Which continuous-monitoring action best justifies
Monitoring live accuracy against realized outcomes with defined drift thresholds both detects the shift and provides the evidence to trigger retraining.
The question
Six months after deployment, a retailer's demand-forecasting model shows steadily rising error as shopping patterns shift, while its offline test metrics at release remain unchanged. The vendor contract allows quarterly retraining at extra cost, and finance wants proof before approving spend. Which continuous-monitoring action best justifies and triggers the maintenance decision?
Preparing for AIGP? Take the free 5-min readiness quiz →
- Track live prediction accuracy against realized outcomes and set drift thresholds that, once breached, trigger a scheduled retraining and revalidation cycle. ✓Correct because continuous monitoring of production performance against actual results, tied to predefined thresholds, is what detects drift and objectively triggers the retraining the objective calls for.
- Rerun the original release-time offline test suite every month and retrain the model only if those fixed benchmark scores begin to noticeably decline over successive quarters.Plausible because offline testing is familiar, but a static release benchmark cannot detect real-world drift, and the stem states those metrics have not moved despite live degradation.
- Schedule retraining automatically every quarter regardless of observed performance, since the vendor contract already budgets for quarterly refreshes.Plausible as a maintenance schedule, but calendar-only retraining wastes spend when stable and lags when drift is fast, and it gives finance no evidence-based trigger.
- Commission an external audit of the model's development documentation to confirm the original training methodology was methodologically sound and well recorded.Plausible as governance assurance, but a documentation audit examines past build quality rather than monitoring live behavior, so it neither detects drift nor justifies the retraining decision.
The trap
Assuming stable offline benchmark scores mean the deployed model is still performing well in production. How to remember it
Monitoring live accuracy against realized outcomes with defined drift thresholds both detects the shift and provides the evidence to trigger retraining.
How many of these would you get right?
One of 1581 AIGP questions on Certsqill. Take a free five-minute check and see your score per domain — not one number, but which section to open tonight.
Test your AIGP readiness — freeMore Understanding How to Govern AI Deployment and Use questions
- Deliver role-based user training so staff can correctly: Which action most directly closes the →
- Red teaming, in which skilled testers deliberately attempt: Which activity is the most fit-for-purpose choice? →
- Record the event in an incident log capturing cause: Which documentation action best satisfies this governance →
- All 424 Understanding How to Govern AI Deployment and Use questions →
Part of the Certsqill AIGP question bank · Understanding How to Govern AI Deployment and Use ·
Every answer, right and wrong, comes with its own explanation.