Conduct targeted security and privacy testing: What evidence gap is decisive?
Performance validation and security testing answer different questions; benchmark success does not resolve demonstrated disclosure risk.
The question
An internal employee-assistance model passes response-quality benchmarks on representative requests. A separate test reveals that crafted prompts can expose confidential workplace information. Management asks whether strong benchmark performance validates the system for deployment. What evidence gap is decisive?
Preparing for AIGP? Take the free 5-min readiness quiz →
- Expand the benchmark with more ordinary employee questionsMore ordinary questions may refine usefulness evidence, but they do not address the demonstrated confidentiality attack.
- Conduct targeted security and privacy testing of disclosure scenarios ✓The observed prompt-based disclosure requires security and privacy evidence, which ordinary performance benchmarks do not provide.
- Compare response quality across additional employee subgroupsSubgroup quality analysis is valuable, but it cannot establish resistance to crafted disclosure prompts.
- Approve deployment with monitoring because benchmark performance establishes acceptable operational reliabilityMonitoring may support later oversight, but benchmark results do not resolve a known confidentiality vulnerability before deployment.
The trap
A good quality score cannot answer a security question; match each risk claim to evidence designed to test it. How to remember it
Performance validation and security testing answer different questions; benchmark success does not resolve demonstrated disclosure risk.
How many of these would you get right?
One of 1581 AIGP questions on Certsqill. Take a free five-minute check and see your score per domain — not one number, but which section to open tonight.
Test your AIGP readiness — freeMore Understanding How to Govern AI Development questions
- Measure class errors and capacity before cost-sensitive: Which next step best supports threshold selection? →
- Test re-identification and residual disclosure risks: What is the most defensible response? →
- Clarify criteria: What evidence should it prioritize before interpreting model performance? →
- All 426 Understanding How to Govern AI Development questions →
Part of the Certsqill AIGP question bank · Understanding How to Govern AI Development ·
Every answer, right and wrong, comes with its own explanation.