Red teaming, where testers adversarially probe the model: Which activity is being described?
Red teaming adversarially probes a model to elicit harmful or policy-violating outputs and expose weaknesses before real adversaries exploit them.
The question
After release, a governance program schedules periodic activities to assess a generative model's safety, including one where a specialized group deliberately crafts adversarial prompts to make the model produce harmful or policy-violating outputs. Which activity is being described?
Preparing for AIGP? Take the free 5-min readiness quiz →
- Threat modeling, where the team systematically enumerates potential attackers, assets and attack paths to reason about risks before any testing is performed.Plausible because threat modeling is a related security activity, but it is an analytical mapping of risks rather than the hands-on adversarial probing described.
- Conformity assessment, where an evaluation confirms the system satisfies the applicable regulatory requirements before it may lawfully be placed on the market.Plausible because conformity assessment is a real gate, but it verifies regulatory compliance rather than adversarially attacking the model to elicit harmful outputs.
- Red teaming, where testers adversarially probe the model to elicit harmful, unsafe or policy-violating behavior and expose weaknesses before adversaries do. ✓Correct because red teaming is the adversarial exercise that deliberately attempts to induce harmful or policy-violating outputs to surface weaknesses.
- Drift detection, where monitoring compares live input and output distributions against a baseline to flag when the model's environment has shifted over time.Plausible because drift detection is important monitoring, but it watches for distributional change rather than actively provoking harmful behavior.
The trap
Confusing red teaming (hands-on adversarial probing) with threat modeling (analytical enumeration of risks). How to remember it
Red teaming adversarially probes a model to elicit harmful or policy-violating outputs and expose weaknesses before real adversaries exploit them.
How many of these would you get right?
One of 1000 AIGP questions on Certsqill. Take a free five-minute check and see your score per domain — not one number, but which section to open tonight.
Test your AIGP readiness — freeMore Understanding How to Govern AI Development questions
- Track live performance and input distributions: Which practice best fulfills this objective? →
- A description of the incident: Which set of contents best fulfills the objective of managing and documenting →
- Lack of quality data: Which contributing factor most accurately explains the failure? →
- All 135 Understanding How to Govern AI Development questions →
Part of the Certsqill AIGP question bank · Understanding How to Govern AI Development ·
Every answer, right and wrong, comes with its own explanation.