AWS AI Practitioner Guidelines for Responsible AI: 156 practice questions
12 of the 156 Guidelines for Responsible AI questions in the Certsqill AWS AI Practitioner bank, shown in full below. Each one carries an explanation for every option, not just the correct one — the wrong answers are where the marks go.
Preparing for AWS AI Practitioner? Take the free 5-min readiness check →
1. Measure recall overall and across relevant subgroups: Which evaluation approach best addresses responsible per
- Measure recall overall and across relevant subgroups, emphasizing missed-alert costs. ✓Recall captures detected alerts, while subgroup analysis reveals whether missed-alert performance differs across relevant populations.
- Remove alert-related group information so evaluation cannot reveal unequal outcomes.Removing attributes can prevent subgroup analysis and does not eliminate proxy information or unequal performance.
- Use balanced class counts as proof that the classifier is fair and reliable.Balanced counts do not prove equal performance, label quality, or acceptable consequences across subgroups.
- Report overall accuracy only because it summarizes all classification outcomes in one number.Overall accuracy can hide missed alerts, especially when alert classes are imbalanced or operational costs differ.
Prioritize recall for costly missed alerts and compare performance across relevant subgroups rather than relying on overall accuracy.
2. Audit data representation and label quality: Which approach is best?
- Balance the number of examples per group and declare fairness once aggregate quality remains acceptable.Balanced counts do not guarantee representative content, comparable outcomes, or fair labels across groups.
- Remove demographic fields before evaluation so the model cannot use them directly.Removing fields does not remove proxy signals and prevents measuring whether outcomes differ across customer groups.
- Publish the model’s architecture and weights so customers can determine whether representation is fair.Technical transparency alone does not measure representation, subgroup outcomes, label quality, or service experiences.
- Audit data representation and label quality, evaluate outcomes by relevant group, and monitor feedback over time. ✓This approach examines representation, labels, subgroup impacts, and changing real-world experiences rather than relying on one proxy.
Assess representation, label quality, subgroup outcomes, and ongoing feedback; no single data or transparency action proves fairness.
3. Assess proxy features and subgroup outcomes before: What is the best approach before relying on the prioritiza
- Remove protected fields, then deploy if aggregate accuracy remains stable.Removing explicit attributes does not remove correlated proxies or demonstrate equitable outcomes.
- Replace the model with a larger foundation model and compare only its overall score.Model size does not inherently remove proxy relationships, biased data, or unequal subgroup performance.
- Optimize aggregate accuracy because employees make the final offer decisions.Aggregate performance can conceal unequal errors or outcomes even when people review recommendations.
- Assess proxy features and subgroup outcomes before deployment, with continued human oversight. ✓Proxy analysis, subgroup evaluation, and oversight address whether neutral-looking inputs produce unequal prioritization.
Assess correlated features and subgroup outcomes with ongoing oversight; removing explicit attributes or changing model size is insufficient.
4. Measure subgroup errors and weigh fairness against: What should the company prioritize?
- Remove subgroup identifiers from evaluation data.Removing identifiers prevents meaningful subgroup measurement and does not remove disparities or proxy effects.
- Apply one global threshold change until the model meets its overall accuracy target.A global threshold change may alter errors without addressing the specific subgroup disparity or defining an acceptable tradeoff.
- Measure subgroup errors and weigh fairness against educational harm. ✓Subgroup metrics reveal unequal impacts, and decision makers can weigh fairness against the consequences of different errors.
- Report overall accuracy while postponing subgroup analysis.Aggregate accuracy can hide unequal outcomes, so postponing subgroup analysis may leave important harm undiscovered.
Overall accuracy can hide subgroup harm; measure subgroup errors and assess fairness against educational consequences.
5. Define escalation criteria for uncertain or consequential: Which approach best fits?
- Define escalation criteria for uncertain or consequential cases and document human review outcomes. ✓Explicit escalation criteria preserve human control where uncertainty or impact warrants judgment, while documented outcomes support monitoring and improvement.
- Escalate only applications containing protected-attribute fields, regardless of model confidence.Sensitive fields alone do not identify every uncertain or harmful case, and many consequential decisions may lack those fields.
- Send every application to human reviewers and use the model only for record storage.Reviewing every case may remove useful automation rather than creating proportionate escalation for uncertain or consequential cases.
- Allow the model to make final decisions whenever its overall accuracy exceeds a target.Overall accuracy does not identify uncertainty or guarantee that high-impact individual decisions are appropriate without human oversight.
Use documented, risk-based escalation criteria so uncertain or consequential scholarship recommendations receive meaningful human review.
6. Evaluate representative changed inputs before release: What is the best responsible-AI practice?
- Evaluate representative changed inputs before release and monitor behavior after launch. ✓Pre-release challenge testing and post-release monitoring reveal whether changed inputs reduce robustness or create new harmful behavior.
- Use a larger model so changed inputs cannot affect response quality.Larger models can still fail on distribution changes, unfamiliar terminology, or unsafe requests and therefore require evaluation.
- Test only the original validation examples because they represent the approved workflow.Original examples cannot reveal failures caused by new terminology, regional variation, or unusual inputs after deployment.
- Retrain automatically whenever any new wording appears in customer messages.New wording may be harmless, and automatic retraining on unchecked data can introduce errors or undesirable patterns.
Test realistic changed inputs before launch and monitor production behavior rather than assuming previous validation or model size ensures robustness.
7. Investigate invented rules as veracity failures: Which distinction should guide the investigation?
- Investigate invented rules as privacy failures and dialect disparities as overfitting.Neither category directly describes the reported symptoms; privacy and overfitting require different evidence.
- Investigate both symptoms primarily as quality problems and use the same corrective review.Both affect quality, but they represent different risks and require different evidence and mitigation questions.
- Investigate invented rules as veracity failures and dialect disparities as fairness failures. ✓Unsupported claims concern veracity, while patterned disadvantage across dialect groups concerns fairness.
- Investigate invented rules as bias and dialect disparities as general accuracy issues.Invented content is not necessarily subgroup bias, and dialect disparities require explicit fairness analysis.
Classify invented claims as veracity concerns and patterned dialect disadvantage as a fairness concern.
8. Configure a denied-topic safeguard for the prohibited: Which control best addresses this requirement?
- Grant fewer identity permissions to users who ask about the prohibited category.Identity permissions control access, not whether generated content concerns a particular conversational topic.
- Increase the model temperature so prohibited requests receive more varied refusals.Temperature changes response variation but does not establish a policy boundary for a prohibited subject category.
- Configure a denied-topic safeguard for the prohibited discussion category. ✓Denied topics are designed to identify and restrict configured subject areas, including requests that may use varied wording.
- Use a general word filter that blocks selected terms in every request.Word filtering may miss indirect phrasing and does not express the broader semantic category the company wants to prohibit.
A denied-topic safeguard directly expresses a prohibited subject category; word filters, generation settings, and identity permissions address different concerns.
9. Mask detected sensitive values and monitor for missed: Which approach best fits?
- Remove assistant permissions while leaving permitted response content unchanged.Authorization may restrict access but does not mask sensitive values in otherwise permitted responses.
- Encrypt stored representations so generated text cannot disclose identification numbers.Encryption of stored representations does not guarantee detection or masking during generation.
- Mask detected sensitive values and monitor for missed detections. ✓Masking reduces exposure of detected values, while monitoring acknowledges that detection is imperfect.
- Rely on training so the model remembers which customer values are private.Training does not provide dependable, request-specific detection or masking of sensitive output.
Mask detected sensitive information and monitor for misses because detection is not perfect.
10. Treat filtering as one layer and separately assess privacy: What is the best response?
- Use filters for offensive language and treat the remaining feedback as suitable for decisions.Remaining feedback may contain sensitive information or other risks, and offensive-language filtering does not establish fairness or accuracy.
- Reject AI assistance because employment decisions require no automated support.Bounded assistance with safeguards and human accountability may be appropriate; rejecting all AI exceeds the evidence presented.
- Treat filtering as one layer and separately assess privacy, impacts, accuracy, and governance. ✓Filtering covers selected safeguard categories and should be combined with broader evaluation appropriate to employment-related impacts.
- Accept filtering as evidence that employment recommendations are unbiased.Content filters address configured risks but do not establish fairness, decision validity, or unbiased recommendations.
Content filtering is one safeguard, not proof of fairness or safe employment decisions; assess privacy, impacts, accuracy, and governance separately.
11. It checks whether the answer relates to the supplied: Which statement is most accurate?
- It authorizes access to any internal policy or employee record used during retrieval.Grounding checks evaluate content relationships and do not grant identity permissions or data access.
- It prevents prompt injection whenever the answer is supported by retrieved policy text.Grounding checks do not guarantee prevention of every injection or unsafe behavior.
- It checks whether the answer relates to the supplied source and query, not universal truth. ✓Contextual grounding evaluates the answer’s relationship to provided context and the query.
- It proves the answer is factually correct whenever a source was retrieved.Retrieved material does not guarantee correct interpretation, complete evidence, or universal truth.
Contextual grounding checks relation to supplied sources and queries; it does not prove universal truth or grant access.
12. Right-size the model against quality needs and measure: Which approach best addresses sustainable model select
- Choose a model based only on its lowest operating price.Price alone does not capture resource consumption, quality, latency, or whether the model satisfies the business requirement.
- Select the largest available model to maximize response quality.A larger model can improve capability, but unnecessary capacity may increase resource use without meeting a defined additional need.
- Right-size the model against quality needs and measure its resource use. ✓Sustainable selection balances required quality with resource consumption, rather than choosing capability or price in isolation.
- Replace model evaluation with a general sustainability certification.A certification may provide context, but it cannot replace evaluating the model against workload-specific quality and resource requirements.
Select the smallest suitable capability while measuring resource use against quality and business requirements.
144 more Guidelines for Responsible AI questions
The remaining 144 questions in this domain are part of the full AWS AI Practitioner bank — 1116 questions, every option explained. Start with the free five-minute check and see your score per domain.
Test your AWS AI Practitioner readiness — freeOther AWS AI Practitioner domains
- Applications of Foundation Models — 311 questions →
- Fundamentals of GenAI — 269 questions →
- Fundamentals of AI and ML — 223 questions →
- Security, Compliance, and Governance for AI Solutions — 157 questions →
- All 1116 AWS AI Practitioner questions →