AWS 4: Operating, Monitoring, and Securing ML and AI Solutions: 264 practice questions
12 of the 264 4: Operating, Monitoring, and Securing ML and AI Solutions questions in the Certsqill AWS bank, shown in full below. Each one carries an explanation for every option, not just the correct one — the wrong answers are where the marks go.
Preparing for AWS? Take the free 5-min readiness check →
1. Publish endpoint and tool-failure counts as CloudWatch: Which TWO choices meet the requirements?
Input feature distribution: shifted substantially
Prediction error rate: unavailable because labels arrive monthly
Endpoint errors: increased during tool-assisted inspections
The team requires one mechanism to detect operational failures promptly and another to detect feature-distribution changes. Which TWO choices meet the requirements? Select TWO.
Select two. More than one option is correct — every correct one is ticked below.
- Retrain automatically whenever either observation changes, without recording separate failure and drift signals.Observations need diagnosis and separate monitoring signals; automatic retraining is not an inherent response to every alarm.
- Create a CloudWatch alarm directly on an unparsed JSON log field without publishing a metric.CloudWatch alarms require metrics; arbitrary structured log fields need metric filters or explicit metric publication first.
- Publish endpoint and tool-failure counts as CloudWatch metrics and alarm when operational thresholds are exceeded. ✓CloudWatch metrics and alarms provide prompt detection of measurable endpoint and orchestration failures.
- Run a scheduled drift-detection pipeline comparing recent input features with the training baseline. ✓A drift pipeline directly compares current feature distributions with a baseline when timely labels are unavailable.
- Use monthly Bedrock Model Evaluation results as the primary alert for immediate endpoint tool failures.Model evaluation is not a prompt operational failure detector, especially when evaluation labels arrive monthly.
Use CloudWatch metrics for operational failures and a scheduled baseline comparison for feature drift.
2. Publish separate loop and truncation metrics: What monitoring design should be implemented?
- Publish separate loop and truncation metrics, with request-correlated logs for investigation. ✓Separate metrics support independent thresholds, while correlated request logs preserve diagnostic context.
- Use model-quality drift monitoring to identify tool loops before instrumenting tool execution metrics and request logs.Model-quality metrics do not directly measure orchestration loops or truncation, and telemetry is needed to investigate those events.
- Alarm on total requests and review individual assistant requests only after users report failures.Total request volume does not distinguish loops from truncations and delays detection until users notice failures.
- Store tool responses with request identifiers but combine loop and truncation counts into one daily metric.Identifiers help tracing, but one combined metric prevents separate thresholds and obscures which failure is increasing.
Publish separate loop and truncation metrics with correlated request logs and independent alarms.
3. Compare recent feature distributions with a training: Which monitoring approach best detects potentially harmf
- Compare current device and region counts with operational capacity thresholds each day.Capacity thresholds can reveal infrastructure pressure, but they do not compare production feature distributions with the training population.
- Compare recent feature distributions with a training baseline using scheduled data-drift monitoring. ✓Scheduled baseline comparisons can detect changes in device and region distributions before ground-truth outcomes are available.
- Run an A/B experiment to determine whether production feature proportions match the original training distribution before changing the model.A/B testing compares model variants or user experiences; it is not the appropriate mechanism for comparing production inputs with a training baseline.
- Wait for labels and measure recommendation accuracy after outcomes are reported.Waiting for labels delays detection and measures model quality rather than current input-distribution change.
Schedule data-drift checks against the training baseline before labels arrive.
4. Assign stable cohorts to model variants: Which approach is most appropriate?
- Use input-drift reports to select the better model.Input-distribution drift does not measure transcription accuracy or average handling time between models.
- Assign stable cohorts to model variants. ✓Stable cohort assignment supports a controlled A/B comparison without switching a request between models.
- Race both models and return the faster response.Request racing selects for latency rather than fairly comparing accuracy and handling-time outcomes.
- Replace the model weekly and compare period totals.Comparing unrelated time periods confounds model performance with changing traffic, customers, and operating conditions.
Assign stable cohorts to separate model variants and compare their outcome metrics.
5. Publish per-tenant loop and tool-failure metrics: Which approach best satisfies both requirements?
- Store structured agent events in CloudWatch Logs and configure alarms directly against arbitrary JSON fields.CloudWatch alarms require metrics; arbitrary JSON log fields need metric filters or embedded metric publication first.
- Enable data-quality monitoring on the agent endpoint and retrain whenever a coordination violation appears.Data-quality monitoring addresses input distributions, while coordination failures require agent telemetry and an explicitly configured response.
- Create a dashboard showing raw traces and ask operators to restart agents when loops are visible.Dashboards support observation but do not provide tenant-specific automated detection or remediation for recurring coordination failures.
- Publish per-tenant loop and tool-failure metrics, then use CloudWatch alarms to invoke a remediation workflow. ✓Published metrics let CloudWatch alarm on measurable agent failures and invoke tenant-aware automation through configured actions.
Publish agent coordination metrics and alarm on them to automate tenant-specific detection and remediation.
6. Run Amazon Bedrock automatic model evaluations using: Select TWO approaches that best satisfy these requiremen
Select two. More than one option is correct — every correct one is ticked below.
- Use SageMaker automatic model tuning to select the model with the highest training objective.Automatic model tuning optimizes specified training hyperparameters, not comparative foundation-model response quality or safety judgments.
- Run Amazon Bedrock automatic model evaluations using representative prompt datasets and defined quality metrics. ✓Automatic Bedrock evaluations provide repeatable, scalable comparisons across models using curated datasets and selected evaluation metrics.
- Use SageMaker Model Monitor to compare foundation-model responses without endpoint data capture or baselines.Model Monitor relies on captured inference data and baselines, and it does not replace Bedrock response-quality evaluations.
- Run Amazon Bedrock human-based model evaluations with qualified reviewers assessing safety and factual nuance. ✓Human evaluation captures qualitative safety and factual judgments that automated metrics may miss in regulated use cases.
- Use CloudWatch CPU utilization as the primary measure of summarization quality across model deployments.CPU utilization measures infrastructure behavior and cannot assess factuality, safety, relevance, or response quality.
Combine automatic Bedrock evaluations for scale with human evaluations for nuanced safety and factual assessment.
7. Use a GPU-backed SageMaker inference instance family: Which instance selection is most appropriate?
- Use a compute-optimized CPU instance family because document processing is primarily text-based.Text inputs do not eliminate the measured GPU-memory requirement imposed by the transformer’s inference workload.
- Use managed Spot Training instances for the production endpoint to reduce inference cost.Managed Spot Training applies to training jobs and interruption-tolerant workloads, not production endpoint hosting.
- Use a GPU-backed SageMaker inference instance family with sufficient memory for the transformer. ✓A measured GPU-memory bottleneck and predictable low-latency requirement justify a suitably sized GPU-backed inference family.
- Use a smaller memory-optimized CPU instance and increase request batching to reduce latency.Additional batching cannot resolve a GPU-memory bottleneck and can increase latency when predictable response times are required.
Choose a GPU-backed inference family sized for the measured transformer memory requirement and latency target.
8. AWS X-Ray distributed tracing with instrumentation across: Which tool should they configure first?
- CloudWatch Logs with application messages written independently by each downstream service.Independent logs can contain useful details but require manual correlation and do not provide distributed service timing automatically.
- AWS X-Ray distributed tracing with instrumentation across the agent and downstream services. ✓X-Ray traces correlate a request across instrumented services and expose downstream segments contributing to end-to-end latency.
- AWS CloudTrail event history to identify the service responsible for response-time variation.CloudTrail records API activity for auditing and governance, not detailed request-path latency across application services.
- Amazon CloudWatch billing reports to locate the downstream call causing latency spikes.Billing reports describe costs and usage charges, not per-request timing or downstream call dependencies.
Configure X-Ray distributed tracing to correlate each request with downstream service segments and latency.
9. Combine endpoint and published quality metrics: Which design best meets the requirement?
- Use AWS Budgets to display endpoint latency, errors, and quality violations.AWS Budgets monitors spending thresholds, not application latency, request errors, or model-quality metrics.
- Combine endpoint and published quality metrics in CloudWatch, with alarms. ✓CloudWatch dashboards can combine endpoint metrics with separately published quality metrics, and CloudWatch alarms can monitor operational thresholds.
- Review endpoint logs in periodic reports.Periodic reports are not a unified operational dashboard and cannot provide timely automated alarms for latency and errors.
- Schedule Model Monitor reports as the complete operational dashboard.Model Monitor reports quality findings but do not replace a dashboard for request volume, latency, errors, and operational alarms.
Combine endpoint and published quality metrics in CloudWatch, then configure alarms.
10. Configure target-tracking autoscaling with a small: Which configuration is best?
- Use asynchronous inference and allow requests to queue until capacity becomes available.Queuing can reduce infrastructure cost but conflicts with the interactive low-latency response requirement.
- Configure target-tracking autoscaling with a small baseline and a bounded maximum instance count. ✓Autoscaling matches capacity to demand while preserving baseline availability and limiting expansion during traffic surges.
- Provision enough instances for the daily peak and keep them running continuously.Peak provisioning meets latency needs but wastes capacity during predictable periods of lower demand.
- Set the endpoint capacity to zero whenever traffic falls below the average rate.Removing all serving capacity risks cold-start or availability delays and does not provide controlled autoscaling behavior.
Use target-tracking autoscaling with baseline capacity and a bounded maximum to balance latency, availability, and cost.
11. Create an AWS Budget with threshold notifications: Select TWO cost-management approaches.
Select two. More than one option is correct — every correct one is ticked below.
- Use Service Quotas to impose an exact monthly dollar limit on inference.Service Quotas manage service capacity limits, not arbitrary monthly spending amounts.
- Use automatic model tuning to enforce production inference spending limits.Automatic model tuning optimizes training configurations and does not enforce production billing limits.
- Create an AWS Budget with threshold notifications for designated owners. ✓AWS Budgets sends notifications when configured actual or forecasted spending thresholds are reached.
- Apply activated cost allocation tags to inference resources for department reporting. ✓Activated cost allocation tags support attribution of eligible resource costs for department chargeback.
- Use latency alarms as the monthly department spending threshold.Latency alarms monitor performance and do not attribute spending or define financial thresholds.
Use cost allocation tags for chargeback and AWS Budgets for threshold alerts.
12. Use SageMaker managed Spot Training with checkpointing: Which purchasing option is most suitable?
- Use SageMaker on-demand training instances for every nightly job.On-demand training improves interruption resistance but ignores the stated tolerance for interruptions and cost-optimization opportunity.
- Use SageMaker managed Spot Training with checkpointing enabled. ✓Managed Spot Training reduces training cost for interruption-tolerant jobs and checkpointing supports recovery after interruptions.
- Purchase a long-term inference capacity commitment for the nightly training workload.Inference capacity commitments address serving resources and do not provide the interruption-aware training behavior required here.
- Run each training job on a continuously running real-time inference endpoint.Inference endpoints are designed for serving predictions and are unsuitable as a cost-effective replacement for training jobs.
Managed Spot Training fits interruption-tolerant nightly jobs because S3 checkpoints enable recovery while reducing training cost.
252 more 4: Operating, Monitoring, and Securing ML and AI Solutions questions
The remaining 252 questions in this domain are part of the full AWS bank — 1100 questions, every option explained. Start with the free five-minute check and see your score per domain.
Test your AWS readiness — freeOther AWS domains
- 1: Data Preparation for ML and AI — 308 questions →
- 2: ML Model and Foundation Model (FM) Development — 264 questions →
- 3: Deployment and Orchestration of ML and AI Workflows — 264 questions →
- All 1100 AWS questions →
- AWS certification: requirements, cost and exam format →