AWS 4: Operating, Monitoring practice questions
48 hours only — 15% off every course with code SAVE15. Browse courses →48h · 15% off all courses · code SAVE15 →
Certifications Tools Flashcards Career Paths Exam Guides Blog Pricing For Teams About

AWS 4: Operating, Monitoring, and Securing ML and AI Solutions: 264 practice questions

AWS 264 questions 12 shown free

12 of the 264 4: Operating, Monitoring, and Securing ML and AI Solutions questions in the Certsqill AWS bank, shown in full below. Each one carries an explanation for every option, not just the correct one — the wrong answers are where the marks go.

Preparing for AWS? Take the free 5-min readiness check →

1. Publish endpoint and tool-failure counts as CloudWatch: Which TWO choices meet the requirements?

Easy
An image-inspection endpoint shows the following weekly observations:

Input feature distribution: shifted substantially
Prediction error rate: unavailable because labels arrive monthly
Endpoint errors: increased during tool-assisted inspections

The team requires one mechanism to detect operational failures promptly and another to detect feature-distribution changes. Which TWO choices meet the requirements? Select TWO.

Select two. More than one option is correct — every correct one is ticked below.

  1. Retrain automatically whenever either observation changes, without recording separate failure and drift signals.
    Observations need diagnosis and separate monitoring signals; automatic retraining is not an inherent response to every alarm.
  2. Create a CloudWatch alarm directly on an unparsed JSON log field without publishing a metric.
    CloudWatch alarms require metrics; arbitrary structured log fields need metric filters or explicit metric publication first.
  3. Publish endpoint and tool-failure counts as CloudWatch metrics and alarm when operational thresholds are exceeded. ✓
    CloudWatch metrics and alarms provide prompt detection of measurable endpoint and orchestration failures.
  4. Run a scheduled drift-detection pipeline comparing recent input features with the training baseline. ✓
    A drift pipeline directly compares current feature distributions with a baseline when timely labels are unavailable.
  5. Use monthly Bedrock Model Evaluation results as the primary alert for immediate endpoint tool failures.
    Model evaluation is not a prompt operational failure detector, especially when evaluation labels arrive monthly.
The trap
This confuses offline quality assessment with real-time operational monitoring. This assumes alarms natively evaluate any raw log content. This treats monitoring detection as sufficient justification for retraining.

Use CloudWatch metrics for operational failures and a scheduled baseline comparison for feature drift.

2. Publish separate loop and truncation metrics: What monitoring design should be implemented?

Easy
An internal research assistant occasionally loops through a tool and sometimes receives truncated tool responses. The team needs separate alerts for loop frequency and truncation frequency, with request identifiers for investigation. What monitoring design should be implemented?
  1. Publish separate loop and truncation metrics, with request-correlated logs for investigation. ✓
    Separate metrics support independent thresholds, while correlated request logs preserve diagnostic context.
  2. Use model-quality drift monitoring to identify tool loops before instrumenting tool execution metrics and request logs.
    Model-quality metrics do not directly measure orchestration loops or truncation, and telemetry is needed to investigate those events.
  3. Alarm on total requests and review individual assistant requests only after users report failures.
    Total request volume does not distinguish loops from truncations and delays detection until users notice failures.
  4. Store tool responses with request identifiers but combine loop and truncation counts into one daily metric.
    Identifiers help tracing, but one combined metric prevents separate thresholds and obscures which failure is increasing.
The trap
This substitutes general traffic volume for workflow-specific anomaly signals. This preserves correlation while sacrificing the required diagnostic distinction. This confuses model outcome monitoring with workflow execution monitoring.

Publish separate loop and truncation metrics with correlated request logs and independent alarms.

3. Compare recent feature distributions with a training: Which monitoring approach best detects potentially harmf

Easy
A recommendation model receives daily subscriber events. The training distribution was stable, but a new subscription plan changes the proportions of device types and regions. Ground-truth outcomes will not be available for several weeks. Which monitoring approach best detects potentially harmful input changes now?
  1. Compare current device and region counts with operational capacity thresholds each day.
    Capacity thresholds can reveal infrastructure pressure, but they do not compare production feature distributions with the training population.
  2. Compare recent feature distributions with a training baseline using scheduled data-drift monitoring. ✓
    Scheduled baseline comparisons can detect changes in device and region distributions before ground-truth outcomes are available.
  3. Run an A/B experiment to determine whether production feature proportions match the original training distribution before changing the model.
    A/B testing compares model variants or user experiences; it is not the appropriate mechanism for comparing production inputs with a training baseline.
  4. Wait for labels and measure recommendation accuracy after outcomes are reported.
    Waiting for labels delays detection and measures model quality rather than current input-distribution change.
The trap
Confuses delayed model-quality monitoring with immediate data-drift monitoring. Treats population composition as a scaling signal. Applies experimentation to a distribution-monitoring question.

Schedule data-drift checks against the training baseline before labels arrive.

4. Assign stable cohorts to model variants: Which approach is most appropriate?

Easy
A team is comparing two transcript-analysis models for an audio support application. Requests must be assigned consistently to one model during the experiment, and the team wants to compare transcription accuracy and average handling time before selecting a winner. Which approach is most appropriate?
  1. Use input-drift reports to select the better model.
    Input-distribution drift does not measure transcription accuracy or average handling time between models.
  2. Assign stable cohorts to model variants. ✓
    Stable cohort assignment supports a controlled A/B comparison without switching a request between models.
  3. Race both models and return the faster response.
    Request racing selects for latency rather than fairly comparing accuracy and handling-time outcomes.
  4. Replace the model weekly and compare period totals.
    Comparing unrelated time periods confounds model performance with changing traffic, customers, and operating conditions.
The trap
This confuses racing responses with controlled experimentation. This lacks simultaneous controlled cohorts. This substitutes drift monitoring for outcome comparison.

Assign stable cohorts to separate model variants and compare their outcome metrics.

5. Publish per-tenant loop and tool-failure metrics: Which approach best satisfies both requirements?

Easy
A multi-tenant analytics agent sometimes loops on tool calls, but operators need tenant-specific alerts and automatic remediation. The team already emits structured telemetry and wants to avoid inspecting raw logs manually. Which approach best satisfies both requirements?
  1. Store structured agent events in CloudWatch Logs and configure alarms directly against arbitrary JSON fields.
    CloudWatch alarms require metrics; arbitrary JSON log fields need metric filters or embedded metric publication first.
  2. Enable data-quality monitoring on the agent endpoint and retrain whenever a coordination violation appears.
    Data-quality monitoring addresses input distributions, while coordination failures require agent telemetry and an explicitly configured response.
  3. Create a dashboard showing raw traces and ask operators to restart agents when loops are visible.
    Dashboards support observation but do not provide tenant-specific automated detection or remediation for recurring coordination failures.
  4. Publish per-tenant loop and tool-failure metrics, then use CloudWatch alarms to invoke a remediation workflow. ✓
    Published metrics let CloudWatch alarm on measurable agent failures and invoke tenant-aware automation through configured actions.
The trap
This confuses searchable log content with native alarm metrics. This treats orchestration failures as model-data drift and assumes alarms automatically retrain models. This mistakes visualization for automation and leaves operators responsible for continuous intervention.

Publish agent coordination metrics and alarm on them to automate tenant-specific detection and remediation.

6. Run Amazon Bedrock automatic model evaluations using: Select TWO approaches that best satisfy these requiremen

Easy
An ML platform compares three foundation models for summarization and regulated advice. Teams need scalable, repeatable quality measurements and separate expert judgment for safety and factual nuance. Evaluation datasets and reviewers are available. Select TWO approaches that best satisfy these requirements.

Select two. More than one option is correct — every correct one is ticked below.

  1. Use SageMaker automatic model tuning to select the model with the highest training objective.
    Automatic model tuning optimizes specified training hyperparameters, not comparative foundation-model response quality or safety judgments.
  2. Run Amazon Bedrock automatic model evaluations using representative prompt datasets and defined quality metrics. ✓
    Automatic Bedrock evaluations provide repeatable, scalable comparisons across models using curated datasets and selected evaluation metrics.
  3. Use SageMaker Model Monitor to compare foundation-model responses without endpoint data capture or baselines.
    Model Monitor relies on captured inference data and baselines, and it does not replace Bedrock response-quality evaluations.
  4. Run Amazon Bedrock human-based model evaluations with qualified reviewers assessing safety and factual nuance. ✓
    Human evaluation captures qualitative safety and factual judgments that automated metrics may miss in regulated use cases.
  5. Use CloudWatch CPU utilization as the primary measure of summarization quality across model deployments.
    CPU utilization measures infrastructure behavior and cannot assess factuality, safety, relevance, or response quality.
The trap
This confuses hyperparameter optimization with foundation-model evaluation. This assumes production drift monitoring is a general-purpose FM benchmarking mechanism. This confuses resource performance with AI-specific output performance.

Combine automatic Bedrock evaluations for scale with human evaluations for nuanced safety and factual assessment.

7. Use a GPU-backed SageMaker inference instance family: Which instance selection is most appropriate?

Medium
A regulated document-processing endpoint runs a transformer whose measured inference bottleneck is GPU memory. Requests require predictable, low latency, and the organization accepts higher instance cost to avoid queuing. Which instance selection is most appropriate?
  1. Use a compute-optimized CPU instance family because document processing is primarily text-based.
    Text inputs do not eliminate the measured GPU-memory requirement imposed by the transformer’s inference workload.
  2. Use managed Spot Training instances for the production endpoint to reduce inference cost.
    Managed Spot Training applies to training jobs and interruption-tolerant workloads, not production endpoint hosting.
  3. Use a GPU-backed SageMaker inference instance family with sufficient memory for the transformer. ✓
    A measured GPU-memory bottleneck and predictable low-latency requirement justify a suitably sized GPU-backed inference family.
  4. Use a smaller memory-optimized CPU instance and increase request batching to reduce latency.
    Additional batching cannot resolve a GPU-memory bottleneck and can increase latency when predictable response times are required.
The trap
This prioritizes input modality over measured compute bottlenecks. This confuses a training purchasing option with an inference instance choice. This assumes batching compensates for insufficient accelerator memory.

Choose a GPU-backed inference family sized for the measured transformer memory requirement and latency target.

8. AWS X-Ray distributed tracing with instrumentation across: Which tool should they configure first?

Medium
A streaming sensor agent invokes several services and occasionally exceeds its response-time objective. Engineers need to identify which downstream call causes delay and correlate that call with the originating request. Which tool should they configure first?
  1. CloudWatch Logs with application messages written independently by each downstream service.
    Independent logs can contain useful details but require manual correlation and do not provide distributed service timing automatically.
  2. AWS X-Ray distributed tracing with instrumentation across the agent and downstream services. ✓
    X-Ray traces correlate a request across instrumented services and expose downstream segments contributing to end-to-end latency.
  3. AWS CloudTrail event history to identify the service responsible for response-time variation.
    CloudTrail records API activity for auditing and governance, not detailed request-path latency across application services.
  4. Amazon CloudWatch billing reports to locate the downstream call causing latency spikes.
    Billing reports describe costs and usage charges, not per-request timing or downstream call dependencies.
The trap
This treats uncorrelated log lines as equivalent to distributed traces. This confuses control-plane auditing with application performance analysis. This confuses resource expenditure analysis with request tracing.

Configure X-Ray distributed tracing to correlate each request with downstream service segments and latency.

9. Combine endpoint and published quality metrics: Which design best meets the requirement?

Medium
A customer-service workflow uses an inference endpoint. Operations wants one dashboard showing request volume, error rate, p95 latency, and model-quality violations, with alarms for latency and errors. Which design best meets the requirement?
  1. Use AWS Budgets to display endpoint latency, errors, and quality violations.
    AWS Budgets monitors spending thresholds, not application latency, request errors, or model-quality metrics.
  2. Combine endpoint and published quality metrics in CloudWatch, with alarms. ✓
    CloudWatch dashboards can combine endpoint metrics with separately published quality metrics, and CloudWatch alarms can monitor operational thresholds.
  3. Review endpoint logs in periodic reports.
    Periodic reports are not a unified operational dashboard and cannot provide timely automated alarms for latency and errors.
  4. Schedule Model Monitor reports as the complete operational dashboard.
    Model Monitor reports quality findings but do not replace a dashboard for request volume, latency, errors, and operational alarms.
The trap
This substitutes delayed manual review for active monitoring. This treats specialized quality monitoring as complete service monitoring. This confuses financial monitoring with service-performance monitoring.

Combine endpoint and published quality metrics in CloudWatch, then configure alarms.

10. Configure target-tracking autoscaling with a small: Which configuration is best?

Medium
A support team uses an interactive document-search endpoint. Demand varies sharply during business hours, but responses must remain low latency and available during brief traffic surges. The team wants to avoid paying for peak capacity continuously. Which configuration is best?
  1. Use asynchronous inference and allow requests to queue until capacity becomes available.
    Queuing can reduce infrastructure cost but conflicts with the interactive low-latency response requirement.
  2. Configure target-tracking autoscaling with a small baseline and a bounded maximum instance count. ✓
    Autoscaling matches capacity to demand while preserving baseline availability and limiting expansion during traffic surges.
  3. Provision enough instances for the daily peak and keep them running continuously.
    Peak provisioning meets latency needs but wastes capacity during predictable periods of lower demand.
  4. Set the endpoint capacity to zero whenever traffic falls below the average rate.
    Removing all serving capacity risks cold-start or availability delays and does not provide controlled autoscaling behavior.
The trap
This ignores the stated requirement to avoid continuous peak-capacity cost. This prioritizes utilization over the required user-facing latency. This confuses aggressive scale-in with reliable capacity management.

Use target-tracking autoscaling with baseline capacity and a bounded maximum to balance latency, availability, and cost.

11. Create an AWS Budget with threshold notifications: Select TWO cost-management approaches.

Medium
A fraud-scoring platform needs department-level chargeback and alerts when monthly inference spending approaches an approved threshold. Alerts should notify owners, while actual request controls remain in the application. Select TWO cost-management approaches.

Select two. More than one option is correct — every correct one is ticked below.

  1. Use Service Quotas to impose an exact monthly dollar limit on inference.
    Service Quotas manage service capacity limits, not arbitrary monthly spending amounts.
  2. Use automatic model tuning to enforce production inference spending limits.
    Automatic model tuning optimizes training configurations and does not enforce production billing limits.
  3. Create an AWS Budget with threshold notifications for designated owners. ✓
    AWS Budgets sends notifications when configured actual or forecasted spending thresholds are reached.
  4. Apply activated cost allocation tags to inference resources for department reporting. ✓
    Activated cost allocation tags support attribution of eligible resource costs for department chargeback.
  5. Use latency alarms as the monthly department spending threshold.
    Latency alarms monitor performance and do not attribute spending or define financial thresholds.
The trap
Confuses operational monitoring with cost management. Confuses capacity quotas with financial controls. Confuses training optimization with cost governance.

Use cost allocation tags for chargeback and AWS Budgets for threshold alerts.

12. Use SageMaker managed Spot Training with checkpointing: Which purchasing option is most suitable?

Medium
A retailer retrains a demand model nightly. Each job can tolerate interruption because checkpoints are written to Amazon S3, and completion by the next morning matters more than exact start time. Exhibit: jobs run 2–5 hours; interruptions have occurred previously without data loss. Which purchasing option is most suitable?
  1. Use SageMaker on-demand training instances for every nightly job.
    On-demand training improves interruption resistance but ignores the stated tolerance for interruptions and cost-optimization opportunity.
  2. Use SageMaker managed Spot Training with checkpointing enabled. ✓
    Managed Spot Training reduces training cost for interruption-tolerant jobs and checkpointing supports recovery after interruptions.
  3. Purchase a long-term inference capacity commitment for the nightly training workload.
    Inference capacity commitments address serving resources and do not provide the interruption-aware training behavior required here.
  4. Run each training job on a continuously running real-time inference endpoint.
    Inference endpoints are designed for serving predictions and are unsuitable as a cost-effective replacement for training jobs.
The trap
This chooses reliability beyond the explicit requirement while overlooking checkpoint support. This confuses model serving infrastructure with managed training capacity. This applies a serving purchase option to a training workload.

Managed Spot Training fits interruption-tolerant nightly jobs because S3 checkpoints enable recovery while reducing training cost.

252 more 4: Operating, Monitoring, and Securing ML and AI Solutions questions

The remaining 252 questions in this domain are part of the full AWS bank — 1100 questions, every option explained. Start with the free five-minute check and see your score per domain.

Test your AWS readiness — free

Other AWS domains

Part of the Certsqill AWS question bank · 4: Operating, Monitoring, and Securing ML and AI Solutions · Every answer, right and wrong, comes with its own explanation.