AWS: 1100 practice questions with explanations
48 hours only — 15% off every course with code SAVE15. Browse courses →48h · 15% off all courses · code SAVE15 →
Certifications Tools Flashcards Career Paths Exam Guides Blog Pricing For Teams About

AWS practice questions: 1100 questions with full explanations

4 domains 1100 questions 170 min exam
Questions on the exam
about 85 — vendor indicates, no fixed count published
Time allowed
170 minutes format →

1100 practice questions for AWS Certified Machine Learning Engineer – Associate MLA-C02, grouped by exam domain. Every question below shows all four options, which one is correct, and why each of the other three is not — the wrong answers are where most candidates lose marks.

Not sure where you stand? Take the free 5-min AWS readiness check →

AWS certification: requirements, cost and exam format → ·  AWS exam format →  · 

Questions by domain

Sample questions

Store encrypted Parquet objects in S3 and scan only: Which approach best meets these requirements?

1: Data Preparation for ML and AI Easy
A retailer must retain five years of training data in its existing encrypted Amazon S3 data lake. Training jobs repeatedly read only selected columns directly from S3. The team wants to reduce bytes scanned without provisioning another storage service. Which approach best meets these requirements?
  1. Store encrypted Parquet files on EFS for shared training access.
    EFS supports shared filesystem access, but the stated durable analytical object workload is better suited to S3.
  2. Store encrypted CSV files in S3 and scan every field during training.
    CSV is row-oriented, and scanning every field prevents efficient column pruning for selective analytical queries.
  3. Store encrypted Parquet objects in S3 and scan only required columns. ✓
    S3 provides durable object storage, Parquet supports column pruning, and server-side encryption protects the objects.
  4. Store encrypted Parquet files on EBS volumes attached to training instances.
    EBS is attached block storage and is less suitable than S3 for a durable, shared, historical object dataset.
The trap
Overlooks columnar storage and selective reads. Chooses block storage for an object-lake workload. Confuses shared filesystem access with the required object-storage pattern.

All 308 1: Data Preparation for ML and AI questions →

Supervised fine-tuning on the labeled multilingual cases: What should it evaluate next?

2: ML Model and Foundation Model (FM) Development Easy
A company classifies multilingual support cases into stable categories. Prompt instructions, labeled few-shot examples and retrieval of category definitions have already failed its held-out consistency threshold. It has sufficient representative labeled cases and a foundation model that supports fine-tuning. What should it evaluate next?
  1. The previously tested few-shot prompt without new examples.
    Repeating an unchanged failed prompt does not introduce the task adaptation requested by the scenario.
  2. Retrieval of the same unchanged category definitions.
    The stem states that this approach was already evaluated and failed the required consistency threshold.
  3. A larger output-token allowance for each classification.
    A higher output allowance does not directly teach the desired category boundaries or terminology.
  4. Supervised fine-tuning on the labeled multilingual cases. ✓
    Labeled task examples can adapt classification behavior after the stated prompting and retrieval approaches failed validation.
The trap
The stem states that this approach was already evaluated and failed the required consistency threshold. Repeating an unchanged failed prompt does not introduce the task adaptation requested by the scenario. A higher output allowance does not directly teach the desired category boundaries or terminology.

All 264 2: ML Model and Foundation Model (FM) Development questions →

Use a multi-container inference pipeline: Which strategy should it choose?

3: Deployment and Orchestration of ML and AI Workflows Easy
A production line inspects images continuously. Each request must pass through an image resize and normalization step before a defect-classification model, and the combined response must remain synchronous. The team wants one managed endpoint with an ordered preprocessing-and-inference path. Which strategy should it choose?
  1. Deploy one real-time classifier and duplicate preprocessing logic in every client application.
    Client-side preprocessing does not provide one managed ordered path and can create inconsistent transformations.
  2. Use a multi-container inference pipeline. ✓
    A multi-container inference pipeline sequences compatible preprocessing and inference containers within one managed endpoint.
  3. Connect separate asynchronous preprocessing and classification endpoints through a queued workflow.
    Separate asynchronous stages do not satisfy the required synchronous combined response.
  4. Place preprocessing and classification models in a multi-model endpoint for sequential execution.
    Multi-model endpoints select among models; they do not inherently sequence preprocessing and classification containers.
The trap
This moves required deployment behavior outside the endpoint. This confuses model sharing with pipeline orchestration. This ignores the production line's synchronous decision requirement.

All 264 3: Deployment and Orchestration of ML and AI Workflows questions →

Publish separate loop and truncation metrics: What monitoring design should be implemented?

4: Operating, Monitoring, and Securing ML and AI Solutions Easy
An internal research assistant occasionally loops through a tool and sometimes receives truncated tool responses. The team needs separate alerts for loop frequency and truncation frequency, with request identifiers for investigation. What monitoring design should be implemented?
  1. Publish separate loop and truncation metrics, with request-correlated logs for investigation. ✓
    Separate metrics support independent thresholds, while correlated request logs preserve diagnostic context.
  2. Use model-quality drift monitoring to identify tool loops before instrumenting tool execution metrics and request logs.
    Model-quality metrics do not directly measure orchestration loops or truncation, and telemetry is needed to investigate those events.
  3. Alarm on total requests and review individual assistant requests only after users report failures.
    Total request volume does not distinguish loops from truncations and delays detection until users notice failures.
  4. Store tool responses with request identifiers but combine loop and truncation counts into one daily metric.
    Identifiers help tracing, but one combined metric prevents separate thresholds and obscures which failure is increasing.
The trap
This substitutes general traffic volume for workflow-specific anomaly signals. This preserves correlation while sacrificing the required diagnostic distinction. This confuses model outcome monitoring with workflow execution monitoring.

All 264 4: Operating, Monitoring, and Securing ML and AI Solutions questions →

Distribute objects across additional prefixes and adjust: What should the engineer do first?

1: Data Preparation for ML and AI Easy
A multilingual classification pipeline reads thousands of objects from Amazon S3. Recently, ingestion duration increased sharply, but object sizes and worker CPU utilization remain normal. CloudWatch shows many retries and throttling responses against one S3 prefix. What should the engineer do first?
  1. Convert all objects to EFS files without changing the ingestion readers.
    Changing storage services alone does not diagnose the concentrated request pattern or guarantee compatible reader behavior.
  2. Distribute objects across additional prefixes and adjust the reader concurrency gradually. ✓
    Spreading requests across prefixes and tuning concurrency addresses concentrated request load while validating ingestion behavior.
  3. Increase the model endpoint instance count before investigating ingestion metrics.
    Endpoint capacity does not resolve S3 prefix throttling occurring before data reaches the model.
  4. Increase worker memory because throttling responses indicate insufficient processing memory.
    Throttling responses identify service request pressure, not a worker-memory shortage, especially with normal CPU and object sizes.
The trap
Confuses inference capacity with upstream object-storage request throttling. Treats migration as a substitute for identifying and correcting the capacity bottleneck. Misinterprets service throttling as local compute resource exhaustion.

All 308 1: Data Preparation for ML and AI questions →

Train a supervised image classifier with an explanation: Which approach best meets these requirements?

2: ML Model and Foundation Model (FM) Development Easy
A factory inspects 12,000 product images daily. Defects are rare, labeled examples are available, and inspectors require a clear reason for every rejection. The solution must run on a low-cost production endpoint and classify only predefined defect categories. Which approach best meets these requirements?
  1. Use unsupervised anomaly detection.
    Anomaly detection finds unusual examples but does not directly classify the predefined defect categories represented by the labels.
  2. Use retrieval over product manuals for image decisions.
    Retrieval supplies document context but does not provide reliable category-specific visual defect classification or rejection evidence.
  3. Use a generative image model to describe defects.
    Image generation or description is unnecessary for fixed-label classification and can add cost and latency.
  4. Train a supervised image classifier with an explanation method such as saliency maps for its predefined defect labels. ✓
    Supervised classification uses the labeled categories, while an explanation method supports inspection decisions and economical inference.
The trap
It overlooks the available labeled categories. It confuses retrieved knowledge with visual classification. It treats generative image capability as equivalent to inspection classification.

All 264 2: ML Model and Foundation Model (FM) Development questions →

Use an asynchronous inference endpoint with configured: Which strategy is most appropriate?

3: Deployment and Orchestration of ML and AI Workflows Easy
An internal research assistant accepts document-analysis requests that can take several minutes and may contain large payloads. Users do not require an immediate HTTP response; they receive completion notifications later. The team wants managed model inference without keeping a client connection open. Which strategy is most appropriate?
  1. Use a real-time endpoint and wait synchronously.
    Synchronous real-time inference keeps the client waiting and is unsuitable for delayed completion.
  2. Use serverless inference with a substantially increased client timeout for each document.
    A longer timeout does not provide the queued request and deferred-result pattern required here.
  3. Use an asynchronous inference endpoint with configured output handling and notifications. ✓
    Asynchronous inference queues supported long-running or large requests and delivers results through configured output handling.
  4. Submit each document as a separate batch transform job and notify users after completion.
    Batch transform is intended for offline datasets, not individual delayed requests requiring endpoint-style submission.
The trap
This confuses interactive inference with queued inference. This treats per-request processing as scheduled bulk processing. This mistakes timeout configuration for asynchronous architecture.

All 264 3: Deployment and Orchestration of ML and AI Workflows questions →

Compare recent feature distributions with a training: Which monitoring approach best detects potentially harmf

4: Operating, Monitoring, and Securing ML and AI Solutions Easy
A recommendation model receives daily subscriber events. The training distribution was stable, but a new subscription plan changes the proportions of device types and regions. Ground-truth outcomes will not be available for several weeks. Which monitoring approach best detects potentially harmful input changes now?
  1. Compare current device and region counts with operational capacity thresholds each day.
    Capacity thresholds can reveal infrastructure pressure, but they do not compare production feature distributions with the training population.
  2. Compare recent feature distributions with a training baseline using scheduled data-drift monitoring. ✓
    Scheduled baseline comparisons can detect changes in device and region distributions before ground-truth outcomes are available.
  3. Run an A/B experiment to determine whether production feature proportions match the original training distribution before changing the model.
    A/B testing compares model variants or user experiences; it is not the appropriate mechanism for comparing production inputs with a training baseline.
  4. Wait for labels and measure recommendation accuracy after outcomes are reported.
    Waiting for labels delays detection and measures model quality rather than current input-distribution change.
The trap
Confuses delayed model-quality monitoring with immediate data-drift monitoring. Treats population composition as a scaling signal. Applies experimentation to a distribution-monitoring question.

All 264 4: Operating, Monitoring, and Securing ML and AI Solutions questions →

AWS exam: the facts

How many questions are on the AWS exam?

Around 85. The vendor does not publish a fixed count for AWS, so this is the figure it indicates rather than a guaranteed number.

How long is the AWS exam?

170 minutes. Across 85 questions that is about 120 seconds per question.

What topics does the AWS exam cover?

4 domains: 1: Data Preparation for ML and AI, 2: ML Model and Foundation Model (FM) Development, 3: Deployment and Orchestration of ML and AI Workflows, 4: Operating, Monitoring, and Securing ML and AI Solutions. Weights: 1: Data Preparation for ML and AI 28%, 2: ML Model and Foundation Model (FM) Development 24%, 3: Deployment and Orchestration of ML and AI Workflows 24%, 4: Operating, Monitoring, and Securing ML and AI Solutions 24%.

How many AWS practice questions does Certsqill have?

1100, spread across 4 exam domains. Every one shows all options, which is correct, and why each of the others is not.

Would you pass AWS today?

Five minutes, and you get a score per domain — not one number, but which section to open tonight.

Test your AWS readiness — free
Certsqill AWS question bank · 1100 questions across 4 domains · Every answer, right and wrong, comes with its own explanation.