AWS 2: ML Model and Foundation Model (FM) practice questions
48 hours only — 15% off every course with code SAVE15. Browse courses →48h · 15% off all courses · code SAVE15 →
Certifications Tools Flashcards Career Paths Exam Guides Blog Pricing For Teams About

AWS 2: ML Model and Foundation Model (FM) Development: 264 practice questions

AWS 264 questions 12 shown free

12 of the 264 2: ML Model and Foundation Model (FM) Development questions in the Certsqill AWS bank, shown in full below. Each one carries an explanation for every option, not just the correct one — the wrong answers are where the marks go.

Preparing for AWS? Take the free 5-min readiness check →

1. Choose Model B for responses requiring internal: Select TWO choices that satisfy the constraints.

Easy
A retailer evaluates Amazon Bedrock models for demand-planning explanations. The solution must answer questions using frequently changing inventory policies and must support an agent that can call an internal planning tool. Exhibit: Model A supports text generation only. Model B supports text generation and tool integration through AgentCore Harness. Select TWO choices that satisfy the constraints.

Select two. More than one option is correct — every correct one is ticked below.

  1. Choose Model B for responses requiring internal planning-tool actions. ✓
    Model B meets the stated tool-integration requirement through AgentCore Harness for agentic planning workflows.
  2. Select the least expensive model regardless of tool support because retrieval changes model behavior automatically.
    Cost alone cannot satisfy a mandatory tool capability, and retrieval does not create unsupported model integration automatically.
  3. Store policies only in the model prompt during initial deployment and avoid later retrieval.
    Initial prompt content becomes stale as policies change and does not meet the frequently changing knowledge requirement.
  4. Retrieve current inventory policies at request time instead of relying on model weights alone. ✓
    Retrieval supplies changing knowledge without requiring repeated model modification or relying on stale embedded information.
  5. Choose Model A and fine-tune it weekly on policy documents to provide tool execution.
    Fine-tuning can adapt behavior but does not automatically provide live tool integration or guarantee current policy knowledge.
The trap
The misconception is treating fine-tuning as both a live knowledge store and a tool-execution mechanism. The misconception is assuming retrieval compensates for missing capabilities. The misconception is treating deployment-time context as continuously current knowledge.

Use the model with required tool integration and retrieve changing policies at request time rather than fine-tuning for freshness.

2. Supervised fine-tuning on the labeled multilingual cases: What should it evaluate next?

Easy
A company classifies multilingual support cases into stable categories. Prompt instructions, labeled few-shot examples and retrieval of category definitions have already failed its held-out consistency threshold. It has sufficient representative labeled cases and a foundation model that supports fine-tuning. What should it evaluate next?
  1. The previously tested few-shot prompt without new examples.
    Repeating an unchanged failed prompt does not introduce the task adaptation requested by the scenario.
  2. Retrieval of the same unchanged category definitions.
    The stem states that this approach was already evaluated and failed the required consistency threshold.
  3. A larger output-token allowance for each classification.
    A higher output allowance does not directly teach the desired category boundaries or terminology.
  4. Supervised fine-tuning on the labeled multilingual cases. ✓
    Labeled task examples can adapt classification behavior after the stated prompting and retrieval approaches failed validation.
The trap
The stem states that this approach was already evaluated and failed the required consistency threshold. Repeating an unchanged failed prompt does not introduce the task adaptation requested by the scenario. A higher output allowance does not directly teach the desired category boundaries or terminology.

Evaluate supervised fine-tuning using representative examples and independent held-out validation.

3. Train a supervised image classifier with an explanation: Which approach best meets these requirements?

Easy
A factory inspects 12,000 product images daily. Defects are rare, labeled examples are available, and inspectors require a clear reason for every rejection. The solution must run on a low-cost production endpoint and classify only predefined defect categories. Which approach best meets these requirements?
  1. Use unsupervised anomaly detection.
    Anomaly detection finds unusual examples but does not directly classify the predefined defect categories represented by the labels.
  2. Use retrieval over product manuals for image decisions.
    Retrieval supplies document context but does not provide reliable category-specific visual defect classification or rejection evidence.
  3. Use a generative image model to describe defects.
    Image generation or description is unnecessary for fixed-label classification and can add cost and latency.
  4. Train a supervised image classifier with an explanation method such as saliency maps for its predefined defect labels. ✓
    Supervised classification uses the labeled categories, while an explanation method supports inspection decisions and economical inference.
The trap
It overlooks the available labeled categories. It confuses retrieved knowledge with visual classification. It treats generative image capability as equivalent to inspection classification.

A supervised image classifier with explanations matches the modality, labels, interpretability, and endpoint-cost requirements.

4. Use a foundation model with retrieval augmented generation: Which approach should be selected?

Easy
A research group needs an internal assistant that answers questions from frequently changing laboratory procedures. Answers must cite current source passages, and the team has only a small set of question-answer examples. The group wants minimal model maintenance and does not need the assistant to learn a new writing style. Which approach should be selected?
  1. Develop and train a custom language model from the laboratory procedures.
    A custom model requires substantial data, engineering, and maintenance compared with managed foundation-model retrieval.
  2. Use a foundation model with retrieval augmented generation over the procedures. ✓
    RAG grounds responses in changing source content without requiring extensive examples or repeated model retraining.
  3. Fine-tune a foundation model on the small question-answer example set.
    Fine-tuning can adapt behavior, but the small dataset and changing knowledge make retrieval more suitable.
  4. Use a fixed pretrained model without supplying laboratory procedures during inference.
    A pretrained model alone cannot reliably cite or reflect frequently changing, organization-specific procedures.
The trap
This mistakes behavior adaptation for maintaining current factual knowledge. This ignores the stated need for minimal maintenance and limited examples. This assumes general pretraining contains the laboratory's current private information.

RAG supplies changing laboratory knowledge and citations while avoiding extensive fine-tuning data and maintenance.

5. Retrieve current catalog and subscriber features: Which architecture is most appropriate?

Easy
A streaming service wants daily recommendations for each subscriber. Recommendations must reflect the subscriber's latest viewing events and the current catalog, while the catalog changes several times each day. The company can tolerate offline preparation but cannot retrain a language model whenever titles or preferences change. Which architecture is most appropriate?
  1. Retrieve current catalog and subscriber features, then generate recommendations with an FM. ✓
    Retrieval supplies current catalog and subscriber information at inference without modifying the foundation model’s weights.
  2. Prompt the FM without retrieving subscriber events or catalog information, relying on its pretrained knowledge of the service.
    Prompt wording alone cannot provide private subscriber events or the service’s frequently changing catalog.
  3. Train a static recommendation model once on the initial catalog snapshot.
    A static model becomes stale as the catalog and subscriber viewing behavior change.
  4. Fine-tune the FM after every catalog and viewing-event update.
    Frequent fine-tuning is costly and unnecessary when changing catalog and subscriber information can be supplied at runtime.
The trap
Treating changing facts as weight-updating training data. Ignoring the explicit freshness requirement. Assuming a general model already knows private and dynamic service data.

Retrieve current catalog and subscriber information before generation instead of retraining model weights.

6. Run the larger model with managed Spot Training to lower: Select TWO actions that satisfy the requirements.

Easy
A team analyzes recorded customer calls. The baseline transcription model achieves 91% word accuracy in 20 minutes and costs 1 unit per training run. A larger model achieves 95% in 90 minutes and costs 5 units. The business requires at least 94% accuracy, permits 100 minutes, and allows a maximum training budget of 5 units. Select TWO actions that satisfy the requirements.

Select two. More than one option is correct — every correct one is ticked below.

  1. Stop the larger model early before its measured accuracy reaches the required threshold.
    Stopping before the measured 95% result provides no evidence that the model will still achieve at least 94% accuracy.
  2. Run the larger model with managed Spot Training to lower its training cost. ✓
    Managed Spot Training is a supported way to reduce SageMaker training costs while retaining the larger model’s configuration; checkpoints can improve recovery from interruptions.
  3. Choose the larger model because it meets the accuracy, time, and budget limits. ✓
    The larger model provides 95% accuracy in 90 minutes and costs 5 units, satisfying every stated limit.
  4. Increase baseline training epochs until the model reaches the required accuracy.
    The scenario provides no evidence that additional epochs can raise the baseline from 91% to at least 94% within the limits.
  5. Choose the baseline model because it trains faster and costs fewer units.
    The baseline model fails the required 94% accuracy threshold despite its lower cost and shorter training time.
The trap
This prioritizes efficiency over a hard quality requirement. This assumes early stopping preserves target performance without validation evidence. This assumes more training guarantees a sufficient capacity improvement.

The larger model meets all hard limits, and managed Spot Training can reduce its training cost.

7. Deploy the smaller model after validating that it meets: Which approach is best?

Medium
A multi-tenant analytics API serves thousands of requests per minute. A large model provides the best answer quality but exceeds the 300-millisecond response target and consumes excessive compute. A smaller model meets the response target at one-third the cost but has lower quality. The business accepts the smaller model if responses meet a defined evaluation threshold. Which approach is best?
  1. Fine-tune the large model to improve quality before deployment.
    Fine-tuning may improve quality but does not address the stated latency and compute-cost constraints.
  2. Use the smaller model without measuring quality because it satisfies the response target.
    Latency and cost alone cannot establish that the smaller model meets the required answer-quality threshold.
  3. Deploy the large model and increase endpoint capacity until latency meets the target.
    Additional capacity might reduce queueing, but the model’s computation and cost remain unsuitable without evidence.
  4. Deploy the smaller model after validating that it meets the minimum quality threshold. ✓
    The smaller model satisfies latency and cost constraints, provided measured quality remains above the accepted threshold.
The trap
This assumes scaling infrastructure alone resolves the large model’s cost and latency tradeoff. This optimizes quality while leaving the dominant deployment constraints unresolved. This treats operational metrics as a substitute for required model evaluation.

Validate the lower-cost model against the accepted quality threshold, then use it to satisfy latency and cost targets.

8. Amazon Textract: Which service should the team use for invoice table extraction?

Medium
An ML platform team receives three requests: extract tables from invoices, transcribe recorded meetings, and detect objects in warehouse images. The team wants managed AWS AI services rather than training or operating custom models. Each request requires the service specialized for its input modality and output. Which service should the team use for invoice table extraction?
  1. Amazon Transcribe
    Amazon Transcribe converts speech to text and does not specialize in extracting invoice tables.
  2. Amazon Comprehend
    Amazon Comprehend provides text analysis, while invoice table extraction requires document-layout processing.
  3. Amazon Rekognition
    Amazon Rekognition analyzes images and video, but it is not the specialized invoice table extractor.
  4. Amazon Textract ✓
    Amazon Textract extracts text, forms, and tables from documents, matching the invoice requirement directly.
The trap
This confuses audio transcription with document structure extraction. This selects a visual analysis service for structured document extraction. This overlooks the distinction between analyzing text and extracting document structure.

Textract is the managed AWS service specialized for extracting invoice text, forms, and tables.

9. Use the SageMaker AI built-in XGBoost algorithm for binary: Which choice is most appropriate?

Medium
A compliance team must classify millions of tabular records into two risk categories. The data is stored in Amazon S3, labels are available, and the team wants a SageMaker AI algorithm without maintaining a custom training container. The evaluation metric is AUC, and the team needs configurable model hyperparameters. Which choice is most appropriate?
  1. Use a text-generation foundation model to generate the risk category for each record.
    A foundation model is unsuitable and unnecessarily complex for structured, labeled tabular classification.
  2. Use SageMaker AI script mode with a custom deep-learning training script.
    Script mode is feasible, but it adds custom-code and container responsibilities unnecessary for this built-in algorithm need.
  3. Use the SageMaker AI built-in XGBoost algorithm for binary classification. ✓
    Built-in XGBoost supports supervised tabular classification and exposes tunable hyperparameters for AUC-oriented evaluation.
  4. Use an unsupervised clustering algorithm to discover the two risk categories.
    Clustering does not use available labels to optimize supervised binary classification performance such as AUC.
The trap
This selects custom implementation despite the explicit desire to avoid custom training containers. This confuses discovering groups with learning labeled risk outcomes. This treats a structured prediction problem as open-ended language generation.

Built-in XGBoost directly supports labeled tabular binary classification and configurable hyperparameters without custom containers.

10. Use a SageMaker AI PyTorch framework container in script: Which configuration best fits?

Medium
A team processes streaming sensor data with a custom PyTorch training loop. The team wants to change only the training script, pass hyperparameters through the SageMaker AI training job, and keep input data in Amazon S3 without building a Docker image. Training data is too large to copy fully before execution. Which configuration best fits?
  1. Use File input mode and copy the complete dataset into the container image.
    Embedding data in an image is operationally unsuitable and does not use the requested S3 streaming approach.
  2. Use a built-in algorithm without supplying the custom PyTorch training loop.
    A built-in algorithm cannot execute the team’s custom PyTorch loop or its specialized training behavior.
  3. Use a SageMaker AI PyTorch framework container in script mode with FastFile input. ✓
    Script mode supplies custom code, while FastFile streams Amazon S3 data through a read-only channel without custom images.
  4. Build and maintain a custom Docker container that embeds the training loop and dataset.
    A custom container can run the code, but it violates the stated simplicity and no-image-maintenance preference.
The trap
This adds container operations when a supported framework container already fits. This confuses training input channels with data baked into container images. This sacrifices required custom logic merely to avoid script configuration.

A supported PyTorch container with script mode and FastFile provides custom code and efficient S3-backed input.

11. Configure the tuning job to optimize validation F1: Select TWO actions that meet the goal.

Medium
A customer-service classifier has three uncertain hyperparameters. The team has a labeled validation set, a measurable F1 objective, and a working SageMaker AI training job. It wants SageMaker AI to search configured ranges automatically while limiting wasted training on poor trials. Select TWO actions that meet the goal.

Select two. More than one option is correct — every correct one is ticked below.

  1. Configure the tuning job to optimize validation F1 reported by each training job. ✓
    AMT requires a measurable objective and can select hyperparameters producing the best reported validation metric.
  2. Tune by changing hyperparameters manually after each training job without defining ranges.
    Manual experimentation does not implement automatic search and provides no configured search space for SageMaker AI.
  3. Create an automatic model tuning job with ranges for the uncertain hyperparameters. ✓
    Automatic model tuning searches specified hyperparameter ranges across training jobs using the selected objective metric.
  4. Optimize customer satisfaction directly without supplying a training metric.
    Automatic tuning optimizes a reported training metric, not an unmeasured business outcome directly.
  5. Start tuning before confirming that the underlying training job runs successfully.
    A working training job is a prerequisite because tuning repeatedly launches that training configuration.
The trap
This confuses iterative experimentation with automatic model tuning. This assumes AMT can evaluate arbitrary business results without a metric signal. This overlooks the need to validate the algorithm and training setup first.

Define hyperparameter ranges and a validation F1 objective so AMT can search and compare training trials.

12. Enable validation-based early stopping and configure: Which approach should be used?

Medium
A support team trains a document-search ranking model on a large Amazon S3 dataset. A pilot run shows validation performance stops improving after several epochs. The training job can be interrupted, and the team wants shorter training while preserving the ability to resume. Exhibit: Training log: epoch 6 validation score 0.88; epoch 7 score 0.88; epoch 8 score 0.879. Which approach should be used?
  1. Enable validation-based early stopping and configure checkpoints for managed Spot Training. ✓
    Early stopping avoids unproductive epochs, while checkpoints let an interrupted Spot job resume from saved progress.
  2. Increase epochs and disable validation checks during training.
    More epochs without validation waste resources and remove the signal needed to detect the observed plateau.
  3. Use managed Spot Training without checkpoints and restart after interruptions.
    Spot Training can reduce cost, but restarting loses completed progress and may extend total completion time.
  4. Replace the job with a larger model solely to shorten training after validation plateaus.
    A larger model changes capacity and resource needs; it does not inherently shorten training or provide interruption recovery.
The trap
Ignores evidence that validation performance has stopped improving. Uses Spot capacity while omitting recovery state. Confuses model capacity with training-time reduction and checkpointing.

Use validation early stopping with checkpoints for managed Spot Training.

252 more 2: ML Model and Foundation Model (FM) Development questions

The remaining 252 questions in this domain are part of the full AWS bank — 1100 questions, every option explained. Start with the free five-minute check and see your score per domain.

Test your AWS readiness — free

Other AWS domains

Part of the Certsqill AWS question bank · 2: ML Model and Foundation Model (FM) Development · Every answer, right and wrong, comes with its own explanation.