AWS 2: ML Model and Foundation Model (FM) Development: 264 practice questions
12 of the 264 2: ML Model and Foundation Model (FM) Development questions in the Certsqill AWS bank, shown in full below. Each one carries an explanation for every option, not just the correct one — the wrong answers are where the marks go.
Preparing for AWS? Take the free 5-min readiness check →
1. Choose Model B for responses requiring internal: Select TWO choices that satisfy the constraints.
Select two. More than one option is correct — every correct one is ticked below.
- Choose Model B for responses requiring internal planning-tool actions. ✓Model B meets the stated tool-integration requirement through AgentCore Harness for agentic planning workflows.
- Select the least expensive model regardless of tool support because retrieval changes model behavior automatically.Cost alone cannot satisfy a mandatory tool capability, and retrieval does not create unsupported model integration automatically.
- Store policies only in the model prompt during initial deployment and avoid later retrieval.Initial prompt content becomes stale as policies change and does not meet the frequently changing knowledge requirement.
- Retrieve current inventory policies at request time instead of relying on model weights alone. ✓Retrieval supplies changing knowledge without requiring repeated model modification or relying on stale embedded information.
- Choose Model A and fine-tune it weekly on policy documents to provide tool execution.Fine-tuning can adapt behavior but does not automatically provide live tool integration or guarantee current policy knowledge.
Use the model with required tool integration and retrieve changing policies at request time rather than fine-tuning for freshness.
2. Supervised fine-tuning on the labeled multilingual cases: What should it evaluate next?
- The previously tested few-shot prompt without new examples.Repeating an unchanged failed prompt does not introduce the task adaptation requested by the scenario.
- Retrieval of the same unchanged category definitions.The stem states that this approach was already evaluated and failed the required consistency threshold.
- A larger output-token allowance for each classification.A higher output allowance does not directly teach the desired category boundaries or terminology.
- Supervised fine-tuning on the labeled multilingual cases. ✓Labeled task examples can adapt classification behavior after the stated prompting and retrieval approaches failed validation.
Evaluate supervised fine-tuning using representative examples and independent held-out validation.
3. Train a supervised image classifier with an explanation: Which approach best meets these requirements?
- Use unsupervised anomaly detection.Anomaly detection finds unusual examples but does not directly classify the predefined defect categories represented by the labels.
- Use retrieval over product manuals for image decisions.Retrieval supplies document context but does not provide reliable category-specific visual defect classification or rejection evidence.
- Use a generative image model to describe defects.Image generation or description is unnecessary for fixed-label classification and can add cost and latency.
- Train a supervised image classifier with an explanation method such as saliency maps for its predefined defect labels. ✓Supervised classification uses the labeled categories, while an explanation method supports inspection decisions and economical inference.
A supervised image classifier with explanations matches the modality, labels, interpretability, and endpoint-cost requirements.
4. Use a foundation model with retrieval augmented generation: Which approach should be selected?
- Develop and train a custom language model from the laboratory procedures.A custom model requires substantial data, engineering, and maintenance compared with managed foundation-model retrieval.
- Use a foundation model with retrieval augmented generation over the procedures. ✓RAG grounds responses in changing source content without requiring extensive examples or repeated model retraining.
- Fine-tune a foundation model on the small question-answer example set.Fine-tuning can adapt behavior, but the small dataset and changing knowledge make retrieval more suitable.
- Use a fixed pretrained model without supplying laboratory procedures during inference.A pretrained model alone cannot reliably cite or reflect frequently changing, organization-specific procedures.
RAG supplies changing laboratory knowledge and citations while avoiding extensive fine-tuning data and maintenance.
5. Retrieve current catalog and subscriber features: Which architecture is most appropriate?
- Retrieve current catalog and subscriber features, then generate recommendations with an FM. ✓Retrieval supplies current catalog and subscriber information at inference without modifying the foundation model’s weights.
- Prompt the FM without retrieving subscriber events or catalog information, relying on its pretrained knowledge of the service.Prompt wording alone cannot provide private subscriber events or the service’s frequently changing catalog.
- Train a static recommendation model once on the initial catalog snapshot.A static model becomes stale as the catalog and subscriber viewing behavior change.
- Fine-tune the FM after every catalog and viewing-event update.Frequent fine-tuning is costly and unnecessary when changing catalog and subscriber information can be supplied at runtime.
Retrieve current catalog and subscriber information before generation instead of retraining model weights.
6. Run the larger model with managed Spot Training to lower: Select TWO actions that satisfy the requirements.
Select two. More than one option is correct — every correct one is ticked below.
- Stop the larger model early before its measured accuracy reaches the required threshold.Stopping before the measured 95% result provides no evidence that the model will still achieve at least 94% accuracy.
- Run the larger model with managed Spot Training to lower its training cost. ✓Managed Spot Training is a supported way to reduce SageMaker training costs while retaining the larger model’s configuration; checkpoints can improve recovery from interruptions.
- Choose the larger model because it meets the accuracy, time, and budget limits. ✓The larger model provides 95% accuracy in 90 minutes and costs 5 units, satisfying every stated limit.
- Increase baseline training epochs until the model reaches the required accuracy.The scenario provides no evidence that additional epochs can raise the baseline from 91% to at least 94% within the limits.
- Choose the baseline model because it trains faster and costs fewer units.The baseline model fails the required 94% accuracy threshold despite its lower cost and shorter training time.
The larger model meets all hard limits, and managed Spot Training can reduce its training cost.
7. Deploy the smaller model after validating that it meets: Which approach is best?
- Fine-tune the large model to improve quality before deployment.Fine-tuning may improve quality but does not address the stated latency and compute-cost constraints.
- Use the smaller model without measuring quality because it satisfies the response target.Latency and cost alone cannot establish that the smaller model meets the required answer-quality threshold.
- Deploy the large model and increase endpoint capacity until latency meets the target.Additional capacity might reduce queueing, but the model’s computation and cost remain unsuitable without evidence.
- Deploy the smaller model after validating that it meets the minimum quality threshold. ✓The smaller model satisfies latency and cost constraints, provided measured quality remains above the accepted threshold.
Validate the lower-cost model against the accepted quality threshold, then use it to satisfy latency and cost targets.
8. Amazon Textract: Which service should the team use for invoice table extraction?
- Amazon TranscribeAmazon Transcribe converts speech to text and does not specialize in extracting invoice tables.
- Amazon ComprehendAmazon Comprehend provides text analysis, while invoice table extraction requires document-layout processing.
- Amazon RekognitionAmazon Rekognition analyzes images and video, but it is not the specialized invoice table extractor.
- Amazon Textract ✓Amazon Textract extracts text, forms, and tables from documents, matching the invoice requirement directly.
Textract is the managed AWS service specialized for extracting invoice text, forms, and tables.
9. Use the SageMaker AI built-in XGBoost algorithm for binary: Which choice is most appropriate?
- Use a text-generation foundation model to generate the risk category for each record.A foundation model is unsuitable and unnecessarily complex for structured, labeled tabular classification.
- Use SageMaker AI script mode with a custom deep-learning training script.Script mode is feasible, but it adds custom-code and container responsibilities unnecessary for this built-in algorithm need.
- Use the SageMaker AI built-in XGBoost algorithm for binary classification. ✓Built-in XGBoost supports supervised tabular classification and exposes tunable hyperparameters for AUC-oriented evaluation.
- Use an unsupervised clustering algorithm to discover the two risk categories.Clustering does not use available labels to optimize supervised binary classification performance such as AUC.
Built-in XGBoost directly supports labeled tabular binary classification and configurable hyperparameters without custom containers.
10. Use a SageMaker AI PyTorch framework container in script: Which configuration best fits?
- Use File input mode and copy the complete dataset into the container image.Embedding data in an image is operationally unsuitable and does not use the requested S3 streaming approach.
- Use a built-in algorithm without supplying the custom PyTorch training loop.A built-in algorithm cannot execute the team’s custom PyTorch loop or its specialized training behavior.
- Use a SageMaker AI PyTorch framework container in script mode with FastFile input. ✓Script mode supplies custom code, while FastFile streams Amazon S3 data through a read-only channel without custom images.
- Build and maintain a custom Docker container that embeds the training loop and dataset.A custom container can run the code, but it violates the stated simplicity and no-image-maintenance preference.
A supported PyTorch container with script mode and FastFile provides custom code and efficient S3-backed input.
11. Configure the tuning job to optimize validation F1: Select TWO actions that meet the goal.
Select two. More than one option is correct — every correct one is ticked below.
- Configure the tuning job to optimize validation F1 reported by each training job. ✓AMT requires a measurable objective and can select hyperparameters producing the best reported validation metric.
- Tune by changing hyperparameters manually after each training job without defining ranges.Manual experimentation does not implement automatic search and provides no configured search space for SageMaker AI.
- Create an automatic model tuning job with ranges for the uncertain hyperparameters. ✓Automatic model tuning searches specified hyperparameter ranges across training jobs using the selected objective metric.
- Optimize customer satisfaction directly without supplying a training metric.Automatic tuning optimizes a reported training metric, not an unmeasured business outcome directly.
- Start tuning before confirming that the underlying training job runs successfully.A working training job is a prerequisite because tuning repeatedly launches that training configuration.
Define hyperparameter ranges and a validation F1 objective so AMT can search and compare training trials.
12. Enable validation-based early stopping and configure: Which approach should be used?
- Enable validation-based early stopping and configure checkpoints for managed Spot Training. ✓Early stopping avoids unproductive epochs, while checkpoints let an interrupted Spot job resume from saved progress.
- Increase epochs and disable validation checks during training.More epochs without validation waste resources and remove the signal needed to detect the observed plateau.
- Use managed Spot Training without checkpoints and restart after interruptions.Spot Training can reduce cost, but restarting loses completed progress and may extend total completion time.
- Replace the job with a larger model solely to shorten training after validation plateaus.A larger model changes capacity and resource needs; it does not inherently shorten training or provide interruption recovery.
Use validation early stopping with checkpoints for managed Spot Training.
252 more 2: ML Model and Foundation Model (FM) Development questions
The remaining 252 questions in this domain are part of the full AWS bank — 1100 questions, every option explained. Start with the free five-minute check and see your score per domain.
Test your AWS readiness — freeOther AWS domains
- 1: Data Preparation for ML and AI — 308 questions →
- 3: Deployment and Orchestration of ML and AI Workflows — 264 questions →
- 4: Operating, Monitoring, and Securing ML and AI Solutions — 264 questions →
- All 1100 AWS questions →
- AWS certification: requirements, cost and exam format →