AWS 3: Deployment and Orchestration of ML and AI Workflows: 264 practice questions
12 of the 264 3: Deployment and Orchestration of ML and AI Workflows questions in the Certsqill AWS bank, shown in full below. Each one carries an explanation for every option, not just the correct one — the wrong answers are where the marks go.
Preparing for AWS? Take the free 5-min readiness check →
1. Use SageMaker AI batch transform for Workload B: Exhibit: Workload A: interactive requests, GPU model Workload
Exhibit:
Workload A: interactive requests, GPU model
Workload B: nightly S3 input, no interactive response
Select TWO.
Select two. More than one option is correct — every correct one is ticked below.
- Use a real-time endpoint for Workload B and keep it running overnight.A continuously running endpoint can process requests, but it is unnecessary for a scheduled offline dataset.
- Use a serverless inference endpoint for Workload A's GPU model.Serverless inference is not an appropriate assumption for hosting this GPU-dependent workload.
- Use SageMaker AI batch transform for Workload B. ✓Batch transform processes an offline dataset without requiring an always-on interactive endpoint.
- Use batch transform for Workload A and poll the output after each request.Batch transform is designed for datasets, not interactive per-request inference requiring responsive GPU predictions.
- Use a SageMaker AI real-time endpoint for Workload A. ✓Real-time endpoints provide interactive inference and can host GPU-backed models for low-latency case classification.
Use real-time inference for interactive GPU requests and batch transform for the scheduled offline dataset.
2. Use a multi-container inference pipeline: Which strategy should it choose?
- Deploy one real-time classifier and duplicate preprocessing logic in every client application.Client-side preprocessing does not provide one managed ordered path and can create inconsistent transformations.
- Use a multi-container inference pipeline. ✓A multi-container inference pipeline sequences compatible preprocessing and inference containers within one managed endpoint.
- Connect separate asynchronous preprocessing and classification endpoints through a queued workflow.Separate asynchronous stages do not satisfy the required synchronous combined response.
- Place preprocessing and classification models in a multi-model endpoint for sequential execution.Multi-model endpoints select among models; they do not inherently sequence preprocessing and classification containers.
Use a multi-container inference pipeline for ordered synchronous stages.
3. Use an asynchronous inference endpoint with configured: Which strategy is most appropriate?
- Use a real-time endpoint and wait synchronously.Synchronous real-time inference keeps the client waiting and is unsuitable for delayed completion.
- Use serverless inference with a substantially increased client timeout for each document.A longer timeout does not provide the queued request and deferred-result pattern required here.
- Use an asynchronous inference endpoint with configured output handling and notifications. ✓Asynchronous inference queues supported long-running or large requests and delivers results through configured output handling.
- Submit each document as a separate batch transform job and notify users after completion.Batch transform is intended for offline datasets, not individual delayed requests requiring endpoint-style submission.
Use asynchronous inference for queued, long-running requests with later results.
4. Use Amazon Bedrock on-demand model inference: Which deployment option best fits?
- Use Amazon Bedrock on-demand model inference for the scheduled generation job. ✓Bedrock on-demand inference avoids customer-managed hosting and suits a scheduled workload using an available foundation model.
- Deploy the foundation model as a SageMaker AI batch transform artifact.Batch transform is for compatible hosted model artifacts and is not the stated Bedrock foundation-model deployment path.
- Create a continuously running SageMaker AI real-time endpoint for the foundation model.A continuously running endpoint adds hosting management and is unnecessary for this scheduled generation workload.
- Purchase provisioned throughput before measuring the predictable workload's capacity needs.Provisioned throughput can suit committed sustained usage, but purchasing it prematurely conflicts with the stated simple managed option.
Bedrock on-demand inference provides managed access to the selected foundation model without customer-managed hosting.
5. Deploy the Docker image to a SageMaker AI real-time: Which approach should it use?
- Upload the Docker image directly as an Amazon Bedrock foundation model.Bedrock model import requires supported model architectures and formats, not an arbitrary custom inference Docker image.
- Convert the model to a SageMaker built-in algorithm without changing its inference interface.Built-in algorithms require their supported training and inference contracts, unlike this custom packaged classifier.
- Deploy the Docker image to a SageMaker AI real-time endpoint. ✓SageMaker AI real-time endpoints can host custom inference containers and serve synchronous predictions from externally built models.
- Run the container only in a scheduled SageMaker AI batch transform job.Batch transform can host compatible containers, but scheduled processing does not satisfy synchronous application predictions.
A SageMaker AI real-time endpoint supports custom containers and synchronous inference for models built outside AWS.
6. Create an agent version and expose an alias: Which TWO configurations best satisfy these requirements?
Select TWO.
Select two. More than one option is correct — every correct one is ticked below.
- Have the client call the agent's test alias directly in production.The test alias is intended for testing and does not provide the stable versioned deployment integration required.
- Associate one shared knowledge base without tenant metadata filtering.A shared unfiltered knowledge base can expose one tenant's information to another tenant.
- Place tenant identifiers only in the user prompt and let the foundation model authorize requests.Prompt text is not a reliable authorization boundary and cannot enforce tenant isolation for billing operations.
- Create an agent version and expose an alias for the application to invoke. ✓An alias provides a stable application target while allowing tested agent versions to change behind that target.
- Configure an action group backed by a Lambda function that validates tenant context and authorizes billing API calls. ✓An action group integrates external APIs, while Lambda can validate tenant context and enforce authorization before executing operations.
Authorize tenant-scoped tools in an action group and expose a versioned agent through a stable alias.
7. Use provisioned throughput for the selected foundation: Which deployment choice is most appropriate?
- Run daily batch inference and cache results for all team requests.Batch results cannot serve unpredictable interactive team requests or satisfy strict response-time requirements.
- Use provisioned throughput for the selected foundation model. ✓Provisioned throughput allocates committed model capacity appropriate for stable sustained usage and strict response-time objectives.
- Use a serverless SageMaker endpoint with no provisioned concurrency.Serverless endpoint behavior does not provide the requested dedicated foundation-model capacity commitment.
- Use on-demand foundation-model inference for all requests.On-demand inference avoids commitments but does not provide the dedicated predictable capacity requested for sustained traffic.
Provisioned throughput fits stable sustained usage when the organization explicitly accepts committed dedicated foundation-model capacity.
8. Use a Bedrock knowledge base with department: Which configuration best meets these requirements?
- Use vector similarity alone and choose the newest result afterward.Similarity and post-retrieval recency selection do not reliably exclude unauthorized departments before generation.
- Retrieve globally and instruct the model to ignore unauthorized passages.Prompt instructions do not reliably enforce retrieval authorization or prevent sensitive passages from entering the model context.
- Use a Bedrock knowledge base with department and effective-date metadata filters, then generate answers with retrieved citations. ✓Metadata filters constrain retrieval before generation, while retrieved source passages support citations and effective-date selection.
- Store documents without metadata and filter displayed citations in the client.Client-side filtering occurs after generation and cannot prevent unauthorized content from influencing the answer.
Filter retrieval by metadata before generation and use retrieved passages for citations.
9. Schedule SageMaker Serverless Inference provisioned: Which solution best satisfies both requirements?
- Run a real-time endpoint with a fixed instance count throughout each day.A fixed endpoint avoids cold starts but continues consuming provisioned capacity overnight unless separately stopped.
- Schedule SageMaker Serverless Inference provisioned concurrency for shifts. ✓Scheduled provisioned concurrency keeps serverless inference warm during shifts and can be removed overnight.
- Submit batch transform jobs after each shift and analyze the accumulated sensor records later.Batch Transform cannot provide timely analysis during the production shifts.
- Use Serverless Inference without provisioned concurrency throughout the day.On-demand serverless capacity can introduce cold starts during the shifts.
Schedule serverless provisioned concurrency for shifts and remove it overnight.
10. Use Step Functions to invoke CloudFormation: Which approach is best?
- Use Step Functions to invoke CloudFormation, wait for completion, read stack outputs, and pass them to subsequent tasks. ✓Step Functions can orchestrate CloudFormation operations, preserve workflow state, and pass generated outputs between integrated service tasks.
- Use AWS Lambda to create the stack and repeatedly poll CloudFormation until completion.Lambda polling can work but requires custom state handling, retry logic, and execution management that the orchestration service provides natively.
- Use CloudFormation nested stacks and manually copy outputs into workflow parameters.Nested stacks manage infrastructure relationships but manual output copying is not an automated communication mechanism.
- Use SageMaker Pipelines to provision arbitrary CloudFormation stacks and automatically expose every stack output as a pipeline parameter.SageMaker Pipelines orchestrates ML steps, but it does not automatically expose arbitrary CloudFormation outputs as pipeline parameters.
Step Functions provides durable service orchestration, CloudFormation integration, output propagation, retries, and cleanup sequencing.
11. Deploy the separate images as containers in a SageMaker: Which TWO actions best satisfy these requirements?
Select two. More than one option is correct — every correct one is ticked below.
- Run both components in one image and select different Python environments at container startup.One image couples dependencies and deployment versions, making independent updates and reliable rollback more difficult.
- Deploy the separate images as containers in a SageMaker multi-container endpoint. ✓Multi-container endpoints support multiple framework-specific containers and can invoke containers sequentially or individually.
- Use SageMaker FastFile mode to mount the container images from an Amazon S3 bucket.FastFile exposes S3 training data read-only; it does not replace container image storage or isolate inference dependencies.
- Build separate versioned container images for the parser and embedding components, then publish them to Amazon ECR. ✓Separate versioned images isolate dependencies and allow independent release, rollback, and vulnerability scanning through Amazon ECR.
- Store the Dockerfiles in Amazon S3 and invoke them directly during endpoint deployment.S3 can store source files, but SageMaker deployment requires a built container image available through a supported registry.
Separate ECR images isolate dependencies, while a multi-container endpoint supports independent component execution within one managed endpoint.
12. Deploy the endpoint with VpcConfig using subnet-a: Which configuration should be used?
- Place the endpoint in private subnets and allow inbound traffic from 0.0.0.0/0.The endpoint would be private but the unrestricted security-group rule violates least-privilege network access.
- Deploy the endpoint in the application subnets without specifying a security group.Private subnet placement alone does not provide the intended explicit traffic control through the endpoint security group.
- Leave the endpoint public and restrict InvokeEndpoint permission to the application IAM role.IAM controls authorization but does not make a public endpoint network-private or prevent internet-level reachability.
- Deploy the endpoint with VpcConfig using subnet-a, subnet-b, and sg-endpoint. ✓VpcConfig places the endpoint network interfaces in selected private subnets and applies the security group controlling application access.
Configure SageMaker VpcConfig with private subnets and the endpoint security group so only routed application traffic can reach inference.
252 more 3: Deployment and Orchestration of ML and AI Workflows questions
The remaining 252 questions in this domain are part of the full AWS bank — 1100 questions, every option explained. Start with the free five-minute check and see your score per domain.
Test your AWS readiness — freeOther AWS domains
- 1: Data Preparation for ML and AI — 308 questions →
- 2: ML Model and Foundation Model (FM) Development — 264 questions →
- 4: Operating, Monitoring, and Securing ML and AI Solutions — 264 questions →
- All 1100 AWS questions →
- AWS certification: requirements, cost and exam format →