AWS 3: Deployment and Orchestration of ML practice questions
48 hours only — 15% off every course with code SAVE15. Browse courses →48h · 15% off all courses · code SAVE15 →
Certifications Tools Flashcards Career Paths Exam Guides Blog Pricing For Teams About

AWS 3: Deployment and Orchestration of ML and AI Workflows: 264 practice questions

AWS 264 questions 12 shown free

12 of the 264 3: Deployment and Orchestration of ML and AI Workflows questions in the Certsqill AWS bank, shown in full below. Each one carries an explanation for every option, not just the correct one — the wrong answers are where the marks go.

Preparing for AWS? Take the free 5-min readiness check →

1. Use SageMaker AI batch transform for Workload B: Exhibit: Workload A: interactive requests, GPU model Workload

Easy
A multilingual case-classification service has two workloads. The first requires interactive predictions from a GPU-backed model. The second classifies a nightly S3 dataset and can finish before morning. The team wants managed SageMaker AI targets appropriate to each workload.

Exhibit:
Workload A: interactive requests, GPU model
Workload B: nightly S3 input, no interactive response

Select TWO.

Select two. More than one option is correct — every correct one is ticked below.

  1. Use a real-time endpoint for Workload B and keep it running overnight.
    A continuously running endpoint can process requests, but it is unnecessary for a scheduled offline dataset.
  2. Use a serverless inference endpoint for Workload A's GPU model.
    Serverless inference is not an appropriate assumption for hosting this GPU-dependent workload.
  3. Use SageMaker AI batch transform for Workload B. ✓
    Batch transform processes an offline dataset without requiring an always-on interactive endpoint.
  4. Use batch transform for Workload A and poll the output after each request.
    Batch transform is designed for datasets, not interactive per-request inference requiring responsive GPU predictions.
  5. Use a SageMaker AI real-time endpoint for Workload A. ✓
    Real-time endpoints provide interactive inference and can host GPU-backed models for low-latency case classification.
The trap
This treats an offline processing mechanism as an interactive serving target. This incorrectly generalizes serverless elasticity to GPU model hosting. This ignores the workload's batch nature and adds an unsuitable serving pattern.

Use real-time inference for interactive GPU requests and batch transform for the scheduled offline dataset.

2. Use a multi-container inference pipeline: Which strategy should it choose?

Easy
A production line inspects images continuously. Each request must pass through an image resize and normalization step before a defect-classification model, and the combined response must remain synchronous. The team wants one managed endpoint with an ordered preprocessing-and-inference path. Which strategy should it choose?
  1. Deploy one real-time classifier and duplicate preprocessing logic in every client application.
    Client-side preprocessing does not provide one managed ordered path and can create inconsistent transformations.
  2. Use a multi-container inference pipeline. ✓
    A multi-container inference pipeline sequences compatible preprocessing and inference containers within one managed endpoint.
  3. Connect separate asynchronous preprocessing and classification endpoints through a queued workflow.
    Separate asynchronous stages do not satisfy the required synchronous combined response.
  4. Place preprocessing and classification models in a multi-model endpoint for sequential execution.
    Multi-model endpoints select among models; they do not inherently sequence preprocessing and classification containers.
The trap
This moves required deployment behavior outside the endpoint. This confuses model sharing with pipeline orchestration. This ignores the production line's synchronous decision requirement.

Use a multi-container inference pipeline for ordered synchronous stages.

3. Use an asynchronous inference endpoint with configured: Which strategy is most appropriate?

Easy
An internal research assistant accepts document-analysis requests that can take several minutes and may contain large payloads. Users do not require an immediate HTTP response; they receive completion notifications later. The team wants managed model inference without keeping a client connection open. Which strategy is most appropriate?
  1. Use a real-time endpoint and wait synchronously.
    Synchronous real-time inference keeps the client waiting and is unsuitable for delayed completion.
  2. Use serverless inference with a substantially increased client timeout for each document.
    A longer timeout does not provide the queued request and deferred-result pattern required here.
  3. Use an asynchronous inference endpoint with configured output handling and notifications. ✓
    Asynchronous inference queues supported long-running or large requests and delivers results through configured output handling.
  4. Submit each document as a separate batch transform job and notify users after completion.
    Batch transform is intended for offline datasets, not individual delayed requests requiring endpoint-style submission.
The trap
This confuses interactive inference with queued inference. This treats per-request processing as scheduled bulk processing. This mistakes timeout configuration for asynchronous architecture.

Use asynchronous inference for queued, long-running requests with later results.

4. Use Amazon Bedrock on-demand model inference: Which deployment option best fits?

Easy
A subscription service generates recommendations once each morning. The workload is predictable and does not require custom model hosting, while the team wants to avoid managing inference instances. The selected foundation model is available through Amazon Bedrock and supports the required generation task. Which deployment option best fits?
  1. Use Amazon Bedrock on-demand model inference for the scheduled generation job. ✓
    Bedrock on-demand inference avoids customer-managed hosting and suits a scheduled workload using an available foundation model.
  2. Deploy the foundation model as a SageMaker AI batch transform artifact.
    Batch transform is for compatible hosted model artifacts and is not the stated Bedrock foundation-model deployment path.
  3. Create a continuously running SageMaker AI real-time endpoint for the foundation model.
    A continuously running endpoint adds hosting management and is unnecessary for this scheduled generation workload.
  4. Purchase provisioned throughput before measuring the predictable workload's capacity needs.
    Provisioned throughput can suit committed sustained usage, but purchasing it prematurely conflicts with the stated simple managed option.
The trap
This assumes every foundation-model workload needs a dedicated endpoint. This confuses a capacity commitment with the least operationally complex deployment choice. This treats Bedrock model access as a SageMaker model artifact.

Bedrock on-demand inference provides managed access to the selected foundation model without customer-managed hosting.

5. Deploy the Docker image to a SageMaker AI real-time: Which approach should it use?

Easy
A data-science team trained an audio transcript classifier outside AWS and packaged it as a Docker image with model artifacts in Amazon S3. The service needs synchronous predictions from an application, and the team wants AWS-managed endpoint deployment while retaining its custom inference code. Which approach should it use?
  1. Upload the Docker image directly as an Amazon Bedrock foundation model.
    Bedrock model import requires supported model architectures and formats, not an arbitrary custom inference Docker image.
  2. Convert the model to a SageMaker built-in algorithm without changing its inference interface.
    Built-in algorithms require their supported training and inference contracts, unlike this custom packaged classifier.
  3. Deploy the Docker image to a SageMaker AI real-time endpoint. ✓
    SageMaker AI real-time endpoints can host custom inference containers and serve synchronous predictions from externally built models.
  4. Run the container only in a scheduled SageMaker AI batch transform job.
    Batch transform can host compatible containers, but scheduled processing does not satisfy synchronous application predictions.
The trap
This assumes any external container can become a Bedrock model. This mistakes container portability for compatibility with a built-in algorithm. This ignores the required online invocation pattern.

A SageMaker AI real-time endpoint supports custom containers and synchronous inference for models built outside AWS.

6. Create an agent version and expose an alias: Which TWO configurations best satisfy these requirements?

Easy
A multi-tenant analytics assistant must answer tenant-specific questions and call each tenant's billing API. The application team requires tool calls to carry tenant context, server-side authorization before execution, and a stable application integration point while agent versions evolve. Which TWO configurations best satisfy these requirements?

Select TWO.

Select two. More than one option is correct — every correct one is ticked below.

  1. Have the client call the agent's test alias directly in production.
    The test alias is intended for testing and does not provide the stable versioned deployment integration required.
  2. Associate one shared knowledge base without tenant metadata filtering.
    A shared unfiltered knowledge base can expose one tenant's information to another tenant.
  3. Place tenant identifiers only in the user prompt and let the foundation model authorize requests.
    Prompt text is not a reliable authorization boundary and cannot enforce tenant isolation for billing operations.
  4. Create an agent version and expose an alias for the application to invoke. ✓
    An alias provides a stable application target while allowing tested agent versions to change behind that target.
  5. Configure an action group backed by a Lambda function that validates tenant context and authorizes billing API calls. ✓
    An action group integrates external APIs, while Lambda can validate tenant context and enforce authorization before executing operations.
The trap
This delegates security decisions to probabilistic model behavior. This confuses data association with tenant-aware access control. This mistakes a testing endpoint for a production deployment alias.

Authorize tenant-scoped tools in an action group and expose a versioned agent through a stable alias.

7. Use provisioned throughput for the selected foundation: Which deployment choice is most appropriate?

Medium
An ML platform serves several teams through a shared foundation-model application. Traffic is steady throughout business hours, response-time targets are strict, and usage is expected to remain stable for the next year. The organization accepts a capacity commitment to obtain predictable dedicated model capacity. Which deployment choice is most appropriate?
  1. Run daily batch inference and cache results for all team requests.
    Batch results cannot serve unpredictable interactive team requests or satisfy strict response-time requirements.
  2. Use provisioned throughput for the selected foundation model. ✓
    Provisioned throughput allocates committed model capacity appropriate for stable sustained usage and strict response-time objectives.
  3. Use a serverless SageMaker endpoint with no provisioned concurrency.
    Serverless endpoint behavior does not provide the requested dedicated foundation-model capacity commitment.
  4. Use on-demand foundation-model inference for all requests.
    On-demand inference avoids commitments but does not provide the dedicated predictable capacity requested for sustained traffic.
The trap
This ignores the explicit stability and capacity-commitment requirements. This substitutes elastic hosting for committed model capacity. This assumes all users can accept stale precomputed responses.

Provisioned throughput fits stable sustained usage when the organization explicitly accepts committed dedicated foundation-model capacity.

8. Use a Bedrock knowledge base with department: Which configuration best meets these requirements?

Medium
A regulated organization uses RAG to answer questions from policy documents. Each document contains department and effective-date metadata. Answers must cite source passages, exclude documents from other departments, and prefer the currently effective policy. Which configuration best meets these requirements?
  1. Use vector similarity alone and choose the newest result afterward.
    Similarity and post-retrieval recency selection do not reliably exclude unauthorized departments before generation.
  2. Retrieve globally and instruct the model to ignore unauthorized passages.
    Prompt instructions do not reliably enforce retrieval authorization or prevent sensitive passages from entering the model context.
  3. Use a Bedrock knowledge base with department and effective-date metadata filters, then generate answers with retrieved citations. ✓
    Metadata filters constrain retrieval before generation, while retrieved source passages support citations and effective-date selection.
  4. Store documents without metadata and filter displayed citations in the client.
    Client-side filtering occurs after generation and cannot prevent unauthorized content from influencing the answer.
The trap
It treats generation instructions as an access-control mechanism. It confuses relevance and recency with retrieval authorization. It assumes removing a citation removes information already supplied to the model.

Filter retrieval by metadata before generation and use retrieved passages for citations.

9. Schedule SageMaker Serverless Inference provisioned: Which solution best satisfies both requirements?

Medium
A factory analyzes sensor streams with predictable surges during two daily production shifts. The model must avoid cold-start delay during shifts, while compute should scale to zero overnight. Which solution best satisfies both requirements?
  1. Run a real-time endpoint with a fixed instance count throughout each day.
    A fixed endpoint avoids cold starts but continues consuming provisioned capacity overnight unless separately stopped.
  2. Schedule SageMaker Serverless Inference provisioned concurrency for shifts. ✓
    Scheduled provisioned concurrency keeps serverless inference warm during shifts and can be removed overnight.
  3. Submit batch transform jobs after each shift and analyze the accumulated sensor records later.
    Batch Transform cannot provide timely analysis during the production shifts.
  4. Use Serverless Inference without provisioned concurrency throughout the day.
    On-demand serverless capacity can introduce cold starts during the shifts.
The trap
This confuses automatic scaling with guaranteed warm capacity. This overlooks the zero-capacity overnight requirement. This treats a latency-sensitive stream as an offline workload.

Schedule serverless provisioned concurrency for shifts and remove it overnight.

10. Use Step Functions to invoke CloudFormation: Which approach is best?

Medium
A customer-service workflow must create a temporary SageMaker processing stack, pass its generated S3 output location to a later classification step, and delete the stack after completion. Operations wants one auditable orchestration workflow rather than custom polling code. Which approach is best?
  1. Use Step Functions to invoke CloudFormation, wait for completion, read stack outputs, and pass them to subsequent tasks. ✓
    Step Functions can orchestrate CloudFormation operations, preserve workflow state, and pass generated outputs between integrated service tasks.
  2. Use AWS Lambda to create the stack and repeatedly poll CloudFormation until completion.
    Lambda polling can work but requires custom state handling, retry logic, and execution management that the orchestration service provides natively.
  3. Use CloudFormation nested stacks and manually copy outputs into workflow parameters.
    Nested stacks manage infrastructure relationships but manual output copying is not an automated communication mechanism.
  4. Use SageMaker Pipelines to provision arbitrary CloudFormation stacks and automatically expose every stack output as a pipeline parameter.
    SageMaker Pipelines orchestrates ML steps, but it does not automatically expose arbitrary CloudFormation outputs as pipeline parameters.
The trap
This assumes custom polling is preferable when managed workflow integrations already provide durable orchestration. This confuses infrastructure composition with runtime communication between provisioning and workflow steps. This overstates native coupling between SageMaker Pipelines and arbitrary infrastructure stack outputs.

Step Functions provides durable service orchestration, CloudFormation integration, output propagation, retries, and cleanup sequencing.

11. Deploy the separate images as containers in a SageMaker: Which TWO actions best satisfy these requirements?

Medium
A support team is containerizing a document-search application. The parser and embedding service require different framework versions, deployments must support independent updates, and failed releases must be reversible. Which TWO actions best satisfy these requirements? Select TWO.

Select two. More than one option is correct — every correct one is ticked below.

  1. Run both components in one image and select different Python environments at container startup.
    One image couples dependencies and deployment versions, making independent updates and reliable rollback more difficult.
  2. Deploy the separate images as containers in a SageMaker multi-container endpoint. ✓
    Multi-container endpoints support multiple framework-specific containers and can invoke containers sequentially or individually.
  3. Use SageMaker FastFile mode to mount the container images from an Amazon S3 bucket.
    FastFile exposes S3 training data read-only; it does not replace container image storage or isolate inference dependencies.
  4. Build separate versioned container images for the parser and embedding components, then publish them to Amazon ECR. ✓
    Separate versioned images isolate dependencies and allow independent release, rollback, and vulnerability scanning through Amazon ECR.
  5. Store the Dockerfiles in Amazon S3 and invoke them directly during endpoint deployment.
    S3 can store source files, but SageMaker deployment requires a built container image available through a supported registry.
The trap
This mistakes runtime environment selection for true component lifecycle isolation. This confuses container build source storage with an executable container image repository. This confuses a training-data input mode with container image delivery.

Separate ECR images isolate dependencies, while a multi-container endpoint supports independent component execution within one managed endpoint.

12. Deploy the endpoint with VpcConfig using subnet-a: Which configuration should be used?

Medium
A fraud-scoring application runs in private application subnets and invokes a SageMaker real-time endpoint. The endpoint must remain unreachable from the public internet. The following configuration is available: private subnets have routes to the application, and the endpoint security group allows inbound traffic from the application security group. Which configuration should be used? Exhibit: Endpoint VPC: unavailable; Application VPC: 10.20.0.0/16; Private subnets: subnet-a, subnet-b; Application security group: sg-app; Endpoint security group: sg-endpoint.
  1. Place the endpoint in private subnets and allow inbound traffic from 0.0.0.0/0.
    The endpoint would be private but the unrestricted security-group rule violates least-privilege network access.
  2. Deploy the endpoint in the application subnets without specifying a security group.
    Private subnet placement alone does not provide the intended explicit traffic control through the endpoint security group.
  3. Leave the endpoint public and restrict InvokeEndpoint permission to the application IAM role.
    IAM controls authorization but does not make a public endpoint network-private or prevent internet-level reachability.
  4. Deploy the endpoint with VpcConfig using subnet-a, subnet-b, and sg-endpoint. ✓
    VpcConfig places the endpoint network interfaces in selected private subnets and applies the security group controlling application access.
The trap
This confuses identity authorization with network isolation. This assumes subnet placement replaces security-group configuration. This treats private subnet placement as sufficient despite an overly broad inbound rule.

Configure SageMaker VpcConfig with private subnets and the endpoint security group so only routed application traffic can reach inference.

252 more 3: Deployment and Orchestration of ML and AI Workflows questions

The remaining 252 questions in this domain are part of the full AWS bank — 1100 questions, every option explained. Start with the free five-minute check and see your score per domain.

Test your AWS readiness — free

Other AWS domains

Part of the Certsqill AWS question bank · 3: Deployment and Orchestration of ML and AI Workflows · Every answer, right and wrong, comes with its own explanation.