AWS 1: Data Preparation for ML and AI practice questions
48 hours only — 15% off every course with code SAVE15. Browse courses →48h · 15% off all courses · code SAVE15 →
Certifications Tools Flashcards Career Paths Exam Guides Blog Pricing For Teams About

AWS 1: Data Preparation for ML and AI: 308 practice questions

AWS 308 questions 12 shown free

12 of the 308 1: Data Preparation for ML and AI questions in the Certsqill AWS bank, shown in full below. Each one carries an explanation for every option, not just the correct one — the wrong answers are where the marks go.

Preparing for AWS? Take the free 5-min readiness check →

1. Consume transaction records continuously and write feature: Which TWO ingestion designs meet these requirement

Easy
A fraud system receives transaction events in Kinesis Data Streams and customer profiles as Parquet files in Amazon S3. Transaction features must reach an existing SageMaker Feature Store online store within seconds. All changed profiles must be refreshed hourly. Which TWO ingestion designs meet these requirements?

Select two. More than one option is correct — every correct one is ticked below.

  1. Refresh the online profile features from the complete S3 dataset each Sunday.
    A weekly refresh uses a supported ingestion pattern but leaves changed profiles stale beyond the stated hourly deadline.
  2. Consume transaction records continuously and write feature updates with PutRecord. ✓
    A continuously running consumer can transform transaction events and update online feature records within the required latency.
  3. Deliver transaction records to S3 and process the accumulated files every night.
    Nightly processing is executable but cannot update transaction features within seconds of receiving each event.
  4. Run a daily Athena aggregation of transactions and load its results into the online store.
    Daily aggregated features can be ingested, but that schedule cannot satisfy the seconds-level update requirement.
  5. Run an hourly Processing job that converts profile rows into feature records. ✓
    The scheduled job reads the Parquet dataset, transforms its rows, and ingests feature records on the required hourly cadence.
The trap
Confusing eventual batch availability with low-latency feature freshness. Ignoring the refresh interval when choosing a technically valid ingestion pattern. Selecting a batch aggregation workflow for event-by-event online freshness.

Continuous transaction ingestion and hourly profile processing satisfy the two independent freshness requirements without confusing files with feature records.

2. Store encrypted Parquet objects in S3 and scan only: Which approach best meets these requirements?

Easy
A retailer must retain five years of training data in its existing encrypted Amazon S3 data lake. Training jobs repeatedly read only selected columns directly from S3. The team wants to reduce bytes scanned without provisioning another storage service. Which approach best meets these requirements?
  1. Store encrypted Parquet files on EFS for shared training access.
    EFS supports shared filesystem access, but the stated durable analytical object workload is better suited to S3.
  2. Store encrypted CSV files in S3 and scan every field during training.
    CSV is row-oriented, and scanning every field prevents efficient column pruning for selective analytical queries.
  3. Store encrypted Parquet objects in S3 and scan only required columns. ✓
    S3 provides durable object storage, Parquet supports column pruning, and server-side encryption protects the objects.
  4. Store encrypted Parquet files on EBS volumes attached to training instances.
    EBS is attached block storage and is less suitable than S3 for a durable, shared, historical object dataset.
The trap
Overlooks columnar storage and selective reads. Chooses block storage for an object-lake workload. Confuses shared filesystem access with the required object-storage pattern.

Encrypted Parquet in S3 supports durable storage, column pruning, and server-side encryption.

3. Distribute objects across additional prefixes and adjust: What should the engineer do first?

Easy
A multilingual classification pipeline reads thousands of objects from Amazon S3. Recently, ingestion duration increased sharply, but object sizes and worker CPU utilization remain normal. CloudWatch shows many retries and throttling responses against one S3 prefix. What should the engineer do first?
  1. Convert all objects to EFS files without changing the ingestion readers.
    Changing storage services alone does not diagnose the concentrated request pattern or guarantee compatible reader behavior.
  2. Distribute objects across additional prefixes and adjust the reader concurrency gradually. ✓
    Spreading requests across prefixes and tuning concurrency addresses concentrated request load while validating ingestion behavior.
  3. Increase the model endpoint instance count before investigating ingestion metrics.
    Endpoint capacity does not resolve S3 prefix throttling occurring before data reaches the model.
  4. Increase worker memory because throttling responses indicate insufficient processing memory.
    Throttling responses identify service request pressure, not a worker-memory shortage, especially with normal CPU and object sizes.
The trap
Confuses inference capacity with upstream object-storage request throttling. Treats migration as a substitute for identifying and correcting the capacity bottleneck. Misinterprets service throttling as local compute resource exhaustion.

Retries and throttling on one prefix indicate concentrated S3 request pressure; redistribute objects and tune concurrency.

4. Use Kinesis Data Streams with each camera assigned: Which design meets these requirements?

Easy
Inspection cameras publish image events to a stream. Events from each camera must be consumed in that camera’s emission order; there is no required ordering across cameras. The application needs replay after consumer restarts and parallel processing across camera groups. Which design meets these requirements?
  1. Write images to EBS and scan the volume periodically.
    EBS does not provide ordered stream records, retention-aware consumer progress, or managed consumer scaling.
  2. Use Kinesis Data Streams with each camera assigned a stable partition key. ✓
    A consistent partition key preserves ordering for a camera within its shard; checkpoints support replay and other shards allow parallel consumption.
  3. Store images in EFS and have workers watch for new files.
    EFS provides shared files, not shard ordering, stream retention, or consumer checkpointing.
  4. Place images in S3 and repeatedly list the prefix.
    S3 polling adds discovery delay and does not provide native stream ordering or consumer checkpoint semantics.
The trap
Mistakes block storage for a streaming service. Uses object polling instead of stream ingestion. Confuses shared file visibility with streaming semantics.

Kinesis Data Streams supplies ordered, retained records and scalable consumers.

5. Apache Parquet: Which format should be used for the training dataset?

Easy
An internal research assistant trains periodically on millions of documents stored in Amazon S3. Training jobs usually select a subset of columns from structured metadata, while analysts occasionally inspect complete records. The team wants to reduce scan volume without sacrificing a broadly supported interchange format. Which format should be used for the training dataset?
  1. Apache Parquet ✓
    Parquet is columnar, enabling efficient column pruning when training jobs read only selected metadata fields.
  2. JSON Lines
    JSON Lines is flexible for records but remains less efficient than Parquet for repeated selective analytical scans.
  3. Apache ORC with every field duplicated as a separate file
    ORC can support columnar analytics, but duplicating every field into separate files adds unnecessary complexity and storage management.
  4. CSV
    CSV is row-oriented, so selecting a few columns generally still requires reading complete rows from each file.
The trap
Assumes a simple text format automatically provides columnar scan efficiency. Prioritizes document flexibility over the stated column-selection access pattern. Adds file fragmentation that is not required to obtain columnar scan benefits.

Parquet supports column pruning, reducing data read when training jobs select only portions of structured metadata.

6. Use a left join from preferences to purchase history: Select TWO actions.

Easy
A subscription service creates daily recommendations by combining customer preferences from DynamoDB with purchase history in Amazon S3. Customer identifiers differ in capitalization and formatting between sources. The pipeline must retain customers with preferences even when no purchase exists, and must prevent duplicate customer rows. Select TWO actions.

Select two. More than one option is correct — every correct one is ticked below.

  1. Append both datasets by row position and select the first record for each output row.
    Appending by position does not establish customer relationships and can associate unrelated records incorrectly.
  2. Use a left join from preferences to purchase history and resolve duplicate purchase records. ✓
    A left join preserves customers lacking purchases, while deduplication prevents multiple output rows per customer.
  3. Use an inner join and retain only customers appearing in both sources.
    An inner join removes customers without purchases, violating the requirement to retain all preference records.
  4. Normalize identifier casing and formatting before joining the datasets. ✓
    Standardizing join keys makes logically identical customer identifiers match across sources despite capitalization and formatting differences.
  5. Join on the original identifiers and let unmatched records receive null recommendations.
    Unnormalized identifiers cause false mismatches, so valid cross-source customer relationships remain unresolved.
The trap
Confuses matching completeness with preserving the primary customer population. Assumes missing matches represent absent purchases rather than inconsistent key formatting. Treats positional coincidence as a reliable relational key.

Normalize keys before joining, then use a left join with deduplication to preserve customers and produce one row each.

7. Use an Amazon OpenSearch Service vector index configured: Which configuration is most appropriate?

Medium
A company stores embeddings for transcribed support calls and needs similarity search across millions of vectors. The embedding model produces 768-dimensional vectors, and queries require approximate nearest-neighbor retrieval with cosine similarity. Which configuration is most appropriate?
  1. Create a 512-dimensional index and truncate each 768-dimensional embedding before insertion.
    Truncation changes the representation and conflicts with the required embedding dimensionality and model semantics.
  2. Use an Amazon OpenSearch Service vector index configured for 768 dimensions and cosine similarity. ✓
    OpenSearch supports vector indexing, and matching dimensions and cosine similarity align the index with generated embeddings.
  3. Use an exact Euclidean-distance scan because similarity configuration does not affect retrieval quality.
    Distance metric affects nearest-neighbor results, and exhaustive scans do not meet the scalable approximate-search requirement.
  4. Store vectors as JSON objects in Amazon S3 and compare every object during each query.
    S3 stores objects but does not automatically provide scalable approximate nearest-neighbor vector search over those objects.
The trap
Assumes object storage inherently supplies vector indexing and similarity retrieval. Treats dimensionality mismatch as harmless preprocessing rather than an index compatibility error. Ignores both the specified cosine metric and the need for scalable indexed retrieval.

Configure a vector index with the model’s 768 dimensions and cosine similarity for scalable approximate retrieval.

8. Store source objects in S3 with tenant-scoped IAM access: Which design should be used?

Medium
A multi-tenant analytics service accepts text reports, images, and audio recordings. Files must remain independently retrievable, tenants require logical isolation, and analysts periodically run batch processing over selected tenant datasets. The service does not require a shared mounted filesystem or millisecond feature lookups. Which design should be used? Tenant access must be enforced independently of object naming.
  1. Store all binary content as rows in a DynamoDB table with one item per byte.
    Byte-level item storage creates inefficient fragmentation and does not suit independent large-object management.
  2. Place every tenant’s files in one EBS volume and separate access through directory names.
    EBS is attached block storage and does not naturally provide scalable shared object retrieval for multiple tenants.
  3. Convert images and audio into CSV rows before storing them in Amazon S3.
    CSV is unsuitable for preserving arbitrary binary media as independently retrievable native objects.
  4. Store source objects in S3 with tenant-scoped IAM access and linked structured metadata. ✓
    S3 retains the original files, linked metadata supports indexing, and tenant-scoped permissions enforce access rather than relying on names alone.
The trap
Confuses item-oriented database storage with object storage for heterogeneous files. Uses block storage for a multi-tenant object repository and treats directories as sufficient isolation. Applies tabular serialization to diverse binary content without a stated transformation requirement.

S3 natively stores heterogeneous objects and supports tenant-oriented organization and batch retrieval without shared filesystem requirements.

9. Configure both stores and stream hourly updates by calling: Which configuration is appropriate?

Medium
Several teams share a SageMaker Feature Store feature group containing customer activity features. Online inference needs the latest value with low-latency reads, while weekly training requires the complete historical record. The platform team wants one logical feature group and will ingest hourly updates. Which configuration is appropriate?
  1. Configure only the online store and export its latest records weekly for training.
    The online store retains only the latest records, so it cannot provide the complete historical record required for training.
  2. Configure only the offline store and query it directly for millisecond online inference.
    The offline store supports historical training and batch workloads, not low-latency online feature reads.
  3. Configure both stores and stream hourly updates by calling PutRecord. ✓
    The online store serves current values, the offline store retains history, and PutRecord supports streaming feature updates.
  4. Create separate online-only feature groups for teams and duplicate hourly updates manually.
    Separate online-only groups add duplication and still fail to provide the shared historical offline store needed for weekly training.
The trap
Confuses latest-value serving storage with historical retention. Reverses the intended roles of the two Feature Store modes. Addresses team separation while omitting historical retention.

Use both stores and stream hourly updates with PutRecord for current serving and historical training.

10. Run a SageMaker batch transform job using the S3 input: Which AWS approach should be used?

Medium
A regulated organization must classify 40 million archived documents each night. The model is already packaged, inference is asynchronous, and no always-on endpoint is required. Input documents reside in Amazon S3, and predictions must be written back to S3 for audit review. Which AWS approach should be used?
  1. Deploy a persistent SageMaker real-time endpoint and invoke it once for every document.
    A persistent endpoint is unnecessary for scheduled asynchronous inference and adds ongoing serving infrastructure.
  2. Upload all documents into a Feature Store online store before invoking the model.
    Feature Store online storage is intended for current feature retrieval, not large document batch inference.
  3. Run a SageMaker batch transform job using the S3 input and output locations. ✓
    Batch Transform processes large S3 datasets without a persistent endpoint and writes corresponding inference outputs to S3.
  4. Run each document through a Lambda function and store individual predictions in EBS.
    This creates unnecessary per-document orchestration and EBS is not the required durable shared output destination.
The trap
Chooses online serving despite an explicitly batch-oriented workload with no persistent endpoint requirement. Confuses feature serving with bulk document inference input. Replaces managed batch distribution with fragmented invocation and unsuitable output storage.

SageMaker Batch Transform is designed for large S3 datasets, asynchronous inference, and output written to S3.

11. Ingest streaming updates by calling the synchronous: Which TWO actions satisfy these requirements?

Medium
A team receives streaming readings from industrial sensors and needs the latest values for millisecond-scale inference while retaining historical values for nightly training. The feature definitions must be reusable by multiple authorized teams, and updates should be available as soon as detected. Which TWO actions satisfy these requirements? Select TWO.

Select two. More than one option is correct — every correct one is ticked below.

  1. Store feature definitions only in each consuming model’s training notebook.
    Notebook-local definitions reduce sharing and can create inconsistent transformations between training and inference.
  2. Run a nightly Processing job that batches all sensor readings into the feature group.
    Nightly batching delays feature availability and therefore fails the requirement for updates as soon as they are detected.
  3. Ingest streaming updates by calling the synchronous PutRecord API. ✓
    PutRecord supports streaming ingestion and makes newly detected feature values available without waiting for batch processing.
  4. Create a SageMaker Feature Store feature group with both online and offline stores. ✓
    The online store supports latest low-latency values, while the offline store preserves historical records for training.
  5. Configure only an offline store and query historical records during every inference request.
    An offline store is intended for training and batch inference, not low-latency real-time feature retrieval.
The trap
This confuses historical feature storage with the online serving requirement. This treats batch ingestion as equivalent to streaming ingestion. This overlooks Feature Store’s reusable feature-group and metadata capabilities.

Use both Feature Store stores and PutRecord so current values serve inference while history supports training.

12. Use SageMaker Batch Transform with SplitType set to Line: Which configuration should the engineer choose?

Medium
A customer-service platform receives large transcript files in Amazon S3 each evening. The team needs asynchronous inference without a persistent endpoint, must process records independently, and wants output records aligned with their input order. The files are line-delimited and contain no embedded newline characters. Which configuration should the engineer choose? Exhibit: Input format: CSV; files: many; endpoint: unnecessary; ordering: required.
  1. Set SplitType to None so each complete file is sent in one request.
    Whole-file requests can exceed payload expectations and do not exploit the stated line-delimited record structure.
  2. Deploy a persistent real-time endpoint and send each transcript in an individual request.
    A persistent endpoint adds unnecessary operational overhead when the workload is scheduled, asynchronous, and stored in S3.
  3. Use SageMaker Batch Transform with SplitType set to Line and write results to Amazon S3. ✓
    Batch Transform handles large S3 datasets without a persistent endpoint, and line splitting preserves corresponding record order.
  4. Use a streaming transform with AssembleWith set to Line to merge all files.
    AssembleWith concatenates records within output files; it does not merge separate input files or replace Batch Transform.
The trap
This confuses real-time serving with scheduled bulk inference. This ignores the requirement to process independent records efficiently. This mistakes an output formatting parameter for a transformation service.

Batch Transform with line splitting fits asynchronous S3 inference and maintains input-to-output record ordering.

296 more 1: Data Preparation for ML and AI questions

The remaining 296 questions in this domain are part of the full AWS bank — 1100 questions, every option explained. Start with the free five-minute check and see your score per domain.

Test your AWS readiness — free

Other AWS domains

Part of the Certsqill AWS question bank · 1: Data Preparation for ML and AI · Every answer, right and wrong, comes with its own explanation.