AWS 1: Data Ingestion and Transformation practice questions
48 hours only — 15% off every course with code SAVE15. Browse courses →48h · 15% off all courses · code SAVE15 →
Certifications Tools Flashcards Career Paths Exam Guides Blog Pricing For Teams About

AWS 1: Data Ingestion and Transformation: 306 practice questions

AWS 306 questions 12 shown free

12 of the 306 1: Data Ingestion and Transformation questions in the Certsqill AWS bank, shown in full below. Each one carries an explanation for every option, not just the correct one — the wrong answers are where the marks go.

Preparing for AWS? Take the free 5-min readiness check →

1. Split each device across subkeys and resequence: Which change best addresses the problem?

Easy
A fleet publishes telemetry to a provisioned Kinesis data stream. One device produces 1.2 MB/s, exceeding the 1 MB/s write limit of a shard. Consumers require ordered records per device, and global ordering is unnecessary. The application can assign monotonically increasing device sequence values and buffer records for resequencing. Which change best addresses the problem?
Exhibit: partition key = deviceId; shard capacity = 1 MB/s writes.
  1. Add consumers that read the hot shard before accepting more device records.
    Consumer coordination does not redistribute writes or increase the hot shard's capacity.
  2. Split each device across subkeys and resequence by its device sequence value. ✓
    Subkeys distribute one hot device across shards, while downstream resequencing preserves device order without requiring global ordering.
  3. Use sequence numbers as partition keys for all records.
    Kinesis assigns sequence numbers after ingestion, so producers cannot use them to route incoming records.
  4. Increase PutRecords batch size until the hot shard accepts the device.
    PutRecords improves request efficiency but does not bypass a shard's write-throughput limit.
The trap
Confuses service-assigned metadata with producer-controlled partitioning. Confuses batching with additional shard capacity. Treats a consumer-side change as producer-side load balancing.

Split the hot device across subkeys and resequence its records downstream using device sequence values.

2. Configure an AWS Glue job to read the S3 objects: Which option best fits?

Easy
A financial institution receives one compressed CSV file per hour in Amazon S3. Each file contains transactions for a complete hour, and processing may begin after the file arrives. The team wants a managed batch ingestion mechanism that can read the files, transform them, and write Parquet output for downstream analytics. Which option best fits?
  1. Create a Kinesis data stream and publish each CSV line as a separate streaming record.
    Kinesis could transport records, but it adds unnecessary streaming complexity to complete hourly batch files.
  2. Use an Amazon MSK consumer group to poll the S3 bucket for completed files.
    MSK consumers process Kafka records; they do not natively poll S3 as a batch-file ingestion mechanism.
  3. Configure DynamoDB Streams to discover newly created S3 objects and process their contents.
    DynamoDB Streams reports DynamoDB item changes and does not provide notifications for S3 object creation.
  4. Configure an AWS Glue job to read the S3 objects, transform rows, and write Parquet. ✓
    AWS Glue jobs provide managed batch processing for S3 files and can transform records into analytics-friendly output formats.
The trap
This treats a Kafka consumer as a generic managed file reader. This assumes DynamoDB Streams is a general-purpose object-storage event source. This confuses a periodic file workload with continuously arriving streaming data.

An AWS Glue job directly processes hourly S3 files as batch input and writes transformed Parquet results.

3. Enable job bookmarks and use a stable landing prefix: Which configuration is most appropriate?

Easy
A logistics Glue job processes daily delivery files from an S3 landing prefix. A retry must not reprocess unchanged input, but corrected files with new object versions must be processed. The job writes transformed results to a separate output prefix. Which configuration is most appropriate?
  1. Delete the output prefix before every run and disable input tracking.
    Deleting outputs does not identify previously processed inputs and can remove valid historical delivery results.
  2. Store the last successful run timestamp only in the output object name.
    Output naming alone does not make Glue aware of which source objects were successfully processed.
  3. Enable job bookmarks and use a stable landing prefix for incremental input tracking. ✓
    Glue job bookmarks track previously processed input and allow newly changed or added source objects to be considered.
  4. Use a fixed filename filter and overwrite every transformed object on each run.
    Fixed filenames and repeated overwrites can miss corrected objects or cause unnecessary reprocessing of unchanged input.
The trap
This confuses output cleanup with source-progress tracking; bookmarks do not clean outputs. This treats a naming convention as an ingestion checkpoint mechanism. This assumes naming conventions provide reliable incremental processing semantics.

Glue job bookmarks track processed input objects, while output management remains a separate responsibility.

4. Persist tokens: Which approach is best?

Easy
A customer-support pipeline calls a vendor API that returns tickets in pages of 200 records. The API supplies a continuation token, occasionally returns throttling responses, and can return the same page after a retry. The pipeline must retrieve all tickets without duplicates and resume after interruption. Which approach is best?
  1. Persist tokens, back off on throttling, and deduplicate ticket IDs. ✓
    Persisted tokens support resumption, backoff handles throttling, and ticket IDs make repeated pages idempotent.
  2. Use parallel requests without coordination to finish quickly.
    Uncoordinated concurrency can exceed vendor limits and create overlapping pages.
  3. Request maximum pages and restart from the beginning after failures.
    Restarting from the beginning increases duplication and does not provide reliable checkpoint recovery.
  4. Use response timestamps as cursors and ignore repeated records.
    Response timestamps are not the documented pagination cursor, and ignoring duplicates can corrupt data.
The trap
It treats parallelism as universally beneficial despite rate limits and resumability requirements. It substitutes local metadata for the vendor's continuation mechanism. It assumes larger pages and restarts solve checkpointing and duplicate delivery.

Persist the vendor token, back off on throttling, and deduplicate ticket IDs.

5. Use Apache Airflow in Amazon MWAA to define dependencies: Select TWO actions that satisfy these requirements.

Easy
A healthcare reporting platform runs two independent Glue jobs every night and also runs a three-step workflow whose steps must execute in order. The organization wants managed scheduling, retries for workflow tasks, and no custom polling code. Select TWO actions that satisfy these requirements.

Select two. More than one option is correct — every correct one is ticked below.

  1. Place a time-based schedule on every workflow step independently.
    Independent schedules can overlap or run out of order, so they cannot reliably enforce step dependencies.
  2. Use an S3 notification for each job and rely on object creation timing to enforce workflow order.
    S3 notifications signal object events but do not inherently model ordered workflow dependencies or task retries.
  3. Create a Kinesis stream and use record arrival to schedule each nightly transformation.
    Kinesis transports streaming records and does not provide the requested time-based batch workflow orchestration.
  4. Use Apache Airflow in Amazon MWAA to define dependencies and retries for the three-step workflow. ✓
    Airflow DAGs express task dependencies and retry behavior for multi-step workflows managed through Amazon MWAA.
  5. Use EventBridge schedules to start each independent nightly Glue job. ✓
    EventBridge provides managed time-based rules suitable for starting independent scheduled Glue jobs.
The trap
This assumes matching times provide orchestration semantics. This substitutes streaming ingestion for a scheduler and workflow engine. This mistakes event arrival for explicit orchestration and dependency management.

EventBridge schedules independent jobs, while Airflow manages ordered multi-step workflows with retries.

6. Route S3 object events to EventBridge and filter the rule: Which configuration best meets the requirements?

Medium
A media company stores uploaded video metadata in an S3 bucket. A Lambda function must start transcoding only for objects under the raw/ prefix with a .json suffix. The design must support multiple event consumers and tolerate duplicate notifications. Which configuration best meets the requirements?
  1. Configure Lambda to scan the entire bucket every minute and identify new JSON objects.
    Periodic full-bucket scans add unnecessary work and do not provide efficient event-driven processing for matching objects.
  2. Use a Kinesis partition key based on the suffix to detect S3 object creation.
    Kinesis partition keys route stream records and do not detect S3 object creation events by themselves.
  3. Route S3 object events to EventBridge and filter the rule by bucket, prefix, and suffix. ✓
    EventBridge filtering targets matching S3 events and supports routing the same event pattern to multiple consumers.
  4. Trigger Lambda for every object event and filter the prefix inside the transcoding code.
    Application filtering invokes the function for irrelevant objects and increases processing without improving event selection.
The trap
This confuses receiving all notifications with filtering at the event-routing layer. This replaces precise event filtering with expensive polling. This treats a stream routing field as an object-storage event trigger.

An EventBridge rule filters S3 events before invoking consumers, while downstream processing remains idempotent for duplicates.

7. Use a Kinesis event source mapping and make reconciliation: Which design is most appropriate?

Medium
A billing team uses Kinesis Data Streams to deliver reconciliation records to Lambda. A transient database failure causes the function to fail after processing part of a batch. The team needs records retried without losing unprocessed records and must prevent duplicate financial updates. Which design is most appropriate?
  1. Increase the Kinesis PutRecords batch size to prevent database failures during consumption.
    Producer batching affects writes to Kinesis and does not resolve downstream database failures during Lambda processing.
  2. Have the producer invoke Lambda directly and delete stream records after invocation.
    Direct invocation bypasses the stream’s managed polling and does not provide reliable retry behavior for failed processing.
  3. Configure Lambda to acknowledge every batch before writing reconciliation results.
    Acknowledging before durable writes can lose records when downstream processing fails after acknowledgment.
  4. Use a Kinesis event source mapping and make reconciliation writes idempotent. ✓
    The event source mapping polls Kinesis for Lambda, while idempotent writes safely handle retries after partial batch failure.
The trap
This confuses producer throughput controls with consumer failure handling. This confuses publishing to Kinesis with invoking a consumer function. This reverses the safe order of processing and acknowledgment.

Lambda event source mapping consumes Kinesis with retry behavior, while idempotent writes protect against duplicate processing.

8. Route the private subnets through a NAT gateway: Which configuration should the data engineer request?

Medium
A Glue job in a VPC connects to a manufacturing database whose firewall accepts only approved public source addresses. The job must reach the database through private subnets, and the network team will permit only one stable address. Which configuration should the data engineer request?
  1. Assign public addresses to each Glue worker and allowlist every worker address.
    Worker-level public addresses are not the intended stable egress control and would violate the requirement for one manageable address.
  2. Allowlist the Glue job security-group identifier in the manufacturing firewall.
    An external database firewall generally receives network source addresses, not AWS security-group identifiers.
  3. Allowlist the private subnet CIDR and route directly to the database without address translation.
    A private subnet CIDR does not provide the required public source path, and the external firewall is specified to accept public addresses.
  4. Route the private subnets through a NAT gateway with an Elastic IP and allowlist that address. ✓
    A NAT gateway provides outbound translation, and its Elastic IP supplies the single stable public source address for the database firewall.
The trap
Confuses an AWS security-group reference with an externally visible address. Assumes private source addresses are directly visible and routable externally. Assumes individual Glue worker addresses are stable and appropriate for allowlisting.

Use a NAT gateway with an Elastic IP to provide stable public egress from private Glue subnets.

9. Use a shared limiter: Which change is best?

Medium
A shared data platform calls a vendor API with a documented limit of 100 requests per minute. Four scheduled jobs currently issue bursts simultaneously, causing throttling and failed loads. The vendor permits retries, and completeness is more important than finishing each burst immediately. Which change is best?
  1. Increase each job's independent concurrency.
    Independent concurrency increases aggregate bursts and makes throttling more likely.
  2. Disable retries and accept failed loads to avoid extra capacity use.
    Disabling retries sacrifices completeness even though temporary throttling can be recovered from.
  3. Use a shared limiter. ✓
    A shared limiter controls aggregate traffic across all jobs and keeps calls within the vendor's limit.
  4. Retry throttled requests immediately until successful.
    Immediate retries intensify bursts and can prolong throttling.
The trap
It assumes per-job speed improvements remain safe when the limit is shared. It mistakes persistence without delay for effective rate-limit recovery. It optimizes request count by violating the explicit completeness requirement.

A shared limiter controls aggregate traffic across the jobs.

10. Use multiple consumer applications so each downstream: Select TWO design choices that satisfy the distribution

Medium
Several partner applications publish order events to a central Kinesis stream. Three independent consumers need the same events: fraud detection, billing, and analytics. Fraud detection requires dedicated read throughput, while each consumer must process the stream independently. Select TWO design choices that satisfy the distribution requirements.

Select two. More than one option is correct — every correct one is ticked below.

  1. Use one consumer to process all events and synchronously call the three downstream systems.
    A single coupled consumer creates shared failure and scaling behavior rather than independent downstream processing.
  2. Batch more records per PutRecords request to provide fraud detection dedicated read capacity.
    Producer batching reduces request overhead but does not create dedicated consumer read throughput.
  3. Assign the same partition key to every partner so all consumers receive each event.
    Partition keys determine shard routing, not which consumer applications receive records; one key can also create a hot shard.
  4. Use multiple consumer applications so each downstream workload reads the stream independently. ✓
    Kinesis supports multiple applications consuming the same stream independently, enabling separate processing and checkpoints.
  5. Register fraud detection as an enhanced fan-out consumer for dedicated read throughput. ✓
    Enhanced fan-out provides a registered consumer dedicated read throughput instead of competing through shared consumer reads.
The trap
This confuses fan-out with sequentially invoking several destinations from one processor. This confuses producer write batching with enhanced consumer fan-out. This mistakes shard placement for consumer distribution.

Independent consumer applications provide fan-out, and enhanced fan-out reserves dedicated read throughput for fraud detection.

11. Use Kinesis retention and consumer checkpoints so: Which design best supports replayability?

Medium
An inventory pipeline receives updates through Kinesis Data Streams. Consumers may be unavailable for several hours, and operations requires replaying records after correcting transformation logic. The business accepts duplicate delivery if downstream processing is idempotent. Which design best supports replayability?
  1. Use PutRecords batching and delete each successful batch after downstream processing completes.
    Batching improves producer throughput, but Kinesis records remain governed by retention and are not deleted by consumers.
  2. Use Kinesis retention and consumer checkpoints so processing can resume or restart from an earlier sequence position. ✓
    Retention preserves records, while checkpoints identify progress and permit consumers to replay records after failures or code corrections.
  3. Use an S3 event notification for each update and replay events by resending notifications when processing fails.
    S3 notifications can be duplicated but do not themselves provide ordered, durable replay positions or complete event-history management.
  4. Use a single partition key for all inventory updates and rely on sequence numbers as permanent record indexes.
    A single key can create a hot shard, and sequence numbers are not permanent indexes for logically separated datasets.
The trap
This confuses request batching with durable replay storage and assumes consumers control record deletion. This mistakes ordering metadata for an independent replay catalog and ignores partition-key capacity concentration. This treats notifications as an event log rather than using durable source records with tracked processing positions.

Kinesis retention and checkpoints preserve records and processing positions, enabling controlled replay after consumer failures or logic corrections.

12. Persist totals and processed event IDs: Which approach is most appropriate?

Medium
A marketing attribution job processes events from several channels. Each event must be enriched with the latest campaign dimension and then accumulated into daily customer totals. A retry must not double-count prior contributions. Exhibit: `event -> lookup campaign -> add amount to customer-day total`. Which approach is most appropriate?
  1. Keep campaign lookups in memory only.
    Memory-only state can disappear across retries or workers and does not protect accumulated totals.
  2. Recompute each event independently without storing prior totals.
    Stateless event processing cannot know whether a retry already contributed to the customer-day total.
  3. Persist totals and processed event IDs, applying each event once. ✓
    Durable totals and event identifiers allow updates to be conditional and retries to avoid double-counting.
  4. Increase the batch size so all customer-day events finish in one invocation.
    A larger batch does not provide durable progress or idempotency after partial failure.
The trap
It confuses temporary enrichment state with durable aggregation state. It confuses deterministic transformation with stateful accumulation. It assumes execution boundaries eliminate retries.

Persist totals and processed event IDs so accumulation is idempotent.

294 more 1: Data Ingestion and Transformation questions

The remaining 294 questions in this domain are part of the full AWS bank — 900 questions, every option explained. Start with the free five-minute check and see your score per domain.

Test your AWS readiness — free

Other AWS domains

Part of the Certsqill AWS question bank · 1: Data Ingestion and Transformation · Every answer, right and wrong, comes with its own explanation.