AWS 1: Data Ingestion and Transformation: 306 practice questions
12 of the 306 1: Data Ingestion and Transformation questions in the Certsqill AWS bank, shown in full below. Each one carries an explanation for every option, not just the correct one — the wrong answers are where the marks go.
Preparing for AWS? Take the free 5-min readiness check →
1. Split each device across subkeys and resequence: Which change best addresses the problem?
Exhibit: partition key = deviceId; shard capacity = 1 MB/s writes.
- Add consumers that read the hot shard before accepting more device records.Consumer coordination does not redistribute writes or increase the hot shard's capacity.
- Split each device across subkeys and resequence by its device sequence value. ✓Subkeys distribute one hot device across shards, while downstream resequencing preserves device order without requiring global ordering.
- Use sequence numbers as partition keys for all records.Kinesis assigns sequence numbers after ingestion, so producers cannot use them to route incoming records.
- Increase PutRecords batch size until the hot shard accepts the device.PutRecords improves request efficiency but does not bypass a shard's write-throughput limit.
Split the hot device across subkeys and resequence its records downstream using device sequence values.
2. Configure an AWS Glue job to read the S3 objects: Which option best fits?
- Create a Kinesis data stream and publish each CSV line as a separate streaming record.Kinesis could transport records, but it adds unnecessary streaming complexity to complete hourly batch files.
- Use an Amazon MSK consumer group to poll the S3 bucket for completed files.MSK consumers process Kafka records; they do not natively poll S3 as a batch-file ingestion mechanism.
- Configure DynamoDB Streams to discover newly created S3 objects and process their contents.DynamoDB Streams reports DynamoDB item changes and does not provide notifications for S3 object creation.
- Configure an AWS Glue job to read the S3 objects, transform rows, and write Parquet. ✓AWS Glue jobs provide managed batch processing for S3 files and can transform records into analytics-friendly output formats.
An AWS Glue job directly processes hourly S3 files as batch input and writes transformed Parquet results.
3. Enable job bookmarks and use a stable landing prefix: Which configuration is most appropriate?
- Delete the output prefix before every run and disable input tracking.Deleting outputs does not identify previously processed inputs and can remove valid historical delivery results.
- Store the last successful run timestamp only in the output object name.Output naming alone does not make Glue aware of which source objects were successfully processed.
- Enable job bookmarks and use a stable landing prefix for incremental input tracking. ✓Glue job bookmarks track previously processed input and allow newly changed or added source objects to be considered.
- Use a fixed filename filter and overwrite every transformed object on each run.Fixed filenames and repeated overwrites can miss corrected objects or cause unnecessary reprocessing of unchanged input.
Glue job bookmarks track processed input objects, while output management remains a separate responsibility.
4. Persist tokens: Which approach is best?
- Persist tokens, back off on throttling, and deduplicate ticket IDs. ✓Persisted tokens support resumption, backoff handles throttling, and ticket IDs make repeated pages idempotent.
- Use parallel requests without coordination to finish quickly.Uncoordinated concurrency can exceed vendor limits and create overlapping pages.
- Request maximum pages and restart from the beginning after failures.Restarting from the beginning increases duplication and does not provide reliable checkpoint recovery.
- Use response timestamps as cursors and ignore repeated records.Response timestamps are not the documented pagination cursor, and ignoring duplicates can corrupt data.
Persist the vendor token, back off on throttling, and deduplicate ticket IDs.
5. Use Apache Airflow in Amazon MWAA to define dependencies: Select TWO actions that satisfy these requirements.
Select two. More than one option is correct — every correct one is ticked below.
- Place a time-based schedule on every workflow step independently.Independent schedules can overlap or run out of order, so they cannot reliably enforce step dependencies.
- Use an S3 notification for each job and rely on object creation timing to enforce workflow order.S3 notifications signal object events but do not inherently model ordered workflow dependencies or task retries.
- Create a Kinesis stream and use record arrival to schedule each nightly transformation.Kinesis transports streaming records and does not provide the requested time-based batch workflow orchestration.
- Use Apache Airflow in Amazon MWAA to define dependencies and retries for the three-step workflow. ✓Airflow DAGs express task dependencies and retry behavior for multi-step workflows managed through Amazon MWAA.
- Use EventBridge schedules to start each independent nightly Glue job. ✓EventBridge provides managed time-based rules suitable for starting independent scheduled Glue jobs.
EventBridge schedules independent jobs, while Airflow manages ordered multi-step workflows with retries.
6. Route S3 object events to EventBridge and filter the rule: Which configuration best meets the requirements?
- Configure Lambda to scan the entire bucket every minute and identify new JSON objects.Periodic full-bucket scans add unnecessary work and do not provide efficient event-driven processing for matching objects.
- Use a Kinesis partition key based on the suffix to detect S3 object creation.Kinesis partition keys route stream records and do not detect S3 object creation events by themselves.
- Route S3 object events to EventBridge and filter the rule by bucket, prefix, and suffix. ✓EventBridge filtering targets matching S3 events and supports routing the same event pattern to multiple consumers.
- Trigger Lambda for every object event and filter the prefix inside the transcoding code.Application filtering invokes the function for irrelevant objects and increases processing without improving event selection.
An EventBridge rule filters S3 events before invoking consumers, while downstream processing remains idempotent for duplicates.
7. Use a Kinesis event source mapping and make reconciliation: Which design is most appropriate?
- Increase the Kinesis PutRecords batch size to prevent database failures during consumption.Producer batching affects writes to Kinesis and does not resolve downstream database failures during Lambda processing.
- Have the producer invoke Lambda directly and delete stream records after invocation.Direct invocation bypasses the stream’s managed polling and does not provide reliable retry behavior for failed processing.
- Configure Lambda to acknowledge every batch before writing reconciliation results.Acknowledging before durable writes can lose records when downstream processing fails after acknowledgment.
- Use a Kinesis event source mapping and make reconciliation writes idempotent. ✓The event source mapping polls Kinesis for Lambda, while idempotent writes safely handle retries after partial batch failure.
Lambda event source mapping consumes Kinesis with retry behavior, while idempotent writes protect against duplicate processing.
8. Route the private subnets through a NAT gateway: Which configuration should the data engineer request?
- Assign public addresses to each Glue worker and allowlist every worker address.Worker-level public addresses are not the intended stable egress control and would violate the requirement for one manageable address.
- Allowlist the Glue job security-group identifier in the manufacturing firewall.An external database firewall generally receives network source addresses, not AWS security-group identifiers.
- Allowlist the private subnet CIDR and route directly to the database without address translation.A private subnet CIDR does not provide the required public source path, and the external firewall is specified to accept public addresses.
- Route the private subnets through a NAT gateway with an Elastic IP and allowlist that address. ✓A NAT gateway provides outbound translation, and its Elastic IP supplies the single stable public source address for the database firewall.
Use a NAT gateway with an Elastic IP to provide stable public egress from private Glue subnets.
9. Use a shared limiter: Which change is best?
- Increase each job's independent concurrency.Independent concurrency increases aggregate bursts and makes throttling more likely.
- Disable retries and accept failed loads to avoid extra capacity use.Disabling retries sacrifices completeness even though temporary throttling can be recovered from.
- Use a shared limiter. ✓A shared limiter controls aggregate traffic across all jobs and keeps calls within the vendor's limit.
- Retry throttled requests immediately until successful.Immediate retries intensify bursts and can prolong throttling.
A shared limiter controls aggregate traffic across the jobs.
10. Use multiple consumer applications so each downstream: Select TWO design choices that satisfy the distribution
Select two. More than one option is correct — every correct one is ticked below.
- Use one consumer to process all events and synchronously call the three downstream systems.A single coupled consumer creates shared failure and scaling behavior rather than independent downstream processing.
- Batch more records per PutRecords request to provide fraud detection dedicated read capacity.Producer batching reduces request overhead but does not create dedicated consumer read throughput.
- Assign the same partition key to every partner so all consumers receive each event.Partition keys determine shard routing, not which consumer applications receive records; one key can also create a hot shard.
- Use multiple consumer applications so each downstream workload reads the stream independently. ✓Kinesis supports multiple applications consuming the same stream independently, enabling separate processing and checkpoints.
- Register fraud detection as an enhanced fan-out consumer for dedicated read throughput. ✓Enhanced fan-out provides a registered consumer dedicated read throughput instead of competing through shared consumer reads.
Independent consumer applications provide fan-out, and enhanced fan-out reserves dedicated read throughput for fraud detection.
11. Use Kinesis retention and consumer checkpoints so: Which design best supports replayability?
- Use PutRecords batching and delete each successful batch after downstream processing completes.Batching improves producer throughput, but Kinesis records remain governed by retention and are not deleted by consumers.
- Use Kinesis retention and consumer checkpoints so processing can resume or restart from an earlier sequence position. ✓Retention preserves records, while checkpoints identify progress and permit consumers to replay records after failures or code corrections.
- Use an S3 event notification for each update and replay events by resending notifications when processing fails.S3 notifications can be duplicated but do not themselves provide ordered, durable replay positions or complete event-history management.
- Use a single partition key for all inventory updates and rely on sequence numbers as permanent record indexes.A single key can create a hot shard, and sequence numbers are not permanent indexes for logically separated datasets.
Kinesis retention and checkpoints preserve records and processing positions, enabling controlled replay after consumer failures or logic corrections.
12. Persist totals and processed event IDs: Which approach is most appropriate?
- Keep campaign lookups in memory only.Memory-only state can disappear across retries or workers and does not protect accumulated totals.
- Recompute each event independently without storing prior totals.Stateless event processing cannot know whether a retry already contributed to the customer-day total.
- Persist totals and processed event IDs, applying each event once. ✓Durable totals and event identifiers allow updates to be conditional and retries to avoid double-counting.
- Increase the batch size so all customer-day events finish in one invocation.A larger batch does not provide durable progress or idempotency after partial failure.
Persist totals and processed event IDs so accumulation is idempotent.
294 more 1: Data Ingestion and Transformation questions
The remaining 294 questions in this domain are part of the full AWS bank — 900 questions, every option explained. Start with the free five-minute check and see your score per domain.
Test your AWS readiness — freeOther AWS domains
- 2: Data Store Management — 234 questions →
- 3: Data Operations and Support — 198 questions →
- 4: Data Security and Governance — 162 questions →
- All 900 AWS questions →
- AWS certification: requirements, cost and exam format →