AWS: 900 practice questions with explanations
48 hours only — 15% off every course with code SAVE15. Browse courses →48h · 15% off all courses · code SAVE15 →
Certifications Tools Flashcards Career Paths Exam Guides Blog Pricing For Teams About

AWS practice questions: 900 questions with full explanations

4 domains 900 questions 130 min exam
Questions on the exam
about 65 — vendor indicates, no fixed count published
Time allowed
130 minutes format →

900 practice questions for AWS Certified Data Engineer – Associate, grouped by exam domain. Every question below shows all four options, which one is correct, and why each of the other three is not — the wrong answers are where most candidates lose marks.

Not sure where you stand? Take the free 5-min AWS readiness check →

AWS certification: requirements, cost and exam format → ·  AWS exam format →  · 

Questions by domain

Sample questions

Split each device across subkeys and resequence: Which change best addresses the problem?

1: Data Ingestion and Transformation Easy
A fleet publishes telemetry to a provisioned Kinesis data stream. One device produces 1.2 MB/s, exceeding the 1 MB/s write limit of a shard. Consumers require ordered records per device, and global ordering is unnecessary. The application can assign monotonically increasing device sequence values and buffer records for resequencing. Which change best addresses the problem?
Exhibit: partition key = deviceId; shard capacity = 1 MB/s writes.
  1. Add consumers that read the hot shard before accepting more device records.
    Consumer coordination does not redistribute writes or increase the hot shard's capacity.
  2. Split each device across subkeys and resequence by its device sequence value. ✓
    Subkeys distribute one hot device across shards, while downstream resequencing preserves device order without requiring global ordering.
  3. Use sequence numbers as partition keys for all records.
    Kinesis assigns sequence numbers after ingestion, so producers cannot use them to route incoming records.
  4. Increase PutRecords batch size until the hot shard accepts the device.
    PutRecords improves request efficiency but does not bypass a shard's write-throughput limit.
The trap
Confuses service-assigned metadata with producer-controlled partitioning. Confuses batching with additional shard capacity. Treats a consumer-side change as producer-side load balancing.

All 306 1: Data Ingestion and Transformation questions →

Amazon Athena querying the S3 data lake: Which service is the best primary query store?

2: Data Store Management Medium
A finance team stores immutable transaction files in Amazon S3. Analysts run occasional SQL joins across years of data, while a dashboard needs a refreshed daily aggregate. Exhibit: files are Parquet; queries are ad hoc; no row-level updates are required. Which service is the best primary query store?
  1. Amazon RDS for PostgreSQL containing every historical transaction row.
    RDS adds relational administration and storage loading for a workload already suited to querying S3 files.
  2. Amazon DynamoDB with a partition key for transaction date.
    DynamoDB suits known-key operational access, not broad ad hoc joins across years of immutable files.
  3. Amazon Athena querying the S3 data lake. ✓
    Athena queries Parquet files in S3 without requiring a persistent warehouse for occasional analytical SQL workloads.
  4. Amazon Kinesis Data Streams retaining the transaction files for analyst queries.
    Kinesis provides streaming ingestion and retention, not a general-purpose SQL store for historical file analytics.
The trap
Assumes transactional relational storage is automatically best for analytics. Confuses streaming transport with durable analytical storage. Chooses a key-value store for scan-heavy analytical queries.

All 234 2: Data Store Management questions →

Use Step Functions: Which design best meets the requirement?

3: Data Operations and Support Easy
A logistics pipeline must run after each delivery file arrives, then load curated data only after validation succeeds. The team wants managed orchestration and explicit failure handling. Exhibit: Event: object-created; validation: failed; load: started. Which design best meets the requirement?
  1. Run a crawler before a scheduled ETL job.
    Crawler completion and ETL completion are separate operations, and scheduling does not guarantee the required validation dependency.
  2. Schedule validation and loading together.
    Simultaneous scheduled starts do not ensure validation succeeds before loading begins.
  3. Trigger the load directly from S3 and validate afterward.
    Direct loading bypasses the required validation gate and can publish invalid data.
  4. Use Step Functions. ✓
    A Step Functions workflow can start from the object event, validate the file, branch on the result, and prevent loading after failure.
The trap
Confuses catalog discovery with orchestration. Treats scheduling as dependency management. Assumes object arrival proves data validity.

All 198 3: Data Operations and Support questions →

Allow TCP 5432 inbound on DbSG from AppSG: Which update should restore access?

4: Data Security and Governance Hard
A support integration runs in application security group AppSG and connects to a database in DbSG. Exhibit: DbSG currently allows inbound TCP 5432 from 10.0.4.0/24. App instances moved to another private subnet, and the connection now times out. Network policy prohibits broad CIDR access. Which update should restore access?
  1. Allow TCP 5432 inbound on DbSG from 0.0.0.0/0.
    This permits database access from every IPv4 source and violates the stated restriction against broad access.
  2. Allow TCP 5432 inbound on AppSG from DbSG.
    Inbound rules on the client group do not authorize the database’s inbound connection request.
  3. Attach the database security group to the application instances.
    Sharing a security group does not replace the required inbound rule on the database security group.
  4. Allow TCP 5432 inbound on DbSG from AppSG. ✓
    Referencing the application security group permits database traffic from current application members without broad subnet-based access.
The trap
Confuses connectivity restoration with least-privilege network authorization. Confuses group association with traffic authorization. Reverses the direction of the required security-group rule.

All 162 4: Data Security and Governance questions →

Configure an AWS Glue job to read the S3 objects: Which option best fits?

1: Data Ingestion and Transformation Easy
A financial institution receives one compressed CSV file per hour in Amazon S3. Each file contains transactions for a complete hour, and processing may begin after the file arrives. The team wants a managed batch ingestion mechanism that can read the files, transform them, and write Parquet output for downstream analytics. Which option best fits?
  1. Create a Kinesis data stream and publish each CSV line as a separate streaming record.
    Kinesis could transport records, but it adds unnecessary streaming complexity to complete hourly batch files.
  2. Use an Amazon MSK consumer group to poll the S3 bucket for completed files.
    MSK consumers process Kafka records; they do not natively poll S3 as a batch-file ingestion mechanism.
  3. Configure DynamoDB Streams to discover newly created S3 objects and process their contents.
    DynamoDB Streams reports DynamoDB item changes and does not provide notifications for S3 object creation.
  4. Configure an AWS Glue job to read the S3 objects, transform rows, and write Parquet. ✓
    AWS Glue jobs provide managed batch processing for S3 files and can transform records into analytics-friendly output formats.
The trap
This treats a Kafka consumer as a generic managed file reader. This assumes DynamoDB Streams is a general-purpose object-storage event source. This confuses a periodic file workload with continuously arriving streaming data.

All 306 1: Data Ingestion and Transformation questions →

Use DynamoDB with tracking number as the partition key: Which storage configuration is most appropriate?

2: Data Store Management Medium
A delivery application frequently retrieves a parcel by tracking number and occasionally updates its status. It does not perform joins, and access must remain low-latency as the dataset grows. Which storage configuration is most appropriate?
  1. Use Amazon Redshift with tracking number as a distribution key.
    Redshift targets analytical workloads and is unsuitable as the primary low-latency operational store for individual updates.
  2. Use Amazon S3 objects named with tracking numbers and scan prefixes for status.
    S3 object access lacks the convenient item update and indexed lookup behavior required by the application.
  3. Use Amazon Athena over JSON files partitioned by tracking number.
    Athena is designed for analytical queries and introduces unnecessary query execution overhead for frequent point operations.
  4. Use DynamoDB with tracking number as the partition key. ✓
    DynamoDB provides scalable key-based reads and updates when tracking number identifies each parcel item.
The trap
Uses an analytical warehouse for transactional key-value access. Assumes object names provide database-style mutable record access. Confuses serverless SQL analytics with operational point reads.

All 234 2: Data Store Management questions →

Inspect DAG dependencies and scheduler logs: Which investigation is most appropriate first?

3: Data Operations and Support Easy
A customer-support integration runs in Amazon MWAA. A DAG suddenly fails before its first task, while other DAGs continue running. The scheduler is healthy, and the failed task has no application log entries. Which investigation is most appropriate first?
  1. Inspect DAG dependencies and scheduler logs. ✓
    A pre-task failure commonly indicates dependency, parsing, or scheduling issues, so DAG state and scheduler evidence should be checked first.
  2. Increase retries without reviewing logs or dependencies.
    Retries do not correct invalid DAG dependencies or parsing problems and can obscure the original cause.
  3. Reset the integration's Glue bookmarks.
    Glue bookmarks track source-processing state and do not diagnose an MWAA DAG that fails before task execution.
  4. Restart downstream applications.
    Downstream restarts do not explain a DAG failure before task execution and may disrupt healthy systems.
The trap
Applies an ingestion-state remedy to an orchestration failure. Treats downstream symptoms instead of isolating the workflow boundary. Assumes the failure is transient.

All 198 3: Data Operations and Support questions →

Store it in Secrets Manager with rotation enabled: Which solution best satisfies these requirements?

4: Data Security and Governance Easy
A media analytics pipeline connects to a database nightly. The database password must rotate automatically, application code cannot contain static credentials, and brief connection retries during rotation are acceptable. Which solution best satisfies these requirements?
  1. Create an IAM user and embed its access keys in the pipeline configuration.
    IAM access keys are static AWS credentials and do not manage the database password.
  2. Place the password in an S3 object and grant the pipeline read access.
    S3 stores the value but does not by itself rotate database credentials or coordinate refreshed connections.
  3. Store it in Secrets Manager with rotation enabled. ✓
    Secrets Manager stores the password centrally, supports configured rotation, and allows runtime retrieval of the current value.
  4. Store the password in an encrypted configuration file deployed with the pipeline.
    Encryption protects the file but does not provide managed rotation or eliminate the need to distribute decryption material.
The trap
Confuses encryption at rest with secret lifecycle management. Assumes protected storage performs credential rotation. Confuses AWS API credentials with database credentials.

All 162 4: Data Security and Governance questions →

AWS exam: the facts

How many questions are on the AWS exam?

Around 65. The vendor does not publish a fixed count for AWS, so this is the figure it indicates rather than a guaranteed number.

How long is the AWS exam?

130 minutes. Across 65 questions that is about 120 seconds per question.

What topics does the AWS exam cover?

4 domains: 1: Data Ingestion and Transformation, 2: Data Store Management, 3: Data Operations and Support, 4: Data Security and Governance. Weights: 1: Data Ingestion and Transformation 34%, 2: Data Store Management 26%, 3: Data Operations and Support 22%, 4: Data Security and Governance 18%.

How many AWS practice questions does Certsqill have?

900, spread across 4 exam domains. Every one shows all options, which is correct, and why each of the others is not.

Would you pass AWS today?

Five minutes, and you get a score per domain — not one number, but which section to open tonight.

Test your AWS readiness — free
Certsqill AWS question bank · 900 questions across 4 domains · Every answer, right and wrong, comes with its own explanation.