AWS 2: Data Store Management: 234 practice questions
12 of the 234 2: Data Store Management questions in the Certsqill AWS bank, shown in full below. Each one carries an explanation for every option, not just the correct one — the wrong answers are where the marks go.
Preparing for AWS? Take the free 5-min readiness check →
1. Amazon Athena querying the S3 data lake: Which service is the best primary query store?
- Amazon RDS for PostgreSQL containing every historical transaction row.RDS adds relational administration and storage loading for a workload already suited to querying S3 files.
- Amazon DynamoDB with a partition key for transaction date.DynamoDB suits known-key operational access, not broad ad hoc joins across years of immutable files.
- Amazon Athena querying the S3 data lake. ✓Athena queries Parquet files in S3 without requiring a persistent warehouse for occasional analytical SQL workloads.
- Amazon Kinesis Data Streams retaining the transaction files for analyst queries.Kinesis provides streaming ingestion and retention, not a general-purpose SQL store for historical file analytics.
Athena directly queries the existing Parquet data in S3 for occasional analytical joins.
2. Use DynamoDB with tracking number as the partition key: Which storage configuration is most appropriate?
- Use Amazon Redshift with tracking number as a distribution key.Redshift targets analytical workloads and is unsuitable as the primary low-latency operational store for individual updates.
- Use Amazon S3 objects named with tracking numbers and scan prefixes for status.S3 object access lacks the convenient item update and indexed lookup behavior required by the application.
- Use Amazon Athena over JSON files partitioned by tracking number.Athena is designed for analytical queries and introduces unnecessary query execution overhead for frequent point operations.
- Use DynamoDB with tracking number as the partition key. ✓DynamoDB provides scalable key-based reads and updates when tracking number identifies each parcel item.
Use DynamoDB with tracking number as the partition key for scalable point reads and updates.
3. Amazon MemoryDB for Redis: Which service best fits?
- Amazon Redshift.Redshift is an analytical warehouse and is not designed for low-latency session key retrieval with expiration.
- Amazon Kinesis Data Streams.Kinesis transports ordered records temporarily and does not provide direct key-based session retrieval or expiration semantics.
- Amazon MemoryDB for Redis. ✓MemoryDB provides fast in-memory key-value access and supports expiration for session-oriented application data.
- Amazon Athena querying S3 session files.Athena query execution is inappropriate for frequent low-latency point retrieval of expiring session data.
Amazon MemoryDB provides low-latency key-value access suitable for expiring customer sessions.
4. AWS Transfer Family configured for SFTP delivery into: Which TWO services should be integrated?
Select two. More than one option is correct — every correct one is ticked below.
- AWS Transfer Family configured for SFTP delivery into Amazon S3. ✓Transfer Family provides managed SFTP access for partners and can deliver uploaded files to S3.
- Amazon Kinesis Data Streams configured to expose an SFTP endpoint and replicate tables.Kinesis transports streaming records but does not provide an SFTP endpoint or database migration workflow by itself.
- Amazon Athena configured as the relational database replication engine.Athena queries data in place and does not perform full-load or change-data-capture database replication.
- AWS Database Migration Service for full load and ongoing CDC replication. ✓DMS supports an initial full load followed by change data capture to minimize downtime during database migration.
- Amazon EventBridge scheduled rules configured to copy database transaction logs.EventBridge schedules invoke targets but do not implement relational full-load and CDC replication semantics.
Use DMS for database full load and CDC, and Transfer Family for managed partner SFTP uploads to S3.
5. Use Redshift Spectrum with external tables to query the S3: Which mechanism should the data engineer implement
- Query an RDS database with a Redshift federated query.Federated queries access supported remote databases, not clickstream files stored in S3.
- Refresh a Redshift materialized view from the S3 files.A materialized view stores query results and requires refreshes, so it does not provide direct current access without copying data.
- Load the files into Redshift with COPY.COPY creates a warehouse copy, conflicting with the requirement to keep the files in S3 without copying the full dataset.
- Use Redshift Spectrum with external tables to query the S3 files alongside Redshift tables. ✓Redshift Spectrum queries external S3 data directly and supports joins with Redshift tables without loading the complete dataset.
Use Redshift Spectrum external tables for direct S3 queries and warehouse joins.
6. Begin a transaction: Which approach should the engineer use?
- Read the invoices in a transaction.A transaction alone does not necessarily prevent another worker from changing rows unless the read uses an appropriate lock.
- Read each invoice normally, then run its update in a separate autocommit statement.The initial read does not hold a row lock until the later update, so another worker can change the invoice between those operations.
- Use the primary database for reconciliation.Using the primary selects the writable database but does not by itself coordinate the workers or lock selected rows.
- Begin a transaction, select the unpaid invoices with row-level locks such as FOR UPDATE, perform the updates, and commit only after all reconciliation work completes. ✓A transaction containing a row-locking read and the subsequent updates prevents competing workers from changing the selected rows until commit or rollback.
Use row-level locks inside one transaction covering the read, update, and commit.
7. Create an Apache Iceberg table in the Glue Catalog: Which choice meets these requirements?
- Use Athena partition projection for corrections and historical snapshots.Partition projection computes partition locations and does not implement record transactions or table snapshots.
- Query independent Parquet files through an external table.Parquet supplies a file format but does not itself provide transactional updates, snapshots, or coordinated concurrent writes.
- Create an Apache Iceberg table in the Glue Catalog. ✓Iceberg manages S3 files as a table and provides record operations, snapshots, schema evolution, and concurrent-write protection.
- Register the files in a business catalog without creating an open table, then query them as transactional data.A business catalog improves discovery but does not add transactions, snapshots, or schema evolution to raw files.
Use an Apache Iceberg table backed by the AWS Glue Catalog.
8. Use an IVF index that groups vectors into lists: Which vector index type best matches these requirements?
- Use an IVF index that groups vectors into lists and searches selected lists. ✓IVF reduces search work by grouping vectors into lists and probing selected lists, enabling tunable recall and latency.
- Use an HNSW index that builds a graph of navigable vector connections.HNSW uses graph navigation and can provide strong recall, but its graph structures commonly require more memory.
- Skip indexing because consistent embedding dimensions automatically make similarity searches efficient.Consistent dimensions are necessary for valid comparisons but do not organize vectors or reduce nearest-neighbor search work.
- Use a relational B-tree index on each embedding dimension independently.B-tree indexes on separate dimensions do not provide effective nearest-neighbor search for high-dimensional vector similarity.
IVF partitions vectors into lists and searches selected lists, providing a controllable recall, latency, and memory tradeoff.
9. Query the cataloged table with Athena so queries read: Select TWO actions.
Select two. More than one option is correct — every correct one is ticked below.
- Query the cataloged table with Athena so queries read the original S3 objects. ✓Athena uses catalog metadata to interpret and query the source files without copying them into the catalog.
- Copy every daily file into the Glue Data Catalog to make it queryable.The Glue Data Catalog stores metadata and does not contain copies of the underlying partner files.
- Configure an AWS Glue crawler to inspect the S3 prefix and update table metadata. ✓A Glue crawler reads the source location, discovers schema information, and updates technical metadata in the Glue Data Catalog.
- Use a SageMaker business catalog entry as the physical schema for Athena automatically.Business catalog entries support discovery and governance but do not automatically create Athena technical table metadata.
- Enable partition projection to discover changing CSV columns and infer their schema.Partition projection computes configured partition values and locations; it does not discover changing file schemas.
Use a Glue crawler for source schema discovery and Athena to query the catalog metadata over the original S3 files.
10. Create a Glue table for the S3 files: Which design should be implemented?
- Copy the S3 file contents into the Glue Data Catalog as table data.The Glue Data Catalog stores metadata and locations, not copies of the underlying files.
- Create an Athena view without defining a table for the S3 objects, expecting the view to provide their missing location and schema.An Athena view references queryable tables and cannot replace the required source metadata definition.
- Create a Glue table for the S3 files. ✓A Glue Data Catalog table records the S3 location and technical schema that Athena uses for SQL queries.
- Create only a SageMaker Catalog business asset and expect Athena to infer the physical schema.A business asset supports discovery and governance but does not automatically become an Athena technical table definition.
Create a Glue Data Catalog table pointing to the existing S3 location.
11. Run an AWS Glue crawler over the S3 prefix and configure: Which action best satisfies the requirement?
- Enable Athena partition projection and expect it to infer newly added JSON columns.Partition projection can calculate partition locations but does not inspect files or infer evolving JSON columns.
- Create a SageMaker Catalog glossary term for each JSON field.Glossary terms describe business concepts and do not discover physical JSON schemas or register Glue partitions.
- Create a materialized view before files arrive so Glue can infer their future schema.Materialized views store query results and cannot inspect future files or populate Glue table definitions.
- Run an AWS Glue crawler over the S3 prefix and configure it to update the catalog. ✓A crawler inspects source files, infers schema information, and can populate or update catalog tables and partitions.
A Glue crawler inspects evolving JSON files and populates technical schema and partition metadata for Athena.
12. Enable Athena partition projection with date ranges: Which configuration should be used?
The partition pattern is `year/month/day`, future folders are predictable, and all access is through Athena.
- Enable Athena partition projection with date ranges and a matching S3 path template. ✓Partition projection lets Athena calculate predictable date partitions and locations during queries, avoiding manual partition registration.
- Create a business catalog asset for every daily folder and query those assets with Athena.Business catalog assets support discovery and governance but do not provide Athena partition metadata or replace table configuration.
- Run a crawler after every folder arrival and manually inspect each discovered partition.A crawler can automatically discover and register partitions, but repeated crawler runs and manual inspection add unnecessary operational work when the partition pattern is predictable.
- Use an Iceberg table property to make ordinary Hive-style folders self-register in Athena.Iceberg metadata and Athena partition projection are different mechanisms; an Iceberg property does not manage partitions for an ordinary Hive-style table.
Use Athena partition projection for predictable date-based S3 paths.
222 more 2: Data Store Management questions
The remaining 222 questions in this domain are part of the full AWS bank — 900 questions, every option explained. Start with the free five-minute check and see your score per domain.
Test your AWS readiness — freeOther AWS domains
- 1: Data Ingestion and Transformation — 306 questions →
- 3: Data Operations and Support — 198 questions →
- 4: Data Security and Governance — 162 questions →
- All 900 AWS questions →
- AWS certification: requirements, cost and exam format →