AWS 2: Data Store Management: 234 practice questions
48 hours only — 15% off every course with code SAVE15. Browse courses →48h · 15% off all courses · code SAVE15 →
Certifications Tools Flashcards Career Paths Exam Guides Blog Pricing For Teams About

AWS 2: Data Store Management: 234 practice questions

AWS 234 questions 12 shown free

12 of the 234 2: Data Store Management questions in the Certsqill AWS bank, shown in full below. Each one carries an explanation for every option, not just the correct one — the wrong answers are where the marks go.

Preparing for AWS? Take the free 5-min readiness check →

1. Amazon Athena querying the S3 data lake: Which service is the best primary query store?

Medium
A finance team stores immutable transaction files in Amazon S3. Analysts run occasional SQL joins across years of data, while a dashboard needs a refreshed daily aggregate. Exhibit: files are Parquet; queries are ad hoc; no row-level updates are required. Which service is the best primary query store?
  1. Amazon RDS for PostgreSQL containing every historical transaction row.
    RDS adds relational administration and storage loading for a workload already suited to querying S3 files.
  2. Amazon DynamoDB with a partition key for transaction date.
    DynamoDB suits known-key operational access, not broad ad hoc joins across years of immutable files.
  3. Amazon Athena querying the S3 data lake. ✓
    Athena queries Parquet files in S3 without requiring a persistent warehouse for occasional analytical SQL workloads.
  4. Amazon Kinesis Data Streams retaining the transaction files for analyst queries.
    Kinesis provides streaming ingestion and retention, not a general-purpose SQL store for historical file analytics.
The trap
Assumes transactional relational storage is automatically best for analytics. Confuses streaming transport with durable analytical storage. Chooses a key-value store for scan-heavy analytical queries.

Athena directly queries the existing Parquet data in S3 for occasional analytical joins.

2. Use DynamoDB with tracking number as the partition key: Which storage configuration is most appropriate?

Medium
A delivery application frequently retrieves a parcel by tracking number and occasionally updates its status. It does not perform joins, and access must remain low-latency as the dataset grows. Which storage configuration is most appropriate?
  1. Use Amazon Redshift with tracking number as a distribution key.
    Redshift targets analytical workloads and is unsuitable as the primary low-latency operational store for individual updates.
  2. Use Amazon S3 objects named with tracking numbers and scan prefixes for status.
    S3 object access lacks the convenient item update and indexed lookup behavior required by the application.
  3. Use Amazon Athena over JSON files partitioned by tracking number.
    Athena is designed for analytical queries and introduces unnecessary query execution overhead for frequent point operations.
  4. Use DynamoDB with tracking number as the partition key. ✓
    DynamoDB provides scalable key-based reads and updates when tracking number identifies each parcel item.
The trap
Uses an analytical warehouse for transactional key-value access. Assumes object names provide database-style mutable record access. Confuses serverless SQL analytics with operational point reads.

Use DynamoDB with tracking number as the partition key for scalable point reads and updates.

3. Amazon MemoryDB for Redis: Which service best fits?

Medium
A customer-support application must retrieve frequently accessed session data by customer ID with very low latency. Sessions can expire automatically, and the data does not require relational joins or durable reporting. Which service best fits?
  1. Amazon Redshift.
    Redshift is an analytical warehouse and is not designed for low-latency session key retrieval with expiration.
  2. Amazon Kinesis Data Streams.
    Kinesis transports ordered records temporarily and does not provide direct key-based session retrieval or expiration semantics.
  3. Amazon MemoryDB for Redis. ✓
    MemoryDB provides fast in-memory key-value access and supports expiration for session-oriented application data.
  4. Amazon Athena querying S3 session files.
    Athena query execution is inappropriate for frequent low-latency point retrieval of expiring session data.
The trap
Selects a warehouse for an application cache-like workload. Treats a streaming service as an application key-value database. Confuses analytical file queries with interactive session access.

Amazon MemoryDB provides low-latency key-value access suitable for expiring customer sessions.

4. AWS Transfer Family configured for SFTP delivery into: Which TWO services should be integrated?

Medium
A healthcare organization must migrate an on-premises relational database to AWS with minimal downtime. Historical rows must be copied first, then ongoing changes replicated until cutover. Separately, an approved partner must upload files over SFTP into Amazon S3. Which TWO services should be integrated? Select TWO.

Select two. More than one option is correct — every correct one is ticked below.

  1. AWS Transfer Family configured for SFTP delivery into Amazon S3. ✓
    Transfer Family provides managed SFTP access for partners and can deliver uploaded files to S3.
  2. Amazon Kinesis Data Streams configured to expose an SFTP endpoint and replicate tables.
    Kinesis transports streaming records but does not provide an SFTP endpoint or database migration workflow by itself.
  3. Amazon Athena configured as the relational database replication engine.
    Athena queries data in place and does not perform full-load or change-data-capture database replication.
  4. AWS Database Migration Service for full load and ongoing CDC replication. ✓
    DMS supports an initial full load followed by change data capture to minimize downtime during database migration.
  5. Amazon EventBridge scheduled rules configured to copy database transaction logs.
    EventBridge schedules invoke targets but do not implement relational full-load and CDC replication semantics.
The trap
Combines unrelated streaming and file-transfer assumptions. Confuses an analytical query service with a migration engine. Mistakes orchestration triggers for database migration functionality.

Use DMS for database full load and CDC, and Transfer Family for managed partner SFTP uploads to S3.

5. Use Redshift Spectrum with external tables to query the S3: Which mechanism should the data engineer implement

Medium
A media company stores clickstream files in Amazon S3 and needs Redshift analysts to join them with warehouse tables each day. The files remain in S3, analysts need current data, and copying the full dataset into Redshift is undesirable. Which mechanism should the data engineer implement?
  1. Query an RDS database with a Redshift federated query.
    Federated queries access supported remote databases, not clickstream files stored in S3.
  2. Refresh a Redshift materialized view from the S3 files.
    A materialized view stores query results and requires refreshes, so it does not provide direct current access without copying data.
  3. Load the files into Redshift with COPY.
    COPY creates a warehouse copy, conflicting with the requirement to keep the files in S3 without copying the full dataset.
  4. Use Redshift Spectrum with external tables to query the S3 files alongside Redshift tables. ✓
    Redshift Spectrum queries external S3 data directly and supports joins with Redshift tables without loading the complete dataset.
The trap
Uses a valid ingestion method that violates the storage constraint. Selects a real remote-query feature for the wrong source type. Confuses cached results with direct external access.

Use Redshift Spectrum external tables for direct S3 queries and warehouse joins.

6. Begin a transaction: Which approach should the engineer use?

Medium
During subscription billing reconciliation, two workers may update the same invoice rows. Each worker must read selected unpaid invoices, modify them, and prevent another worker from changing those rows until the transaction completes. The database is Amazon RDS for PostgreSQL. Which approach should the engineer use?
  1. Read the invoices in a transaction.
    A transaction alone does not necessarily prevent another worker from changing rows unless the read uses an appropriate lock.
  2. Read each invoice normally, then run its update in a separate autocommit statement.
    The initial read does not hold a row lock until the later update, so another worker can change the invoice between those operations.
  3. Use the primary database for reconciliation.
    Using the primary selects the writable database but does not by itself coordinate the workers or lock selected rows.
  4. Begin a transaction, select the unpaid invoices with row-level locks such as FOR UPDATE, perform the updates, and commit only after all reconciliation work completes. ✓
    A transaction containing a row-locking read and the subsequent updates prevents competing workers from changing the selected rows until commit or rollback.
The trap
Confuses write availability with concurrency control. Omits the required row-locking operation. Treats separate statements as one protected read-modify-write transaction.

Use row-level locks inside one transaction covering the read, update, and commit.

7. Create an Apache Iceberg table in the Glue Catalog: Which choice meets these requirements?

Medium
A manufacturer stores quality measurements in Amazon S3. Analysts need record-level corrections, schema evolution, and time-travel queries while multiple ingestion jobs write concurrently. The design should use an open table format rather than treating individual Parquet files as an independent dataset. Which choice meets these requirements?
  1. Use Athena partition projection for corrections and historical snapshots.
    Partition projection computes partition locations and does not implement record transactions or table snapshots.
  2. Query independent Parquet files through an external table.
    Parquet supplies a file format but does not itself provide transactional updates, snapshots, or coordinated concurrent writes.
  3. Create an Apache Iceberg table in the Glue Catalog. ✓
    Iceberg manages S3 files as a table and provides record operations, snapshots, schema evolution, and concurrent-write protection.
  4. Register the files in a business catalog without creating an open table, then query them as transactional data.
    A business catalog improves discovery but does not add transactions, snapshots, or schema evolution to raw files.
The trap
Confuses a file format with an open table format. Mistakes partition metadata optimization for table management. Treats business metadata as storage semantics.

Use an Apache Iceberg table backed by the AWS Glue Catalog.

8. Use an IVF index that groups vectors into lists: Which vector index type best matches these requirements?

Medium
An enterprise platform must search millions of fixed-dimension embeddings. The workload is read-heavy, permits approximate nearest-neighbor results, and has limited memory. The team can tune how many candidate groups are searched to trade recall for latency. Which vector index type best matches these requirements?
  1. Use an IVF index that groups vectors into lists and searches selected lists. ✓
    IVF reduces search work by grouping vectors into lists and probing selected lists, enabling tunable recall and latency.
  2. Use an HNSW index that builds a graph of navigable vector connections.
    HNSW uses graph navigation and can provide strong recall, but its graph structures commonly require more memory.
  3. Skip indexing because consistent embedding dimensions automatically make similarity searches efficient.
    Consistent dimensions are necessary for valid comparisons but do not organize vectors or reduce nearest-neighbor search work.
  4. Use a relational B-tree index on each embedding dimension independently.
    B-tree indexes on separate dimensions do not provide effective nearest-neighbor search for high-dimensional vector similarity.
The trap
This overlooks the scenario's limited-memory and list-probing tradeoff. This applies scalar indexing to a multidimensional similarity problem. This confuses compatible vector shape with an actual search acceleration structure.

IVF partitions vectors into lists and searches selected lists, providing a controllable recall, latency, and memory tradeoff.

9. Query the cataloged table with Athena so queries read: Select TWO actions.

Medium
A partner delivers daily CSV files to a known Amazon S3 prefix. Analysts must consume newly arrived files using discoverable table metadata, while the files must remain in their original location. The schema can change occasionally, and the solution should avoid manually defining every new file. Select TWO actions.

Select two. More than one option is correct — every correct one is ticked below.

  1. Query the cataloged table with Athena so queries read the original S3 objects. ✓
    Athena uses catalog metadata to interpret and query the source files without copying them into the catalog.
  2. Copy every daily file into the Glue Data Catalog to make it queryable.
    The Glue Data Catalog stores metadata and does not contain copies of the underlying partner files.
  3. Configure an AWS Glue crawler to inspect the S3 prefix and update table metadata. ✓
    A Glue crawler reads the source location, discovers schema information, and updates technical metadata in the Glue Data Catalog.
  4. Use a SageMaker business catalog entry as the physical schema for Athena automatically.
    Business catalog entries support discovery and governance but do not automatically create Athena technical table metadata.
  5. Enable partition projection to discover changing CSV columns and infer their schema.
    Partition projection computes configured partition values and locations; it does not discover changing file schemas.
The trap
This confuses partition-location generation with schema discovery. This mistakes a metadata catalog for a data repository. This confuses business discovery metadata with the technical catalog consumed by Athena.

Use a Glue crawler for source schema discovery and Athena to query the catalog metadata over the original S3 files.

10. Create a Glue table for the S3 files: Which design should be implemented?

Medium
An inventory team needs a technical description of curated S3 files that Athena can reference repeatedly. The data engineer must record the location, columns, and file format without moving the files, and analysts should query the resulting table through SQL. Which design should be implemented?
  1. Copy the S3 file contents into the Glue Data Catalog as table data.
    The Glue Data Catalog stores metadata and locations, not copies of the underlying files.
  2. Create an Athena view without defining a table for the S3 objects, expecting the view to provide their missing location and schema.
    An Athena view references queryable tables and cannot replace the required source metadata definition.
  3. Create a Glue table for the S3 files. ✓
    A Glue Data Catalog table records the S3 location and technical schema that Athena uses for SQL queries.
  4. Create only a SageMaker Catalog business asset and expect Athena to infer the physical schema.
    A business asset supports discovery and governance but does not automatically become an Athena technical table definition.
The trap
Treats business catalog metadata as a technical schema. Confuses metadata storage with data storage. Assumes a view can supply missing source metadata.

Create a Glue Data Catalog table pointing to the existing S3 location.

11. Run an AWS Glue crawler over the S3 prefix and configure: Which action best satisfies the requirement?

Hard
Marketing attribution files arrive under one S3 prefix in JSON format. New files may add fields, and date-based folders appear regularly. The team needs Glue to discover the evolving schema and register tables and partitions for downstream Athena queries. Which action best satisfies the requirement?
  1. Enable Athena partition projection and expect it to infer newly added JSON columns.
    Partition projection can calculate partition locations but does not inspect files or infer evolving JSON columns.
  2. Create a SageMaker Catalog glossary term for each JSON field.
    Glossary terms describe business concepts and do not discover physical JSON schemas or register Glue partitions.
  3. Create a materialized view before files arrive so Glue can infer their future schema.
    Materialized views store query results and cannot inspect future files or populate Glue table definitions.
  4. Run an AWS Glue crawler over the S3 prefix and configure it to update the catalog. ✓
    A crawler inspects source files, infers schema information, and can populate or update catalog tables and partitions.
The trap
This treats a stored query result as a schema-discovery mechanism. This confuses partition management with schema discovery. This substitutes business vocabulary for technical metadata discovery.

A Glue crawler inspects evolving JSON files and populates technical schema and partition metadata for Athena.

12. Enable Athena partition projection with date ranges: Which configuration should be used?

Hard
A retail table stores orders under predictable S3 paths such as `orders/year=2026/month=09/day=20/`. New daily folders are added automatically. Athena queries always filter by date, and only Athena needs the table. The team wants to avoid registering each new partition manually. Which configuration should be used?

The partition pattern is `year/month/day`, future folders are predictable, and all access is through Athena.
  1. Enable Athena partition projection with date ranges and a matching S3 path template. ✓
    Partition projection lets Athena calculate predictable date partitions and locations during queries, avoiding manual partition registration.
  2. Create a business catalog asset for every daily folder and query those assets with Athena.
    Business catalog assets support discovery and governance but do not provide Athena partition metadata or replace table configuration.
  3. Run a crawler after every folder arrival and manually inspect each discovered partition.
    A crawler can automatically discover and register partitions, but repeated crawler runs and manual inspection add unnecessary operational work when the partition pattern is predictable.
  4. Use an Iceberg table property to make ordinary Hive-style folders self-register in Athena.
    Iceberg metadata and Athena partition projection are different mechanisms; an Iceberg property does not manage partitions for an ordinary Hive-style table.
The trap
Confuses Iceberg table metadata with partition projection. Confuses business discovery assets with technical table metadata. This is executable but less suitable than projection for predictable Athena-only access.

Use Athena partition projection for predictable date-based S3 paths.

222 more 2: Data Store Management questions

The remaining 222 questions in this domain are part of the full AWS bank — 900 questions, every option explained. Start with the free five-minute check and see your score per domain.

Test your AWS readiness — free

Other AWS domains

Part of the Certsqill AWS question bank · 2: Data Store Management · Every answer, right and wrong, comes with its own explanation.