AWS 3: Data Operations and Support: 198 practice questions
12 of the 198 3: Data Operations and Support questions in the Certsqill AWS bank, shown in full below. Each one carries an explanation for every option, not just the correct one — the wrong answers are where the marks go.
Preparing for AWS? Take the free 5-min readiness check →
1. Use Step Functions: Which design best meets the requirement?
- Run a crawler before a scheduled ETL job.Crawler completion and ETL completion are separate operations, and scheduling does not guarantee the required validation dependency.
- Schedule validation and loading together.Simultaneous scheduled starts do not ensure validation succeeds before loading begins.
- Trigger the load directly from S3 and validate afterward.Direct loading bypasses the required validation gate and can publish invalid data.
- Use Step Functions. ✓A Step Functions workflow can start from the object event, validate the file, branch on the result, and prevent loading after failure.
Use Step Functions to sequence validation, conditional loading, and failure handling.
2. Inspect DAG dependencies and scheduler logs: Which investigation is most appropriate first?
- Inspect DAG dependencies and scheduler logs. ✓A pre-task failure commonly indicates dependency, parsing, or scheduling issues, so DAG state and scheduler evidence should be checked first.
- Increase retries without reviewing logs or dependencies.Retries do not correct invalid DAG dependencies or parsing problems and can obscure the original cause.
- Reset the integration's Glue bookmarks.Glue bookmarks track source-processing state and do not diagnose an MWAA DAG that fails before task execution.
- Restart downstream applications.Downstream restarts do not explain a DAG failure before task execution and may disrupt healthy systems.
Check DAG dependencies, parsing state, and scheduler evidence before changing data-processing state.
3. Use an SDK client with pagination: Which implementation is most appropriate?
- Use one SDK response and write its records to the report.A single response may be only the first page, so this can silently omit records.
- Retry throttled calls and skip records that continue failing.Retries address throttling, but silently skipping failed records violates the requirement to record them for review.
- Use an SDK client with pagination, backoff, and record-level error capture. ✓The application must follow pagination tokens, handle throttling, and isolate failed records to produce complete and auditable results.
- Use an SDK paginator and let failed records abort the entire batch.Pagination supports completeness, but aborting the batch does not record failed records for review or preserve successful processing.
Explicitly implement pagination, throttling handling, and record-level failure capture.
4. Configure Athena partition projection for the date: Which feature should improve partition discovery for Athen
- Use partition projection and expect other query engines to use its settings.Partition projection is an Athena query feature; other engines use standard partition metadata instead.
- Configure Athena partition projection for the date and region columns. ✓Partition projection lets Athena calculate predictable partition values without retrieving partition metadata from the catalog.
- Create a view that lists all partitions before each query.Views do not configure partition projection and cannot eliminate catalog metadata retrieval for the underlying table.
- Enable projection on a table with unrestricted, unpredictable partition values.Projection requires configured ranges, patterns, or enumerated values; unpredictable values cannot be safely calculated.
Projection calculates predictable partitions in Athena, reducing catalog lookups while retaining partition pruning for constrained queries.
5. Retry throttled requests with bounded backoff and record: Select TWO actions that meet these requirements.
Select two. More than one option is correct — every correct one is ticked below.
- Retry throttled requests with bounded backoff and record individual invoice errors. ✓Bounded backoff addresses transient throttling, while error recording preserves rejected invoices for controlled replay.
- Follow pagination tokens until the API indicates no additional page. ✓Following continuation tokens ensures the reconciliation job examines the partner's complete invoice result set.
- Stop the entire job when the first invoice is rejected.A single rejected invoice should be isolated so accepted invoices and later pages continue processing.
- Retry every invoice indefinitely until the partner accepts it.Indefinite retries can stall reconciliation and do not distinguish permanent invoice rejection from transient throttling.
- Read only the first page because the API client handles completeness automatically.API clients do not universally consume every page, so application code must follow pagination explicitly.
Complete API consumption requires pagination; resilient reconciliation requires bounded throttling retries and durable per-invoice failure capture.
6. Profile, standardize, flag, and report: Which approach is best?
- Profile, standardize, flag, and report. ✓This creates a repeatable quality assessment while preserving questionable records for later decisions.
- Delete missing measurements and convert units during queries.Deletion loses potentially useful records, while deferred conversion can produce inconsistent interpretations.
- Load raw files into production and infer units during transformations.Downstream inference can produce inconsistent results and does not provide an initial quality assessment.
- Deduplicate entire rows and discard invalid timestamps.Whole-row deduplication can miss business-key duplicates, and discarding records loses quality evidence.
Profile and standardize data, preserve questionable records, and create explicit quality flags.
7. Use partition projection with a bounded date range: Which configuration best improves these queries?
- Enable projection and expect SHOW PARTITIONS to list every projected date.Projected partitions are calculated at query time and do not necessarily appear in catalog partition listings.
- Remove the date predicate because projection automatically reads only populated dates.Projection can include nonexistent locations, so removing the predicate may make Athena consider many unnecessary projected values.
- Use partition projection with a bounded date range and retain the date predicate. ✓Projection avoids catalog partition lookups, while the date predicate enables pruning; nonexistent projected dates simply return no rows.
- Use projection while querying the table through Amazon Redshift Spectrum.Partition projection is usable through Athena; Redshift Spectrum uses standard partition metadata for the table.
Bounded projection reduces metadata lookup, while the date filter preserves pruning and limits empty projected candidates.
8. Invoke Lambda from S3 events: Which design is most appropriate?
- Transform the object, delete it, and then record that processing completed.Deleting before durable completion state can cause data loss or ambiguous retries if the function fails between operations.
- Schedule Lambda to rescan the prefix and transform every discovered object.Repeated scans can reprocess existing objects unless additional durable state and output controls are added.
- Invoke Lambda from S3 events, claim each object with a conditional idempotency record, and mark success only after the output is committed. ✓A conditional claim prevents concurrent duplicate work, while a success state recorded after output completion allows retries after failures without treating an unfinished object as complete.
- Invoke Lambda from S3 events and assume each notification is delivered once.Notifications can be duplicated, so exactly-once delivery cannot be assumed.
Use event-driven Lambda with conditional idempotency state and post-success completion recording.
9. Create an EventBridge rule matching InventoryChanged: Which mechanism best meets these requirements?
- Send all warehouse events to the target and filter in application code.This is executable but causes unnecessary invocations and moves required routing logic downstream.
- Create an EventBridge rule matching InventoryChanged and route matches to the target. ✓An EventBridge event-pattern rule filters by detail type and routes matching events with their event payload.
- Create a scheduled EventBridge invocation for periodic synchronization.A schedule is time-based and does not react to event content or route only matching events.
- Use an S3 notification rule to match the warehouse event detail type.S3 notifications address S3 events and do not provide detail-type matching for arbitrary warehouse events.
Use an EventBridge event-pattern rule to match InventoryChanged and route the original event.
10. Aggregate attribution by campaign and week before: Select TWO actions that meet these requirements.
Select two. More than one option is correct — every correct one is ticked below.
- Aggregate attribution by campaign and week before presenting the dashboard. ✓Pre-aggregation at the dashboard grain simplifies visuals and ensures campaign-week metrics use deliberate grouping semantics.
- Use a QuickSight view and assume it materializes results automatically.A view defines query logic but does not automatically materialize or cache its results for dashboard performance.
- Refresh SPICE only when a user opens the dashboard, regardless of source changes.Opening a dashboard does not necessarily refresh cached data, so planned refresh scheduling is required for predictable currency.
- Create a QuickSight dataset from the Athena source and configure scheduled SPICE refreshes. ✓QuickSight provides governed dashboards, while scheduled SPICE refreshes support interactive analysis without requiring real-time source reads.
- Query every raw S3 event directly for each dashboard interaction.Repeated raw scans can increase latency and work, conflicting with the requirement for fast interactive governed visuals.
QuickSight with scheduled SPICE refreshes supports governed interactive dashboards, while campaign-week aggregation provides the required reporting grain.
11. Enable job bookmarks and rewind the bookmark only when: Which mechanism should the data engineer choose?
- Disable job bookmarks and scan the complete S3 prefix during every run.Disabling bookmarks causes every run to rescan prior inputs, increasing duplicate-processing risk and requiring manual output management.
- Reset the job bookmark after every successful run to establish a fresh checkpoint.Resetting after each run removes progress tracking, so subsequent runs can reread all previously processed files.
- Enable job bookmarks and rewind the bookmark only when a historical backfill is required. ✓Bookmarks track processed S3 inputs, while rewinding permits controlled reprocessing from an earlier successful job run.
- Pause job bookmarks permanently and manually delete prior target objects before each scheduled run.Pause processes a selected range without updating state, and deleting targets does not provide dependable source progress tracking.
Enable Glue job bookmarks for incremental processing, using bookmark rewind only for deliberate historical backfills.
12. Configure partition projection properties on the base: Which action is best?
- Register every daily partition manually and remove the date predicate from dashboard queries.Manual registration preserves metadata lookups, while removing date predicates prevents effective partition pruning and increases scanned data.
- Configure projection properties on the table and expect Redshift Spectrum to use them automatically.Partition projection is Athena-specific; other readers use standard catalog or metastore partition metadata instead.
- Configure partition projection properties on the base table with the valid date range and path template. ✓Partition projection lets Athena calculate predictable partitions, avoid GetPartitions calls, and still prune partitions matching query predicates.
- Create a view over the table and configure projection properties only on that view.Athena does not use view properties for partition projection; projection must be configured on referenced base tables.
Configure partition projection on the base table because Athena can calculate predictable partitions without catalog lookups.
186 more 3: Data Operations and Support questions
The remaining 186 questions in this domain are part of the full AWS bank — 900 questions, every option explained. Start with the free five-minute check and see your score per domain.
Test your AWS readiness — freeOther AWS domains
- 1: Data Ingestion and Transformation — 306 questions →
- 2: Data Store Management — 234 questions →
- 4: Data Security and Governance — 162 questions →
- All 900 AWS questions →
- AWS certification: requirements, cost and exam format →