AWS 5: Incident and Event Response: 119 practice questions
12 of the 119 5: Incident and Event Response questions in the Certsqill AWS bank, shown in full below. Each one carries an explanation for every option, not just the correct one — the wrong answers are where the marks go.
Preparing for AWS? Take the free 5-min readiness check →
1. Configure an Amazon SQS standard DLQ on the EventBridge: Select TWO actions that satisfy these requirements.
Select two. More than one option is correct — every correct one is ticked below.
- Use the EventBridge DLQ to capture every carrier API failure after the workflow has accepted the event.An EventBridge target DLQ captures delivery failures, not failures occurring after the Step Functions workflow accepts the event.
- Configure only Step Functions Catch handling and omit the EventBridge target DLQ because workflow errors are retained automatically.Step Functions error handling does not retain events that EventBridge fails to deliver to the target.
- Replace the EventBridge DLQ with an Amazon SQS FIFO queue to guarantee ordered failure retention.EventBridge target DLQs support standard SQS queues, not FIFO queues, regardless of ordering requirements.
- Configure an Amazon SQS standard DLQ on the EventBridge rule target. ✓An EventBridge target DLQ retains events that EventBridge cannot successfully deliver after applicable retry behavior.
- Add Step Functions Retry and Catch handling for transient carrier API errors and route exhausted failures explicitly. ✓Step Functions Retry and Catch control workflow error handling after EventBridge has successfully delivered the event.
Configure an EventBridge standard SQS DLQ and independent Step Functions Retry/Catch handling.
2. Use an SQS event source mapping: Which design should the team choose?
- Write a Lambda poller that receives each message, processes it, and deletes it after success.This is executable, but custom polling violates the managed-polling requirement and requires the team to implement scaling and retry handling.
- Publish each job to SNS and invoke Lambda asynchronously, relying on Lambda’s service retries alone.This replaces the requested SQS consumer workflow and does not provide SQS visibility and redrive handling for the existing queue.
- Use an EventBridge rule with a target DLQ, then invoke Lambda directly for each submitted job.An EventBridge target DLQ records delivery failures, not failures after Lambda accepts and processes an SQS job.
- Use an SQS event source mapping, align visibility with processing time, and configure an SQS redrive policy. ✓The event source mapping provides managed polling and scaling. Visibility settings allow retries, and the redrive policy moves repeatedly failed messages to a durable DLQ.
Use an SQS event source mapping with aligned visibility and an SQS redrive policy.
3. Start an AWS AppConfig deployment using a gradual: Which solution should the team implement?
- Use Systems Manager Automation to edit the AppConfig configuration directly on every host without an approval or rollback step.Direct host edits bypass AppConfig distribution controls and do not provide the required monitored gradual rollback behavior.
- Start an AWS AppConfig deployment using a gradual deployment strategy and associate the billing alarms for rollback monitoring. ✓AppConfig gradual deployments distribute configuration progressively and can roll back when monitored alarms indicate deployment failure.
- Write the pricing rule to Parameter Store and require each application instance to restart before reading it.Parameter Store can store configuration, but restart-based retrieval lacks AppConfig gradual rollout and alarm-triggered rollback.
- Update the application container image through a rolling deployment and store the pricing rule in an environment variable.Container rolling deployment couples configuration to application release and does not provide AppConfig alarm-monitored configuration rollback.
Deploy through AWS AppConfig using a gradual strategy with CloudWatch alarm monitoring and rollback.
4. Use AWS AppConfig deployment strategies with gradual: Which design best satisfies these requirements?
- Deploy flag changes through AWS CloudFormation stack updates and use CloudFormation rollback for application errors.CloudFormation rollback handles stack update failures, not gradual runtime configuration distribution or application alarm thresholds.
- Store each flag version in Amazon S3 and use EventBridge to invoke Lambda for immediate replacement.This provides event-driven replacement but lacks AppConfig gradual rollout monitoring and native configuration rollback evidence.
- Use AWS AppConfig deployment strategies with gradual growth, CloudWatch alarms, and automatic rollback on alarm activation. ✓AppConfig gradually distributes configuration, monitors deployment alarms, and rolls back configuration when distribution or monitoring fails.
- Bake each flag version into an Amazon Machine Image and use CodeDeploy in-place deployments across instances.This changes application images and uses instance deployments, violating the requirement to change flags without redeploying binaries.
AWS AppConfig provides gradual configuration rollout, alarm monitoring, and automatic rollback without rebuilding application binaries.
5. Configure each regional Config rule to invoke that runbook: Select TWO complementary configuration changes.
Select two. More than one option is correct — every correct one is ticked below.
- Configure each regional Config rule to invoke that runbook automatically with bounded concurrency and an error threshold. ✓The Config remediation association supplies the trigger and rate controls; the separately defined runbook supplies safe change logic.
- Run CloudFormation drift detection and record differences without invoking any correction workflow.Detection alone reports differences but leaves the noncompliant group unchanged.
- Implement an idempotent Automation runbook that rechecks live rules, preserves newer compliant changes and records each result. ✓Fresh state checks avoid overwriting an already corrected resource; conditional operations and execution records support safe retries and audits.
- Invoke an unrestricted Lambda function for every finding, immediately replace each group’s rules, and rely on Lambda retries for failed changes.Lambda can modify security groups, but this design can overwrite the newer manual correction and lacks bounded fleet execution and a defined stop threshold.
- Run a nightly Maintenance Window that applies the desired rules to every security group and records the commands in a central log.The scheduled bulk operation is auditable but can be delayed and can overwrite a newer administrator correction.
Use Config-triggered Automation that rechecks current state, is idempotent and auditable, and enforces concurrency and error limits.
6. Check the deployment group's platform: Which investigation is most appropriate?
- Review alarm history and immediately redeploy the last successful revision to every instance before collecting deployment logs.Normal alarms do not rule out deployment failures, and redeploying first can obscure the original evidence.
- Inspect lifecycle-event logs and agent status, then compare revisions on instances launched during deployment.These checks help, but they omit deployment-group target status and the specific failed lifecycle event.
- Run CloudFormation drift detection on the deployment group, then replace instances that differ from the template.Drift detection concerns CloudFormation-managed properties and does not diagnose CodeDeploy lifecycle failures or application revisions.
- Check the deployment group's platform, target health, and failed lifecycle event; inspect agent logs and verify Auto Scaling revision updates. ✓The failed lifecycle event and agent logs identify the cause, while target health and Auto Scaling checks establish fleet consistency.
Correlate CodeDeploy lifecycle events, target health, agent logs, and Auto Scaling revision state.
7. The task security group likely blocks inbound traffic: Which diagnosis should guide the next investigation?
- The network ACL is necessarily blocking return traffic because security groups cannot affect load-balancer health checks.Security groups directly control task ingress, while network ACLs are only one possible network-layer cause.
- The container image is invalid because ECS tasks reached RUNNING before the deployment failed.RUNNING confirms task startup, but the evidence points to network reachability between the load balancer and task.
- The task security group likely blocks inbound traffic from the load balancer security group on the health-check port. ✓Stateful security groups must allow the load balancer source on the health-check port for target connectivity.
- CodePipeline must have published an incorrect artifact because target health checks do not evaluate network connectivity.Application Load Balancer health checks explicitly test target connectivity and response behavior after task startup.
The recent task security-group change and health-check timeouts indicate missing load-balancer-to-task ingress.
8. The target DLQ covers delivery failures: Which conclusion is most accurate?
- Replace the standard target DLQ with a FIFO queue so accepted Lambda failures can be retained in order.EventBridge target DLQs support standard SQS queues, and accepted Lambda failures are outside this DLQ path.
- Grant the Lambda execution role permission to send to the DLQ because Lambda must record its own processing failures there.The configured EventBridge target DLQ is written by EventBridge for delivery failures; Lambda processing failures require Lambda asynchronous failure handling.
- EventBridge retries the target until Lambda completes provisioning successfully, so the queue policy needs no change.EventBridge retry policy governs target delivery, not processing after Lambda accepts the invocation.
- The target DLQ covers delivery failures; configure Lambda asynchronous failure handling for errors after acceptance. ✓EventBridge considers delivery successful when Lambda accepts the invocation. Later Lambda processing failures require Lambda retry and failure handling.
The EventBridge DLQ covers delivery failures; Lambda asynchronous handling covers failures after acceptance.
9. Configure bounded Step Functions retries and catches: Select TWO designs.
Select two. More than one option is correct — every correct one is ticked below.
- Use an EventBridge target DLQ to redrive SQS messages after the workflow accepts them.EventBridge target DLQs handle EventBridge delivery failures, not messages accepted from SQS and processed by a workflow.
- Configure bounded Step Functions retries and catches, and make entitlement updates idempotent. ✓Step Functions Retry and Catch policies provide bounded workflow recovery, while idempotence limits duplicate external effects.
- Retry every workflow failure indefinitely so transient outages cannot permanently remove an order event.Unlimited workflow retries can retain poison messages indefinitely and prevent bounded recovery or redrive.
- Set SQS visibility to 30 seconds so stalled messages quickly return for another consumer.Thirty seconds is shorter than the possible three-minute processing time and can cause premature concurrent delivery.
- Set SQS visibility longer than the maximum processing time and configure a redrive policy for repeated failures. ✓A sufficient visibility timeout prevents premature reappearance, while the redrive policy isolates poison messages after bounded receive attempts.
Use bounded Step Functions retries with idempotent effects, plus SQS visibility alignment and redrive.
10. Allow the deployment principal to use kms: Which corrective action addresses the decisive failure?
- Allow the deployment principal to use kms:GenerateDataKey on the customer managed key. ✓AppConfig deployments using a customer managed key require the initiating IAM principal to have kms:GenerateDataKey permission.
- Replace the customer managed key with an AWS managed key and retain the existing deployment permissions without investigating the denied KMS action.Changing encryption design is unnecessary and does not directly remediate the reported permission failure.
- Grant the unchanged AppConfig service role permission to rewrite the encrypted S3 object before starting the deployment.The decisive requirement is the caller's KMS permission; AppConfig does not need to rewrite the source object for this correction.
- Add only kms:Decrypt to the deployment principal because AppConfig generates the data key independently.The reported denial is for kms:GenerateDataKey, which must be allowed for the initiating principal.
Allow the initiating principal to call kms:GenerateDataKey on the customer managed key.
11. The recreated target role’s trust policy does not allow: Which diagnosis is most likely?
- The runbook action role needs permission to assume the EventBridge target role before capacity can change.The EventBridge service assumes the target role; the runbook action role independently performs resource modifications.
- The Automation document must be converted into a Lambda function because EventBridge cannot target Systems Manager Automation.Systems Manager Automation is a supported EventBridge target when the target role is configured correctly.
- The Auto Scaling group lacks a CloudWatch alarm because EventBridge cannot invoke Automation without alarm state.EventBridge can invoke Automation directly; a CloudWatch alarm is not required when the rule receives the event.
- The recreated target role’s trust policy does not allow events.amazonaws.com to assume it. ✓EventBridge must assume its configured target role before invoking Systems Manager Automation; permission policies alone are insufficient.
An assume-role failure after recreating the target role indicates a missing EventBridge service principal in its trust policy.
12. Verify the ECS task execution role can authenticate to ECR: Which investigation should be performed first?
- Modify the application container’s task role to grant ECR permissions because ECS uses that role during image retrieval.ECS uses the task execution role for image retrieval; the application task role is used after containers start.
- Replace the ECR image with a new tag because ECS service deployment circuit breakers cannot handle private repositories.ECS supports private ECR images, and changing tags does not address the stated execution-role or network evidence.
- Increase the ECS service desired count so additional tasks can eventually pull the image successfully.Increasing desired count cannot correct authorization or network failures and may increase stopped-task churn.
- Verify the ECS task execution role can authenticate to ECR and retrieve image layers through the private network path. ✓Image-pull failures commonly result from missing execution-role ECR permissions or unavailable private connectivity to ECR dependencies.
Check the replacement task execution role and private ECR connectivity before changing capacity or application roles.
107 more 5: Incident and Event Response questions
The remaining 107 questions in this domain are part of the full AWS bank — 849 questions, every option explained. Start with the free five-minute check and see your score per domain.
Test your AWS readiness — freeOther AWS domains
- 1: SDLC Automation — 187 questions →
- 2: Configuration Management and IaC — 144 questions →
- 6: Security and Compliance — 144 questions →
- 3: Resilient Cloud Solutions — 128 questions →
- 4: Monitoring and Logging — 127 questions →
- All 849 AWS questions →
- AWS certification: requirements, cost and exam format →