AWS 5: Incident and Event Response: 119 practice questions
48 hours only — 15% off every course with code SAVE15. Browse courses →48h · 15% off all courses · code SAVE15 →
Certifications Tools Flashcards Career Paths Exam Guides Blog Pricing For Teams About

AWS 5: Incident and Event Response: 119 practice questions

AWS 119 questions 12 shown free

12 of the 119 5: Incident and Event Response questions in the Certsqill AWS bank, shown in full below. Each one carries an explanation for every option, not just the correct one — the wrong answers are where the marks go.

Preparing for AWS? Take the free 5-min readiness check →

1. Configure an Amazon SQS standard DLQ on the EventBridge: Select TWO actions that satisfy these requirements.

Medium
A logistics transaction service receives shipment events through a custom EventBridge event bus and invokes a Step Functions workflow. The workflow occasionally fails because a downstream carrier API is unavailable. Failed deliveries from EventBridge must be retained, and workflow failures must be distinguishable from delivery failures. Operators need automatic retries for transient workflow errors and a durable queue for events EventBridge cannot deliver. Select TWO actions that satisfy these requirements.

Select two. More than one option is correct — every correct one is ticked below.

  1. Use the EventBridge DLQ to capture every carrier API failure after the workflow has accepted the event.
    An EventBridge target DLQ captures delivery failures, not failures occurring after the Step Functions workflow accepts the event.
  2. Configure only Step Functions Catch handling and omit the EventBridge target DLQ because workflow errors are retained automatically.
    Step Functions error handling does not retain events that EventBridge fails to deliver to the target.
  3. Replace the EventBridge DLQ with an Amazon SQS FIFO queue to guarantee ordered failure retention.
    EventBridge target DLQs support standard SQS queues, not FIFO queues, regardless of ordering requirements.
  4. Configure an Amazon SQS standard DLQ on the EventBridge rule target. ✓
    An EventBridge target DLQ retains events that EventBridge cannot successfully deliver after applicable retry behavior.
  5. Add Step Functions Retry and Catch handling for transient carrier API errors and route exhausted failures explicitly. ✓
    Step Functions Retry and Catch control workflow error handling after EventBridge has successfully delivered the event.
The trap
This assumes successful workflow error handling replaces transport-level failure retention. This assumes FIFO queues are valid EventBridge target DLQs. This confuses target delivery failure with downstream workflow execution failure.

Configure an EventBridge standard SQS DLQ and independent Step Functions Retry/Catch handling.

2. Use an SQS event source mapping: Which design should the team choose?

Medium
A production container fleet submits asynchronous image jobs to Amazon SQS. Each job invokes Lambda, which calls a third-party image service that may be unavailable for several minutes. The design must retain messages during transient failures, move permanently invalid jobs aside after bounded attempts, preserve an operator-visible record, and avoid custom polling. Which design should the team choose?
  1. Write a Lambda poller that receives each message, processes it, and deletes it after success.
    This is executable, but custom polling violates the managed-polling requirement and requires the team to implement scaling and retry handling.
  2. Publish each job to SNS and invoke Lambda asynchronously, relying on Lambda’s service retries alone.
    This replaces the requested SQS consumer workflow and does not provide SQS visibility and redrive handling for the existing queue.
  3. Use an EventBridge rule with a target DLQ, then invoke Lambda directly for each submitted job.
    An EventBridge target DLQ records delivery failures, not failures after Lambda accepts and processes an SQS job.
  4. Use an SQS event source mapping, align visibility with processing time, and configure an SQS redrive policy. ✓
    The event source mapping provides managed polling and scaling. Visibility settings allow retries, and the redrive policy moves repeatedly failed messages to a durable DLQ.
The trap
Confuses target-delivery failure with consumer-processing failure. Treats asynchronous Lambda retries as equivalent to SQS redrive. Uses custom consumer logic when an event source mapping is required.

Use an SQS event source mapping with aligned visibility and an SQS redrive policy.

3. Start an AWS AppConfig deployment using a gradual: Which solution should the team implement?

Medium
An event-driven billing system stores feature configuration in AWS AppConfig. A new pricing rule must roll out gradually because incorrect values could overcharge customers. The service already publishes CloudWatch metrics and has alarms for billing errors. The deployment must automatically stop and roll back if those alarms breach, while the configuration remains managed outside application redeployments. Which solution should the team implement?
  1. Use Systems Manager Automation to edit the AppConfig configuration directly on every host without an approval or rollback step.
    Direct host edits bypass AppConfig distribution controls and do not provide the required monitored gradual rollback behavior.
  2. Start an AWS AppConfig deployment using a gradual deployment strategy and associate the billing alarms for rollback monitoring. ✓
    AppConfig gradual deployments distribute configuration progressively and can roll back when monitored alarms indicate deployment failure.
  3. Write the pricing rule to Parameter Store and require each application instance to restart before reading it.
    Parameter Store can store configuration, but restart-based retrieval lacks AppConfig gradual rollout and alarm-triggered rollback.
  4. Update the application container image through a rolling deployment and store the pricing rule in an environment variable.
    Container rolling deployment couples configuration to application release and does not provide AppConfig alarm-monitored configuration rollback.
The trap
This assumes configuration can safely be changed independently on hosts without a controlled deployment. This replaces a configuration deployment with an application deployment. This treats centralized storage as equivalent to monitored configuration deployment.

Deploy through AWS AppConfig using a gradual strategy with CloudWatch alarm monitoring and rollback.

4. Use AWS AppConfig deployment strategies with gradual: Which design best satisfies these requirements?

Medium
A healthcare application stores feature flags in AWS AppConfig. A critical flag must reach production gradually, automatically stop when error metrics breach thresholds, and return to the last working version without redeploying application binaries. The team also requires deployment status and rollback evidence for audits. Which design best satisfies these requirements?
  1. Deploy flag changes through AWS CloudFormation stack updates and use CloudFormation rollback for application errors.
    CloudFormation rollback handles stack update failures, not gradual runtime configuration distribution or application alarm thresholds.
  2. Store each flag version in Amazon S3 and use EventBridge to invoke Lambda for immediate replacement.
    This provides event-driven replacement but lacks AppConfig gradual rollout monitoring and native configuration rollback evidence.
  3. Use AWS AppConfig deployment strategies with gradual growth, CloudWatch alarms, and automatic rollback on alarm activation. ✓
    AppConfig gradually distributes configuration, monitors deployment alarms, and rolls back configuration when distribution or monitoring fails.
  4. Bake each flag version into an Amazon Machine Image and use CodeDeploy in-place deployments across instances.
    This changes application images and uses instance deployments, violating the requirement to change flags without redeploying binaries.
The trap
Confuses infrastructure template rollback with monitored runtime configuration deployment. Treats runtime configuration as immutable compute code and loses AppConfig rollout controls. Assumes custom orchestration provides the same controlled rollout and rollback guarantees as AppConfig.

AWS AppConfig provides gradual configuration rollout, alarm monitoring, and automatic rollback without rebuilding application binaries.

5. Configure each regional Config rule to invoke that runbook: Select TWO complementary configuration changes.

Medium
A multi-Region customer portal uses AWS Config to detect public security groups. Each Region must remediate only noncompliant groups, avoid overwriting an administrator’s newer correction, and limit simultaneous changes during an incident. The remediation must be auditable and stop after a defined error threshold. Exhibit: Config evaluates a group as NON_COMPLIANT; its latest recorded state is five minutes old, while a manual change occurred two minutes ago. Select TWO complementary configuration changes.

Select two. More than one option is correct — every correct one is ticked below.

  1. Configure each regional Config rule to invoke that runbook automatically with bounded concurrency and an error threshold. ✓
    The Config remediation association supplies the trigger and rate controls; the separately defined runbook supplies safe change logic.
  2. Run CloudFormation drift detection and record differences without invoking any correction workflow.
    Detection alone reports differences but leaves the noncompliant group unchanged.
  3. Implement an idempotent Automation runbook that rechecks live rules, preserves newer compliant changes and records each result. ✓
    Fresh state checks avoid overwriting an already corrected resource; conditional operations and execution records support safe retries and audits.
  4. Invoke an unrestricted Lambda function for every finding, immediately replace each group’s rules, and rely on Lambda retries for failed changes.
    Lambda can modify security groups, but this design can overwrite the newer manual correction and lacks bounded fleet execution and a defined stop threshold.
  5. Run a nightly Maintenance Window that applies the desired rules to every security group and records the commands in a central log.
    The scheduled bulk operation is auditable but can be delayed and can overwrite a newer administrator correction.
The trap
Prioritizes immediate execution without stale-state protection or blast-radius controls. Uses periodic convergence where event-driven conditional remediation is required. Drift reporting is not remediation.

Use Config-triggered Automation that rechecks current state, is idempotent and auditable, and enforces concurrency and error limits.

6. Check the deployment group's platform: Which investigation is most appropriate?

Hard
An internal operations team reports that a CodeDeploy deployment to Amazon EC2 failed after several instances entered the replacement phase. The deployment dashboard shows a generic failure, but application availability alarms remained normal. Some instances were launched by an Auto Scaling group during deployment. The team needs the decisive failure cause and must determine whether the remaining fleet is running the intended revision before retrying. Which investigation is most appropriate?
  1. Review alarm history and immediately redeploy the last successful revision to every instance before collecting deployment logs.
    Normal alarms do not rule out deployment failures, and redeploying first can obscure the original evidence.
  2. Inspect lifecycle-event logs and agent status, then compare revisions on instances launched during deployment.
    These checks help, but they omit deployment-group target status and the specific failed lifecycle event.
  3. Run CloudFormation drift detection on the deployment group, then replace instances that differ from the template.
    Drift detection concerns CloudFormation-managed properties and does not diagnose CodeDeploy lifecycle failures or application revisions.
  4. Check the deployment group's platform, target health, and failed lifecycle event; inspect agent logs and verify Auto Scaling revision updates. ✓
    The failed lifecycle event and agent logs identify the cause, while target health and Auto Scaling checks establish fleet consistency.
The trap
Relies on instance evidence without correlating deployment state. Confuses infrastructure drift with deployment evidence. Treats application alarms as complete deployment diagnostics.

Correlate CodeDeploy lifecycle events, target health, agent logs, and Auto Scaling revision state.

7. The task security group likely blocks inbound traffic: Which diagnosis should guide the next investigation?

Hard
A partner-facing API platform runs on Amazon ECS behind an Application Load Balancer. A release succeeds in CodePipeline and tasks enter RUNNING, but the deployment fails health checks and traffic remains on the old task set. Container logs show the process listening on the configured port. The service’s target group reports repeated health-check timeouts, and the task security group was changed during a recent network hardening effort. Which diagnosis should guide the next investigation?
  1. The network ACL is necessarily blocking return traffic because security groups cannot affect load-balancer health checks.
    Security groups directly control task ingress, while network ACLs are only one possible network-layer cause.
  2. The container image is invalid because ECS tasks reached RUNNING before the deployment failed.
    RUNNING confirms task startup, but the evidence points to network reachability between the load balancer and task.
  3. The task security group likely blocks inbound traffic from the load balancer security group on the health-check port. ✓
    Stateful security groups must allow the load balancer source on the health-check port for target connectivity.
  4. CodePipeline must have published an incorrect artifact because target health checks do not evaluate network connectivity.
    Application Load Balancer health checks explicitly test target connectivity and response behavior after task startup.
The trap
Treats ECS lifecycle state as proof that load-balancer health checks can connect. Incorrectly dismisses security groups and assumes a stateless network ACL is the sole cause. Confuses successful pipeline artifact delivery with successful service reachability.

The recent task security-group change and health-check timeouts indicate missing load-balancer-to-task ingress.

8. The target DLQ covers delivery failures: Which conclusion is most accurate?

Hard
A SaaS application uses an EventBridge rule to invoke a Lambda function for enterprise tenant provisioning. EventBridge metrics show successful invocations, but some tenants are missing after the function returns an error while processing requests. The team configured an EventBridge target dead-letter queue and expects failed tenant operations to appear there. Permissions allow EventBridge to send messages to the queue. Which conclusion is most accurate?
  1. Replace the standard target DLQ with a FIFO queue so accepted Lambda failures can be retained in order.
    EventBridge target DLQs support standard SQS queues, and accepted Lambda failures are outside this DLQ path.
  2. Grant the Lambda execution role permission to send to the DLQ because Lambda must record its own processing failures there.
    The configured EventBridge target DLQ is written by EventBridge for delivery failures; Lambda processing failures require Lambda asynchronous failure handling.
  3. EventBridge retries the target until Lambda completes provisioning successfully, so the queue policy needs no change.
    EventBridge retry policy governs target delivery, not processing after Lambda accepts the invocation.
  4. The target DLQ covers delivery failures; configure Lambda asynchronous failure handling for errors after acceptance. ✓
    EventBridge considers delivery successful when Lambda accepts the invocation. Later Lambda processing failures require Lambda retry and failure handling.
The trap
Extends delivery retries into application execution. Confuses EventBridge's DLQ writer with the Lambda execution role. Confuses queue type requirements with the failure boundary.

The EventBridge DLQ covers delivery failures; Lambda asynchronous handling covers failures after acceptance.

9. Configure bounded Step Functions retries and catches: Select TWO designs.

Hard
A global media service consumes order events through Amazon SQS and invokes a Step Functions workflow for entitlement updates. Processing normally takes 40 seconds but occasionally reaches 3 minutes. Duplicate events are possible, and poison messages must not block unrelated orders indefinitely. The team wants bounded retries, reliable redrive, and no premature message reappearance while a consumer is still working. Select TWO designs.

Select two. More than one option is correct — every correct one is ticked below.

  1. Use an EventBridge target DLQ to redrive SQS messages after the workflow accepts them.
    EventBridge target DLQs handle EventBridge delivery failures, not messages accepted from SQS and processed by a workflow.
  2. Configure bounded Step Functions retries and catches, and make entitlement updates idempotent. ✓
    Step Functions Retry and Catch policies provide bounded workflow recovery, while idempotence limits duplicate external effects.
  3. Retry every workflow failure indefinitely so transient outages cannot permanently remove an order event.
    Unlimited workflow retries can retain poison messages indefinitely and prevent bounded recovery or redrive.
  4. Set SQS visibility to 30 seconds so stalled messages quickly return for another consumer.
    Thirty seconds is shorter than the possible three-minute processing time and can cause premature concurrent delivery.
  5. Set SQS visibility longer than the maximum processing time and configure a redrive policy for repeated failures. ✓
    A sufficient visibility timeout prevents premature reappearance, while the redrive policy isolates poison messages after bounded receive attempts.
The trap
Applies EventBridge delivery controls to an SQS consumer path. Confuses rapid retry with safe visibility management. Confuses durable recovery with unbounded execution.

Use bounded Step Functions retries with idempotent effects, plus SQS visibility alignment and redrive.

10. Allow the deployment principal to use kms: Which corrective action addresses the decisive failure?

Hard
A shared developer platform deploys encrypted freeform configuration through AWS AppConfig. Deployments previously succeeded, but a new deployment fails immediately with an AccessDenied error for kms:GenerateDataKey. The configuration is stored in Amazon S3 and encrypted with a customer managed KMS key. The deployment operator can read the S3 object, and the AppConfig service role is unchanged. Which corrective action addresses the decisive failure?
  1. Allow the deployment principal to use kms:GenerateDataKey on the customer managed key. ✓
    AppConfig deployments using a customer managed key require the initiating IAM principal to have kms:GenerateDataKey permission.
  2. Replace the customer managed key with an AWS managed key and retain the existing deployment permissions without investigating the denied KMS action.
    Changing encryption design is unnecessary and does not directly remediate the reported permission failure.
  3. Grant the unchanged AppConfig service role permission to rewrite the encrypted S3 object before starting the deployment.
    The decisive requirement is the caller's KMS permission; AppConfig does not need to rewrite the source object for this correction.
  4. Add only kms:Decrypt to the deployment principal because AppConfig generates the data key independently.
    The reported denial is for kms:GenerateDataKey, which must be allowed for the initiating principal.
The trap
Confuses decryption authorization with data-key generation authorization. Confuses source-object handling with the denied KMS operation. Uses an avoidable storage and encryption redesign instead of correcting caller authorization.

Allow the initiating principal to call kms:GenerateDataKey on the customer managed key.

11. The recreated target role’s trust policy does not allow: Which diagnosis is most likely?

Hard
A logistics transaction service emits an EventBridge event when a production queue exceeds its threshold. The rule targets a Systems Manager Automation runbook that increases worker capacity. The first event starts successfully, but later events remain in FAILED status with an assume-role error. The runbook’s action role has permission to modify the Auto Scaling group, and the automation document itself passes validation. The organization recently recreated the EventBridge target role. Which diagnosis is most likely?
  1. The runbook action role needs permission to assume the EventBridge target role before capacity can change.
    The EventBridge service assumes the target role; the runbook action role independently performs resource modifications.
  2. The Automation document must be converted into a Lambda function because EventBridge cannot target Systems Manager Automation.
    Systems Manager Automation is a supported EventBridge target when the target role is configured correctly.
  3. The Auto Scaling group lacks a CloudWatch alarm because EventBridge cannot invoke Automation without alarm state.
    EventBridge can invoke Automation directly; a CloudWatch alarm is not required when the rule receives the event.
  4. The recreated target role’s trust policy does not allow events.amazonaws.com to assume it. ✓
    EventBridge must assume its configured target role before invoking Systems Manager Automation; permission policies alone are insufficient.
The trap
Reverses the trust relationship between EventBridge and the runbook execution role. Confuses the event source condition with a separate alarm-based invocation design. Assumes EventBridge supports only compute targets.

An assume-role failure after recreating the target role indicates a missing EventBridge service principal in its trust policy.

12. Verify the ECS task execution role can authenticate to ECR: Which investigation should be performed first?

Hard
A production container fleet uses an ECS service with tasks launched from a private Amazon ECR repository. After a task-definition revision, new tasks repeatedly stop during startup while existing tasks remain healthy. ECS service events report that the task cannot pull the image. The task execution role was recently replaced, and the VPC endpoints for Amazon ECR and Amazon S3 remain unchanged. Which investigation should be performed first?
  1. Modify the application container’s task role to grant ECR permissions because ECS uses that role during image retrieval.
    ECS uses the task execution role for image retrieval; the application task role is used after containers start.
  2. Replace the ECR image with a new tag because ECS service deployment circuit breakers cannot handle private repositories.
    ECS supports private ECR images, and changing tags does not address the stated execution-role or network evidence.
  3. Increase the ECS service desired count so additional tasks can eventually pull the image successfully.
    Increasing desired count cannot correct authorization or network failures and may increase stopped-task churn.
  4. Verify the ECS task execution role can authenticate to ECR and retrieve image layers through the private network path. ✓
    Image-pull failures commonly result from missing execution-role ECR permissions or unavailable private connectivity to ECR dependencies.
The trap
Blames deployment mechanics instead of checking the image-pull authorization path. Confuses task role permissions with task execution role permissions. Treats a deterministic image-pull failure as insufficient capacity.

Check the replacement task execution role and private ECR connectivity before changing capacity or application roles.

107 more 5: Incident and Event Response questions

The remaining 107 questions in this domain are part of the full AWS bank — 849 questions, every option explained. Start with the free five-minute check and see your score per domain.

Test your AWS readiness — free

Other AWS domains

Part of the Certsqill AWS question bank · 5: Incident and Event Response · Every answer, right and wrong, comes with its own explanation.