AWS Solutions Architect Professional practice questions
48 hours only — 15% off every course with code SAVE15. Browse courses →48h · 15% off all courses · code SAVE15 →
Certifications Tools Flashcards Career Paths Exam Guides Blog Pricing For Teams About

AWS Solutions Architect Professional Design for New Solutions: 297 practice questions

AWS Solutions Architect Professional 297 questions 12 shown free

12 of the 297 Design for New Solutions questions in the Certsqill AWS Solutions Architect Professional bank, shown in full below. Each one carries an explanation for every option, not just the correct one — the wrong answers are where the marks go.

Preparing for AWS Solutions Architect Professional? Take the free 5-min readiness check →

1. Use blue-green environments with expand-contract schema: Which architecture best satisfies both requirements?

Medium
A financial services group is releasing a customer-facing payment service. Deployment must support rollback without rebuilding the previous version, and database changes must support old and new application versions during traffic transition. Rolling deployments currently mix versions and have caused schema errors. The pipeline can provision a separate environment and run health checks, but rollback actions must be configured. Which architecture best satisfies both requirements?
  1. Continue rolling deployment and run a CloudFormation change set before each release.
    A change set previews infrastructure changes, but rolling deployment still mixes versions and does not ensure schema compatibility.
  2. Use a canary release while applying an irreversible schema change before compatibility testing.
    Canary traffic limits exposure, but an irreversible schema change can prevent reliable rollback to the previous application.
  3. Replace the old environment in place and restore its previous AMI if errors occur.
    In-place replacement can leave partial changes and does not provide a separately validated environment or safe schema transition.
  4. Use blue-green environments with expand-contract schema changes, validate green, then shift traffic. ✓
    A separate validated environment enables traffic reversal, while expand-contract changes keep the schema compatible with both application versions.
The trap
Infrastructure preview does not solve version mixing. Irreversible schema changes undermine rollback. AMI restoration is not an atomic application rollback.

Use blue-green deployment with expand-contract schema evolution.

2. Commit only to the measured baseline with a Compute: Which approach is most appropriate?

Easy
An energy provider is launching a compute-heavy analytics service. Forecasting is uncertain: the steady baseline may increase, but seasonal bursts are unpredictable and the team does not want unused commitments if adoption is slow. Workloads run on eligible EC2 usage, and the company can tolerate normal on-demand pricing for capacity above the baseline. Finance asks for a purchasing approach that balances savings with elasticity rather than maximizing a discount based on an unverified forecast. Which approach is most appropriate?
  1. Use a budget threshold as a guaranteed hard spending cap for all compute usage.
    Budgets monitor configured thresholds and actions but do not universally guarantee an immediate service-wide spend cap.
  2. Run all analytics on Spot Instances without fallback capacity or interruption handling.
    Spot can reduce cost for interruptible workloads, but using it exclusively conflicts with workloads needing dependable baseline capacity.
  3. Commit only to the measured baseline with a Compute Savings Plan and keep uncertain peaks on demand. ✓
    A baseline commitment can reduce eligible steady usage cost while on-demand capacity preserves flexibility for uncertain growth.
  4. Purchase commitments for the highest projected seasonal demand before production measurements exist.
    Overcommitting against an uncertain forecast risks paying for unused commitment when adoption or seasonal demand is lower.
The trap
Budget alerts are not universal cost enforcement. Spot interruptions require resilient fallback design. Forecast uncertainty makes peak commitment risky.

Commit to measured baseline usage and retain on-demand elasticity for uncertain peaks.

3. Configure a DLQ and redrive policy: Which action is most appropriate?

Medium
A multinational enterprise processes order events with an SQS Standard queue. Consumers record an idempotency key in a database before acknowledging each message, so repeated deliveries do not repeat the business effect. During a downstream outage, queue age rises sharply, and several malformed messages repeatedly become visible after their visibility timeout. The team has enabled consumer Auto Scaling based only on average CPU utilization. Leadership asks which remaining design issue most threatens recovery and operational stability. Which action is most appropriate?
  1. Remove database idempotency because Standard queues already suppress duplicate deliveries.
    Standard queues are at-least-once and can redeliver messages, so consumer idempotency remains necessary for safe side effects.
  2. Increase visibility timeout indefinitely so malformed messages cannot return to the queue.
    A longer timeout delays retries but can hide failures and does not isolate messages that cannot ever process successfully.
  3. Change the queue to FIFO and assume deduplication makes downstream processing exactly once.
    FIFO provides ordering and bounded deduplication semantics, but it does not make external side effects exactly once automatically.
  4. Configure a DLQ and redrive policy, and scale consumers using backlog or queue-age signals. ✓
    A DLQ isolates repeatedly failing messages, while backlog-aware scaling addresses growing work that average CPU can miss.
The trap
Queue deduplication is not transactional side-effect protection. Timeout extension does not resolve poison messages. Standard queues do not suppress every duplicate.

Use a DLQ for poison messages and scale from backlog or queue age rather than CPU alone.

4. Perform a measured restore rehearsal through dependencies: Which TWO actions should the architect require befo

Medium
An industrial manufacturer is moving a transactional workload between Regions. The business requires a four-hour recovery time objective, protection from logical corruption, and a rollback point before production cutover. The team has rehearsed restoration once but has not measured dependency recovery or application validation. Which TWO actions should the architect require before approving the migration sequence? Select TWO.

Select two. More than one option is correct — every correct one is ticked below.

  1. Switch DNS when replication reports healthy.
    Replication health does not establish dependency readiness, application recovery time, logical-corruption protection, or a rollback boundary.
  2. Perform a measured restore rehearsal through dependencies and application validation. ✓
    A measured rehearsal establishes whether the four-hour recovery objective is achievable and verifies dependency readiness and application correctness.
  3. Create and retain an immutable point-in-time source backup through destination validation. ✓
    An isolated immutable point-in-time copy protects against logical corruption and preserves a pre-cutover recovery point while the destination is validated.
  4. Delete the source immediately after cutover.
    Immediate deletion removes the rollback boundary before destination validation and increases exposure to migration defects or logical errors.
  5. Rely on an untested backup policy for rollback.
    A backup policy may retain recovery material, but an untested restore does not establish recovery time, application correctness, or an operational rollback sequence.
The trap
Healthy replication is not proof of recoverability. Deleting the source makes rollback unavailable too early. Recovery material without a tested procedure does not prove the objective.

Measure full recovery and retain an immutable pre-cutover recovery point.

5. Enforce authorization in the application: Which option best addresses the limitation?

Hard
A telecommunications firm exposes account-management APIs through an internet-facing application load balancer. Security proposes AWS WAF rules that block requests lacking a customer identifier and concludes that unauthorized account access is prevented. The application supports authenticated sessions and must enforce tenant ownership for every request. The company wants the smallest architectural correction without replacing its edge service. Which option best addresses the limitation?
  1. Enforce authorization in the application. ✓
    Application authorization can evaluate authenticated identity, tenant ownership, requested object, and business permissions for every operation.
  2. Use signed CloudFront URLs for API account requests.
    Signed URLs control configured delivery access but do not replace application identity, session validation, or tenant authorization.
  3. Add more WAF rules for every account path.
    WAF filters request characteristics but cannot reliably determine authenticated identity, tenant ownership, or business authorization decisions.
  4. Enable GuardDuty findings to deny unauthorized account requests.
    GuardDuty detects supported threats; it does not automatically enforce application permissions or authorize individual API operations.
The trap
Filtering request patterns does not enforce object-level authorization. Signed delivery authorization is not full application authorization. Threat detection does not provide per-request authorization enforcement.

WAF can filter traffic, but only application authorization can enforce identity and tenant ownership.

6. Use the Aurora reader endpoint for read connections: Which component is missing?

Easy
A managed service provider operates an Aurora cluster for several customer portals. Monitoring shows writer CPU below 35%, reader CPU below 40%, and database latency rising only when portal traffic increases. Connection logs show that 96% of read sessions terminate on the writer, although two Aurora Replicas are healthy. Applications currently use the cluster endpoint for all database operations. The provider wants to distribute new read connections without changing the database engine or adding a cache. Which component is missing?
  1. Add an RDS Proxy while leaving all application reads on the cluster endpoint.
    RDS Proxy can pool and manage connections, but leaving reads on the cluster endpoint does not provide the required reader routing.
  2. Add more Aurora Replicas while leaving applications connected to the cluster endpoint.
    Additional replicas do not receive read sessions when applications continue targeting the writer-oriented cluster endpoint.
  3. Use the Aurora reader endpoint for read connections. ✓
    The reader endpoint distributes new read connections among available Aurora Replicas instead of directing every connection to the writer.
  4. Export portal read data to S3 and query the objects instead of using Aurora readers.
    S3 object storage does not replace the portals’ transactional relational queries or provide Aurora connection distribution.
The trap
Connection pooling does not change the selected Aurora endpoint. More capacity is ineffective when the connection target is wrong. Object storage is not a relational read-routing component.

Direct read sessions to Aurora’s reader endpoint; healthy replicas cannot help while applications target the cluster endpoint.

7. Add canary alarms with automatic rollback: Which change is the smallest complete solution?

Medium
A university research team deploys a new analysis service through an automated pipeline. The team requires a small canary exposure, automatic rollback when error rate increases, and no manual approval during overnight experiments. The proposed pipeline creates a canary deployment but has no configured alarms or rollback action. Existing application metrics are already published as CloudWatch metrics, and the deployment target supports traffic shifting. Which change is the smallest complete solution?
  1. Require an overnight operator to approve rollback manually.
    Manual approval violates the stated unattended operation requirement and does not provide automatic failure response.
  2. Create a CloudFormation change set before deployment.
    A change set previews infrastructure changes but does not monitor application health or automatically roll back traffic.
  3. Increase the canary duration without adding alarms.
    A longer canary supplies exposure time but cannot detect failure or initiate automatic rollback without configured monitoring.
  4. Add canary alarms with automatic rollback. ✓
    Configured health alarms connected to traffic shifting provide the monitoring and automatic rollback missing from the proposed canary.
The trap
Manual intervention contradicts the overnight automation requirement. Previewing resource changes is not deployment health validation. Duration alone cannot detect or reverse faulty deployments.

Connect existing health metrics to canary rollback alarms so traffic shifts can reverse automatically.

8. Use diversified Spot workers: Which architecture best fits these constraints?

Medium
A travel company runs nightly fare-recalculation jobs that can restart from checkpoints and may be interrupted without customer impact. Demand varies substantially, and the company wants the lowest compute cost while retaining enough capacity for deadlines. Job state must survive instance interruption, and the design cannot depend on one undiversified capacity pool. Which architecture best fits these constraints?
  1. Run all jobs on one large On-Demand instance and copy checkpoints only after the nightly schedule completes.
    A single host creates a capacity bottleneck, and delayed checkpoint copying prevents timely recovery after interruption.
  2. Use Lambda workers with reserved concurrency and store checkpoints in ephemeral execution state.
    Reserved concurrency limits simultaneous executions, while ephemeral execution state cannot reliably preserve progress across interruption or replacement.
  3. Use an Auto Scaling group of only On-Demand instances with local checkpoint files.
    Only On-Demand capacity misses the cost objective, and local checkpoint files may be lost when instances are replaced.
  4. Use diversified Spot workers, On-Demand fallback, and durable checkpoints. ✓
    Diversified Spot capacity lowers cost, fallback capacity protects deadlines, and durable checkpoints allow interrupted jobs to resume safely.
The trap
Concurrency limits do not provide durable batch state. Reliable instance pricing does not satisfy cost or durable-state requirements. One host is neither diversified nor resilient for checkpointed work.

Combine diversified Spot capacity, On-Demand fallback, and durable checkpoints for resilient interruptible batch processing.

9. Map additional accounts to stable groups and scale: Which change best resolves the backlog without violating o

Medium
An acquired subsidiary processes payment events through an SQS FIFO queue. Events for each account must remain ordered, while different accounts should process concurrently. Exhibit: 12,000 messages per minute, 40 currently active message-group IDs per minute, 80 active consumers, increasing oldest-message age, stable processing time, and healthy consumers. The account population supports mapping additional accounts to distinct stable message groups without splitting any account across groups. Which change best resolves the backlog without violating ordering?
  1. Map additional accounts to stable groups and scale consumers from backlog. ✓
    Ordering remains intact within each account’s stable group, while more independently active groups expose parallelism. Backlog-based scaling then adds consumers when queued work exists.
  2. Use an SQS Standard queue.
    Standard queues provide best-effort ordering and cannot satisfy the required per-account ordering guarantee.
  3. Add consumers while retaining only the 40 current groups.
    Consumers beyond the available message groups cannot create additional parallel FIFO work.
  4. Increase the visibility timeout for the existing groups.
    Visibility timeout changes redelivery timing but does not increase the number of independently executable message groups.
The trap
Consumer deduplication does not restore queue ordering. Consumer count cannot exceed useful group-level concurrency. Timeout tuning does not remove the concurrency bottleneck.

Use more stable account groups and scale consumers from backlog.

10. Use warm standby with reduced-capacity instances already: Which strategy should the architect recommend?

Hard
A logistics operator must recover a regional order-processing service within 30 minutes, with a recovery point objective of 15 minutes. The current design keeps replicated data and deployment artifacts ready in a second Region, but application instances require 55 minutes to provision and validate during a quarterly exercise. Management will fund a reduced-capacity continuously running environment, provided the measured recovery objective is met. Which strategy should the architect recommend?
  1. Use backup and restore with automated deployment after detecting the regional outage.
    Restoration and deployment would add work to an already measured 55-minute provisioning path, exceeding the RTO.
  2. Use active-active deployment with full production capacity in both Regions continuously.
    Active-active could reduce interruption, but it exceeds the stated reduced-capacity funding tradeoff without a requirement for full duplicate capacity.
  3. Use warm standby with reduced-capacity instances already running. ✓
    A functional reduced-capacity environment avoids the measured 55-minute provisioning delay and is aligned with the approved continuous-capacity budget.
  4. Use pilot light and provision the application tier on demand during regional failure.
    The measured application provisioning and validation time already exceeds the 30-minute recovery objective.
The trap
Pilot-light provisioning is too slow for the measured RTO. Automation does not make an untested, slower recovery path meet the objective. The strategy overprovisions beyond the approved cost constraint.

Use warm standby because measured provisioning exceeds the RTO and reduced continuous capacity is funded.

11. Grant the workload role KMS use and S3 access separately: Which design is most appropriate?

Easy
An online marketplace stores customer documents in S3 and requires encryption with a customer-managed KMS key. A workload role must upload and download objects, while operators may administer the key but must not read documents. The marketplace has experienced outages caused by incomplete cross-account key permissions and wants a low-maintenance correction without weakening separation of duties. Which design is most appropriate?
  1. Grant the workload role KMS use and S3 access separately. ✓
    Separate KMS usage from S3 permissions lets the workload decrypt objects while operators retain key administration without document access.
  2. Rotate the KMS key.
    Rotation does not grant permissions or repair cross-account authorization relationships.
  3. Add an IAM allow while leaving the cross-account KMS key policy unchanged.
    Cross-account KMS use requires compatible key-policy authorization; an IAM allow alone cannot overcome the key policy.
  4. Give operators S3 read access for key troubleshooting.
    This violates separation of duties and unnecessarily grants document access.
The trap
Cryptographic rotation is not an access-control fix. Key troubleshooting does not require document access. IAM permission alone cannot authorize cross-account key use.

Authorize KMS use and S3 data access independently.

12. Keep classic Multi-AZ and add a read replica for reporting: Which architecture should the architect recommend?

Medium
A global retailer runs a relational order system on a classic RDS Multi-AZ DB instance. It requires automatic standby failover for regional hardware failure and read scaling for a reporting workload. The reporting team currently queries the primary, causing contention during business hours. The retailer must preserve transactional semantics and wants the smallest change that addresses both availability and reporting load. Which architecture should the architect recommend?
  1. Keep classic Multi-AZ and add a read replica for reporting. ✓
    The classic Multi-AZ standby supplies high availability, while a read replica supplies separate read capacity for reporting queries.
  2. Add an Aurora reader without migrating the relational database.
    Aurora readers belong to Aurora clusters, so this does not add read capacity to the existing classic RDS deployment.
  3. Remove Multi-AZ and use only a read replica.
    A read replica supports read scale but does not replace synchronous standby high availability for the primary database instance.
  4. Read from the classic Multi-AZ standby.
    A classic Multi-AZ standby supports failover but cannot serve application read traffic for reporting workloads.
The trap
Aurora reader capacity requires an Aurora architecture. Classic standby instances are not readable. Read scaling does not provide equivalent primary HA.

Retain classic Multi-AZ for failover and add a read replica for reporting traffic.

285 more Design for New Solutions questions

The remaining 285 questions in this domain are part of the full AWS Solutions Architect Professional bank — 1024 questions, every option explained. Start with the free five-minute check and see your score per domain.

Test your AWS Solutions Architect Professional readiness — free

Other AWS Solutions Architect Professional domains

Part of the Certsqill AWS Solutions Architect Professional question bank · Design for New Solutions · Every answer, right and wrong, comes with its own explanation.