Sandboxed adversarial traces with approval logs: Before approval, which evidence most directly resolves that
Adversarial traces with approval logs directly test whether hostile content can trigger payment actions despite model instructions.
The question
A bank proposes an agent that extracts invoice data and can initiate payments. Retrieved documents may contain hostile instructions, and the team has not shown whether prompt rules constrain payment tools. Before approval, which evidence most directly resolves that uncertainty?
Preparing for AIGP? Take the free 5-min readiness quiz →
- A general accuracy report covering invoice extractionExtraction accuracy may support the use case, but it leaves the payment-tool authorization uncertainty unresolved.
- Sandboxed adversarial traces with approval logs ✓These traces test untrusted instructions, tool restrictions, and consequential-action approvals while logs reveal observable behavior and containment points.
- A polished demonstration using trusted invoicesA trusted demonstration does not test hostile retrieved content or establish that payment authorization remains outside model-controlled instructions.
- A vendor statement describing intended agent behaviorIntent statements provide limited assurance but cannot demonstrate effective permission boundaries, approval enforcement, or actual action logging.
The trap
For agents, prioritize observed tool behavior, authorization boundaries, approvals, and logs over demonstrations or assurances. How to remember it
Adversarial traces with approval logs directly test whether hostile content can trigger payment actions despite model instructions.
How many of these would you get right?
One of 1581 AIGP questions on Certsqill. Take a free five-minute check and see your score per domain — not one number, but which section to open tonight.
Test your AIGP readiness — freeMore Understanding How to Govern AI Deployment and Use questions
- Test role- and plant-specific retrieval: What evidence should precede approval? →
- A baseline of current service outcomes: What evidence should precede approval of the replacement? →
- Observed review of representative patches: Which evidence best addresses readiness? →
- All 424 Understanding How to Govern AI Deployment and Use questions →
Part of the Certsqill AIGP question bank · Understanding How to Govern AI Deployment and Use ·
Every answer, right and wrong, comes with its own explanation.