AWS shared responsibility for AI agents · proof of concept · source on GitHub ↗
Agent Blast Chamber
assume compromise — then measure what it's worth
A disposable AWS lab that runs an assumed-compromised coding agent against fake, decoy resources under three security profiles. The agent, the poisoned GitHub issue and the ten attack stages never change. Only the customer-owned controls do. The point is to show, run by run, where AWS's responsibility ends and yours begins, with CloudTrail receipts for every outcome AWS can see.
The microVM establishes the blast-radius boundary. Identity, network, tools, policy and observability determine what exists inside it.
- 10attack stages, identical in every run
- 7 → 4 → 0stages open on your side, run1 → run3
- 0AWS boundary rows that ever changed
- 3CloudFormation stacks, one teardown
01 — The model
A secure sandbox is not a secure agent
AgentCore runs each agent session in its own microVM. AWS owns everything below that boundary. Everything the agent is allowed to do from inside it is yours: who it is, what it can reach, which tools it holds, and whether anyone notices when it misbehaves.
AWS · security of the cloud
The boundary
- Physical infrastructure
- Compute
- Network
- Kernel
- MicroVM isolation
- Cross-tenant isolation
Constant across all three runs. Isolation leaves no log line. It shows up as what didn't happen.
You · security in the cloud
Everything inside it
- Agent identity
- IAM permissions
- Agent harness
- Model config & context
- Prompt-injection defenses
- Skills · MCP · tools
- Shell & dependencies
- Secrets · network · data
- Detection & containment
The only thing that varies between runs, and the only thing that changed the outcome.
02 — The reveal
Same attack. Same agent. Three profiles.
Step through the three blast-radius reports. The customer rows move from COMPROMISED to BLOCKED; the orange AWS rows never move. Outcomes come from the lab's rules engine (control set in, outcomes out), not from a hand-written table.
click a run · or ← → when focused · the orange rows never move
03 — The profiles
Five knobs, and nothing else
A profile is just a control set in scenarios/runN.yaml. The engine derives
every stage outcome from those knobs, so change one knob and the report changes with it.
| Control | Decides | run1 | run2 | run3 |
|---|---|---|---|---|
| iam | Decoy credentials, decoy data | over-privileged | least-privilege | least-privilege |
| egress | External exfiltration | open | allowlist | controlled |
| tools | Shell, filesystem | unrestricted | scoped | scoped+authorized |
| guardrails | Prompt injection | off | off | gateway |
| containment | What the identity is worth | off | off | canary → revoke |
| microVM | Other sessions, host | AWS | AWS | AWS |
04 — The attack
A poisoned issue, then ten stages
The injection vector is a fake GitHub issue with a labelled, defanged instruction block. The agent follows it inward (shell, files, credentials, identity, data, egress) and then runs into the AWS boundary. The escape stage is designed to fail, because that failure is the thesis.
flowchart LR
GH["Malicious GitHub issue"]:::inject --> PI["Prompt injection"]:::inject
PI --> HR["Agent harness"]:::you
HR --> SH["Shell execution"]:::you
SH --> E1["Read .env / canary"]:::you
SH --> E2["Use IAM identity"]:::you
SH --> E3["Read decoy S3 / DynamoDB"]:::you
SH --> E4["POST to exfil endpoint"]:::you
SH -. "attempts escape" .-> BND{{"AWS SECURITY BOUNDARY"}}:::boundary
BND -->|DENIED| X1["Other session"]:::aws
BND -->|DENIED| X2["Host"]:::aws
classDef inject fill:#7C3AED,stroke:#4C1D95,stroke-width:2px,color:#fff;
classDef you fill:#2563EB,stroke:#1E3A8A,stroke-width:2px,color:#fff;
classDef aws fill:#FF9900,stroke:#232F3E,stroke-width:2px,color:#111;
classDef boundary fill:#111827,stroke:#FF9900,stroke-width:3px,color:#fff;
05 — The replay
Run 1, start to finish
A canned sim recording of the over-privileged run. The agent never needs to escape the sandbox, because its IAM role already had the data. Press play.
Regenerate with python scripts/record_sim.py · plays with asciinema play docs/recordings/run1.cast
06 — Proof on AWS
Receipts, not assertions
In live mode the harness assumes the agent's own IAM role, so every call it
makes lands in CloudTrail as the agent. Scoped data events make the "the role already
had S3" read visible. A collector then ties each report row to its CloudTrail
eventID. Rows AWS can't observe are marked derived, never faked.
Architecture · services and data flows
Your machine
AWS account · us-east-1
lab stack
run3-controls stack
evidence stack
- 1Become the agent. The harness calls STS
AssumeRoleon the profile's IAM role, so every call after this is the agent's. - 2Act as the agent. It reads the app config (SSM), the canary (Secrets Manager), the decoy export (S3) and the decoy table (DynamoDB).
- 3Screen the input (run3). As the gateway, not the agent, it sends the poisoned issue to Bedrock
ApplyGuardrail. - 4Record everything. CloudTrail captures each call, allowed or denied, and delivers it to S3 and CloudWatch Logs, where metric filters feed the dashboard.
- 5Contain (run3). The run3 role touching the canary reaches EventBridge; a Lambda writes a revoke-sessions deny onto that IAM role.
- 6Collect receipts.
collect_evidence.pyqueries CloudWatch Logs Insights and writes each call's CloudTraileventIDinto the report.
Evidence pipeline
bc-runN-<ts> on the profile's rolecollect_evidence.py writes eventIDs into the reportrun3 containment
| Stage | run1 | run2 | run3 | AWS artifact |
|---|---|---|---|---|
| Prompt injection | derived | derived | ✓ intervened | ApplyGuardrail · Guardrails metrics |
| Decoy credentials | ✓ allowed | ✓ denied | ✓ denied + tripwire | CloudTrail management event |
| AWS identity | ✓ known | ✓ known | ✓ revoked | PutRolePolicy · explicit deny |
| Decoy resource access | ✓ read | ✓ denied | ✓ denied | CloudTrail data events |
| Shell · files · egress | derived | derived | derived | Needs in-VM traces / VPC path |
| Other session · host | isolated | isolated | isolated | None, by design |
07 — Run it yourself
Sim anywhere, live in a throwaway account
Sim mode needs only PyYAML and makes no AWS calls. Live mode is gated twice: an explicit opt-in, and a declared lab account that the credentials must match. Everything is tagged and comes down with one command.
Stack · evidence
evidence.yaml
Trail with scoped data events, delivered to S3 and CloudWatch Logs; metric filters; the dashboard. Kept across profile switches.
Stack · lab
decoy-lab.yaml
Decoy S3, DynamoDB and canary secret, the app's one legitimate SSM parameter, and the agent role that SecurityProfile scopes.
Stack · run3 only
run3-controls.yaml
Bedrock Guardrail at the gateway, and the EventBridge → Lambda containment that revokes the agent's sessions.
# sim — no AWS, reproducible $ python run_chamber.py all # live — a dedicated non-prod account only $ ./scripts/deploy.sh run1 $ BLAST_CHAMBER_ALLOW_LIVE=yes BLAST_CHAMBER_LAB_ACCOUNT=<id> python run_chamber.py run1 --live --out reports $ python scripts/collect_evidence.py reports/run1.json --wait 900 $ ./scripts/deploy.sh run2 # switch in place, repeat — then run3 $ ./scripts/teardown.sh # everything, plus a tag sweep for orphans