AWS shared responsibility for AI agents · proof of concept · source on GitHub ↗

Agent Blast Chamber

assume compromise — then measure what it's worth

A disposable AWS lab that runs an assumed-compromised coding agent against fake, decoy resources under three security profiles. The agent, the poisoned GitHub issue and the ten attack stages never change. Only the customer-owned controls do. The point is to show, run by run, where AWS's responsibility ends and yours begins, with CloudTrail receipts for every outcome AWS can see.

The microVM establishes the blast-radius boundary. Identity, network, tools, policy and observability determine what exists inside it.

01 — The model

A secure sandbox is not a secure agent

AgentCore runs each agent session in its own microVM. AWS owns everything below that boundary. Everything the agent is allowed to do from inside it is yours: who it is, what it can reach, which tools it holds, and whether anyone notices when it misbehaves.

YOUsecurity in the agent
Prompt · context · untrusted inputissues, READMEs, tool output: anything the agent readsrun3
Model configurationsystem prompt, model choice
Guardrails & policyat the gateway, outside the agentrun3
Agent harnesswhat it may runrun3
Tools · MCP · skillsscoped + authorizedrun3
Shell & dependencieswhat's installed in the VM
Agent identity & IAM permissionsthe execution role: what "owning the VM" is worthrun2
Secretscanary, credentialsrun2
Network egressallowlistrun2
DataS3, DynamoDBrun2
Observability · detection · containmentcanary tripwire → revoke the sessionrun3
Language runtime & base imageYOU with your own container image · AWS with direct code deploy
Trust boundary · the microVMAWS's side of the line. The lab asserts it in every run and never attacks it
AWSsecurity of the sandbox
MicroVM isolationone VM per session
Session isolationother agents' sessions
Cross-tenant isolationother customers
Kernel · hypervisor · host
Compute
Storage
Networking
Physical infrastructureRegions · Availability Zones · edge locations
The classic model, restacked for agents. A run2 or run3 tag marks the lab profile that turns that layer's control on. Every tagged layer sits above the boundary.

AWS · security of the cloud

The boundary

  • Physical infrastructure
  • Compute
  • Network
  • Kernel
  • MicroVM isolation
  • Cross-tenant isolation

Constant across all three runs. Isolation leaves no log line. It shows up as what didn't happen.

You · security in the cloud

Everything inside it

  • Agent identity
  • IAM permissions
  • Agent harness
  • Model config & context
  • Prompt-injection defenses
  • Skills · MCP · tools
  • Shell & dependencies
  • Secrets · network · data
  • Detection & containment

The only thing that varies between runs, and the only thing that changed the outcome.

02 — The reveal

Same attack. Same agent. Three profiles.

Step through the three blast-radius reports. The customer rows move from COMPROMISED to BLOCKED; the orange AWS rows never move. Outcomes come from the lab's rules engine (control set in, outcomes out), not from a hand-written table.

7
open on your side
AWS boundary — HELD ✓ (all runs)
BROAD
blast radius

click a run · or ← → when focused · the orange rows never move

03 — The profiles

Five knobs, and nothing else

A profile is just a control set in scenarios/runN.yaml. The engine derives every stage outcome from those knobs, so change one knob and the report changes with it.

ControlDecidesrun1run2run3
iamDecoy credentials, decoy dataover-privilegedleast-privilegeleast-privilege
egressExternal exfiltrationopenallowlistcontrolled
toolsShell, filesystemunrestrictedscopedscoped+authorized
guardrailsPrompt injectionoffoffgateway
containmentWhat the identity is worthoffoffcanary → revoke
microVMOther sessions, hostAWSAWSAWS

04 — The attack

A poisoned issue, then ten stages

The injection vector is a fake GitHub issue with a labelled, defanged instruction block. The agent follows it inward (shell, files, credentials, identity, data, egress) and then runs into the AWS boundary. The escape stage is designed to fail, because that failure is the thesis.

flowchart LR
    GH["Malicious GitHub issue"]:::inject --> PI["Prompt injection"]:::inject
    PI --> HR["Agent harness"]:::you
    HR --> SH["Shell execution"]:::you
    SH --> E1["Read .env / canary"]:::you
    SH --> E2["Use IAM identity"]:::you
    SH --> E3["Read decoy S3 / DynamoDB"]:::you
    SH --> E4["POST to exfil endpoint"]:::you
    SH -. "attempts escape" .-> BND{{"AWS SECURITY BOUNDARY"}}:::boundary
    BND -->|DENIED| X1["Other session"]:::aws
    BND -->|DENIED| X2["Host"]:::aws
    classDef inject fill:#7C3AED,stroke:#4C1D95,stroke-width:2px,color:#fff;
    classDef you fill:#2563EB,stroke:#1E3A8A,stroke-width:2px,color:#fff;
    classDef aws fill:#FF9900,stroke:#232F3E,stroke-width:2px,color:#111;
    classDef boundary fill:#111827,stroke:#FF9900,stroke-width:3px,color:#fff;

05 — The replay

Run 1, start to finish

A canned sim recording of the over-privileged run. The agent never needs to escape the sandbox, because its IAM role already had the data. Press play.

agent-blast-chamber — run1 (sim)

Regenerate with python scripts/record_sim.py · plays with asciinema play docs/recordings/run1.cast

06 — Proof on AWS

Receipts, not assertions

In live mode the harness assumes the agent's own IAM role, so every call it makes lands in CloudTrail as the agent. Scoped data events make the "the role already had S3" read visible. A collector then ties each report row to its CloudTrail eventID. Rows AWS can't observe are marked derived, never faked.

Architecture · services and data flows

Your machine

operatorrun_chamber.pyharness · gateway123
operatorcollect_evidence.pyCloudTrail receipts6

AWS account · us-east-1

lab stack

security · identityAWS STSAssumeRole → bc-runN session1
security · identityIAM roleagent-exec-runN, scoped per profile15
managementSSM Parameter Storethe app's one legitimate call2
security · identitySecrets Managercanary secret (tripwire)25
storageAmazon S3decoy bucket, fake export2
databaseAmazon DynamoDBdecoy customer table2

run3-controls stack

machine learningBedrock Guardrailsprompt-attack filter at the gateway3
integrationAmazon EventBridgecanary tripwire rule5
computeAWS Lambdacontainment: revoke sessions5

evidence stack

managementAWS CloudTrailmanagement + scoped data events45
storageAmazon S3trail bucket4
managementCloudWatch Logstrail + Lambda log groups46
managementCloudWatchmetric filters + dashboard4
  1. 1Become the agent. The harness calls STS AssumeRole on the profile's IAM role, so every call after this is the agent's.
  2. 2Act as the agent. It reads the app config (SSM), the canary (Secrets Manager), the decoy export (S3) and the decoy table (DynamoDB).
  3. 3Screen the input (run3). As the gateway, not the agent, it sends the poisoned issue to Bedrock ApplyGuardrail.
  4. 4Record everything. CloudTrail captures each call, allowed or denied, and delivers it to S3 and CloudWatch Logs, where metric filters feed the dashboard.
  5. 5Contain (run3). The run3 role touching the canary reaches EventBridge; a Lambda writes a revoke-sessions deny onto that IAM role.
  6. 6Collect receipts. collect_evidence.py queries CloudWatch Logs Insights and writes each call's CloudTrail eventID into the report.

Evidence pipeline

sts:AssumeRole
Act as the agent
Session bc-runN-<ts> on the profile's role
cloudtrail
Mgmt + data events
Scoped to the decoy bucket & table
cloudwatch logs
Metric filters
Reads, denials, canary hits, split by run
dashboard
Three runs, one view
Bars + the per-call evidence table
logs insights
Receipts
collect_evidence.py writes eventIDs into the report

run3 containment

bedrock
Gateway guardrail
ApplyGuardrail screens the issue before the agent sees it
secretsmanager
Canary touched
Any GetSecretValue by the run3 role
eventbridge
Tripwire
Read-only management events enabled
lambda
Revoke sessions
Deny * for tokens issued before now
proof
Explicit deny
The app's own call now fails
Stagerun1run2run3AWS artifact
Prompt injectionderivedderived✓ intervenedApplyGuardrail · Guardrails metrics
Decoy credentials✓ allowed✓ denied✓ denied + tripwireCloudTrail management event
AWS identity✓ known✓ known✓ revokedPutRolePolicy · explicit deny
Decoy resource access✓ read✓ denied✓ deniedCloudTrail data events
Shell · files · egressderivedderivedderivedNeeds in-VM traces / VPC path
Other session · hostisolatedisolatedisolatedNone, by design

07 — Run it yourself

Sim anywhere, live in a throwaway account

Sim mode needs only PyYAML and makes no AWS calls. Live mode is gated twice: an explicit opt-in, and a declared lab account that the credentials must match. Everything is tagged and comes down with one command.

Stack · evidence

evidence.yaml

Trail with scoped data events, delivered to S3 and CloudWatch Logs; metric filters; the dashboard. Kept across profile switches.

Stack · lab

decoy-lab.yaml

Decoy S3, DynamoDB and canary secret, the app's one legitimate SSM parameter, and the agent role that SecurityProfile scopes.

Stack · run3 only

run3-controls.yaml

Bedrock Guardrail at the gateway, and the EventBridge → Lambda containment that revokes the agent's sessions.

# sim — no AWS, reproducible
$ python run_chamber.py all

# live — a dedicated non-prod account only
$ ./scripts/deploy.sh run1
$ BLAST_CHAMBER_ALLOW_LIVE=yes BLAST_CHAMBER_LAB_ACCOUNT=<id> python run_chamber.py run1 --live --out reports
$ python scripts/collect_evidence.py reports/run1.json --wait 900
$ ./scripts/deploy.sh run2   # switch in place, repeat — then run3
$ ./scripts/teardown.sh      # everything, plus a tag sweep for orphans