Mass deletion
Recursive deletes outside the project, wiped home or system directories, emptied trash and backups.
rm -rf · del /s · Remove-Item -RecurseAgentRisk Lab runs on your machine, stops an AI agent’s destructive actions before they execute, and tells you the moment it does.
The headline above is the first thing the agent’s command reaches. A selection sweeps across the words “safe actions of agents.” and the letters begin to fall. The guard holds the command, the letters freeze in mid-air, then rewind into place, and “safe” is underlined.
09:41. The agent is doing its job in the billing API project, reading and writing files safely.
The agent reports the disk is 97 percent full and blocking the build. The largest directory is the home folder at 412 GB, so it reaches for rm -rf ~/. The command is not malicious. It is missing one piece of context.
The guard holds the command at the syscall boundary, before it touches disk. It does not run.
AgentRisk Lab plays the command forward in a sandbox and measures the result: 14,302 files predicted lost, including 14 SSH keys, a production database with no backup, and 41 uncommitted changes across 3 repos.
The predicted loss runs back to zero. The real files were never touched. Files lost: 0.
You get one alert with the exact command and the rule it matched. You choose Deny.
The agent is refused and tries a narrower command, rm -rf node_modules/.cache. The guard passes it. About 1.2 GB is freed, and nothing else changes.
Files lost: 0, across 14,302 files, 3 repos and 1 production database. Held, measured, un-happened.
Agents now run shell commands, edit databases and push to production on your behalf. Most of them ask first — until one doesn’t. A prompt is a request. An interlock is a guarantee.
Try it
This is a simulation of the interlock. Pick an action an agent might attempt and watch what happens before it executes.
$ agent “clean up the repo and ship the fix”
Pick an action the agent might attempt ↓
Simulation. Real rules and decisions run in the local daemon.
What it stops
Reads and safe writes pass without friction. Anything you can’t undo gets held for a human.
Recursive deletes outside the project, wiped home or system directories, emptied trash and backups.
rm -rf · del /s · Remove-Item -RecurseDrops, truncates and unscoped deletes or updates on anything that isn’t a throwaway test database.
DROP · TRUNCATE · DELETE without WHEREForce-pushes to protected branches, hard resets over uncommitted work, deleted remote branches and tags.
push --force · reset --hard · branch -DReads of keys, tokens and .env files combined with an outbound request to an unknown host.
.env → curl · ~/.ssh · cloud credsPiping downloads into a shell, installing unpinned packages from new registries, editing install hooks.
curl | sh · npm i unknownDestroying stacks, volumes, buckets and clusters, or widening access policies, from a CLI or an API key.
terraform destroy · aws s3 rbHow it works
A small background service sits between the agent and your system. Tool calls, shell commands, file writes and network requests pass through it before they execute.
Each action is checked against your policy and a set of built-in destructive-action detectors. The check is deterministic and local — no model has to agree for a block to hold.
Risky actions are stopped, not logged after the fact. You get an immediate alert with the exact command, the agent that issued it and one tap to allow once or deny.
Policy
Start with sensible defaults. Tighten per project with a plain file that lives in your repo and goes through review like any other code.
# agentrisk.policy.yml default: allow-reads, hold-destructive protect: paths: ["~/", "/etc", "./.git", "./infra"] branches: ["main", "release/*"] databases: ["prod-*"] allow: - "npm test" - "git push origin feature/*" notify: on: [block, repeated-attempt] channels: [desktop, phone, slack] quiet_hours: never # blocks always break through
Questions
No. AgentRisk Lab is in active development and this page is its front door. Early-access members get the first builds and a say in which agents and platforms come first.
The goal is any agent that runs commands or tools on your computer: terminal coding agents, IDE assistants and MCP-based agents. The first supported set will be decided with early-access members.
That is the central design constraint. The guard runs as a separate process from the agent, with its own permissions, and changes to its policy require you, not the agent.
Decisions are made locally. Alerts only leave the machine through channels you configure, and they contain the blocked action, not your files.
Safe actions pass straight through. Only held actions wait, and they wait for you — which is the point.
Pricing isn’t set. Early-access members will hear first, before anything is charged.
Early access
Tell us what you run agents on. We’ll email when there’s a build to try — and not otherwise.