AgentRiskLab

Early access · in active development

We care about safe actions of agents.

AgentRisk Lab runs on your machine, stops an AI agent’s destructive actions before they execute, and tells you the moment it does.

The film: an agent’s command, played forward and held

The headline above is the first thing the agent’s command reaches. A selection sweeps across the words “safe actions of agents.” and the letters begin to fall. The guard holds the command, the letters freeze in mid-air, then rewind into place, and “safe” is underlined.

09:41. The agent is doing its job in the billing API project, reading and writing files safely.

The agent reports the disk is 97 percent full and blocking the build. The largest directory is the home folder at 412 GB, so it reaches for rm -rf ~/. The command is not malicious. It is missing one piece of context.

The guard holds the command at the syscall boundary, before it touches disk. It does not run.

AgentRisk Lab plays the command forward in a sandbox and measures the result: 14,302 files predicted lost, including 14 SSH keys, a production database with no backup, and 41 uncommitted changes across 3 repos.

The predicted loss runs back to zero. The real files were never touched. Files lost: 0.

You get one alert with the exact command and the rule it matched. You choose Deny.

The agent is refused and tries a narrower command, rm -rf node_modules/.cache. The guard passes it. About 1.2 GB is freed, and nothing else changes.

Files lost: 0, across 14,302 files, 3 repos and 1 production database. Held, measured, un-happened.

Agents now run shell commands, edit databases and push to production on your behalf. Most of them ask first — until one doesn’t. A prompt is a request. An interlock is a guarantee.

Try it

Try to make it run something it shouldn’t.

This is a simulation of the interlock. Pick an action an agent might attempt and watch what happens before it executes.

agent-session · ~/projects/billing-api Guard armed

$ agent “clean up the repo and ship the fix”

Pick an action the agent might attempt ↓

!
AgentRisk Lab

Simulation. Real rules and decisions run in the local daemon.

What it stops

Irreversible by default is the line.

Reads and safe writes pass without friction. Anything you can’t undo gets held for a human.

Filesystem

Mass deletion

Recursive deletes outside the project, wiped home or system directories, emptied trash and backups.

rm -rf · del /s · Remove-Item -Recurse
Database

Destructive queries

Drops, truncates and unscoped deletes or updates on anything that isn’t a throwaway test database.

DROP · TRUNCATE · DELETE without WHERE
Git

History rewrites

Force-pushes to protected branches, hard resets over uncommitted work, deleted remote branches and tags.

push --force · reset --hard · branch -D
Secrets

Credential exfiltration

Reads of keys, tokens and .env files combined with an outbound request to an unknown host.

.env → curl · ~/.ssh · cloud creds
Supply chain

Blind execution

Piping downloads into a shell, installing unpinned packages from new registries, editing install hooks.

curl | sh · npm i unknown
Cloud

Infrastructure teardown

Destroying stacks, volumes, buckets and clusters, or widening access policies, from a CLI or an API key.

terraform destroy · aws s3 rb

How it works

Three steps, in this order, every time.

  1. Intercept

    A small background service sits between the agent and your system. Tool calls, shell commands, file writes and network requests pass through it before they execute.

  2. Judge

    Each action is checked against your policy and a set of built-in destructive-action detectors. The check is deterministic and local — no model has to agree for a block to hold.

  3. Hold and notify

    Risky actions are stopped, not logged after the fact. You get an immediate alert with the exact command, the agent that issued it and one tap to allow once or deny.

Local-firstDecisions are made on your machine, offline.
Fail closedIf the guard can’t decide, the action waits.
Agent-agnosticBuilt for terminal agents, IDE agents and MCP servers.
Full audit trailEvery allow, block and override is recorded.

Policy

Rules you can read in one sitting.

Start with sensible defaults. Tighten per project with a plain file that lives in your repo and goes through review like any other code.

Questions

Straight answers.

Is it available today?

No. AgentRisk Lab is in active development and this page is its front door. Early-access members get the first builds and a say in which agents and platforms come first.

Which agents will it work with?

The goal is any agent that runs commands or tools on your computer: terminal coding agents, IDE assistants and MCP-based agents. The first supported set will be decided with early-access members.

Can an agent just turn it off?

That is the central design constraint. The guard runs as a separate process from the agent, with its own permissions, and changes to its policy require you, not the agent.

Does my code or data leave my machine?

Decisions are made locally. Alerts only leave the machine through channels you configure, and they contain the blocked action, not your files.

Will it slow my agents down?

Safe actions pass straight through. Only held actions wait, and they wait for you — which is the point.

What does it cost?

Pricing isn’t set. Early-access members will hear first, before anything is charged.

Early access

Be on the list before the first agent goes rogue.

Tell us what you run agents on. We’ll email when there’s a build to try — and not otherwise.