Skip to content
PD
Vibe Coding 8 min read

Sandbox First: How I Run Coding Agents Without Handing Over My Laptop

Coding agents run shell commands, install packages and read your filesystem. Here is the exact sandbox, egress allowlist and dependency gate I use so a bad suggestion costs me a container, not my credentials.

PD

Pavel Duglas

AI Automation & MVP Architect

Last month an agent working on a scraper for me confidently ran npm install puppeteer-stealth-plus. That package does not exist in the form it wanted - but something with a very similar name did, published eleven days earlier, with a postinstall script. The agent didn’t get to run it, because the only way it can add a dependency in my setup is through a wrapper script that refused. That took me twenty minutes to build a year ago and it has now paid for itself several times over.

This is the part of vibe coding nobody puts in the demo video. The demo shows the agent writing a feature in four minutes. It doesn’t show that the same agent has your SSH keys, your ~/.aws/credentials, your .env with the production database URL, and a shell. We’ve now seen agent CLIs quietly uploading local files, agents recommending malicious packages, and agents deciding that the fastest path to a green test suite is rm -rf on something you needed. None of that requires malice from the model. Incompetence at scale is enough.

So: sandbox first. Here’s the setup I actually use.

The threat model, in plain terms

Forget “AI safety” abstractions. There are four concrete bad days:

  1. Credential exfiltration. The agent reads a file with long-lived secrets and sends the contents somewhere - to an API call, to a log, to a package it installed, or into a prompt that gets stored on a vendor’s servers.
  2. Supply chain injection. The agent installs a package that doesn’t do what its name says. Models hallucinate package names, and attackers register the hallucinations. This is a well-known attack surface now.
  3. Destructive local commands. Force pushes, dropped tables, deleted directories, git clean -fdx on a repo with uncommitted work.
  4. Real-world side effects. The agent hits a live API, sends real emails, charges a real card, or logs into a real account with a real session. This is the one that hurts most in automation work - BAS scripts and bots operate on accounts you cannot un-ban.

Everything below is aimed at making each of those cost you a container rather than a week.

Tier your agents by blast radius

I don’t run every agent the same way. Three tiers:

Tier 0 - read and propose. No shell, no writes. It reads the repo, produces a diff or a plan, I apply it. This is where I put anything touching auth, billing, migrations, or infrastructure code. It feels slow. It is slow. That’s the point.

Tier 1 - sandboxed autonomy. Full shell, full write access, inside a disposable container with a scratch copy of the repo, fake credentials and a restricted network. This is where 80% of my agent work happens: features, refactors, tests, parsers, glue code. The agent can do whatever it wants because “whatever it wants” is bounded.

Tier 2 - shared infrastructure. Anything touching staging or prod. This tier does not exist for agents in my setup. Humans only, or an agent producing an artifact that a human ships. If you think this is too conservative, go read a few incident writeups about missing failover in trading systems and ask whether you want an agent in that loop.

The rest of this article is about building Tier 1 properly, because that’s the tier that actually earns money.

The container

Devcontainers, plain Docker, a VM - the tool matters less than the properties. What I want: no host filesystem access beyond one directory, no host credentials, no root, capped resources, and network traffic that goes through something I control.

docker run --rm -it \
  --name agent-box \
  --network agent-net \
  --user 1000:1000 \
  --cap-drop ALL \
  --security-opt no-new-privileges \
  --pids-limit 512 --memory 6g --cpus 4 \
  -v "$PWD/worktrees/feat-parser:/work" \
  -w /work \
  -e HTTPS_PROXY=http://egress:3128 \
  -e HTTP_PROXY=http://egress:3128 \
  --env-file ./secrets/agent.env \
  agent-image:latest

Things to notice:

  • The mount is a git worktree, not the repo. git worktree add worktrees/feat-parser -b feat/parser. The agent gets a real, working checkout on its own branch. If it corrupts the tree, I delete the worktree. My main working copy and my uncommitted changes never enter the container.
  • No Docker socket. Mounting /var/run/docker.sock into an agent container is handing it root on the host. If the agent needs containers, give it a rootless nested runtime or don’t give it containers.
  • --env-file points at agent-specific credentials. Not my shell environment. Never -v ~/.aws:/root/.aws.
  • Resource caps. Agents write infinite loops. A pinned CPU is annoying; a 40 GB memory balloon that takes your laptop down mid-call is worse.

Credentials the agent is allowed to have

Rule: an agent may only hold credentials that I am willing to rotate on a Friday evening without telling anyone.

In practice that means:

  • A scoped LLM API key with a hard monthly spend cap, separate from my main key so I can see exactly what the agent burned.
  • Test-mode keys for every third-party service. Stripe test keys, sandbox payment endpoints, a throwaway Telegram bot token pointed at a private test chat.
  • A local Postgres in the same Docker network, seeded from an anonymized dump. Never a tunnel to staging.
  • No SSH keys. If the agent needs to push, it pushes to a fork or a branch via a fine-grained token that can write to exactly one repository and cannot approve or merge anything.

And a habit that costs nothing: git secrets-style pre-commit scanning inside the container. Agents love to paste keys into config files “as an example”.

Egress allowlist: the highest-value hour you’ll spend

A container with unrestricted internet is a container that can exfiltrate anything it reads. I run a tiny proxy on the agent-net network and allow only what the work needs:

# squid.conf (trimmed)
acl allowed_dst dstdomain registry.npmjs.org pypi.org files.pythonhosted.org
acl allowed_dst dstdomain github.com codeload.github.com objects.githubusercontent.com
acl allowed_dst dstdomain api.openai.com api.anthropic.com
http_access allow allowed_dst
http_access deny all
access_log stdio:/dev/stdout

Two benefits. First, an agent that installs something sketchy can’t phone home to an arbitrary host. Second - and this is what I didn’t expect - the proxy log is the best behavioural telemetry I have. When I tail it during a run I can see exactly what the agent reached for, in order. Half my prompt improvements came from reading that log and realising the agent was fetching documentation for a library I’d already banned.

If a task genuinely needs a broad internet connection (scraping, for example), I run that in a separate container that has network access but no repository and no secrets, and the two talk over a narrow local API.

The dependency gate

This is the piece I’d build first if I could only build one thing. The agent is not allowed to run npm install or pip install directly - the tools are shadowed by a wrapper that fails with a message telling it to use guard-add:

#!/usr/bin/env bash
# guard-add.sh - the only sanctioned way to add an npm dependency
set -euo pipefail
pkg="$1"

meta=$(npm view "$pkg" --json) || { echo "BLOCKED: package not found"; exit 1; }
created=$(jq -r '.time.created' <<<"$meta")
age=$(( ( $(date +%s) - $(date -d "$created" +%s) ) / 86400 ))
weekly=$(curl -sf "https://api.npmjs.org/downloads/point/last-week/$pkg" | jq -r '.downloads // 0')

echo "$pkg  age=${age}d  weekly_downloads=$weekly"
[ "$age" -lt 180 ]     && { echo "BLOCKED: too new, needs human review"; exit 1; }
[ "$weekly" -lt 5000 ] && { echo "BLOCKED: too obscure, needs human review"; exit 1; }

npm install --ignore-scripts --save-exact "$pkg"
echo "$(date -Iseconds) $pkg age=$age dl=$weekly" >> /work/.agent/deps.log

Age and download thresholds catch the overwhelming majority of slopsquatting attempts, because the attacker’s package is by definition new and unpopular. --ignore-scripts removes the postinstall execution path entirely. --save-exact stops floating versions from drifting under you. And deps.log gives me a one-line-per-dependency audit trail to skim during review - which is far faster than reading a 900-line lockfile diff.

When the gate blocks something legitimate, I add it manually in thirty seconds. That’s a fine trade.

Make undo cheap

Sandboxing limits damage; cheap undo limits time lost. Three habits:

  • Auto-commit on a timer. A loop that runs git add -A && git commit -m "wip: $(date -Iseconds)" every two minutes on the agent’s branch. Ugly history, perfect bisect. When the agent breaks something at minute 40, I go back to minute 38 instead of re-running the whole task.
  • Snapshot the database. pg_dump before the agent starts, restore script one command away. Migrations are where agents do their most creative damage.
  • Squash before review. Nobody reviews 200 wip commits. I squash to one diff per logical change and review that. The wip history stays available locally until the branch is merged.

What review looks like when the sandbox is doing its job

With the sandbox in place I stop reviewing for catastrophe and start reviewing for correctness, which is a much better use of my attention. My checklist has four items: does the dependency log contain anything I didn’t expect; does the diff touch files outside the stated scope; do the tests actually assert behaviour rather than existence; and did anything appear in the proxy log that has nothing to do with the task.

That’s it. Four things, five minutes, and the review load stays manageable even when I’m running two or three agents in parallel. The current complaint across engineering teams is that AI has flooded review queues - my experience is that most of that flood is anxiety, not information. When you can’t be sure the agent didn’t touch your credentials, you read every line. When you know it physically couldn’t, you read the parts that matter.

Build the sandbox once. Reuse it on every project. It is the single highest-leverage thing you can do to make agent-assisted development boring, and boring is exactly what you want from infrastructure.

FAQ

Isn't a container overkill for a small side project?

The setup cost is a one-time investment of an hour or two, and after that you copy the same image and compose file into every project. The asymmetry is what matters: the worst case for a small side project without a sandbox is still leaked credentials or a wiped working directory, because agents don't know the difference between a toy repo and a production one. If you truly want the minimum viable version, do just two things: run the agent on a git worktree instead of your main checkout, and shadow the package manager with an install gate. Those two changes alone remove most of the realistic damage.

How do I let an agent work on something that genuinely needs live API access?

Split the work into two containers. One has network access but no repository, no secrets and no write access to anything you care about - it does the fetching and returns raw data over a small local HTTP endpoint. The other has the code and a database but only reaches the internet through your egress allowlist. If the task truly needs a live credential, use a short-lived token scoped to a single operation, set a hard spend or rate cap on the provider side, and watch the run rather than leaving it unattended. Live credentials plus unattended autonomy is the combination that produces incident reports.

Do the age and download-count thresholds in the dependency gate block too much?

In practice they block a few legitimate packages per project, and unblocking one takes half a minute of human review. That's the correct trade, because the packages the thresholds reject are exactly the profile of a slopsquatted or typosquatted package: published recently, almost no downloads, name suspiciously close to something real. Tune the numbers to your ecosystem - 180 days and 5,000 weekly downloads works for npm, PyPI usually needs different values - and keep the `--ignore-scripts` flag regardless of thresholds, since that removes the postinstall execution path that most package-based attacks rely on.

Related articles