AI SecurityHow-To & HardeningNews

How to Sandbox AI Agents: Package Proxies, Egress Allowlists and Kill Switches

By OnCloudSec Research Team · Published Oct 6, 2026 · 1 min read

The OpenAI–Hugging Face incident showed that AI agents can find and exploit weaknesses in their own sandboxes. Organizations running agents — even simple ones — should treat sandbox design as a security control. Here is a practical approach.

Principle 1: Default-deny network egress

  • Agents get no internet access unless a task requires it.
  • When access is needed, use an allowlist of specific domains or endpoints enforced at the network layer (Azure Firewall, AWS Network Firewall, DNS firewall, Entra Internet Access), not just in agent instructions.
  • Log every outbound connection.

Principle 2: Harden the package proxy

The reported escape path ran through a package proxy. Treat proxies and registries as part of the sandbox perimeter:

  • Patch them like internet-facing systems.
  • Restrict them to retrieving approved packages from approved upstreams.
  • Prevent them from being used as general-purpose proxies.
  • Isolate proxies used by agents from those used by production systems.

Principle 3: Isolate agents from each other

Separate runs and agents shouldn't share writable storage, registries or message channels unless designed to. Monitor shared resources for unexpected use.

Principle 4: Scoped, short-lived identities

  • Each agent gets its own identity (Entra Agent ID, IAM role, managed identity) with minimal permissions.
  • No standing credentials inside the sandbox; issue short-lived tokens per task.
  • Never place production credentials in evaluation environments.

Principle 5: Kill switches

  • A tested mechanism to stop agents immediately: revoke identities, cut network access, halt compute.
  • Automatic triggers on anomalies (unexpected egress, privilege changes, volume spikes).

Principle 6: Monitor behavior, not just outputs

Log tool calls, network activity and resource changes. Review anomalies quickly.

Verify

Red-team your sandbox: can an agent reach the internet, other agents or production? Test before attackers — or your agents — do.

Sources

  1. Source
ai agent sandboxingOpenAI agents breach Hugging Face2026

More on this story