Sandbox backends

Note

Experimental: this can change or be removed in a minor release of this provider. See Stable and experimental features.

OpenSandbox (self-hosted remote)

OpenSandboxBackend runs sandboxes through an OpenSandbox server. The server may use Docker or Kubernetes; Airflow workers only use its HTTP API and do not need access to the container runtime.

Install the SDK extra:

pip install "apache-airflow-providers-common-ai[opensandbox]"

Use a generic Airflow connection, resolved lazily on first use:

from airflow.providers.common.ai.sandbox import OpenSandboxBackend
from airflow.providers.common.ai.toolsets import SandboxToolset

SandboxToolset(OpenSandboxBackend(opensandbox_conn_id="opensandbox_default"))

The connection host is required; port is optional, schema defaults to http, and password carries the API key when required. Extras may set request_timeout (default 30 seconds) and use_server_proxy (default true). Set opensandbox_conn_id=None to let the SDK read OPEN_SANDBOX_DOMAIN and OPEN_SANDBOX_API_KEY.

SandboxSpec.env is sent at creation. A default spec sends a deny-all network policy; allow_egress_to becomes explicit allow rules. With block_network=False, the backend omits network policy entirely so a deployment without the egress sidecar can still run an intentionally open sandbox. Deny/allowlist policy requires the sidecar, so the backend reads the enforced policy back after creation and destroys the sandbox if it does not match the requested spec. allow_egress_to_cidrs is refused: OpenSandbox only enforces CIDR targets in dns+nft mode, and the Python SDK does not expose that enforcement mode on policy read-back, so this backend cannot prove the address-layer restriction is active.

Every sandbox carries created-by: airflow metadata and an airflow-sandbox-* name for attribution and cleanup. The server enforces a sandbox lifetime (default 3600 seconds). If the SDK event stream stalls, the worker abandons the call after the command budget plus a grace period, destroys the sandbox, and reports sandbox_terminated so the toolset provisions a fresh one. Output is bounded per stream after the SDK yields it; a single newline-free line is the SDK-level exception, because the SDK assembles that line before the backend sees it.

Constructor parameters:

  • image: image used by the server. Default "python:3.12-slim".

  • cpu and memory: resource limits. Defaults "1" and "2Gi".

  • sandbox_timeout: server-side lifetime in seconds. Default 3600.

  • ready_timeout: provisioning/reconnect timeout. Default 120.

  • use_server_proxy: override the connection extra for file and command calls.

The runtime remains a deployment choice. The default Docker runtime shares the host kernel; choose a stronger runtime such as Kata when your threat model needs a VM boundary.

sbx (Docker Sandboxes, local)

SbxSandboxBackend runs each sandbox in a Docker Sandboxes microVM by driving the sbx CLI. Each sandbox is a real microVM with its own kernel.

Warning

Use this backend for local development, not production. Docker Sandboxes is built for running coding agents against a checkout on your own machine, and driving it from an Airflow worker is off-label use. A production worker would need the sbx binary on the host, an authenticated Docker account (sbx login), a one-time sbx policy init, and on Linux, KVM or nested virtualization, which a worker in an unprivileged container cannot provide.

Orphans are not reclaimed. There is no server-side lifetime. If the worker is killed outright, the microVM and its workspace directory survive; sandboxes are named airflow-sandbox-* so an operator can find and remove them.

Installing the CLI is a Deployment Manager prerequisite (brew install docker/tap/sbx or winget install Docker.sbx); the backend needs no Python dependency. The template image must provide GNU coreutils timeout, base64, stat, head, find, mkdir and dirname, which any Debian or Ubuntu based image has.

Constructor parameters:

  • image: Container image for the sandbox. Default "python:3.12-slim".

  • memory: Memory limit in binary units. sbx enforces a 1 GiB minimum. Default "2g".

  • cpus: CPUs to allocate. None (default) uses the sbx default, which is every host CPU.

  • sbx_path: Path to the sbx binary. Default "sbx".

  • create_timeout: Seconds allowed for provisioning; a first-run microVM boot plus an image pull can be slow. Default 600.

  • host_network_policy: What sbx policy is set to on this host. "unknown" (default) makes create refuse any spec asking for a network guarantee this backend cannot make, and since block_network defaults to True that includes a bare SandboxSpec(). Set "deny-all" after running sbx policy init deny-all, or "allow-all" to state that egress is open and pass SandboxSpec(block_network=False) to match.

What differs between the backends

Swapping the backend is one constructor argument, and tool names, spec and prompt do not change. Four behaviours do, so read them before assuming the same Dag behaves identically everywhere:

  • CPU. sbx gives a sandbox every host CPU; Modal defaults to a request of 0.125 of one, so set cpu; OpenSandbox takes cpu as a limit the server enforces.

  • Egress allowlists. sbx enforces allow_egress_to at the host policy layer; Modal matches TLS handshake names, which is weaker and has to be opted into; OpenSandbox enforces it in an egress sidecar, and the backend reads the enforced policy back rather than trusting the create request. allow_egress_to_cidrs is enforced at the address layer on Modal, refused on sbx, and refused by OpenSandbox because its SDK cannot prove that the sidecar is running in the dns+nft mode required for CIDR enforcement.

  • Command timeouts. A timeout destroys an sbx sandbox and its files; Modal and a server-enforced OpenSandbox timeout preserve the sandbox and files. OpenSandbox destroys it only if the command event stream itself stalls past the client-side grace period.

  • Symlinks. write_file through a symlink follows the link on sbx and replaces it on Modal and OpenSandbox.

  • Attaching. A Modal sandbox can be provisioned by one task and used by an agent in another (A sandbox another task owns). An sbx microVM lives on the worker that created it and cannot be reached from another task, and OpenSandbox has no per-sandbox metadata the ownership rules could be kept in, so both refuse SandboxSpec.owner and the toolset refuses attach_to for them.

Bringing your own backend

Any vendor that can create a sandbox, run a command in it and destroy it can plug in. Subclass SandboxBackend in your own package and pass an instance to SandboxToolset.

Three methods are required: create, run_command and destroy. The file operations ship as defaults implemented over run_command, because reading, writing, listing and exporting a file are all expressible as shell commands. The default export_file, behind SandboxToolset(exports=...), copies a file in 4 MiB slices, one command each, and needs stat, tail, head and base64 in the guest. It relies on run_command returning each slice’s output intact, or setting stdout_truncated when it could not. Override it when the vendor can stream a download, as the sbx and OpenSandbox backends do. Override the others only when the vendor has a native file API:

from airflow.providers.common.ai.sandbox import (
    SandboxBackend,
    SandboxExecResult,
    SandboxSpec,
)


class AcmeSandboxBackend(SandboxBackend):
    name = "acme"

    def create(self, *, spec: SandboxSpec | None = None) -> str:
        return acme_sdk.create_sandbox().id

    def run_command(self, sandbox, command, *, timeout, max_output_bytes):
        r = acme_sdk.exec(sandbox, command, timeout=timeout)
        return SandboxExecResult(exit_code=r.exit_code, stdout=r.stdout, stderr=r.stderr)

    def destroy(self, sandbox) -> None:
        acme_sdk.delete_sandbox(sandbox)

    # Optional: inherited from SandboxBackend unless the vendor has
    # something better than shelling out.
    def read_file(self, sandbox, path, *, max_bytes) -> bytes:
        return acme_sdk.download(sandbox, path, limit=max_bytes)

Four rules for an implementation:

  • Constructors run at Dag-parse time, so resolve credentials lazily, on first use.

  • destroy must be idempotent; destroying an already-gone sandbox is not an error.

  • Raise SandboxTerminalError when retrying cannot help and SandboxError when it might. The first fails the task for Airflow to retry; the second becomes a bounded prompt back to the model.

  • If you cannot enforce something the SandboxSpec asks for, raise. Never provision a weaker sandbox than the Dag author asked for.

If your sandboxes can be found again from another process, subclass AttachableSandboxBackend instead, and a @task can provision a sandbox for an agent task to attach to (A sandbox another task owns). It adds two methods, read_tags and write_tags, over whatever key-value metadata the vendor keeps on a sandbox, and the ownership rules are written once on the base class on top of them. Three things the base class relies on: read_tags raises SandboxTerminalError for a sandbox that does not exist or has ended; write_tags replaces the whole set, because releasing a claim is a rewrite without the holder key; and create stamps SandboxSpec.owner under OWNER_TAG, the sandbox’s end time as Unix seconds under EXPIRES_AT_TAG, and the network policy under NETWORK_TAG using encode_network_policy(spec), all importable from airflow.providers.common.ai.sandbox.base. If the vendor lets a caller of your backend set metadata too, make your reserved keys overwrite theirs. Stamping the working directory under WORKDIR_TAG is optional; without it the attaching backend asks the sandbox. The toolset shortens run_command to the remaining lifetime itself; a backend clamps only if its own file operations need it.

Was this entry helpful?