Docker has documented a GitHub Actions workflow that runs an AI coding agent inside a Docker Sandbox microVM. The example uses GitHub Agentic Workflows (gh-aw) to inspect a Java project, run PostgreSQL-backed integration tests with Testcontainers, correct a seeded bug, and propose the change through a draft pull request.
The design is notable because it separates the GitHub Actions runner from the environment where the agent executes. However, availability is time-sensitive: although Docker’s report identifies docker-sbx as an integrated gh-aw runtime in version 0.82.9, the current gh-aw runtime reference marks that runtime as deprecated and says it is expected to be removed in a future release.
What changed
gh-aw lets teams author an agent workflow in Markdown with YAML frontmatter. The tool then compiles that source into a conventional GitHub Actions workflow with a .lock.yml suffix. In the Docker example, .github/workflows/sandbox-explorer.md is the editable source and .github/workflows/sandbox-explorer.lock.yml is the generated workflow.
The sample selects the awf agent and the docker-sbx runtime on ubuntu-24.04. It grants the agent edit and Bash capabilities, enables sudo, and configures a network allowlist containing defaults, github, containers, and java. The workflow is started with workflow_dispatch.
Its GitHub permissions are limited to contents: read and copilot-requests: write, and it uses the Copilot engine. The safe-output settings create a draft pull request, apply a title prefix, block protected files, and limit proposed changes to src/**.
Docker reports that the example ran the agent in a Sandbox microVM, executed Java and PostgreSQL Testcontainers integration tests, fixed a case-insensitive-email bug, and opened the draft pull request. The reported execution time of 11 minutes and 16 seconds is an observation from that sample, not a general performance promise.
How the execution boundary works
There are three distinct layers in this arrangement:
- The GitHub-hosted Actions runner provides the workflow execution environment and GitHub Actions integration.
- The Docker Sandbox microVM provides the environment in which the AI agent runs.
- A private Docker Engine inside the sandbox supports Docker-dependent work such as Testcontainers without giving the agent access to the host runner’s Docker daemon.
According to Docker’s Sandbox security documentation, the sandbox has a separate kernel, filesystem, network boundary, and Docker Engine. The agent has full privileges inside the VM, including the ability to use sudo and install packages, but it cannot access the host Docker daemon except through explicitly shared resources.
That boundary is not the same as unrestricted isolation. A directly shared workspace is read-write and exposes changes to the host. Docker also documents a clone mode with a private in-VM clone and a read-only repository mount, as well as a mountless mode with no host workspace mount. Shared files, network access, credentials, and host integrations therefore remain important parts of the security design.
Credentials, networking, and permissions
Docker documents a credential-isolation model in which a host-side proxy injects authentication headers while raw credential values remain outside the VM. Its documented CI workflow uses a Docker Personal Access Token with at least Read scope, passed through sbx login --username and --password-stdin. Sandboxes can then be created and managed with sbx create, sbx exec or sbx run, and sbx rm.
Docker also documents importing all available secrets with sbx secret import --all or setting selected secrets for CI credentials. The supplied documentation does not establish exactly how those mechanisms map to the credentials used by the gh-aw sample, so teams should not assume that the example represents every supported credential path.
Outbound TCP traffic is deny-by-default in Docker’s documented network model. Allowed destinations are proxied through the host, while direct external UDP and ICMP traffic are blocked. The sample’s allowlist should therefore be treated as configuration for that workflow, not as a universal list for every agent project.
For Copilot-based workflows, GitHub documentation says gh-aw uses GITHUB_TOKEN by default and requires copilot-requests: write. The organization must also allow Copilot CLI usage billed to the organization. The available documentation does not fully establish the requirements or behavior of the alternative COPILOT_GITHUB_TOKEN path.
Runner and platform requirements
The Docker Sandbox installation documentation lists Ubuntu 24.04 or later, a supported 64-bit Intel, AMD, or Arm processor, KVM hardware virtualization, and membership in the kvm group for Linux. Nested virtualization is required when the runner itself operates inside a VM or VDI environment.
The gh-aw runtime reference additionally identifies Linux, reachable Docker, sufficient CPU, memory, and disk, outbound HTTPS access to GitHub and the selected AI provider, access to ghcr.io, and any domains required by setup and network policy. Its description of docker-sbx also lists KVM, nested virtualization, sudo, apt, Docker Hub credentials, and local Docker.
This creates an availability qualification for the reported example. Docker’s article describes a successful run on a GitHub-hosted ubuntu-24.04 runner, while the current runtime documentation describes substantial virtualization and host prerequisites. The supplied sources do not explain how every capability was provided by that hosted run. Developers should not treat the example as proof that every GitHub-hosted runner, GitHub plan, repository type, or runner image supports the integration.
Docker separately documents Sandbox installation on macOS Sonoma 14 or later on Apple silicon and Windows 11 on 64-bit Intel or AMD systems with Windows Hypervisor Platform. Those installation options do not establish that the gh-aw docker-sbx CI runtime works on macOS or Windows. The runtime reference also identifies ARC DinD as incompatible with docker-sbx, rather than as an interchangeable runner configuration.
What you should do
- Check the current gh-aw runtime documentation before starting. The current reference marks
docker-sbxdeprecated and says it will be removed in a future release. - Keep the Markdown workflow as the editable source and regenerate the lock file with
gh aw compileafter changing its frontmatter. Compilation can catch unsupported values and known incompatible combinations. - Confirm Linux, Docker, KVM, nested virtualization,
sudo,apt, resource, credential, and outbound-network prerequisites for the intended runner. - Use the narrowest permissions and network destinations that the workflow needs. Retain a restricted safe-output boundary such as
src/**where it fits the project. - Review every draft pull request and inspect changes when a workspace is shared directly. Files modified by an agent can affect build scripts, CI configuration, hooks, or other code that is later executed.
- Do not combine the documented
docker-sbxruntime with ARC DinD based on the current compatibility guidance.
Availability remains the main caveat
The Docker example demonstrates a useful pattern: give an agent broad freedom inside a microVM while limiting its repository output, network destinations, permissions, and pull-request behavior. It does not establish a universal deployment model or a formal security guarantee.
The historical claim that version 0.82.9 introduced the integration comes from Docker’s report; independent release-history confirmation and a current reproduction of the sample were not established in the supplied sources. The exact enforcement details for safe-output file restrictions and draft pull requests are also undocumented. Most importantly, the current gh-aw reference’s deprecation notice means teams should treat the example as a time-sensitive implementation to evaluate, not as a stable long-term recommendation.



