This specification defines the configuration model, processing rules, and environment semantics for the Agentic Workflow Firewall (AWF). It is the normative reference for:
- the
awfCLI runtime (--config) - tooling that compiles workflows into AWF invocations (e.g.,
gh-aw) - IDE and static-analysis validation via JSON Schema
The machine-readable schema is published alongside this specification at
docs/awf-config.schema.json (live, tracking main) and as a versioned
release asset (e.g.,
https://github.com/github/gh-aw-firewall/releases/download/v0.23.1/awf-config.schema.json).
This document is normative. Informative notes are marked with Note: or placed in blockquotes. All other text is normative unless stated otherwise.
The key words "MUST", "MUST NOT", "REQUIRED", "SHALL", "SHALL NOT", "SHOULD", "SHOULD NOT", "RECOMMENDED", "MAY", and "OPTIONAL" in this document are to be interpreted as described in RFC 2119.
A conforming AWF configuration document is one that:
- is valid JSON or YAML;
- satisfies all constraints defined by
docs/awf-config.schema.json; and - contains no properties beyond those defined by the schema (closed-world assumption).
A conforming AWF implementation MUST accept every conforming configuration document and MUST reject every non-conforming one.
When the user invokes awf --config <path|-> -- <command>, a conforming
implementation MUST execute the following steps in order:
- If
<path>is-, read configuration bytes from standard input. - Determine the serialisation format:
- If
<path>ends with.json, parse as JSON. - If
<path>ends with.yamlor.yml, parse as YAML. - Otherwise, attempt JSON first; if that fails, attempt YAML.
- If
- Validate the parsed document against
docs/awf-config.schema.json. - On validation failure, abort with non-zero exit status (see §7).
- Map configuration fields to CLI-option semantics per §5.
- Apply precedence rules per §3.
The effective value for any configuration parameter SHALL be determined by the following precedence order (highest wins):
- Explicit CLI flags
- Config file (
--config) - AWF internal defaults
Note: This model enables reusable, checked-in configuration files with environment-specific CLI overrides.
The root object of a conforming configuration document MAY contain the following top-level properties. All are OPTIONAL:
| Property | Type | Description |
|---|---|---|
$schema |
string | JSON Schema URI for IDE validation |
network |
object | Network egress configuration |
filesystem |
object | Host filesystem write-boundary configuration (see §4.1) |
apiProxy |
object | API proxy sidecar configuration |
security |
object | Security and isolation settings |
container |
object | Container and Docker settings |
cloudHypervisor |
object | Cloud Hypervisor v53.0 microVM preview settings (see §4.2) |
chroot |
object | Chroot execution overrides for split-filesystem ARC/DinD runners |
dind |
object | Bootstrap helpers for ARC/DinD split runner/daemon filesystems |
runner |
object | Runner topology declaration (standard vs. ARC/DinD) |
environment |
object | Environment variable propagation (see §8) |
logging |
object | Logging and diagnostics |
rateLimiting |
object | Egress rate limiting |
platform |
object | GitHub platform deployment type declaration |
enclaves |
object | Unified private-repository enclave subsystem (see §14) |
Property-level constraints, types, and descriptions are defined
normatively by docs/awf-config.schema.json.
When filesystem.allowWrite is present, AWF MUST mount existing writable host
binds read-only except at the listed guest-visible absolute paths. The list
narrows existing access: it MUST NOT expose a new host path or make an
otherwise read-only mount writable. Every listed path MUST exist within an
existing writable host mount. An empty list makes all non-internal host bind
mounts read-only.
AWF-owned agent log and session-state mounts and required virtual devices remain
writable so the sandbox can operate. filesystem.allowWrite is supported by the
Docker and gVisor compose runtimes and by the Cloud Hypervisor microVM runtime,
where it is enforced by the host mount tree that backs each virtio-fs export
(see docs/cloud-hypervisor-foundation.md).
Cloud Hypervisor has no always-writable internal mounts, so every export it
publishes is subject to the policy, including /tmp/gh-aw and the guest home
directory at /workspace/.awf-home; paths not covered by allowWrite become
read-only. Because every listed path MUST already exist, and because AWF MUST
NOT auto-create or exempt the guest home, a Cloud Hypervisor workload that needs
a writable home MUST have the backing host directory
$GITHUB_WORKSPACE/.awf-home created before AWF starts and MUST then list the
guest path /workspace/.awf-home in allowWrite; otherwise planning fails
because the path does not exist within a writable export. AWF rejects
filesystem.allowWrite with the sbx runtime and with Docker-in-Docker agent
execution.
Any directory the workload writes to at run time — including caller-owned ones
under /tmp/gh-aw such as gh-aw's repo-memory directory /tmp/gh-aw/repo-memory
— MUST appear in allowWrite, or writes to it fail with EROFS. To make that
diagnosable, the Cloud Hypervisor runtime logs the resolved boundary (per export:
writable, read-only, or read-only with named writable paths) before the guest
boots.
The cloudHypervisor surface binds Cloud Hypervisor v53.0, virtiofsd, the
PCI-capable guest kernel, rootfs, and shared AWF guest supervisor to one
release-pinned GitHub-attested manifest. It requires explicit
--cloud-hypervisor-preview opt-in plus
container.containerRuntime: "cloud-hypervisor" to execute a workload. The supported host target is
GitHub-hosted Ubuntu x86_64 runners with KVM. Runner eligibility, missing KVM
access, and unsupported host-policy failures warn and fall back to the standard
Docker backend; invalid configuration and artifact trust, integrity, digest, or
version failures remain fatal. Host eligibility is enforced by
src/cloud-hypervisor/host-eligibility.ts.
See
src/cloud-hypervisor/preflight.ts
for the artifact/host trust-check module,
src/cloud-hypervisor/launcher.ts
for the secure host launcher (network-namespace join, privilege drop, Landlock
filesystem confinement, and seccomp),
src/cloud-hypervisor/manager.ts for
the VM lifecycle, and guest/cloud-hypervisor/
for the guest artifact build/verification pipeline. See
docs/cloud-hypervisor-foundation.md for
the full architecture and security-boundary writeup.
cloudHypervisor.mountPolicy controls host directory exposure:
workspace-onlyis the default and recommended secure posture. It does not infer a mount fromRUNNER_TOOL_CACHEorAGENT_TOOLSDIRECTORY. The workspace and the narrow gh-aw runtime directories (RUNNER_TEMP/gh-awand/tmp/gh-aw) remain eligible when present so generated workflows can run.workspace-and-tool-cacheexplicitly opts into exporting the whole path named byRUNNER_TOOL_CACHE, falling back toAGENT_TOOLSDIRECTORY. AWF requires that value to name an existing real directory and stages it recursively read-only on the host before virtiofsd starts. Cache paths that overlap writable export sources, missing paths, and unknown policy values are errors.
AWF forwards RUNNER_TOOL_CACHE or AGENT_TOOLSDIRECTORY into the guest only
when the corresponding export exists. gh-aw versions that generate a command
which scans RUNNER_TOOL_CACHE must emit
--cloud-hypervisor-mount-policy workspace-and-tool-cache; workflows that do
not need runner-installed tools should retain the default.
Release test artifacts (cloud-hypervisor-test-x86_64) are x86_64
test/preview artifacts built and verified by both the release workflow and
test-cloud-hypervisor.yml,
which also runs the live-KVM parity/security smoke suite
(scripts/ci/cloud-hypervisor-live-smoke.sh) on GitHub-hosted Ubuntu x86_64
runners when triggered by manual dispatch or the cloud-hypervisor-kvm pull
request label. They are published as explicitly named release/workflow
assets, but are not production defaults and are never auto-downloaded. See
docs/cloud-hypervisor-foundation.md
for the complete CI workflow specification and troubleshooting reference.
Releases also publish distinct enclave-script-rootfs.ext4 and
enclave-agent-rootfs.ext4 artifacts with a separately attested enclave
manifest, per-role provenance bundles, and per-role SBOMs. The release setup
script verifies those bindings before exporting unambiguous cached role paths.
These artifacts do not enable Cloud Hypervisor enclave execution by themselves:
the runtime remains terminal until every ADR 0002 host-executor gate is present,
and enclaves[].image remains invalid for runtime: cloud-hypervisor.
Normal execution requires artifactManifestPath,
artifactManifestBundlePath, and artifactReleaseTag. AWF verifies the
Sigstore bundle offline with gh attestation verify, constraining the signer
to github/gh-aw-firewall/.github/workflows/release.yml, before parsing any
manifest field. It then verifies all five local artifact names and SHA-256
digests. The legacy sha256 object is not a trust root.
Same-run development builds may set
developmentAllowUnattestedArtifacts: true only when the process environment
also contains
AWF_CLOUD_HYPERVISOR_DEVELOPMENT_ALLOW_UNATTESTED_ARTIFACTS=1. This
preview-only dual opt-in still requires all five hashes and must not be used
for release artifacts.
This section is normative.
Tools generating AWF invocations (such as gh-aw) SHOULD use the mapping
below. The left side is the configuration-document path; the right side is
the corresponding CLI flag.
Security-sensitive values (API keys, tokens, and credential secrets) MUST be
provided via environment variables, not AWF config documents. Non-sensitive
AWF settings MAY be supplied via config files, including stdin (--config -).
network.allowDomains[]→--allow-domains <csv>network.blockDomains[]→--block-domains <csv>network.dnsServers[]→--dns-servers <csv>network.subnet→--network-subnet <cidr>(relocates theawf-netDocker network; use when the default172.30.0.0/24collides with the host or cluster network, e.g. the OpenShift service CIDR172.30.0.0/16)network.upstreamProxy→--upstream-proxynetwork.isolation→--network-isolation(experimental; enforces egress via Docker network topology instead of host iptables)network.verifySbxEgress→--verify-sbx-egress(fail-closed verification that Docker sbx direct traffic cannot bypass Squid; requires the sbx runtime)network.topologyAttach[]→--topology-attach <name>(repeatable; requiresnetwork.isolation: true)apiProxy.enabled→--enable-api-proxy([DEPRECATED] API proxy is always enabled; this flag is ignored)apiProxy.caCert→--api-proxy-ca-cert <path>(mounts an additional CA certificate into the api-proxy sidecar and setsNODE_EXTRA_CA_CERTSfor upstream TLS verification)apiProxy.enableTokenSteering→--enable-token-steering(maps toAWF_ENABLE_TOKEN_STEERING; omit or set tofalseto opt out)apiProxy.anthropicAutoCache→--anthropic-auto-cacheapiProxy.anthropicCacheTailTtl→--anthropic-cache-tail-ttl <5m|1h>apiProxy.hostedWeb.claude→ (config-only; maps toAWF_CLAUDE_HOSTED_WEB_POLICY— AWF-owned domain policy for Anthropic-hostedweb_search_*/web_fetch_*server tools; see §9.8 Claude Hosted Web Search and Fetch)apiProxy.hostedWeb.codex→ (config-only; maps toAWF_CODEX_HOSTED_WEB_POLICY— AWF-owned domain policy for OpenAI Responsesweb_searchand Codex/v1/alpha/search; see §9.9 Codex/OpenAI Hosted Web Policy)apiProxy.maxEffectiveTokens→ (config-only; no CLI equivalent)apiProxy.maxAiCredits→ (config-only; maps toAWF_MAX_AI_CREDITS)apiProxy.defaultAiCreditsPricing→ (config-only; maps toAWF_DEFAULT_AI_CREDITS_PRICING)apiProxy.providers→ (config-only; maps toAWF_API_PROXY_PROVIDERS)apiProxy.modelMultipliers→--max-model-multiplier <model:multiplier,...>apiProxy.defaultModelMultiplier→ (config-only; maps toAWF_EFFECTIVE_TOKEN_DEFAULT_MODEL_MULTIPLIER)apiProxy.maxTurns→ (config-only; no CLI equivalent)apiProxy.maxRuns→ (deprecated alias formaxTurns; maps toAWF_MAX_RUNS)apiProxy.maxModelMultiplierCap→--max-model-multiplier-cap <number>apiProxy.maxCacheMisses→--max-cache-misses <number>apiProxy.maxPermissionDenied→--max-permission-denied <number>apiProxy.requestedModel→ (config-only; maps toAWF_REQUESTED_MODELfor pre-startup validation)apiProxy.modelFallback→ (config-only; model fallback strategy)apiProxy.fallbackModels→ (config-only; maps toAWF_FALLBACK_MODELS— ordered model IDs retried on 5xx, timeout, or model-not-supported failures)experimental.modelRouting→ (config-only; experimental opt-in required forapiProxy.routing; defaults to off)apiProxy.routing→ (config-only; requiresexperimental.modelRouting: true; task-level routing objective and task conversation input)apiProxy.routing.candidateModels→ (optional glob patterns that limit router/classifier choices; intersected with the model policy; defaults toapiProxy.allowedModels)apiProxy.modelRouter.providerType→ (config-only; maps toCOPILOT_PROVIDER_TYPE)apiProxy.modelRouter.baseUrl→ (config-only; maps toCOPILOT_PROVIDER_BASE_URL)apiProxy.allowedModels→ (config-only; maps toAWF_ALLOWED_MODELS— JSON array of glob patterns; only matching models are permitted)apiProxy.disallowedModels→ (config-only; maps toAWF_DISALLOWED_MODELS— JSON array of glob patterns; matching models are rejected with HTTP 403)apiProxy.models→ (config-only; model alias rewriting)apiProxy.logging.debugTokens→ (config-only; maps toAWF_DEBUG_TOKENS)apiProxy.logging.tokenLogDir→ (config-only; maps toAWF_TOKEN_LOG_DIR)apiProxy.diagnostics.captureBlockedRequests→ (config-only; maps toAWF_CAPTURE_BLOCKED_LLM_REQUESTS)apiProxy.diagnostics.maxCapturedBytes→ (config-only; maps toAWF_MAX_BLOCKED_CAPTURE_BYTES)apiProxy.auth.type→ (config-only; maps toAWF_AUTH_TYPE)apiProxy.auth.provider→ (config-only; maps toAWF_AUTH_PROVIDER)apiProxy.auth.oidcAudience→ (config-only; maps toAWF_AUTH_OIDC_AUDIENCE)apiProxy.auth.azureTenantId→ (config-only; maps toAWF_AUTH_AZURE_TENANT_ID)apiProxy.auth.azureClientId→ (config-only; maps toAWF_AUTH_AZURE_CLIENT_ID)apiProxy.auth.azureScope→ (config-only; maps toAWF_AUTH_AZURE_SCOPE)apiProxy.auth.azureCloud→ (config-only; maps toAWF_AUTH_AZURE_CLOUD)apiProxy.auth.awsRoleArn→ (config-only; maps toAWF_AUTH_AWS_ROLE_ARN)apiProxy.auth.awsRegion→ (config-only; maps toAWF_AUTH_AWS_REGION)apiProxy.auth.awsRoleSessionName→ (config-only; maps toAWF_AUTH_AWS_ROLE_SESSION_NAME)apiProxy.auth.gcpWorkloadIdentityProvider→ (config-only; maps toAWF_AUTH_GCP_WORKLOAD_IDENTITY_PROVIDER)apiProxy.auth.gcpServiceAccount→ (config-only; maps toAWF_AUTH_GCP_SERVICE_ACCOUNT)apiProxy.auth.gcpScope→ (config-only; maps toAWF_AUTH_GCP_SCOPE)apiProxy.auth.anthropicFederationRuleId→ (config-only; maps toAWF_AUTH_ANTHROPIC_FEDERATION_RULE_ID)apiProxy.auth.anthropicOrganizationId→ (config-only; maps toAWF_AUTH_ANTHROPIC_ORGANIZATION_ID)apiProxy.auth.anthropicServiceAccountId→ (config-only; maps toAWF_AUTH_ANTHROPIC_SERVICE_ACCOUNT_ID)apiProxy.auth.anthropicWorkspaceId→ (config-only; maps toAWF_AUTH_ANTHROPIC_WORKSPACE_ID)apiProxy.auth.anthropicTokenUrl→ (config-only; maps toAWF_AUTH_ANTHROPIC_TOKEN_URL)apiProxy.targets.<provider>.host→--<provider>-api-target(exceptantigravity.host, which maps to the Gemini flag below). Accepts a bare hostname or a full URL; a bare hostname or explicithttps://URL both dial the target over HTTPS on port 443 (the default), while an explicithttp://URL dials it in cleartext on port 80 — matching the runner-side allowlist, which already adds anhttp://-scoped allowlist entry for such targets. Custom ports are not supported:http://hostalways dials port 80 andhttps://hostalways dials port 443.)apiProxy.targets.antigravity.host→--gemini-api-targetapiProxy.targets.copilot.extraHeaders→ (config-only; non-sensitive supplemental BYOK headers, maps toAWF_BYOK_EXTRA_HEADERS)apiProxy.targets.copilot.extraBodyFields→ (config-only; non-sensitive supplemental BYOK body fields, maps toAWF_BYOK_EXTRA_BODY_FIELDS)apiProxy.targets.copilot.sessionId→ (config-only; opt-inx-session-idheader /session_idbody field for Copilot BYOK requests, maps toAWF_PROVIDER_SESSION_ID. Never auto-derived fromGITHUB_RUN_ID.)apiProxy.targets.openai.basePath→--openai-api-base-pathapiProxy.targets.openai.authHeader→--openai-api-auth-headerapiProxy.targets.openai.baseUrlEnv→--openai-base-url-env(names a runner environment variable holding a secret OpenAI-compatible base URL; see §9.7 Secret-Backed OpenAI Target)apiProxy.targets.anthropic.basePath→--anthropic-api-base-pathapiProxy.targets.anthropic.authHeader→--anthropic-api-auth-headerapiProxy.targets.gemini.basePath→--gemini-api-base-pathapiProxy.targets.antigravity.basePath→--gemini-api-base-path- When both
apiProxy.targets.antigravityandapiProxy.targets.geminiare set,antigravitytakes precedence per field. apiProxy.targets.vertex.host→--vertex-api-targetapiProxy.targets.vertex.basePath→--vertex-api-base-pathsecurity.legacySecurity→--legacy-securitysecurity.securityMode→--security-mode <strict|compat>([DEPRECATED] Usesecurity.legacySecurityinstead)security.sslBump→--ssl-bumpsecurity.enableDlp→--enable-dlpsecurity.enableHostAccess→--enable-host-accesssecurity.allowHostPorts→--allow-host-portssecurity.allowHostServicePorts→--allow-host-service-portssecurity.difcProxy.host→--difc-proxy-hostsecurity.difcProxy.caCert→--difc-proxy-ca-certcontainer.memoryLimit→--memory-limitcontainer.pidsLimit→--pids-limitcontainer.agentTimeout→--agent-timeoutcontainer.enableDind→--enable-dindcontainer.workDir→--work-dircontainer.containerWorkDir→--container-workdircontainer.images→ (config-only; a closed compiler-authorized manifest of literal, registry-qualifiedtag@sha256:<digest>OCI references. Supported keys aresquid,agent,apiProxy,router,cliProxy,buildTools,dohProxy,enclaveScript,enclaveAgent,enclaveMcpServer, anddindStaging. Every image AWF runs — including consumers outside Docker Compose such as DinD staging,awf predownload --config, and rootless artifact repair — resolves through this manifest, and the effective per-role references are recorded inimage-manifest.json. AWF rejects missing enabled roles and never falls back to the official registry. It cannot be combined with controls that would select a different image:container.imageRegistry,container.imageTag,container.agentImage,container.buildLocal,security.sslBump(requires a locally built Squid image),runner.sysrootImage,dind.stagingImage, or per-enclave image overrides. Registry credentials are intentionally not configured by AWF; use a pre-authenticated Docker daemon.)container.imageRegistry→--image-registrycontainer.imageTag→--image-tagcontainer.skipPull→--skip-pullcontainer.buildLocal→--build-localcontainer.agentImage→--agent-imagecontainer.tty→--ttycontainer.dockerHost→--docker-hostcontainer.dockerHostPathPrefix→--docker-host-path-prefixcontainer.runnerToolCachePath→ (config-only; checked first for optional read-only runner tool cache mount, beforeRUNNER_TOOL_CACHEand/home/runner/work/_toolauto-detection)container.mounts[]→-v, --mount(repeatable; each array entry maps to one Docker volume mount in/host_path:/container_path[:ro|rw]format (both paths must be absolute; host path must exist); in chroot mode, container paths are automatically prefixed with/host)container.containerRuntime→--container-runtime(user-facing runtime name:"gvisor"for an OCI runtime in Compose,"sbx"for a Docker sbx microVM,"cloud-hypervisor"for the explicit Cloud Hypervisor v53.0 workload preview (GitHub-hosted Ubuntu x86_64 KVM runners only; see §4.2), or"nvx"for the explicit NVX/OpenVMM one-shot workload preview (Linux x86_64 KVM-only; see docs/nvx-security-design.md). gVisor translates to"runsc"and injectsextra_hostsfor its DNS workaround. For sbx, Cloud Hypervisor, and NVX, infrastructure stays in Compose while the primary agent runs in a microVM.)filesystem.allowWrite[]→ (config-only; no CLI equivalent; narrows existing writable host binds to the listed guest-visible absolute paths, see §4.1)cloudHypervisor.previewEnabled→--cloud-hypervisor-preview(requirescontainer.containerRuntime: "cloud-hypervisor"or at least oneenclaves[].runtime: "cloud-hypervisor"entry, plus a GitHub-hosted Ubuntu x86_64 KVM runner to execute a workload)cloudHypervisor.mountPolicy→--cloud-hypervisor-mount-policy(workspace-onlyby default; useworkspace-and-tool-cacheonly when the workload needs the runner tool cache)cloudHypervisor.cloudHypervisorBinary→--cloud-hypervisor-binarycloudHypervisor.kernelPath→--cloud-hypervisor-kernelcloudHypervisor.rootfsPath→--cloud-hypervisor-rootfscloudHypervisor.supervisorPath→--cloud-hypervisor-supervisorcloudHypervisor.artifactManifestPath→--cloud-hypervisor-artifact-manifestcloudHypervisor.artifactManifestBundlePath→--cloud-hypervisor-artifact-manifest-bundlecloudHypervisor.artifactReleaseTag→--cloud-hypervisor-artifact-release-tagcloudHypervisor.developmentAllowUnattestedArtifacts→--cloud-hypervisor-development-allow-unattested-artifacts(development only; also requires the matching environment variable and complete legacy hashes)cloudHypervisor.vcpuCount→--cloud-hypervisor-vcpuscloudHypervisor.memoryMib→--cloud-hypervisor-memory-mibcloudHypervisor.apiTimeoutMs→--cloud-hypervisor-api-timeout-mscloudHypervisor.sha256.cloudHypervisor→--cloud-hypervisor-binary-sha256(development bypass only)cloudHypervisor.sha256.virtiofsd→--cloud-hypervisor-virtiofsd-sha256cloudHypervisor.sha256.kernel→--cloud-hypervisor-kernel-sha256cloudHypervisor.sha256.rootfs→--cloud-hypervisor-rootfs-sha256cloudHypervisor.sha256.supervisor→--cloud-hypervisor-supervisor-sha256nvx.previewEnabled→--nvx-preview(requirescontainer.containerRuntime: "nvx"; execution never falls back to Docker, Cloud Hypervisor, or another runtime on failure)nvx.layerPath→--nvx-layer(guest distro layer; required)nvx.artifactManifestPath→--nvx-artifact-manifest(required)nvx.artifactManifestBundlePath→--nvx-artifact-manifest-bundle(required)nvx.signerWorkflow→--nvx-signer-workflownvx.openvmmPath→--nvx-openvmm(required)nvx.kernelPath→--nvx-kernel(required)nvx.initramfsPath→--nvx-initramfs(required)nvx.memoryMib→--nvx-memory-mib(default 512)nvx.memoryMaxBytes→--nvx-memory-max-bytes(default 512 MiB)nvx.pidsMax→--nvx-pids-max(default 128)nvx.scratchBytes→--nvx-scratch-bytesnvx.mountPolicy→--nvx-mount-policy("workspace-only"(default) exports$GITHUB_WORKSPACEinto the guest at/workspaceread-write;"workspace-and-tool-cache"additionally exports the runner tool cache read-only.filesystem.allowWritenarrows the workspace export, andcontainer.containerWorkDirmust resolve inside/workspace.)chroot.binariesSourcePath→ (config-only; mounts a runner-side binaries directory at/tmp/awf-runner-bininside chroot mode and prepends it toPATH)chroot.identity.home→ (config-only; forwarded asAWF_CHROOT_IDENTITY_HOMEand applied after chroot pivot)chroot.identity.user→ (config-only; forwarded asAWF_CHROOT_IDENTITY_USERand applied toUSER/LOGNAMEafter chroot pivot)chroot.identity.uid→ (config-only; forwarded asAWF_CHROOT_IDENTITY_UIDfor chroot user mapping)chroot.identity.gid→ (config-only; forwarded asAWF_CHROOT_IDENTITY_GIDfor chroot user mapping)dind.preStageDirs→ (config-only; enables daemon-side pre-staging of the DinD work directory tree before compose startup)dind.workDir→ (config-only; daemon-visible staging root, default/tmp/gh-aw)dind.stagingImage→ (config-only; image used for short-lived DinD staging containers)dind.stageEngineBinary.path→ (config-only; runner-side engine binary source path for DinD staging)dind.stageEngineBinary.targetPath→ (config-only; daemon-side destination path for staged engine binary)environment.envFile→--env-fileenvironment.envAll→--env-allenvironment.excludeEnv[]→--exclude-env(repeatable)logging.logLevel→--log-levellogging.diagnosticLogs→--diagnostic-logslogging.auditDir→--audit-dirlogging.proxyLogsDir→--proxy-logs-dirlogging.sessionStateDir→--session-state-dirrateLimiting.enabled: false→--no-rate-limitrateLimiting.requestsPerMinute→--rate-limit-rpmrateLimiting.requestsPerHour→--rate-limit-rphrateLimiting.bytesPerMinute→--rate-limit-bytes-pmrateLimiting.maxGithubApiPointsRest→--max-github-api-points-rest(requiressecurity.difcProxy.host)rateLimiting.maxGithubApiPointsGraphql→--max-github-api-points-graphql(requiressecurity.difcProxy.host)rateLimiting.maxNumToolCalls→--max-num-tool-calls(requiresenclaves, see §14.2a)
GitHub API point budgets apply to the entire AWF run and are enforced only by
the protected gh CLI proxy enabled through security.difcProxy.host. REST
read requests consume one point, REST mutations consume five points, and
GraphQL requests consume one point, matching GitHub's secondary-rate-limit
point accounting. A command that would exceed its REST or GraphQL budget is
rejected before it is executed; the CLI proxy writes a
github_api_points_limited structured audit record with the API kind, used
points, configured limit, and remaining budget.
- (no config equivalent) →
--reflect(CLI-only; starts AWF, queries the API proxy/reflectendpoint, and prints its JSON response instead of running a command; mutually exclusive with a command argument) platform.type→ (config-only; maps toAWF_PLATFORM_TYPE)runner.topology→ (config-only; sets runner deployment model —standardorarc-dind; whenarc-dind, enables sysroot staging and emits RUNNER_TOOL_CACHE warnings)runner.sysrootImage→ (config-only; sysroot init-container image forarc-dindtopology; defaults to<container.imageRegistry>/build-tools:<container.imageTag>, wherecontainer.imageRegistrydefaults toghcr.io/github/gh-aw-firewall)enclaves[]→ (config-only; no CLI equivalent, see §14)enclaves[].repos[]→ (config-only; no CLI equivalent, see §14)enclaves[].timeout→ (config-only; no CLI equivalent, see §14)enclaves[].runtime→ (config-only; no CLI equivalent, see §14)enclaves[].image→ (config-only; no CLI equivalent, see §14)enclaves[].memoryLimit→ (config-only; no CLI equivalent, see §14)enclaves[].cpuLimit→ (config-only; no CLI equivalent, see §14)enclaves[].pidsLimit→ (config-only; no CLI equivalent, see §14)enclaves[].tmpfsLimit→ (config-only; no CLI equivalent, see §14)enclaves[].maxOutputBytes→ (config-only; no CLI equivalent, see §14)enclaves[].maxInvocations→ (config-only; no CLI equivalent, see §14)enclaves[].script→ (config-only; no CLI equivalent, see §14)enclaves[].script.maxScriptBytes→ (config-only; no CLI equivalent, see §14)enclaves[].agent→ (config-only; no CLI equivalent, see §14)enclaves[].agent.engine→ (config-only; no CLI equivalent, see §14)enclaves[].agent.profile→ (config-only; no CLI equivalent, see §14)enclaves[].agent.model→ (config-only; no CLI equivalent, see §14)enclaves[].agent.maxTaskBytes→ (config-only; no CLI equivalent, see §14)enclaves[].agent.maxModelRequests→ (config-only; no CLI equivalent, see §14)enclaves[].agent.maxModelTokens→ (config-only; no CLI equivalent, see §14)enclaves[].agent.github.cli→ (config-only; deprecated closed legacy profile, see §14.3)enclaves[].agent.tools.github→ (config-only; closed GitHub MCP tool contract, see §14.3)enclaves[].dynamic→ (config-only; no CLI equivalent; requires compiler handoff; see §14.1a)
When container.dockerHostPathPrefix points at a daemon-visible shared /tmp path, the implementation stages the invoking CLI binary together with /etc/passwd, /etc/group, and the generated chroot /etc/hosts under that shared path so chroot mode can bootstrap on split-filesystem ARC/DinD hosts.
When DinD is detected, AWF preserves the detected DOCKER_HOST value for the agent environment (including MCP servers) so DinD-aware tooling can reach the correct daemon without manual workflow env overrides.
security.allowHostPorts (--allow-host-ports) is accepted together with
security.enableHostAccess (--enable-host-access) in strict security mode
(the default, without --legacy-security), but it does not provide a direct
route to raw-protocol GitHub Actions services: containers. Strict topology
intentionally omits the agent's host.docker.internal mapping and host-access
iptables bypass; use legacy security or a separately verified tunnel for direct
service clients.
security.allowHostServicePorts (--allow-host-service-ports), which relies
on host iptables, remains suppressed in strict mode.
The following CLI flag has no config-file equivalent by design:
-e, --env <KEY=VALUE>— inject a single environment variable into the agent container (repeatable; CLI-only)
A conforming implementation MUST accept --config - to read configuration
from standard input, enabling programmatic and pipeline scenarios.
On parse or validation failure, a conforming implementation MUST:
- exit with a non-zero status code;
- emit a diagnostic message identifying the location and nature of the error; and
- refrain from partial execution of the agent command.
This section is normative.
The agent container's environment is constructed by merging variables from multiple sources. This section defines the merge order and exclusion rules.
Note: For usage guidance, examples, and troubleshooting, see docs/environment.md.
Variables from the following sources are merged in order of increasing precedence. A value set at a higher level MUST override the same-named value from any lower level.
| Level | Source | Description |
|---|---|---|
| 1 (lowest) | AWF-reserved | Proxy routing, DNS, container paths |
| 2 | --env-all |
Inherited host environment (when enabled) |
| 3 | --env-file |
Variables read from a file |
| 4 (highest) | -e / --env |
Explicit CLI key-value pairs |
A conforming implementation MUST set the following variables in the agent
container regardless of user configuration. Values from --env-all and
--env-file MUST NOT override these variables.
| Variable | Value | Purpose |
|---|---|---|
HTTP_PROXY |
http://<squid-ip>:3128 |
Squid forward proxy for HTTP |
HTTPS_PROXY |
http://<squid-ip>:3128 |
Squid forward proxy for HTTPS |
https_proxy |
http://<squid-ip>:3128 |
Lowercase alias (Yarn 4, undici, Corepack) |
NO_PROXY |
localhost,127.0.0.1,::1,... |
Loopback and container IPs bypassing Squid |
SQUID_PROXY_HOST |
squid-proxy |
Proxy hostname (for tools requiring host separately) |
SQUID_PROXY_PORT |
3128 |
Proxy port |
PATH |
(container default) | MUST use the container's PATH, not the host's |
HOME |
(host user's home) | Derived via sudo-aware detection |
Note: Lowercase
http_proxyis intentionally NOT set. Certain curl builds on Ubuntu 22.04 ignore uppercaseHTTP_PROXYfor HTTP URLs (httpoxy mitigation), causing HTTP traffic to fall through to iptables DNAT interception — the intended defense-in-depth behavior.
The following variables MUST be excluded from --env-all and --env-file
passthrough. A conforming implementation MUST NOT inherit them from the host:
| Category | Variables |
|---|---|
| System | PATH, PWD, OLDPWD, SHLVL, _, SUDO_COMMAND, SUDO_USER, SUDO_UID, SUDO_GID |
| Proxy | HTTP_PROXY, HTTPS_PROXY, http_proxy, https_proxy, NO_PROXY, no_proxy, ALL_PROXY, all_proxy, FTP_PROXY, ftp_proxy |
| Actions runtime credentials | ACTIONS_RUNTIME_TOKEN, ACTIONS_RESULTS_URL, ACTIONS_ID_TOKEN_REQUEST_URL, ACTIONS_ID_TOKEN_REQUEST_TOKEN |
| AWF internal controls | AWF_PREFLIGHT_BINARY, AWF_ENSURE_USR_LOCAL_BIN, AWF_GEMINI_ENABLED |
Note: Host proxy variables are read for upstream proxy auto-detection (see
--upstream-proxy) but MUST NOT propagate into the agent container. AWF sets its own proxy variables pointing to Squid.
When --env-all is NOT active, a conforming implementation SHOULD forward
the following host variables into the agent container:
| Category | Variables |
|---|---|
| GitHub authentication | GITHUB_TOKEN, GH_TOKEN, GITHUB_PERSONAL_ACCESS_TOKEN |
| GitHub enterprise | GITHUB_SERVER_URL, GITHUB_API_URL |
| Docker client | DOCKER_HOST, DOCKER_TLS, DOCKER_TLS_VERIFY, DOCKER_CERT_PATH, DOCKER_CONFIG, DOCKER_CONTEXT, DOCKER_API_VERSION, DOCKER_DEFAULT_PLATFORM |
| User environment | USER, XDG_CONFIG_HOME |
When --env-all IS active, all host variables not in the excluded set
(§8.3) SHALL be forwarded, subject to credential isolation rules (§9).
Actions OIDC request variables MUST be forwarded directly to the api-proxy
sidecar when apiProxy.auth.type is github-oidc and MUST NOT be forwarded
to the agent through any environment input path.
Variables passed via -e / --env MUST override values from --env-all
and --env-file.
Reserved proxy routing variables MAY be overridden only via -e / --env.
Other AWF-reserved variables and source credentials protected by credential
isolation (§9) MUST NOT be overridden.
Note: There is no config-file equivalent for
-e/--env. Individual environment variable injection is a runtime concern, not a static configuration concern.
This section is normative.
AWF implements defense-in-depth credential isolation for LLM API keys.
Behavior is governed by the value of apiProxy.enabled.
Note: For architectural diagrams and protocol-level details, see docs/authentication-architecture.md.
A conforming implementation MUST recognize the following environment variables as source credentials — real API keys read from the host:
| Variable | Provider |
|---|---|
OPENAI_API_KEY |
OpenAI |
ANTHROPIC_API_KEY |
Anthropic (Claude) |
COPILOT_GITHUB_TOKEN |
GitHub Copilot — enables sidecar routing to api.githubcopilot.com (CAPI BYOK / offline mode) |
COPILOT_PROVIDER_API_KEY |
GitHub Copilot BYOK provider key (e.g., Azure OpenAI / OpenRouter API key); independently enables sidecar routing — typically combined with COPILOT_PROVIDER_BASE_URL to point at an arbitrary upstream |
GEMINI_API_KEY |
Google Gemini |
GOOGLE_API_KEY |
Google Vertex AI |
The following secondary aliases SHOULD also be recognized:
OPENAI_KEY, CODEX_API_KEY, CLAUDE_API_KEY.
When the API proxy sidecar is enabled, the following rules apply:
-
Source credentials (§9.1) MUST NOT be exposed in the agent container's environment. They SHALL be passed exclusively to the API proxy sidecar.
-
The
--env-allflag MUST NOT reintroduce excluded credentials into the agent environment. -
A conforming implementation MAY inject placeholder values into the agent container for tool compatibility (e.g.,
OPENAI_API_KEY=sk-placeholder-for-api-proxy). Placeholder values are not secrets and MUST NOT be treated as credentials. -
A conforming implementation MUST inject proxy-routing variables so that agent tools reach the sidecar rather than upstream APIs:
Agent variable Value Purpose OPENAI_BASE_URLhttp://172.30.0.30:10000Routes OpenAI calls to sidecar ANTHROPIC_BASE_URLhttp://172.30.0.30:10001Routes Anthropic calls to sidecar COPILOT_API_URLhttp://172.30.0.30:10002Routes Copilot calls to sidecar GOOGLE_GEMINI_BASE_URLhttp://172.30.0.30:10003Routes Gemini calls to sidecar GEMINI_API_BASE_URLhttp://172.30.0.30:10003Alias for compatibility GOOGLE_VERTEX_BASE_URLhttp://172.30.0.30:10004Routes Vertex AI calls to sidecar -
The API proxy sidecar SHALL inject the real credentials into upstream requests. Sidecar port assignments: 10000 (OpenAI), 10001 (Anthropic), 10002 (Copilot), 10003 (Gemini), 10004 (Vertex AI).
-
A conforming implementation MUST forward the following OpenTelemetry variables from the host into the api-proxy sidecar container so that the sidecar can participate in the distributed trace established by the workflow:
Variable Description GH_AW_OTLP_ENDPOINTSJSON array of {url, headers}objects for fan-out export to multiple OTLP collectors. Takes priority overOTEL_EXPORTER_OTLP_ENDPOINT.OTEL_EXPORTER_OTLP_ENDPOINTOTLP/HTTP collector URL. Single-endpoint fallback when GH_AW_OTLP_ENDPOINTSis absent.OTEL_EXPORTER_OTLP_HEADERSComma-separated key=valueauth headers for the OTLP endpoint. Only used withOTEL_EXPORTER_OTLP_ENDPOINT.OTEL_SERVICE_NAMEService name tag. Defaults to awf-api-proxywhen not set.GITHUB_AW_OTEL_TRACE_IDW3C trace-id of the parent workflow trace. GITHUB_AW_OTEL_PARENT_SPAN_IDW3C span-id of the parent workflow span. These variables are NOT forwarded to the agent container via this mechanism; the agent receives OTEL variables through the standard
OTEL_*prefix forwarding described in §8.4.The sidecar selects its exporter using the following priority order:
GH_AW_OTLP_ENDPOINTS(JSON array) — spans are exported concurrently to all listed endpoints (fan-out mode); partial failures on individual endpoints do not block others.OTEL_EXPORTER_OTLP_ENDPOINT(single URL) — legacy single-endpoint mode.- Neither set — the sidecar writes span NDJSON to
/var/log/api-proxy/otel.jsonlas a local fallback.
When
GITHUB_AW_OTEL_TRACE_ID/GITHUB_AW_OTEL_PARENT_SPAN_IDare present and valid hex, each sidecar span is created as a child of the specified parent span, enabling end-to-end distributed tracing from the GitHub Actions workflow through the api-proxy to the LLM provider.
apiProxy.enabled is deprecated and ignored. The API proxy sidecar is always started; there is no disabled mode. Setting apiProxy.enabled: false in a config file is silently ignored for backward compatibility. The CLI flag --enable-api-proxy is similarly ignored; --no-enable-api-proxy is rejected at runtime with an error.
Credential isolation described in §9.2 therefore always applies: source credentials are never forwarded directly to the agent container.
This constraint is normative for tools generating AWF configurations.
Because the API proxy sidecar is always active, source credentials (§9.1) are always excluded from the agent environment and held in the sidecar. A conforming implementation MUST NOT rely on environment.excludeEnv to suppress API keys — the sidecar handles exclusion automatically.
Tools that compile AWF configurations (e.g., gh-aw) MUST ensure that when an LLM agent requires an API key (OpenAI, Anthropic, Gemini, etc.), the real key is held by the sidecar and a placeholder is injected for tool compatibility.
Real credentials forwarded to the agent — GitHub tokens (GITHUB_TOKEN, GH_TOKEN) and any non-LLM credentials — MUST
be protected by the one-shot-token mechanism. Protected tokens are cached
on first access and removed from /proc/self/environ to prevent
environment variable inspection.
The default protected token list is:
COPILOT_GITHUB_TOKEN, GITHUB_TOKEN, GH_TOKEN, GITHUB_API_TOKEN,
GITHUB_PAT, GH_ACCESS_TOKEN, OPENAI_API_KEY, OPENAI_KEY,
ANTHROPIC_API_KEY, ANTHROPIC_AUTH_TOKEN, CLAUDE_API_KEY,
CODEX_API_KEY, COPILOT_PROVIDER_API_KEY
Placeholder compatibility values (§9.2 item 3) are not secrets. However,
provider credential variable names such as ANTHROPIC_AUTH_TOKEN MAY remain
on the protection list as defense-in-depth so unexpectedly forwarded real
credentials are still scrubbed on first read.
When apiProxy.auth.type is set to github-oidc, the API proxy sidecar
exchanges a GitHub Actions OIDC token for a provider-specific access token.
The apiProxy.auth.provider field (default: azure) selects the token
exchange protocol. A conforming implementation MUST:
-
Forward the common OIDC configuration to the sidecar via the following environment variables:
Config path Environment variable Required Default apiProxy.auth.typeAWF_AUTH_TYPE✅ — apiProxy.auth.providerAWF_AUTH_PROVIDERNo azureapiProxy.auth.oidcAudienceAWF_AUTH_OIDC_AUDIENCENo (provider-specific) -
Forward the GitHub Actions OIDC runtime tokens (
ACTIONS_ID_TOKEN_REQUEST_URL,ACTIONS_ID_TOKEN_REQUEST_TOKEN) to the sidecar whenAWF_AUTH_TYPE=github-oidc. These are injected automatically by the Actions runner when the workflow declarespermissions: id-token: write.If OIDC is requested for a provider but these runtime variables are not present in the sidecar environment, the provider adapter MUST fail closed and return an explicit configuration error (rather than falling back to static-key mode).
-
NOT expose the exchanged provider token in the agent container environment. The sidecar SHALL inject it into upstream request headers.
Exchanges the GitHub OIDC JWT for an Azure AD / Microsoft Entra access
token via workload identity federation. The sidecar injects the resulting
token as a Bearer Authorization header on upstream requests.
| Config path | Environment variable | Required | Default |
|---|---|---|---|
apiProxy.auth.azureTenantId |
AWF_AUTH_AZURE_TENANT_ID |
✅ | — |
apiProxy.auth.azureClientId |
AWF_AUTH_AZURE_CLIENT_ID |
✅ | — |
apiProxy.auth.azureScope |
AWF_AUTH_AZURE_SCOPE |
No | https://cognitiveservices.azure.com/.default |
apiProxy.auth.azureCloud |
AWF_AUTH_AZURE_CLOUD |
No | public |
Default OIDC audience: api://AzureADTokenExchange
Note:
azureTenantIdandazureClientIdare required for Azure AD federated credential exchange but MAY be omitted when using managed identity. See docs/api-proxy-sidecar.md for protocol-level details.
Exchanges the GitHub OIDC JWT for temporary AWS credentials via
sts.amazonaws.com AssumeRoleWithWebIdentity. The sidecar uses these
credentials to sign upstream requests to AWS Bedrock using SigV4.
| Config path | Environment variable | Required | Default |
|---|---|---|---|
apiProxy.auth.awsRoleArn |
AWF_AUTH_AWS_ROLE_ARN |
✅ | — |
apiProxy.auth.awsRegion |
AWF_AUTH_AWS_REGION |
✅ | — |
apiProxy.auth.awsRoleSessionName |
AWF_AUTH_AWS_ROLE_SESSION_NAME |
No | awf-oidc-session |
Default OIDC audience: sts.amazonaws.com
Note: AWS Bedrock uses IAM/SigV4 request signing rather than Bearer tokens. This means the sidecar MUST sign the complete request (method, path, headers, body hash) with the temporary credentials — it is not sufficient to inject a single
Authorizationheader.
Exchanges the GitHub OIDC JWT for a GCP access token via the Security
Token Service (sts.googleapis.com), optionally followed by service
account impersonation via iamcredentials.googleapis.com. The sidecar
injects the resulting token as a Bearer Authorization header.
| Config path | Environment variable | Required | Default |
|---|---|---|---|
apiProxy.auth.gcpWorkloadIdentityProvider |
AWF_AUTH_GCP_WORKLOAD_IDENTITY_PROVIDER |
✅ | — |
apiProxy.auth.gcpServiceAccount |
AWF_AUTH_GCP_SERVICE_ACCOUNT |
No | — |
apiProxy.auth.gcpScope |
AWF_AUTH_GCP_SCOPE |
No | https://www.googleapis.com/auth/cloud-platform |
Default OIDC audience: the gcpWorkloadIdentityProvider value
When gcpServiceAccount is provided, the sidecar performs a two-step
exchange:
- Exchange GitHub OIDC JWT for a federated access token via GCP STS
- Impersonate the service account to obtain a short-lived OAuth2 token
When gcpServiceAccount is omitted, only step 1 is performed and the
federated token is used directly. This requires that the federated
principal has direct access grants on the target resource.
Exchanges the GitHub OIDC JWT for an Anthropic Workload Identity Federation
token via Anthropic OAuth token endpoint (default:
https://api.anthropic.com/v1/oauth/token). The sidecar injects
the resulting token as an Authorization header on upstream requests.
| Config path | Environment variable | Required | Default |
|---|---|---|---|
apiProxy.auth.anthropicFederationRuleId |
AWF_AUTH_ANTHROPIC_FEDERATION_RULE_ID |
✅ | — |
apiProxy.auth.anthropicOrganizationId |
AWF_AUTH_ANTHROPIC_ORGANIZATION_ID |
✅ | — |
apiProxy.auth.anthropicServiceAccountId |
AWF_AUTH_ANTHROPIC_SERVICE_ACCOUNT_ID |
✅ | — |
apiProxy.auth.anthropicWorkspaceId |
AWF_AUTH_ANTHROPIC_WORKSPACE_ID |
Conditional¹ | — |
apiProxy.auth.anthropicTokenUrl |
AWF_AUTH_ANTHROPIC_TOKEN_URL |
❌ | https://api.anthropic.com/v1/oauth/token |
¹ AWF_AUTH_ANTHROPIC_WORKSPACE_ID is required when the federation rule covers
multiple workspaces. When the rule is scoped to a single workspace, it may be
omitted.
anthropicTokenUrl is non-sensitive and SHOULD be supplied via AWF config (including stdin config via --config -); env var support exists for compatibility.
Default OIDC audience: https://api.anthropic.com
When security.difcProxy.host is set, GITHUB_TOKEN and GH_TOKEN MUST
be excluded from the agent environment. These tokens SHALL be held
exclusively by the external DIFC proxy.
Under --network-isolation, the credential-bearing cli-proxy sidecar remains on
awf-net only. When security.difcProxy.host is classified as external
(host.docker.internal, an IP address outside awf-net's subnet, or a dotted DNS name), AWF creates a separate,
credential-free cli-proxy-egress relay that is the only CLI-proxy component
dual-homed onto the external bridge; it forwards solely to the configured DIFC
host and port and never receives GITHUB_TOKEN or GH_TOKEN. When the DIFC
proxy is an attached sibling container already reachable on awf-net, no relay
is created.
apiProxy.targets.openai.baseUrlEnv names a runner environment variable
(typically bound to ${{ secrets.* }} by the gh-aw compiler) whose value is the
base URL of a private OpenAI-compatible endpoint. This allows engine: codex
workflows to route through a sensitive endpoint without writing the URL into
workflow source or generated lockfiles.
{
"apiProxy": {
"targets": {
"openai": {
"baseUrlEnv": "CODEX_LB_BASE_URL"
}
}
}
}A conforming implementation:
- MUST read the named variable only in runner-side configuration code, before any container starts.
- MUST require an absolute
https://URL and MUST reject embedded credentials (user:pass@), query strings, fragments, malformed hosts, unsupported schemes, and non-default ports (the sidecar connects on port 443). - MUST derive the host,
host:port, and optional base path from the URL. - MUST add the derived destination to the effective Squid policy without
persisting it in repository configuration, and MUST keep it out of
allowedDomains(it is carried in the sensitive allowlist instead). - MUST configure the OpenAI api-proxy adapter with the derived host
(
OPENAI_API_TARGET) and base path (OPENAI_API_BASE_PATH); the derived values take precedence overapiProxy.targets.openai.host/basePath. - MUST exclude the named variable from the primary agent environment, including
under
--env-all. - MUST redact the URL, host, and
host:portforms from logs, diagnostics, and uploaded audit artifacts (squid.conf,docker-compose.redacted.yml). - MUST fail before agent startup, with an error message that does not contain the value, when the variable is unset or invalid.
The same rules apply to agent and detection phases, which share this configuration path.
This section is normative.
Anthropic's hosted server tools — web_search_YYYYMMDD and
web_fetch_YYYYMMDD — are executed by Anthropic infrastructure, not inside the
agent container. Squid therefore only ever observes the Anthropic API endpoint;
the searched or fetched destination is invisible to the domain ACL and to Squid
access logs, and the retrieved content returns inside an already-allowed API
response. Without additional enforcement this is a semantic egress gap: an
indirect prompt-injection and blind-exfiltration channel driven by
model-generated queries and URLs.
Because the api-proxy sidecar is trusted and already inspects Anthropic Messages
bodies, AWF closes the gap there. apiProxy.hostedWeb.claude is config-only:
there is no CLI flag and no environment alias. AWF serializes the validated
policy into the sidecar as AWF_CLAUDE_HOSTED_WEB_POLICY; that variable is an
implementation detail, is generated only from validated configuration, and MUST
NOT be accepted from an agent-controlled source.
apiProxy:
hostedWeb:
claude:
enabled: true
allowedDomains:
- docs.github.com
- nodejs.org
maxUses: 5Blocklist mode:
apiProxy:
hostedWeb:
claude:
enabled: true
blockedDomains:
- untrusted.example
- ads.exampleProhibit Claude hosted web tools entirely:
apiProxy:
hostedWeb:
claude:
enabled: falseapiProxy.hostedWeb and apiProxy.hostedWeb.claude are closed objects
(additionalProperties: false).
| Property | Type | Required | Semantics |
|---|---|---|---|
enabled |
boolean | yes | false rejects every matching Claude hosted search/fetch tool. true enables enforcement using exactly one configured domain mode. |
allowedDomains |
non-empty array of unique domains | conditional | Provider allowlist upper bound. Mutually exclusive with blockedDomains. |
blockedDomains |
non-empty array of unique domains | conditional | Provider blocklist lower bound. Mutually exclusive with allowedDomains. |
maxUses |
positive integer | no | Maximum uses AWF permits per matching hosted tool. Request values may only reduce it. |
There is deliberately no implicit unrestricted enabled: true state: when
enabled is true, exactly one of allowedDomains or blockedDomains MUST be
present and non-empty. This prevents a success-shaped configuration that appears
protected while forwarding unrestricted hosted web access. Domains MUST NOT be
combined with enabled: false.
Domain values use one canonical syntax: lowercase DNS hostnames with at least
two labels and no URL scheme, path, query, fragment, credentials or port, and no
raw IP address, CIDR, wildcard, localhost-style single-label name, Docker
service alias, or empty string. Duplicates are removed after normalization. A
parent domain also covers its subdomains, matching Anthropic's semantics.
Absence of apiProxy.hostedWeb.claude preserves the pre-existing pass-through
behaviour. Warning: omitting the object does not constrain Claude-hosted
egress. A compiler such as gh-aw can opt into enforcement by always emitting
the object when workflow frontmatter declares a network policy.
The hosted-web list is intentionally explicit rather than derived from
network.allowDomains. The network list may contain provider API endpoints,
internal services, or ecosystem expansions that are inappropriate to disclose to
or authorize for provider-hosted retrieval. Sensitive, secret-derived allowlist
entries (see §9.7) MUST NOT be disclosed to
the provider.
The behaviour is identical for a JSON file, a YAML file, JSON on stdin, and YAML
on stdin. Validation errors from --config - identify stdin as the source and
reject the run before any container starts. A compiler therefore only needs to
emit a stable config document:
{
"network": {
"allowDomains": ["api.github.com", "docs.github.com"]
},
"apiProxy": {
"hostedWeb": {
"claude": {
"enabled": true,
"allowedDomains": ["docs.github.com"]
}
}
}
}and pipe it to awf --config - -- <command>. Compilers do not need to understand
Anthropic request JSON, tool-version names, or api-proxy environment variables;
AWF owns the provider-specific translation and enforcement.
The effective configured policy is selected before request processing:
- an explicit CLI option, if one is added in a later revision;
- the AWF configuration document, including JSON or YAML supplied via
--config -; - AWF's internal default (no enforcement; current compatibility behaviour).
The configured policy is an immutable security boundary. Request-provided tool settings MAY narrow it but MUST NEVER broaden, remove, replace, or disable it.
| Configured policy | Request tool fields | Effective result |
|---|---|---|
enabled: false |
any | Reject (claude_hosted_web_disabled) |
| allowlist | none | Inject configured allowed_domains |
| allowlist | allowed_domains |
Intersection of configured and requested |
| allowlist | allowed_domains with empty intersection |
Reject (claude_hosted_web_empty_intersection) |
| allowlist | blocked_domains |
Reject (claude_hosted_web_policy_conflict) |
| blocklist | none | Inject configured blocked_domains |
| blocklist | blocked_domains |
Union of configured and requested |
| blocklist | allowed_domains |
Reject (claude_hosted_web_policy_conflict) |
| any enabled mode | both allowed_domains and blocked_domains |
Reject (claude_hosted_web_policy_conflict) |
maxUses configured |
max_uses omitted |
Inject configured value |
maxUses configured |
max_uses present |
min(configured, requested) |
| any enabled mode | malformed max_uses (zero, negative, non-integer) |
Reject (claude_hosted_web_max_uses_invalid) |
Cross-mode request restrictions are rejected rather than silently dropped
because Anthropic cannot represent allowed_domains and blocked_domains on the
same tool definition; rejecting is safer and more diagnosable than pretending
both policies were enforced. AWF-generated values replace the corresponding
request fields only after the effective policy has been calculated: a hosted tool
is never silently stripped, and an unprotected hosted tool is never silently
forwarded.
A conforming implementation:
- MUST parse and validate the serialized policy once at sidecar startup and MUST fail startup — not the first request — on an invalid internal policy.
- MUST apply enforcement only to Anthropic Messages request bodies that contain a matching hosted tool. A configured policy MUST NOT mutate an ordinary Messages request, including requests to custom Anthropic-compatible endpoints that implement no hosted server tools.
- MUST recognize versioned tool types with a validated
web_search_YYYYMMDD/web_fetch_YYYYMMDDmatcher — including versions published after this revision — rather than a fixed list of exact names. - MUST fail closed on a value that claims to be a hosted web tool but does not
match the versioned form (
claude_hosted_web_tool_unrecognized), so a future or malformed shape can never be forwarded without policy. - MUST apply enforcement after generic body parsing and after the other Anthropic transforms (model rewriting, prompt-cache optimizations, tool drop, custom transform file) so that no other transform can re-expand or remove the enforced policy.
- MUST enforce every matching tool in a request, not only the first.
- MUST return structured API-proxy errors with the stable
error.codevalues listed above (HTTP 403 for policy rejections, HTTP 400 for malformed tool definitions, domains, ormax_uses), and MUST NOT include prompts, queries, URLs, or request bodies in those errors.
This section is normative.
OpenAI-hosted retrieval executes beyond the AWF network boundary. Squid sees
the permitted OpenAI endpoint, not the searched or fetched destination. AWF
therefore enforces apiProxy.hostedWeb.codex inside the trusted API proxy on:
- Responses requests containing a
web_searchor datedweb_search_YYYY_MM_DDtool; and - every request to
/v1/alpha/search.
The config object is closed and references the same schema, normalization, and
source-precedence contract as apiProxy.hostedWeb.claude: enabled is
required; enabled: true requires exactly one non-empty allowedDomains or
blockedDomains; enabled: false permits neither; and maxUses, when present,
is a positive integer. Domains are normalized and validated identically for
both providers. The two policies may coexist in one config.
apiProxy:
hostedWeb:
claude:
enabled: true
blockedDomains:
- untrusted.example
codex:
enabled: true
allowedDomains:
- docs.github.com
- nodejs.org
maxUses: 5Configuration-source precedence is:
- an explicit CLI option, if one is added in the future;
- the validated AWF document, including JSON/YAML via
--config -; and - compatibility behavior.
The initial implementation is config-only. AWF_CODEX_HOSTED_WEB_POLICY is an
internal transport generated from validated config and cannot independently
override it. Omission preserves pass-through behavior and does not constrain
Codex-hosted egress.
The configured policy is immutable. A request can only narrow it. OpenAI can represent allowed and blocked filters together, so the canonical cross-mode rule is to preserve the request's opposite-mode restriction while applying the configured mode; no restriction is silently removed.
| Config mode | Request filters | Effective filters |
|---|---|---|
| disabled | Any Responses hosted tool or standalone request | Reject with codex_hosted_web_disabled |
| allow | none | Inject configured allowed_domains |
| allow | allowed_domains |
Intersect with configured allowlist; reject an empty result |
| allow | blocked_domains |
Inject configured allowlist and preserve request blocklist |
| block | none | Inject configured blocked_domains |
| block | blocked_domains |
Union with configured blocklist |
| block | allowed_domains |
Preserve request allowlist and inject configured blocklist |
Both request filter arrays, when present, must be non-empty valid domain lists.
On the standalone route, the effective settings.filters is applied again to
each commands.search_query[].domains and commands.image_query[].domains
list. An allowlist is intersected; blocked domains are removed; and an empty
effective query scope rejects with codex_hosted_web_empty_query_scope rather
than falling back to the request-level scope.
Literal HTTP(S) URLs in commands.open[].ref_id,
commands.find[].ref_id, and commands.screenshot[].ref_id are host-checked
against both effective lists. Non-URL provider reference IDs remain valid.
A malformed URL/host rejects; a disallowed host rejects with
codex_hosted_web_url_disallowed.
The most restrictive applicable configured, request, per-query, and URL
scope governs each retrieval. No scope can relocate or widen authority granted
by another. network.allowDomains and network.sensitiveAllowDomains are
never copied into provider policy or request bodies.
Responses external_web_access and indexed_web_access values must be
booleans. Standalone settings.external_web_access accepts booleans or the
current cached, indexed, and live modes. Unknown future values fail closed
with codex_hosted_web_access_invalid; they are never interpreted as
permissive.
For Responses tools, configured maxUses is injected as max_uses, or
resolved as min(configured, requested). Malformed, zero, negative, or
non-integer request values reject. /v1/alpha/search exposes no equivalent
per-request cap, so a configured maxUses rejects that route with
codex_hosted_web_max_uses_unsupported rather than being silently ignored.
Recognized standalone command families are search_query, image_query,
open, click, find, and screenshot. Unknown command shapes and malformed
dated hosted-tool names fail closed. Enforcement runs after existing OpenAI
body/model transforms and immediately before dispatch, so no later transform
can remove or expand the effective policy.
Policy failures use invalid_request_error envelopes with stable
codex_hosted_web_* codes:
| Code | Meaning |
|---|---|
codex_hosted_web_disabled |
Hosted retrieval is prohibited |
codex_hosted_web_filter_invalid / codex_hosted_web_domain_invalid |
Malformed filters or domains |
codex_hosted_web_empty_intersection / codex_hosted_web_empty_query_scope |
Narrowing produced no permitted scope |
codex_hosted_web_url_invalid / codex_hosted_web_url_disallowed |
Literal URL is malformed or outside policy |
codex_hosted_web_access_invalid |
Unknown hosted-access mode |
codex_hosted_web_max_uses_invalid / codex_hosted_web_max_uses_unsupported |
Invalid cap or route cannot enforce it |
codex_hosted_web_tool_unrecognized / codex_hosted_web_command_unrecognized / codex_hosted_web_shape_invalid |
Unknown future or malformed hosted-search shape |
Errors do not include prompts, queries, literal URLs, or request bodies. Serialized internal policy is validated during sidecar startup. Ordinary OpenAI-compatible requests without a hosted-web surface are unchanged.
Compiler integrations only emit the stable config and pipe it to
awf --config - -- <command>; they do not need to understand provider routes,
body fields, tool versions, or internal environment variables.
This section is normative.
When apiProxy.maxEffectiveTokens is configured, the API proxy MUST enforce
a cumulative effective-token budget across all LLM API requests in a single
run. The budget limits total weighted token consumption, not raw token
counts.
Each upstream response's usage object is decomposed into four categories,
each with a fixed weight:
| Category | Weight | Usage field |
|---|---|---|
| Input | 1.0 | input_tokens / prompt_tokens |
| Cache read | 0.1 | cache_read_input_tokens / prompt_tokens_details.cached_tokens |
| Output | 4.0 | output_tokens / completion_tokens |
| Reasoning | 4.0 | reasoning_tokens / completion_tokens_details.reasoning_tokens |
The base weighted tokens for a single response are:
base = (1.0 × input) + (0.1 × cache_read) + (4.0 × output) + (4.0 × reasoning)
When apiProxy.modelMultipliers is configured, each model name MAY have
an associated positive multiplier. The effective tokens for a response are:
effective_tokens = model_multiplier × base_weighted_tokens
If no exact multiplier is configured, AWF MUST attempt to match
apiProxy.modelMultipliers keys against the request model using a hyphen-suffix
prefix match so family keys like claude-opus-4.7 apply to concrete model IDs
like claude-opus-4.7-20260501.
If no exact or prefix match is found, and apiProxy.defaultModelMultiplier is
configured, that default multiplier MUST be used.
Otherwise, if no exact or prefix match is found, the multiplier MUST default to
the highest configured model multiplier. If no model multipliers are configured
at all, the multiplier defaults to 1.
When AWF falls back to the default multiplier because no configured model key matched, it MUST emit a warning log entry.
The API proxy MUST enforce the budget as follows:
-
Accumulation: After each successful upstream response, the proxy extracts the
usageobject, computes effective tokens, and adds them to a running total for the session. -
Pre-request check: Before forwarding each subsequent request to the upstream provider, the proxy checks whether the cumulative total has reached or exceeded
maxEffectiveTokens. -
Rejection: When the budget is reached or exceeded, the proxy MUST reject the request with:
- HTTP status:
403 Forbidden - Content-Type:
application/json - Response body:
{ "error": { "type": "effective_tokens_limit_exceeded", "message": "Maximum effective tokens exceeded (1234.56 / 1000).", "total_effective_tokens": 1234.56, "max_effective_tokens": 1000 } }
- HTTP status:
-
WebSocket rejection: For WebSocket upgrade requests, the proxy MUST reject with
HTTP/1.1 403 Forbiddenand include the same JSON error body before destroying the socket. -
Finality: Once the budget is reached or exceeded, all subsequent requests in the same run MUST be rejected. The budget is not recoverable.
The proxy MUST track when cumulative effective tokens cross the following
percentage thresholds of maxEffectiveTokens:
| Threshold |
|---|
| 80% |
| 90% |
| 95% |
| 99% |
Each threshold MUST be recorded at most once per run.
Token steering is opt-in. It is active only when apiProxy.enableTokenSteering
is true (CLI: --enable-token-steering), which sets AWF_ENABLE_TOKEN_STEERING=true
in the api-proxy sidecar. When disabled (the default), thresholds are still tracked
(for introspection) but no warning messages are injected. Setting the field to
false, or omitting it, opts a workflow out; the env var is only emitted when the
value is true.
When token steering is enabled and a threshold is first crossed, the proxy MUST inject a budget-warning system message into the body of the very next eligible request sent by the agent, then discard the pending message so that it is injected at most once per threshold per run.
The injected message has the format:
[AWF TOKEN WARNING] <threshold-specific text>
| Threshold | Injected text |
|---|---|
| 80% | You have used 80% of your effective token budget. Begin planning to wrap up your current work. |
| 90% | You have used 90% of your effective token budget. Complete your current task and prepare final output. |
| 95% | You have used 95% of your effective token budget. Finalize and submit your work now. |
| 99% | You have used 99% of your effective token budget. You are about to be cut off. Submit immediately. |
If multiple thresholds are crossed simultaneously (e.g. a single large response crosses both 80% and 90%), the proxy MUST inject only the highest crossed threshold on the next request and queue the remaining thresholds for subsequent requests (one per request).
Provider-specific injection rules:
- OpenAI / Copilot — the proxy inserts a
{ "role": "system", "content": "<message>" }entry into themessagesarray immediately after any pre-existing system messages. - Anthropic — the proxy appends the warning to the
systemfield: ifsystemis a string it is concatenated (separated by\n\n); ifsystemis an array of content blocks a{ "type": "text", "text": "<message>" }block is appended; ifsystemis absent it is created as the warning string. - Gemini — the proxy appends a
{ "text": "<message>" }part tosystemInstruction.parts; ifsystemInstructionis absent it is created.
If the request body cannot be parsed as JSON, or if the body format does not match the expected structure, the proxy MUST silently skip injection for that request and NOT re-queue the message.
When token steering is enabled and container.agentTimeout is configured,
the proxy MUST also inject runtime warnings at 80/90/95/99% of elapsed run time
using the same queueing behavior (highest crossed threshold first, then one
pending warning per subsequent request):
[AWF TIME WARNING] <threshold-specific text>
The API proxy exposes a GET /reflect endpoint on every provider port
(10000–10004). Each port returns the same aggregate reflection payload, whose
endpoints array lists all provider adapters. Only the management port
(10000, OpenAI) serves /metrics and the aggregate /health; non-management
ports still serve provider-local /health responses.
maxAiCredits is a positive number. It is supplied via the AWF config file
(including stdin config via --config -) and maps to the
AWF_MAX_AI_CREDITS environment variable injected into the api-proxy
container.
When configured, the proxy MUST enforce this budget in addition to any
configured maxEffectiveTokens budget. Once cumulative AI credits reach or
exceed maxAiCredits, subsequent requests MUST be rejected with HTTP 403
and error type ai_credits_limit_exceeded.
Regardless of maxAiCredits configuration, AWF also enforces a non-overridable
hard cap of 10,000 AI credits. When cumulative AI credits reach this hard
cap, subsequent requests MUST be rejected with HTTP 403 and error type
ai_credits_limit_exceeded, and the error/log payload MUST include
hard_cap: true.
If both limits are present, the effective enforcement threshold is the lower of:
- configured
maxAiCredits - the fixed hard cap (10,000)
Setting maxAiCredits above 10,000 MUST NOT raise the effective limit.
The AI credits guard resolves model names using this lookup order:
- Operator provider overlay — model prices configured under
apiProxy.providers. - Runtime provider metadata — authoritative token prices discovered from the configured provider. Copilot supports this today.
- Curated pricing table — a built-in table of known models with exact pricing.
- Bundled models.dev catalog — a bundled snapshot of the models.dev catalog used as a fallback when the model is not found in the curated table.
Model names are canonicalized before lookup: provider prefixes
(e.g. copilot/) are stripped, and separators (., _, -) are treated
as interchangeable. For example, copilot/claude-sonnet-4.6,
claude_sonnet_4_6, and claude-sonnet-4-6 all resolve to the same pricing
entry.
If none of these sources resolves the model, the defaultAiCreditsPricing fallback
(if configured) is used. If that is also absent, the request is rejected.
Models whose catalog entry carries zero-cost pricing are recognized as known
models with zero AI credit impact, so they are never rejected as "unknown".
Runtime tiered pricing uses the provider's default-tier prompt threshold. When the total input exceeds that threshold, all token categories use the long-context tier. Pricing source, API version, observation time, selected tier, and any provider-advertised promotion are retained in provenance. Promotions are informational only because provider discovery does not prove that a discount applies to a specific request; they never reduce accounting. Failed or empty discovery responses do not replace the last successful runtime snapshot.
Provider overlays use the models.dev provider structure and per-token dollar rates:
apiProxy:
providers:
github-copilot:
models:
custom-model:
cost:
input: "3e-06"
output: "1.5e-05"
cache_read: "3e-07"
cache_write: "3.75e-06"The overlay is passed to both normal and threat-detection API proxy instances
through AWF_API_PROXY_PROVIDERS. Provider aliases github-copilot and
copilot resolve to the Copilot proxy.
defaultAiCreditsPricing is an optional object with input and output
fields (both required, in $/1M tokens), plus optional cachedInput and
cacheWrite fields.
It is supplied via the AWF config file and maps to the
AWF_DEFAULT_AI_CREDITS_PRICING environment variable (JSON string) injected
into the api-proxy container.
When configured, any model not found in the curated built-in pricing table or the bundled models.dev catalog uses these rates as a fallback for AI credits calculation.
When maxAiCredits is active and the proxy encounters a request whose model
cannot be resolved from the curated built-in pricing table or the bundled
models.dev catalog:
-
If
defaultAiCreditsPricingis configured: the fallback rates are used and the request proceeds normally. -
If
defaultAiCreditsPricingis NOT configured: the proxy MUST reject the request with HTTP400and error typeunknown_model_ai_credits. The error payload includes:model: the unresolved model namemessage: human-readable instructions to configureapiProxy.defaultAiCreditsPricing
This fail-closed behavior prevents unaccounted spending from models whose pricing is unknown to the proxy.
Note: Requests without a model field in the body (e.g. non-chat endpoints)
are not subject to this check.
Recognized dynamic selectors. The Copilot auto model (copilot:auto) is
not subject to unknown_model_ai_credits rejection. Because its concrete
runtime model is not known at request time, the proxy accounts it using a
conservative fallback ceiling (the maximum per-token rate across the curated
pricing catalog) rather than rejecting the request or silently under-counting
spend. Token-usage records for these requests set pricing_source and
accounting_policy to dynamic_selector_fallback, pricing_tier to
conservative, fallback_pricing_used to true, and dynamic_selector to
copilot:auto. This dynamic-selector accounting path applies only to the
recognized Copilot auto selector; other unresolved models still follow the
fallback/rejection behavior above.
When AI credits and/or effective tokens are computed, the token-usage.jsonl
records include additional optional fields:
| Field | Type | Description |
|---|---|---|
effective_tokens_this_response |
number | Weighted tokens for this request |
effective_tokens_total |
number | Running total of effective tokens |
model_multiplier |
number | Cost multiplier applied for this model |
ai_credits_this_response |
number | AI credits consumed by this request |
ai_credits_total |
number | Running total of AI credits |
These fields are only present when the respective guard is active.
This section is normative.
When apiProxy.maxTurns is configured, the API proxy MUST enforce an absolute
maximum number of LLM invocations per run.
An invocation is counted each time the proxy receives a successful (2xx)
HTTP response from an upstream LLM provider. Each response increments a
per-run counter by one, regardless of the number of tokens consumed.
The API proxy MUST enforce the max-runs limit as follows:
-
Pre-request check: Before forwarding each request to the upstream provider, the proxy checks whether the invocation count has reached or exceeded
maxTurns. -
Rejection: When the limit is reached or exceeded, the proxy MUST reject the request with:
- HTTP status:
403 Forbidden - Content-Type:
application/json - Response body:
{ "error": { "type": "max_runs_exceeded", "message": "Maximum LLM invocations exceeded (5 / 5).", "invocation_count": 5, "max_runs": 5 } }
- HTTP status:
-
WebSocket rejection: For WebSocket upgrade requests, the proxy MUST reject with
HTTP/1.1 403 Forbiddenand include the same JSON error body before destroying the socket. -
Finality: Once the limit is reached, all subsequent requests in the same run MUST be rejected. The counter is not recoverable.
The /reflect endpoint (available on all provider ports 10000–10004; see
§10.6) MUST include the current max-runs state:
{
"runs": {
"enabled": true,
"max_runs": 5,
"invocation_count": 3,
"remaining_runs": 2
}
}When maxTurns is not configured, the enabled field MUST be false and
max_runs and remaining_runs MUST be null.
This section is normative.
When apiProxy.maxPermissionDenied is configured, the API proxy MUST halt
further LLM requests after the upstream returns a configurable number of
401 or 403 responses, preventing token waste when API credentials are
misconfigured or expired.
A permission error is counted each time the proxy receives an HTTP 401 or
403 response from an upstream LLM provider. Each such response increments
a per-run counter by one.
The API proxy MUST enforce the permission-denied limit as follows:
-
Post-response counting: After receiving a
401or403from upstream, the proxy increments the denied count. -
Pre-request check: Before forwarding each subsequent request to the upstream provider, the proxy checks whether the denied count has reached or exceeded
maxPermissionDenied. -
Rejection: When the limit is reached or exceeded, the proxy MUST reject the request with:
- HTTP status:
403 Forbidden - Content-Type:
application/json - Response body:
{ "error": { "type": "permission_denied_limit_exceeded", "message": "Permission denied limit exceeded (3 / 3). The run has been stopped due to repeated permission errors — check that all API keys and tokens are correctly configured.", "denied_count": 3, "max_permission_denied": 3 } }
- HTTP status:
-
Finality: Once the limit is reached, all subsequent requests in the same run MUST be rejected until the configured limit changes (changing
AWF_MAX_PERMISSION_DENIEDresets the counter).
The /reflect endpoint (available on all provider ports 10000–10004; see
§10.6) MUST include the current permission-denied guard state:
{
"permission_denied": {
"enabled": true,
"max_permission_denied": 3,
"denied_count": 1
}
}When maxPermissionDenied is not configured, the enabled field MUST be
false, max_permission_denied MUST be null, and denied_count MUST be 0.
maxPermissionDenied is a positive integer. It is supplied via the AWF
config file (stdin config) or the --max-permission-denied CLI flag, and
maps to the AWF_MAX_PERMISSION_DENIED environment variable injected into
the api-proxy container.
Example:
apiProxy:
maxPermissionDenied: 3 # stop run after 3 upstream 401/403 responsesThis section is normative.
When apiProxy.maxCacheMisses is configured, the API proxy MUST halt further
LLM requests after the configured number of consecutive responses that had no
prompt-cache hits, preventing runaway token spend caused by a broken or expired
cache (e.g., mismatched cache keys, context window overflow, or prompt drift).
A cache miss is counted for a response when all of the following are true:
- The response is a successful upstream completion (not a proxy-level error).
input_tokens > 0(zero-input responses such as empty tool calls are excluded so they do not inflate the streak counter).cache_read_tokens === 0(no prompt-cache hit occurred).
A cache hit (cache_read_tokens > 0) resets the consecutive miss streak to
zero.
The API proxy MUST enforce the cache-miss limit as follows:
-
Post-response counting: After receiving each successful upstream response, the proxy inspects the normalized token usage and increments or resets the miss streak counter.
-
Pre-request check: Before forwarding each subsequent request to the upstream provider, the proxy checks whether the miss streak has reached or exceeded
maxCacheMisses. -
Rejection: When the limit is reached or exceeded, the proxy MUST reject the request with:
- HTTP status:
403 Forbidden - Content-Type:
application/json - Response body:
{ "error": { "type": "max_cache_misses_exceeded", "message": "Maximum consecutive cache misses exceeded (3 / 3).", "consecutive_cache_misses": 3, "max_cache_misses": 3 } }
- HTTP status:
-
WebSocket rejection: For WebSocket upgrade requests, the proxy MUST reject with
HTTP/1.1 403 Forbiddenand include the same JSON error body before destroying the socket. -
Finality: Once the streak limit is reached, all subsequent requests in the same run MUST be rejected. Changing
AWF_MAX_CACHE_MISSESresets the streak counter.
The /reflect endpoint (available on all provider ports 10000–10004; see
§10.6) MUST include the current cache-miss guard state:
{
"cache_misses": {
"enabled": true,
"max_cache_misses": 3,
"consecutive_cache_misses": 1,
"remaining_cache_misses": 2
}
}When maxCacheMisses is not configured, the enabled field MUST be false,
max_cache_misses MUST be null, consecutive_cache_misses MUST be 0, and
remaining_cache_misses MUST be null.
maxCacheMisses is a positive integer. It is supplied via the AWF config file
(stdin config) or the --max-cache-misses CLI flag, and maps to the
AWF_MAX_CACHE_MISSES environment variable injected into the api-proxy
container.
Example:
apiProxy:
maxCacheMisses: 3 # stop run after 3 consecutive cache missesThis section is normative.
When apiProxy.maxModelMultiplierCap is configured, the API proxy MUST
reject any request whose resolved model multiplier exceeds the cap before
forwarding the request to the upstream provider.
The proxy resolves the effective multiplier for the requested model using the same algorithm as the effective-token guard:
- Exact match: if
apiProxy.modelMultiplierscontains the exact model name, use its multiplier. - Longest-prefix match: if any configured model name is a prefix of the
requested model name (followed by
-), use the multiplier of the longest-matching prefix. - Default: use
apiProxy.defaultModelMultiplierif configured, otherwise default to1.
Before forwarding each POST/PUT/PATCH request to an upstream LLM provider, the proxy MUST:
-
Extract the
modelfield from the request body. -
Resolve the model's effective multiplier (§12.1).
-
If the multiplier exceeds
maxModelMultiplierCap, reject the request with:- HTTP status:
400 Bad Request - Content-Type:
application/json - Response body:
{ "error": { "type": "model_multiplier_cap_exceeded", "message": "Model multiplier cap exceeded: model \"claude-opus-4.7\" has multiplier 27 which exceeds the configured maximum of 5.", "model": "claude-opus-4.7", "model_multiplier": 27, "max_model_multiplier": 5 } }
- HTTP status:
-
If the model field is absent or the multiplier is within the cap, the request MUST be forwarded normally.
maxModelMultiplierCap is a positive number. It is supplied via the AWF
config file (stdin config) and maps to the AWF_MAX_MODEL_MULTIPLIER
environment variable injected into the api-proxy container. The CLI flag
--max-model-multiplier-cap <number> may also be used.
Example:
apiProxy:
maxModelMultiplierCap: 5 # reject any model with multiplier > 5
modelMultipliers:
claude-opus-4.7: 27
gpt-4o: 2This section is normative.
When apiProxy.modelFallback is configured, the API proxy provides automatic
model selection when a requested model is unavailable. The fallback mechanism
ensures requests complete gracefully without requiring explicit agent-side
handling.
Model fallback is controlled via apiProxy.modelFallback:
{
"apiProxy": {
"modelFallback": {
"enabled": true,
"strategy": "middle_power"
}
}
}| Field | Type | Default | Description |
|---|---|---|---|
enabled |
boolean | true |
Enable/disable the fallback mechanism |
strategy |
string | middle_power |
Selection strategy (middle_power is currently the only strategy) |
excludeEngines |
string[] | [] |
Engines for which middle-power fallback is suppressed (e.g. ["openai"]). Excluded engines receive native model-unavailable errors instead of silent rewrites. |
When strategy is middle_power, the proxy selects the median capability-tier
model from the available models for the current provider.
Capability tiers are assigned based on model family and version:
| Provider | Tier 5 | Tier 4 | Tier 3 | Tier 1 |
|---|---|---|---|---|
| Anthropic | claude-opus* |
claude-sonnet* |
claude-haiku* |
(other) |
| OpenAI / Copilot | gpt-5* |
gpt-4*, gpt-4o* |
gpt-3.5* |
(other) |
| Gemini | (reserved) | (reserved) | (reserved) | (all) |
Selection algorithm:
- Sort available models by capability tier (highest first), then lexicographically
- Select the median model from the sorted list
- Log the selection with the reason and full candidate list
Example:
Available: ['gpt-3.5-turbo', 'gpt-5.2', 'gpt-4.1']
Sorted: ['gpt-5.2' (tier 5), 'gpt-4.1' (tier 4), 'gpt-3.5-turbo' (tier 3)]
Median: gpt-4.1 (index 1 of 0-2)
The fallback is activated when:
- Direct match fails: The requested model is not found in the available models list for the provider.
- Family version fallback doesn't apply: For
gpt-5.*models on OpenAI, if a lowergpt-5.*version is available, use that before triggering middle-power fallback. - Alias has no candidates: An alias pattern matched but produced no resolvable models on the current provider.
The fallback is NOT activated when:
- A direct model match is found (return it immediately)
- A family version fallback is available (for
gpt-5.*only) - The fallback is disabled (
enabled: false) - An alias has
fallback: false(see §12.4) - The provider is in the
excludeEngineslist - Copilot engine in standard mode (no BYOK env vars): the Copilot CLI is authoritative for its own model catalogue, so retired/restricted model names should fail fast with a clear upstream error rather than being silently rewritten to a middle-power fallback
- Copilot BYOK that still targets a GitHub Copilot catalog host (for example
api.githubcopilot.com): the catalog remains authoritative, so fallback is still suppressed - Copilot is configured for a BYOK non-
githubcopilottarget (for example Azure OpenAI deployment endpoints), where deployment names are provider-local and must not be rewritten to catalog model IDs
Model aliases now support an extended syntax that permits per-alias fallback control:
Legacy syntax (string array):
{
"models": {
"sonnet": ["copilot/*sonnet*", "openai/*sonnet*"]
}
}Fallback is enabled by default for legacy syntax.
Extended syntax (object with patterns):
{
"models": {
"sonnet": {
"patterns": ["copilot/*sonnet*"],
"fallback": false
}
}
}| Field | Type | Default | Description |
|---|---|---|---|
patterns |
string[] | — | Glob patterns to match against available models |
fallback |
boolean | true |
Enable fallback for this alias if no candidates are found |
When fallback: false, if the alias patterns produce no candidates, the
resolution returns null instead of activating middle-power fallback.
The health endpoint (GET /health) includes a model_fallback field in the
response:
{
"status": "healthy",
"service": "awf-api-proxy",
"model_fallback": {
"enabled": true,
"strategy": "middle_power"
}
}The /reflect endpoint does not include fallback state by design (it is static
per run).
When apiProxy.requestedModel is configured, the API proxy validates at startup
that the specified model is available in at least one provider's model catalogue.
Configuration:
{
"apiProxy": {
"requestedModel": "gpt-4o"
}
}Mapping: apiProxy.requestedModel → AWF_REQUESTED_MODEL (config-only; set by AWF CLI)
Behavior:
- After
fetchStartupModels()completes, the proxy checksAWF_REQUESTED_MODELagainst all cached provider model lists. - If the model is found directly or resolves via model aliases, a confirmation
model_validationlog is emitted. - If the model is NOT found, a
model_unavailable_at_startuperror log is emitted listing available models as a diagnostic aid. - Validation is non-blocking — the proxy continues serving requests regardless of the outcome, so agents that ignore the model hint are not affected.
This enables workflow authors to get clear, early feedback when a retired or misspelled model is specified, rather than waiting for the first API request to fail with an opaque error.
When apiProxy.fallbackModels is configured, the API proxy retries a failed
request with the next model in an ordered list. The middle-power fallback above
picks a model when the request is resolved. This chain applies after the
upstream has rejected the model.
{
"apiProxy": {
"fallbackModels": ["gpt-5.4", "claude-sonnet-4.6"]
}
}Mapping: apiProxy.fallbackModels → AWF_FALLBACK_MODELS (JSON array; a
comma-separated list is also accepted when the variable is set directly)
Behavior:
- Fallback triggers only on model-specific failures:
- any upstream
5xxresponse (504is reported asupstream_timeout) - a connection error or timeout before any upstream response
(
upstream_connection_error) - a
400/404whose body says the model is unsupported, not found, or not accessible, such asmodel_not_supported,model_not_found, ornot accessible via the … endpoint(model_not_supported)
- any upstream
401,403, and429responses never trigger a fallback. A generic400validation error, such as a bad tool schema or a too-long context, is returned to the client unchanged.- The proxy rewrites the request body's
modelfield (OpenAI, Anthropic, Copilot). For Gemini, where the body has nomodel, it rewrites the/models/<model>:<method>path segment instead. The proxy strips a redundant<provider>/prefix on fallback entries. - Entries are tried in order. The proxy skips models already attempted for the request and models rejected by the model-policy, retired-model, multiplier-cap, or budget guards. When the chain is exhausted, the proxy returns the last upstream error to the client.
- Each switch emits a
model_fallbackwarning log withfrom_model,to_model,requested_model,attempt,reason, andstatus. The token-usage record for the successful response holds the model that served the request inmodel, plus amodel_fallbackobject withrequested_model,model,attempt,reason, andstatus. gh-aw can report that model inGH_AW_INFO_MODELand telemetry. - Copilot's existing transient
model not supportedretries and the alias endpoint-blocked candidate retry still run first. The ordered chain applies only after those have been exhausted. - WebSocket (Responses API) upgrades and AWF-internal routing classifier requests are not covered.
Alias resolution MUST only consider provider slots that are actually configured
for the run. Before an alias is expanded (and before aliases are advertised via
/reflect and models.json), the cached model lists of providers that report
configured: false are treated as empty. A provider-scoped pattern such as
copilot/*sonnet* therefore yields no candidate when no Copilot credential is
present, even if a model list was cached earlier in the run.
Configuration is determined from each provider's reflected configured slot,
not from request readiness. A configured OIDC provider remains eligible while
its token is being minted. When configured providers have no model catalogue
yet, aliases scoped only to unconfigured providers are still omitted, while
aliases that can target a configured provider remain advertised until model
data is available.
Without this filter, a Copilot-first alias group would steer every request to a
slot that answers provider_not_configured, producing a 100% call-failure rate
and, for retry-happy clients, a non-terminating retry loop.
A provider_not_configured response is a terminal run-level misconfiguration:
it is returned with HTTP 403 and "retryable": false so clients fail fast.
Only transient OIDC readiness states, such as a token that has not been minted
yet, use HTTP 503 and "retryable": true.
Task-level model routing is experimental and opt-in. To enable it, set
root-level experimental.modelRouting: true alongside the compiler-authored
apiProxy.routing request. An existing config containing apiProxy.routing
without this gate must add the experimental block shown below; otherwise
validation fails rather than silently ignoring the request. The gate alone,
without apiProxy.routing, is valid but does not activate routing. Omitting
the gate (or setting it to false) without a routing request preserves normal
operation without routing infrastructure or environment.
As of this release, both the proxy-side and host-side halves of task-level
routing are wired and shipped on main. The API proxy's routing controller
is wired into the running server
(PR #8966): when
AWF_ROUTING_CONFIG is present, a routing session starts after key
validation and model discovery and selects one model/effort for the run.
The host workflow stages and validates apiProxy.routing input before
the proxy starts (PR #8985):
it writes the task conversation into a private per-run routing directory,
rejects unsupported configurations (non-Linux, non-runc, disabled API proxy,
--keep-containers, DinD/split filesystems, Docker-socket exposure, or an
unpinned router image),
and waits for selection.json before starting the agent.
The selection is advisory, not admitted-only. The agent is seeded with the
selected model, effort, and endpoint, but the proxy does not pin requests to
it: an agent (or a sub-agent that declares its own model:) MAY send any model
that model policy permits, and such a request completes normally. Model choice
is not a containment boundary — allowedModels / disallowedModels
(AWF_ALLOWED_MODELS / AWF_DISALLOWED_MODELS) remain the enforcement
surface that bounds cost and policy, independently of routing, and a request
for an excluded model is still rejected by that policy guard. WebSocket
upgrades are proxied normally while a routing session exists. For each
inference request the proxy logs a model_routing event with
stage: "request", recording the requested and selected provider, model,
effort, and endpoint, routed: "as_selected" or "deviated", and the list of
deviations, so routing quality stays measurable without enforcement.
A genuine routing failure still surfaces as host exit code 78 instead of the
run silently continuing: no selection could be produced (no_route, router
unreachable, contract or configuration errors), or an upstream failure on a
request that used the selected provider and model (a native provider error
code, an SSE error event, or a prematurely closed response). A request that
uses another model is not a routing failure, and neither is its upstream
error or a model-policy rejection of it.
The agent learns the selected model, effort, and endpoint from the API proxy's
GET /reflect routing field (see
api-proxy-sidecar.md); the private selection.json
is not visible to the agent.
experimental:
modelRouting: true
apiProxy:
routing:
provider: copilot
objective:
goal: cost
mode: balanced
task:
conversationFile: /tmp/gh-aw/routing-conversation.json| Field | Allowed values | Description |
|---|---|---|
provider |
copilot (default), openai, anthropic |
Restricts routing to one configured native API-proxy provider; AWF does not switch credentials or translate across providers. |
candidateModels |
non-empty array of model glob patterns | Limits router/classifier choices without widening the request policy; defaults to apiProxy.allowedModels. |
objective.goal |
cost, cost-speed |
Optimization goal used by the router |
objective.mode |
economy, balanced, robust, auto |
Fixed routing profile, or auto classification |
task.conversationFile |
non-empty string | Host path to the task conversation whose description the router classifies |
The task conversation must be written by the workflow host before AWF starts.
Callers are responsible for preparing task-relevant conversation content before
AWF starts. AWF classifies the supplied conversation as-is; it does not parse
workflow prompt markup or remove injected system instructions. If a rendered
prompt contains a system block, the caller should provide the task conversation
without that unrelated block.
The conversation is a JSON array in the router's
conversation format, for example
[{"role":"user","parts":[{"text":"Fix the failing unit test."}]}], with at
least one non-blank user message and at most 1 MiB. Its user messages form the
task description, which the router classifies once per run. The resulting
classification (task type, scope, complexity, and, for mode: auto, the
routing profile) selects the one model and effort the agent is seeded with for
the whole run. The router is not invoked again per request or per sub-agent.
The routing object is closed: objective and task are required, provider
and candidateModels are optional, and unknown properties are rejected.
Omitting provider preserves Copilot routing. OpenAI and Anthropic routing
require the matching provider to be configured for the agent and currently
require the native api.openai.com or api.anthropic.com target; AWF does not
route custom gateways or translate or forward requests across provider
boundaries. A supported routed run also requires a complete
container.images manifest containing digest-pinned references for router
and every other enabled image role. The legacy latest router default is kept
only for resolver compatibility and is not a supported tag-only routed
configuration.
The candidate pool uses models discovered for the selected native provider.
The /reflect model_api_mapping includes maintained routing metadata for
selected model families where provider /models endpoints do not publish
context limits or reasoning-effort support. The initial maintained set covers
OpenAI gpt-5.4 and gpt-5.4-2026-03-05, plus Anthropic
claude-opus-5-5, claude-opus-5, claude-fable-5-1, claude-fable-5,
claude-mythos-5-1, claude-mythos-5, claude-mythos-preview*,
claude-sonnet-5-5, claude-sonnet-5, claude-opus-4-6/4-7/4-8, and
claude-sonnet-4-6. This is deliberately not exhaustive: discovered IDs are
eligible only when endpoint and effort support are known from runtime metadata
or an exact maintained mapping. For example, o3, gpt-5-nano, and
gpt-5.4-mini do not inherit metadata from the GPT-5.4 base model and can yield
no_route. Positive context limits are also maintained in /reflect; without
one, a model is excluded from classifier preflight but may still remain
available to the router.
An explicitly empty effort list (or explicit lack of reasoning-effort support) allows one effortless choice. OpenAI uses its mapped Responses or Chat Completions protocol; Anthropic uses Messages. Unsupported effort values are discarded, and a model with no remaining advertised effort is excluded rather than converted into an effortless choice.
Request guards and alias resolution use the provider-aware
allowedModels / disallowedModels policy. Candidate filtering additionally
uses apiProxy.routing.candidateModels when supplied; those patterns only
narrow the router/classifier pool and never widen the request policy. When
omitted, candidates continue to be derived from allowedModels. Native patterns
such as gpt-* match the native model name; qualified patterns such as
github-copilot/gpt-* match that provider only. Copilot recognizes the existing
copilot, github-copilot, and github provider aliases. Matching remains
case-insensitive with * wildcards, and deny rules take precedence. A
provider-prefixed auto remains subject to dynamic-model verification: under
a denylist it must also match an explicit allow rule.
The API proxy emits structured logging events during model alias resolution. These events are critical for debugging model routing decisions in production.
The following events are emitted as JSON lines to the API proxy's stdout
(captured by Docker logging). They are always active when model aliases
are configured (apiProxy.models):
| Event | Trigger | Key fields |
|---|---|---|
model_resolution |
Every request where a model alias resolves | requested_model, resolved_model, provider, resolution_log[] |
model_rewrite |
Every request where the model field is rewritten | original_model, rewritten_model, provider |
model_fallback_activated |
Fallback strategy selected a replacement | reason, selected, candidates[] |
model_fallback_skipped |
Fallback was available but explicitly suppressed | reason, requested_model |
model_fallback_candidates |
Informational: available fallback models | candidates[], strategy |
These events are written by logRequest() in containers/api-proxy/logging.js.
When apiProxy.logging.debugTokens is true (or AWF_DEBUG_TOKENS=1),
additional diagnostic events are written to token-diag.jsonl in the directory
specified by apiProxy.logging.tokenLogDir (default: /var/log/api-proxy):
| Event | Description |
|---|---|
model_alias_resolution_step |
Each step in the alias resolution chain (input → pattern match → candidate) |
model_alias_rewrite |
Final rewrite decision with before/after model names and matched pattern |
Each diagnostic record follows the token-diag/v<version> schema:
{
"_schema": "token-diag/v0.25.40",
"timestamp": "2025-01-15T10:30:00.000Z",
"event": "model_alias_resolution_step",
"data": {
"alias": "sonnet",
"pattern": "anthropic/*sonnet*",
"candidate": "claude-sonnet-4-5",
"provider": "anthropic"
}
}apiProxy:
models:
sonnet: ["copilot/*sonnet*", "anthropic/*sonnet*"]
logging:
debugTokens: true
tokenLogDir: "/var/log/api-proxy"
diagnostics:
captureBlockedRequests: summary # false | summary | redacted | full
maxCapturedBytes: 250000| Property | Type | Default | Env var | Description |
|---|---|---|---|---|
apiProxy.logging.debugTokens |
boolean | false |
AWF_DEBUG_TOKENS |
Enable diagnostic token/model-alias logging to file |
apiProxy.logging.tokenLogDir |
string | /var/log/api-proxy |
AWF_TOKEN_LOG_DIR |
Directory for token-usage.jsonl and token-diag.jsonl |
apiProxy.diagnostics.captureBlockedRequests |
string | boolean | false |
AWF_CAPTURE_BLOCKED_LLM_REQUESTS |
Capture body-shape info for guard-blocked requests (false/true/summary/redacted/full; true is an alias for summary) |
apiProxy.diagnostics.maxCapturedBytes |
integer | 250000 |
AWF_MAX_BLOCKED_CAPTURE_BYTES |
Max bytes per record in full capture mode |
AWF produces the following structured and unstructured log files at runtime.
All JSONL files use the .jsonl extension.
All AWF JSONL records MUST include the following top-level fields:
timestamp(string, required): ISO 8601 UTC with milliseconds (YYYY-MM-DDTHH:mm:ss.SSSZ).event(string, required): Stable snake_case record discriminator._schema(string, required): Schema identifier in the form<record-type>/v<version>.
Directory: configured by logging.proxyLogsDir (default: <workDir>/squid-logs/)
| File | Format | Description | Always written |
|---|---|---|---|
access.log |
Custom text (firewall_detailed logformat) |
L7 HTTP/HTTPS traffic decisions with timestamps, client IP, domain, status, and decision codes | Yes |
audit.jsonl |
JSONL (audit/v<version> schema) |
Structured version of access log; preferred for programmatic consumption | Yes |
cache.log |
Squid native text | Squid internal diagnostics (startup, shutdown, errors) | Yes |
Directory: configured by apiProxy.logging.tokenLogDir / AWF_TOKEN_LOG_DIR
(default: /var/log/api-proxy/; must be /var/log/api-proxy or a subdirectory to be preserved by AWF's default bind mount)
On the runner, these files are preserved under <logging.proxyLogsDir>/api-proxy-logs/
(or /tmp/api-proxy-logs-<ts>/ when proxyLogsDir is not set). After cleanup AWF logs
Token usage log available at: <path> and, when $GITHUB_ENV is set, exports
AWF_TOKEN_USAGE_LOG=<path> so later workflow steps can locate token-usage.jsonl
without hardcoding a path (see ARC + DinD).
| File | Format | Description | Always written |
|---|---|---|---|
token-usage.jsonl |
JSONL (token-usage/v<version> schema) |
Per-API-call token usage and cost records | Yes (when API proxy is active) |
token-diag.jsonl |
JSONL (token-diag/v<version> schema) |
Diagnostic events: model resolution steps, alias rewrites, token budget decisions | Only when apiProxy.logging.debugTokens: true |
blocked-request-diag.jsonl |
JSONL (blocked-request-diag/v<version> schema) |
Body-shape diagnostics for guard-blocked requests (effective tokens, AI credits, etc.) | Only when apiProxy.diagnostics.captureBlockedRequests is set |
otel.jsonl |
JSONL (OpenTelemetry spans) | Distributed tracing spans; written as local fallback when no OTLP collector is configured | Only when OTEL is active and no collector endpoint set |
Directory: /var/log/cli-proxy/ (or AWF_CLI_PROXY_LOG_DIR)
| File | Format | Description | Always written |
|---|---|---|---|
access.jsonl |
JSONL | CLI proxy request audit records (gh CLI invocations routed through DIFC proxy) | Yes (when CLI proxy is active) |
The API proxy also emits JSON lines to stdout (captured by docker logs).
These are always active and include model resolution events (model_resolution,
model_rewrite, model_fallback_*). Use docker logs awf-api-proxy or
the AWF diagnostic log collection to access them.
Model alias logging was introduced in v0.25.40 (PR #2329). The diagnostic
file mechanism (token-persistence.js) was refactored into a dedicated module
in v0.25.50 but the logging events and their format have been stable since
initial release.
When a guard hard-rails a request (e.g. effective_tokens_limit_exceeded,
ai_credits_limit_exceeded, max_runs_exceeded), the api-proxy can write a
structured diagnostic record to blocked-request-diag.jsonl. This is
opt-in and disabled by default.
Set the environment variable or config key before starting the container:
# Minimal (body-shape only, no content):
AWF_CAPTURE_BLOCKED_LLM_REQUESTS=summary
# Include first 200 chars of each message (for debugging over-large tool results):
AWF_CAPTURE_BLOCKED_LLM_REQUESTS=redacted
# Full body up to AWF_MAX_BLOCKED_CAPTURE_BYTES (default 250 000 bytes):
AWF_CAPTURE_BLOCKED_LLM_REQUESTS=full
AWF_MAX_BLOCKED_CAPTURE_BYTES=250000Or via config YAML:
apiProxy:
diagnostics:
captureBlockedRequests: summary # false | summary | redacted | full
maxCapturedBytes: 250000| Mode | Content | Use case |
|---|---|---|
false (default) |
Nothing written | Production default |
summary |
Counts, sizes, hashes — no content | Safe for normal debugging; identify which message/tool-result was large |
redacted |
Summary + first 200 chars per message | Debug prompt growth without full disclosure |
full |
Full body up to maxCapturedBytes |
Local/private runs only; explicitly document and review |
Each record follows the blocked-request-diag/v<version> schema:
{
"_schema": "blocked-request-diag/v0.26.0",
"timestamp": "2025-01-15T10:30:00.000Z",
"event": "blocked_request_diag",
"capture_mode": "summary",
"request_id": "bc446626-a67b-4a78-a8c3-7293a2bc7306",
"provider": "anthropic",
"path": "/v1/messages",
"guard_type": "effective_tokens_limit_exceeded",
"guard_totals": {
"total_effective_tokens": 27198679,
"max_effective_tokens": 25000000
},
"body_transformed": true,
"inbound_bytes": 184320,
"body_bytes": 185040,
"body_sha256": "a3f2b1c8d9e0f1a2",
"model": "claude-opus-4.7",
"streaming": true,
"message_count": 52,
"tool_result_count": 14,
"message_sizes": [
{ "role": "user", "content_type": "text", "chars": 312, "bytes": 312, "estimated_tokens": 78 },
{ "role": "assistant", "content_type": "text", "chars": 1840, "bytes": 1840, "estimated_tokens": 460 },
{ "role": "user", "content_type": "tool_result", "chars": 94321, "bytes": 94321, "estimated_tokens": 23580, "tool_blocks": 3 }
]
}summarymode captures no message content and is safe for shared/public workflow runs.redactedmode includes short previews; review before attaching to public issues.fullmode captures potentially sensitive prompt and tool-result content. Use only for private runs and rotate or delete the artifact promptly.- The file is written to
AWF_TOKEN_LOG_DIRalongsidetoken-usage.jsonland is governed by the same artifact-retention policy.
The optional top-level enclaves array defines AWF's sole supported private-repository execution surface. It is structurally identical to the gh-aw compiler's enclave frontmatter: every entry declares exactly one script or agent executor, its own non-empty repos list, and entry-level shared controls including timeout, runtime, image, resource limits, and disclosure limits. AWF stages immutable repository seeds on the host, starts one AWF-owned enclave-mcp-server, maintains one shared per-repository ledger for the run, and exposes configured executors only through compiler-launched gh-aw-mcpg.
Dynamic repository-policy entries (dynamic in place of repos on an agent
entry, per docs/adr/0001-agent-enclaves.md) select one canonical repository at
runtime, receive one short-lived github-repository-read-v1 identity, and read
that repository through GitHub MCP without cloning or mounting a seed. They run
only when the gh-aw compiler has started mcpg's
github-repository-delegation-v1 controller and handed AWF its private control
endpoint and capability; otherwise AWF rejects the run. See §14.1a for the
envelope, the handoff, and the identity lifecycle.
enclaves:
- script: {}
repos:
- repo: octo-org/private-service
sensitivity: confidential
timeout: 45
- agent:
model: gpt-5
maxModelRequests: 3
maxModelTokens: 10000
tools:
github:
allowed:
- list_issues
- issue_read
allowedRepos:
- octo-org/private-service
minIntegrity: none
runtime: gvisor
memoryLimit: 256m
maxOutputBytes: 2048
maxInvocations: 3
repos:
- repo: octo-org/private-service
sensitivity: confidential
timeout: 180- Script executor — an entry keyed by
script; launches a no-network, read-only, single-use Python sandbox. An emptyscript: {}object is valid and selects AWF's pinned defaults. - Agent executor — an entry keyed by
agent; launches a bounded single-use Copilot enclave.agent.modelis REQUIRED. Optionalagent.tools.github(or the deprecated legacyagent.github.cli: issues-read-v1marker; the two are mutually exclusive) adds compiler-owned shared mcpg as the only additional peer. - Entry-level controls —
runtime,image,memoryLimit,cpuLimit,pidsLimit,tmpfsLimit,maxOutputBytes, andmaxInvocationsapply to the entry's selected executor.script.maxScriptBytesand agentmaxTaskBytes,maxModelRequests, andmaxModelTokensremain executor-specific. Network and interpreter are AWF-owned invariants, not input fields.
At most one entry MAY exist per executor kind, and each entry MUST declare exactly one executor key. Every entry's repos list is merged into one trusted repository catalog: a repository shared by both entries MUST declare the same sensitivity, because sensitivity fixes one shared per-run information budget that both executors debit.
timeout is a per-invocation wall-clock bound in seconds. It defaults to 30 for script entries and 120 for agent entries, and values above 4740 are rejected. Responses use fixed timing buckets at 100 ms, 1 second, 10 seconds, 60 seconds, 120 seconds, 180 seconds, 240 seconds, 300 seconds, 600 seconds, 1200 seconds, 2400 seconds, and 4800 seconds, followed by a cryptographically random, secret-independent delay from 0 through 1000 ms. The canonical enclave MCP tools use a fixed toolTimeout of 4860 seconds, covering the maximum bucket, response jitter, and a bounded transport allowance.
gvisor requires an exactly registered runsc runtime and never falls back. sbx remains fail-closed for both executors until the audited capability proof lands.
cloud-hypervisor is a reserved enclave runtime value governed by
ADR 0002. AWF preserves the
selection through parsing and validates the shared top-level cloudHypervisor
preview, host, and attested-artifact configuration, but currently fails closed
before launching an enclave. Execution remains disabled until the host executor,
dedicated script and agent rootfs artifacts, workload-specific networking,
resource parity, and durable recovery gates land. The initial scope is static
script and static agent entries only; dynamic entries and custom image
overrides are rejected, and no configuration falls back to another runtime.
An enclave-only cloud-hypervisor selection requires top-level
cloudHypervisor configuration but does not select Cloud Hypervisor for the
primary agent or alter its mounts, TTY, or container runtime.
The agent executor additionally requires enableApiProxy, a configured provider route for its fixed engine/profile, a configured model, and the absence of enableDind. AWF validates those requirements before repository staging.
enclaves:
- agent:
model: gpt-5
dynamic:
allowedOwners:
- octo-org
allowedRepositories: []
sensitivity: confidential
executor: agent
githubPolicy:
version: github-repository-read-v1
tools:
- list_issues
- issue_read
maxRepositories: 4
limits:
timeoutSeconds: 180
memoryLimit: 256m
cpuLimit: "1"
pidsLimit: 128
tmpfsLimit: 256m
maxOutputBytes: 2048
maxTaskBytes: 4096
maxModelRequests: 8
maxModelTokens: 4096
quotas:
maxInvocations: 100
maxOutputBytes: 100000
maxExecutionSeconds: 3600
auditLabels:
- awf-enclave-dynamic
expiresAt: "2030-01-01T00:00:00Z"The dynamic object is byte-for-byte the envelope the gh-aw compiler emits. An agent entry MUST declare exactly one of repos or dynamic, never both, and a script entry MUST NOT declare dynamic. The object is closed (unknown fields are rejected) and every field is REQUIRED:
allowedOwners/allowedRepositories— exact canonical lowercase ASCII owner scopes (owner) and/orowner/reposelectors. Either list MAY be empty, but at least one selector MUST be reachable for the envelope to admit anything. AWF performs no trimming, case folding, Unicode normalization, or URL decoding before matching; a selector that is not already in this exact canonical form is rejected.sensitivity— one ofpublic,trusted,internal,confidential,sealed; fixes the shared per-repository information budget an admitted repository debits (§14.4). An admitted repository opens its balance in the same run-wide ledger the static executors debit, so re-admitting a repository can never refill a budget it has already spent.executor— fixed toagent; any other value is rejected.githubPolicy— fixed to{ version: "github-repository-read-v1", tools: ["list_issues", "issue_read"] }. Any other version, tool set, or additional tool is rejected; this is the sole supported dynamic GitHub policy.maxRepositories— integer1..1000: distinct repositories this envelope may admit for the run.limits— per-invocation trusted bounds, all REQUIRED:timeoutSeconds(1..4740, whole seconds),memoryLimit,cpuLimit,pidsLimit(1..4096),tmpfsLimit,maxOutputBytes(1..8192),maxTaskBytes(1..65536),maxModelRequests(1..64),maxModelTokens(1..32768). These are the same resource and response controls a static agent entry declares at entry level.quotas— run-wide totals debited across every admission under this envelope, all REQUIRED:maxInvocations(1..10000),maxOutputBytes(1..1048576),maxExecutionSeconds(1..86400).auditLabels— a non-empty, unique array of at most 32 opaque labels matching^[A-Za-z0-9][A-Za-z0-9_.:-]{0,127}$. Labels are what AWF and mcpg reconcile dynamic state against at shutdown; they are never repository names or credentials.expiresAt— an absolute ISO-8601 timestamp, never later than the workflow job lifetime; admission at or after this time is denied.
A dynamic entry runs only when the gh-aw compiler has started mcpg's
github-repository-delegation-v1 controller and handed AWF both private
values:
AWF_ENCLAVE_GITHUB_DELEGATION_CONTROL_ENDPOINT— the loopback-only control endpoint, published by Docker on the runner's own127.0.0.1(http://127.0.0.1:<port>/internal/awf-enclave-mcp-control/github-repository-delegation-v1). AWF accepts only the literal loopback hosts127.0.0.1and[::1]; a resolver-dependent name such aslocalhost, a routable or wildcard address, embedded credentials, a query string, a fragment, or any other path is rejected. An omitted or explicit:80is accepted as port 80 (the WHATWG URL parser normalizes both to the same value).AWF_ENCLAVE_GITHUB_DELEGATION_CONTROL_CAPABILITY— the AWF-only 256-bit hex control capability.
A missing, partial, or malformed handoff is a hard failure. AWF never falls back to a static seed catalog, a job-lifetime identity, or a broader policy.
Because the control listener is published on host loopback, only the AWF host
process can reach it through the published port, which is why the control
client runs there rather than in a container. That is a property of the
publication rather than a general routing guarantee: under network isolation the
in-container listener binds 0.0.0.0, and a peer co-attached to a Docker
network with mcpg addresses the container IP directly without traversing the
published port. The control plane is therefore protected by authentication —
every request must carry the AWF-only capability, and mcpg rejects anything else
with 403 delegation_access_denied.
AWF takes custody of both values before any inherited
environment is assembled, stages them into the 0700 enclave private root with
exclusive 0600 files, and never mounts either one into the enclave MCP broker,
the single-use executor, the model sidecar, the general MCP route, or the
delegated data plane. Both variables are also in the primary agent's environment
exclusion set.
The broker routes each enclave_run_agent through AWF's host-side admission
authority over a 0700 request/response directory that is bind-mounted only
into the broker. That channel carries the caller's selector, the exact finite
output-schema hash, and the invocation's settlement; it never carries the
control endpoint, the control capability, the identity handle, the compiler
envelope, mcpg's state path, or its policy generation.
A dynamic-only entry needs no GH_TOKEN/GITHUB_TOKEN, clones no repository,
writes no seed catalog (not even an empty one), and mounts neither /awf/seed
nor a seed map. A run that also declares a separate static entry keeps the full
static staging path unchanged.
Admission runs in src/enclave/dynamic-registry.ts before any repository
content is exposed and before any control call is made:
- Admission is idempotent by
(run, enclave entry, invocation id, canonical repository). A retried request with the same key returns the previously recorded outcome, and joins an in-flight admission rather than reserving capacity a second time. A different repository under an already-bound invocation id is rejected rather than rebinding. maxRepositoriesand all three quotas are reserved synchronously before any asynchronous lookup, so concurrent admissions can never both observe capacity and both commit.maxOutputBytesandmaxExecutionSecondsare reserved at their per-invocation worst case (limits.maxOutputBytes,limits.timeoutSeconds) and replaced by the actual reported usage once the invocation settles. Charges committed after admission stay committed even if the invocation later fails.- Every outcome — malformed selector, policy denial, resolution failure, or success — is delayed to the same fixed timing bucket (§14.3's bucket list) plus secret-independent jitter, so elapsed wall-clock time cannot distinguish failure causes. Every failure returns one non-disclosing canonical denial.
For each admitted invocation AWF calls mcpg's control API, authenticated with
the capability in Authorization, using bounded request/response bodies,
explicit timeouts, and strict JSON validation:
| Operation | Path |
|---|---|
| create or confirm | /internal/awf-enclave-mcp-control/create-or-confirm |
| status | /internal/awf-enclave-mcp-control/status |
| reconcile | /internal/awf-enclave-mcp-control/reconcile |
| revoke | /internal/awf-enclave-mcp-control/revoke |
| revoke by labels | /internal/awf-enclave-mcp-control/revoke-by-labels |
Operation paths are siblings of the controller name in the exported endpoint,
not children of it. requested_ttl is a positive whole number of seconds,
matching the user-configured limits.timeoutSeconds and mcpg's
max_identity_ttl unit; timestamps are RFC 3339.
Every create-or-confirm response is verified before it is trusted: non-empty
handle and executor bearer, an exact repository match against the admitted
selector, tool_policy equal to github-repository-read-v1, tools exactly
list_issues and issue_read, an admitted SHA that matches when one was
requested, and an expiry no later than the requested TTL or the invocation
deadline. Only the executor bearer leaves the AWF host process; the handle stays
in AWF-private state.
At startup AWF calls status, revokes any stale labelled identity, and only
then calls the transactional reconcile. New admissions stay blocked until that
sequence succeeds. Every terminal path — success, agent failure, schema failure,
timeout, cancellation, broker error — settles the reserved quota and revokes the
identity, and revoke-by-labels sweeps the run at teardown. An unresolved
revocation re-blocks admissions and downgrades the invocation to the canonical
error rather than returning a success-shaped result.
admitted_default_branch_sha is optional in both the ADR and mcpg's contract.
AWF has no already-authorized, repository-confined path to resolve a
default-branch SHA before the delegated identity exists, and
github-repository-read-v1 grants only list_issues and issue_read
afterwards. AWF therefore omits the field, never widens a token or tool to
obtain one, and audits every repository read as a live read. A read is
marked pinned only when the control binding actually carries a resolved SHA.
The primary agent reaches private-repository execution only through these MCP tools:
enclave_run_script({
privateRepo: "owner/repo",
schema: <finite disclosure schema>,
script: <bounded UTF-8 Python source>
})
enclave_run_agent({
privateRepo: "owner/repo",
schema: <finite disclosure schema>,
prompt: <bounded UTF-8 task prompt>
})
Both tool schemas are closed (additionalProperties: false). A call can never provide or override images, runtimes, models, engines, profiles, mounts, network settings, credentials, repository catalogs, budgets, timeouts, or any other trusted control.
The primary agent MUST NOT receive a broker socket, wrapper binary, direct server URL, capability token, repository seed, ledger state, or alternate enclave transport.
rateLimiting.maxNumToolCalls (CLI: --max-num-tool-calls <n>) caps the number of enclave tool calls the primary agent may make in one AWF run. It is OPTIONAL and defaults to unlimited: when omitted, enclave-mcp-server keeps no counter, persists no state, and publishes no advisory text. Setting it without any configured enclave is a startup error.
When configured:
- Scope — one run-wide budget shared by
enclave_run_scriptandenclave_run_agent, independent of each entry'smaxInvocationsand of the per-repository disclosure ledger. Dynamic repository admission is performed by the broker inside anenclave_run_agentcall and does not consume additional units; tool calls made inside an enclave agent are bounded bymaxModelRequests/maxModelTokensinstead. - Counting — every attempted, well-formed
tools/callto a published enclave tool consumes one unit before any other admission decision, including calls later rejected as busy, oversized, invalid, or failed. Malformed JSON-RPC requests that never name a published tool do not count. - Advisory — each published tool description, and the
initializeresult'sinstructions, state: "You are allowed to make at most N enclave tool calls in this run. After that, the system will deny any further enclave tool calls." - Denial — once exhausted, calls never reach an executor and return an in-band
isErrorresult whose text is "Max tool call count reached, no more tool calls are allowed. Make a decision based on what you already have in context." The decision depends only on the caller's own call count and trusted configuration, so it discloses no repository information. On the first denied call the broker logs one warning,Max tool call count reached. {"toolName":…,"agentName":<executor kind>,"sessionID":<run id>,"maxToolCalls":N}, so repeated retries cannot flood the logs. - Persistence — the count is persisted (mode
0600) in the broker's private control directory keyed by the AWF run id, so a restarted broker for the same run resumes the count rather than resetting it. - Exemptions — the enclave server publishes no final-answer or structured-result submission tool; the agent's own result/safe-output tools are served elsewhere and are never counted or denied by this cap.
enclave-mcp-server joins only the private awf-enclave-mcp-control network. The compiler launches gh-aw-mcpg, labels it for the run, and passes AWF the private gateway endpoint plus a run-unique capability/identity handoff. The server is reachable only through that gateway.
When the agent executor is enabled, each invocation joins only the dedicated
internal awf-enclave-agent network. Its mandatory peer is the dedicated
enclave API proxy. When GitHub access is configured (agent.tools.github or
the deprecated legacy agent.github.cli: issues-read-v1 marker), the only
additional peer is compiler-owned shared mcpg.
AWF attaches that existing container directly at 172.31.0.40 under the fixed
alias awf-enclave-github-mcp; the enclave uses /mcp/github on port 8080.
Squid, the primary agent, general proxies, safe outputs, and the enclave MCP
server itself remain excluded.
The base compiler handoff from github/gh-aw#50920 and late backend
rediscovery from github/gh-aw-mcpg#10784 are present in mcpg v0.4.15, which
reports MCP Gateway spec 1.16.0. The base floor remains spec 1.15.0 and a
post-v0.4.8 mcpg release. That floor covers static entries only; dynamic
repository admission requires mcpg v0.4.18 or newer for the whole-second
delegation duration encoding (see §14 and docs/enclaves-architecture.md).
GitHub access additionally requires compiler support for mcpg multi-agent
identities and policies, tracked by github/gh-aw#57787. The compiler MUST gate
or pin the first supporting AWF release. Older AWF versions reject the closed
github/tools.github fields; AWF has no compatibility fallback.
While the backend is still starting, mcpg may return retryable HTTP 503 backend_unavailable. AWF retries initialize with bounded backoff until AWF_ENCLAVE_MCP_READINESS_TIMEOUT_MS expires, then fails closed before the primary agent starts.
The gateway has two separate authorization hops. AWF's upstream
Authorization header authenticates mcpg to the AWF-owned enclave server with
AWF_ENCLAVE_MCP_CAPABILITY. mcpg then generates the client-facing gateway
Authorization header from its gateway agent ID/API key; this is the header
present in mcpg's rewritten gateway output consumed by engine config adapters.
That downstream credential is distinct from the AWF capability.
Adapters consuming mcpg's rewritten output MUST keep the client-facing
Authorization value runtime-only: they must not resolve it while generating
configuration or persist the resolved credential under GITHUB_WORKSPACE (or
any other agent-readable path). The AWF upstream contract cannot enforce this
requirement on the downstream output/converter path.
When GitHub access is configured, the compiler supplies the shared gateway
contract (AWF_ENCLAVE_MCP_GATEWAY_CONTAINER, AWF_ENCLAVE_MCP_GATEWAY_ENDPOINT, and
AWF_ENCLAVE_MCP_GATEWAY_IDENTITY) plus a distinct
AWF_ENCLAVE_GITHUB_MCP_AGENT_ID. The compiler configures that identity in
mcpg gateway.agentIds and restricts it with gateway.agentPolicies to the
github server, agent.tools.github.allowed (or the legacy marker's fixed
list_issues/issue_read pair), and the trusted enclave repository catalog
named in agent.tools.github.allowedRepos:
list_issuesissue_readwith methodgetissue_readwith methodget_comments
AWF itself only validates the closed agent.tools.github contract — that
allowed is a non-empty subset of the two supported tools, that
allowedRepos is a non-empty list of exact owner/repository slugs each
present in this same agent entry's own repos list (not the run-wide merged
catalog shared with the script executor), and that minIntegrity (when set)
is one of none, unapproved, approved, or merged — and wires the shared
gateway connection. The legacy agent.github.cli: issues-read-v1 marker has
no allowedRepos field of its own; its repository scope comes entirely from
the compiler-defined trusted catalog baked into that fixed marker, not from
anything AWF validates. Repository and integrity enforcement live entirely in
the compiler-created, enclave-specific mcpg identity; AWF never broadens or
replaces that policy.
AWF stores the enclave identity in a mode-0600 private file, removes it from
the host environment, and gives each invocation a read-only private copy. The
enclave sends the identity directly as Authorization to
http://172.31.0.40:8080/mcp/github. It contains no gh executable and fails
preflight if one is present. Before primary-agent work begins, AWF initializes
the endpoint and fails closed unless tools/list advertises exactly the
configured allowed tools — no more, no fewer.
This mcpg identity lasts for the job rather than one invocation. Its repository policy therefore covers the union of trusted repositories configured for the enclave agent; it is not independently expired, revoked, or narrowed to the repository assigned to a particular invocation. This is an explicit tradeoff of direct shared-mcpg connectivity. AWF's per-invocation process, seed, admission, shared-ledger debit, finite output schema, and timing controls remain in force.
Script and agent calls debit the same live per-repository balance and share one AWF-owned admission lane. Switching executor kinds never resets or forks the ledger.
Repository sensitivity selects the response schema and per-run disclosure policy:
| Sensitivity | Per-run budget | Response schema |
|---|---|---|
trusted |
Unmetered | Structured schemas, including free-form string nodes |
public |
Unmetered | Finite schemas only |
internal |
64 bits | Finite schemas only |
confidential |
8 bits | Finite schemas only |
sealed |
0 bits | No invocation can be admitted |
The trusted class is intended only for repositories whose content may be
returned to the primary agent without confidentiality accounting. It permits
an exact { "type": "string" } schema node at any otherwise valid schema
position. Strings remain bounded by maxOutputBytes and the global 8192-byte
result ceiling. The schema remains strict and structured: floats, optional
fields, extra properties, $ref, recursion, regex schemas, and untagged unions
are unsupported. Every other sensitivity rejects a schema containing a
free-form string node before launching an enclave.
enclave_run_agent necessarily sends repository-derived content to the configured model provider through the dedicated API proxy. The ledger bounds what the calling agent learns; it does not bound what the provider sees.
Legacy bounded smoke and runtime-matrix workflow assets have been removed from the owned surface. Until a unified gh-aw enclave smoke workflow exists, local coverage remains unit-focused:
src/services/enclave-mcp-service.test.tssrc/services/enclave-agent-service.test.tssrc/enclave/script-runner-spec.test.tssrc/enclave/agent-runner-spec.test.tssrc/enclave/manager.test.tssrc/enclave/mcp-server.test.tssrc/enclave/agent-mcp-server.test.ts
See Unified Enclave Architecture for the operator-facing summary.
- RFC 2119 — Key words for use in RFCs to Indicate Requirement Levels
docs/awf-config.schema.json— Machine-readable JSON Schema for configuration documents (normative)
AWF emits structured JSONL artifact files at runtime. Most record types have
a corresponding JSON Schema in the schemas/ directory; opt-in diagnostic
formats are documented inline in this spec instead:
| Schema | JSONL file | Description |
|---|---|---|
schemas/audit.schema.json |
audit.jsonl |
L7 HTTP/HTTPS traffic decisions (allowed/denied) from the Squid proxy |
schemas/token-usage.schema.json |
token-usage.jsonl |
Per-API-call token usage records from the api-proxy sidecar |
schemas/otel-span.schema.json |
otel.jsonl |
OpenTelemetry span records emitted by the local file exporter |
schemas/cli-proxy-access.schema.json |
access.jsonl (cli-proxy) |
CLI proxy request audit records |
| (inline, see §13.2) | token-diag.jsonl |
Model alias resolution steps and diagnostic events (opt-in via apiProxy.logging.debugTokens) |
| (inline, see §13.6) | blocked-request-diag.jsonl |
Body-shape diagnostics for guard-blocked requests (opt-in via apiProxy.diagnostics.captureBlockedRequests) |
Schema files do not carry an independent version. The repository release tag serves as the version:
- The
$idfield in each schema resolves to a stable release download URL. - Each JSONL record includes a
_schemawire-format field encoding the record type and AWF version (e.g.,"_schema": "audit/v0.26.0"). - Consumers SHOULD use a prefix match (
_schema.startsWith("audit/")) rather than an exact match to handle future versions gracefully.
Versioned (release assets):
https://github.com/github/gh-aw-firewall/releases/download/<tag>/awf-config.schema.json
https://github.com/github/gh-aw-firewall/releases/download/<tag>/audit.schema.json
https://github.com/github/gh-aw-firewall/releases/download/<tag>/token-usage.schema.json
https://github.com/github/gh-aw-firewall/releases/download/<tag>/otel-span.schema.json
https://github.com/github/gh-aw-firewall/releases/download/<tag>/cli-proxy-access.schema.json
Latest (main branch):
https://github.com/raw/github/gh-aw-firewall/main/docs/awf-config.schema.json
https://github.com/raw/github/gh-aw-firewall/main/schemas/audit.schema.json
https://github.com/raw/github/gh-aw-firewall/main/schemas/token-usage.schema.json
https://github.com/raw/github/gh-aw-firewall/main/schemas/otel-span.schema.json
https://github.com/raw/github/gh-aw-firewall/main/schemas/cli-proxy-access.schema.json
- docs/arc-dind.md — ARC/DinD split-filesystem architecture, sysroot staging, and end-to-end configuration examples
- docs/environment.md — Usage guide for environment variables
- docs/authentication-architecture.md — Credential isolation architecture and diagrams
- docs/api-proxy-sidecar.md — API proxy sidecar configuration including OIDC authentication for Azure OpenAI
- schemas/README.md — JSONL schema directory with validation examples and versioning policy