Skip to content

Latest commit

 

History

History
2766 lines (2238 loc) · 141 KB

File metadata and controls

2766 lines (2238 loc) · 141 KB

AWF Configuration Specification

Abstract

This specification defines the configuration model, processing rules, and environment semantics for the Agentic Workflow Firewall (AWF). It is the normative reference for:

  • the awf CLI runtime (--config)
  • tooling that compiles workflows into AWF invocations (e.g., gh-aw)
  • IDE and static-analysis validation via JSON Schema

The machine-readable schema is published alongside this specification at docs/awf-config.schema.json (live, tracking main) and as a versioned release asset (e.g., https://github.com/github/gh-aw-firewall/releases/download/v0.23.1/awf-config.schema.json).

Status of This Document

This document is normative. Informative notes are marked with Note: or placed in blockquotes. All other text is normative unless stated otherwise.

1. Conformance

The key words "MUST", "MUST NOT", "REQUIRED", "SHALL", "SHALL NOT", "SHOULD", "SHOULD NOT", "RECOMMENDED", "MAY", and "OPTIONAL" in this document are to be interpreted as described in RFC 2119.

A conforming AWF configuration document is one that:

  1. is valid JSON or YAML;
  2. satisfies all constraints defined by docs/awf-config.schema.json; and
  3. contains no properties beyond those defined by the schema (closed-world assumption).

A conforming AWF implementation MUST accept every conforming configuration document and MUST reject every non-conforming one.

2. Processing Model

When the user invokes awf --config <path|-> -- <command>, a conforming implementation MUST execute the following steps in order:

  1. If <path> is -, read configuration bytes from standard input.
  2. Determine the serialisation format:
    • If <path> ends with .json, parse as JSON.
    • If <path> ends with .yaml or .yml, parse as YAML.
    • Otherwise, attempt JSON first; if that fails, attempt YAML.
  3. Validate the parsed document against docs/awf-config.schema.json.
  4. On validation failure, abort with non-zero exit status (see §7).
  5. Map configuration fields to CLI-option semantics per §5.
  6. Apply precedence rules per §3.

3. Precedence Rules

The effective value for any configuration parameter SHALL be determined by the following precedence order (highest wins):

  1. Explicit CLI flags
  2. Config file (--config)
  3. AWF internal defaults

Note: This model enables reusable, checked-in configuration files with environment-specific CLI overrides.

4. Data Model

The root object of a conforming configuration document MAY contain the following top-level properties. All are OPTIONAL:

Property Type Description
$schema string JSON Schema URI for IDE validation
network object Network egress configuration
filesystem object Host filesystem write-boundary configuration (see §4.1)
apiProxy object API proxy sidecar configuration
security object Security and isolation settings
container object Container and Docker settings
cloudHypervisor object Cloud Hypervisor v53.0 microVM preview settings (see §4.2)
chroot object Chroot execution overrides for split-filesystem ARC/DinD runners
dind object Bootstrap helpers for ARC/DinD split runner/daemon filesystems
runner object Runner topology declaration (standard vs. ARC/DinD)
environment object Environment variable propagation (see §8)
logging object Logging and diagnostics
rateLimiting object Egress rate limiting
platform object GitHub platform deployment type declaration
enclaves object Unified private-repository enclave subsystem (see §14)

Property-level constraints, types, and descriptions are defined normatively by docs/awf-config.schema.json.

4.1 Filesystem write boundary

When filesystem.allowWrite is present, AWF MUST mount existing writable host binds read-only except at the listed guest-visible absolute paths. The list narrows existing access: it MUST NOT expose a new host path or make an otherwise read-only mount writable. Every listed path MUST exist within an existing writable host mount. An empty list makes all non-internal host bind mounts read-only.

AWF-owned agent log and session-state mounts and required virtual devices remain writable so the sandbox can operate. filesystem.allowWrite is supported by the Docker and gVisor compose runtimes and by the Cloud Hypervisor microVM runtime, where it is enforced by the host mount tree that backs each virtio-fs export (see docs/cloud-hypervisor-foundation.md). Cloud Hypervisor has no always-writable internal mounts, so every export it publishes is subject to the policy, including /tmp/gh-aw and the guest home directory at /workspace/.awf-home; paths not covered by allowWrite become read-only. Because every listed path MUST already exist, and because AWF MUST NOT auto-create or exempt the guest home, a Cloud Hypervisor workload that needs a writable home MUST have the backing host directory $GITHUB_WORKSPACE/.awf-home created before AWF starts and MUST then list the guest path /workspace/.awf-home in allowWrite; otherwise planning fails because the path does not exist within a writable export. AWF rejects filesystem.allowWrite with the sbx runtime and with Docker-in-Docker agent execution.

Any directory the workload writes to at run time — including caller-owned ones under /tmp/gh-aw such as gh-aw's repo-memory directory /tmp/gh-aw/repo-memory — MUST appear in allowWrite, or writes to it fail with EROFS. To make that diagnosable, the Cloud Hypervisor runtime logs the resolved boundary (per export: writable, read-only, or read-only with named writable paths) before the guest boots.

4.2 Cloud Hypervisor microVM preview

The cloudHypervisor surface binds Cloud Hypervisor v53.0, virtiofsd, the PCI-capable guest kernel, rootfs, and shared AWF guest supervisor to one release-pinned GitHub-attested manifest. It requires explicit --cloud-hypervisor-preview opt-in plus container.containerRuntime: "cloud-hypervisor" to execute a workload. The supported host target is GitHub-hosted Ubuntu x86_64 runners with KVM. Runner eligibility, missing KVM access, and unsupported host-policy failures warn and fall back to the standard Docker backend; invalid configuration and artifact trust, integrity, digest, or version failures remain fatal. Host eligibility is enforced by src/cloud-hypervisor/host-eligibility.ts. See src/cloud-hypervisor/preflight.ts for the artifact/host trust-check module, src/cloud-hypervisor/launcher.ts for the secure host launcher (network-namespace join, privilege drop, Landlock filesystem confinement, and seccomp), src/cloud-hypervisor/manager.ts for the VM lifecycle, and guest/cloud-hypervisor/ for the guest artifact build/verification pipeline. See docs/cloud-hypervisor-foundation.md for the full architecture and security-boundary writeup.

cloudHypervisor.mountPolicy controls host directory exposure:

  • workspace-only is the default and recommended secure posture. It does not infer a mount from RUNNER_TOOL_CACHE or AGENT_TOOLSDIRECTORY. The workspace and the narrow gh-aw runtime directories (RUNNER_TEMP/gh-aw and /tmp/gh-aw) remain eligible when present so generated workflows can run.
  • workspace-and-tool-cache explicitly opts into exporting the whole path named by RUNNER_TOOL_CACHE, falling back to AGENT_TOOLSDIRECTORY. AWF requires that value to name an existing real directory and stages it recursively read-only on the host before virtiofsd starts. Cache paths that overlap writable export sources, missing paths, and unknown policy values are errors.

AWF forwards RUNNER_TOOL_CACHE or AGENT_TOOLSDIRECTORY into the guest only when the corresponding export exists. gh-aw versions that generate a command which scans RUNNER_TOOL_CACHE must emit --cloud-hypervisor-mount-policy workspace-and-tool-cache; workflows that do not need runner-installed tools should retain the default.

Release test artifacts (cloud-hypervisor-test-x86_64) are x86_64 test/preview artifacts built and verified by both the release workflow and test-cloud-hypervisor.yml, which also runs the live-KVM parity/security smoke suite (scripts/ci/cloud-hypervisor-live-smoke.sh) on GitHub-hosted Ubuntu x86_64 runners when triggered by manual dispatch or the cloud-hypervisor-kvm pull request label. They are published as explicitly named release/workflow assets, but are not production defaults and are never auto-downloaded. See docs/cloud-hypervisor-foundation.md for the complete CI workflow specification and troubleshooting reference.

Releases also publish distinct enclave-script-rootfs.ext4 and enclave-agent-rootfs.ext4 artifacts with a separately attested enclave manifest, per-role provenance bundles, and per-role SBOMs. The release setup script verifies those bindings before exporting unambiguous cached role paths. These artifacts do not enable Cloud Hypervisor enclave execution by themselves: the runtime remains terminal until every ADR 0002 host-executor gate is present, and enclaves[].image remains invalid for runtime: cloud-hypervisor.

Normal execution requires artifactManifestPath, artifactManifestBundlePath, and artifactReleaseTag. AWF verifies the Sigstore bundle offline with gh attestation verify, constraining the signer to github/gh-aw-firewall/.github/workflows/release.yml, before parsing any manifest field. It then verifies all five local artifact names and SHA-256 digests. The legacy sha256 object is not a trust root.

Same-run development builds may set developmentAllowUnattestedArtifacts: true only when the process environment also contains AWF_CLOUD_HYPERVISOR_DEVELOPMENT_ALLOW_UNATTESTED_ARTIFACTS=1. This preview-only dual opt-in still requires all five hashes and must not be used for release artifacts.

5. CLI Mapping

This section is normative.

Tools generating AWF invocations (such as gh-aw) SHOULD use the mapping below. The left side is the configuration-document path; the right side is the corresponding CLI flag.

Security-sensitive values (API keys, tokens, and credential secrets) MUST be provided via environment variables, not AWF config documents. Non-sensitive AWF settings MAY be supplied via config files, including stdin (--config -).

  • network.allowDomains[] → --allow-domains <csv>
  • network.blockDomains[] → --block-domains <csv>
  • network.dnsServers[] → --dns-servers <csv>
  • network.subnet → --network-subnet <cidr> (relocates the awf-net Docker network; use when the default 172.30.0.0/24 collides with the host or cluster network, e.g. the OpenShift service CIDR 172.30.0.0/16)
  • network.upstreamProxy → --upstream-proxy
  • network.isolation → --network-isolation (experimental; enforces egress via Docker network topology instead of host iptables)
  • network.verifySbxEgress → --verify-sbx-egress (fail-closed verification that Docker sbx direct traffic cannot bypass Squid; requires the sbx runtime)
  • network.topologyAttach[] → --topology-attach <name> (repeatable; requires network.isolation: true)
  • apiProxy.enabled → --enable-api-proxy ([DEPRECATED] API proxy is always enabled; this flag is ignored)
  • apiProxy.caCert → --api-proxy-ca-cert <path> (mounts an additional CA certificate into the api-proxy sidecar and sets NODE_EXTRA_CA_CERTS for upstream TLS verification)
  • apiProxy.enableTokenSteering → --enable-token-steering (maps to AWF_ENABLE_TOKEN_STEERING; omit or set to false to opt out)
  • apiProxy.anthropicAutoCache → --anthropic-auto-cache
  • apiProxy.anthropicCacheTailTtl → --anthropic-cache-tail-ttl <5m|1h>
  • apiProxy.hostedWeb.claude → (config-only; maps to AWF_CLAUDE_HOSTED_WEB_POLICY — AWF-owned domain policy for Anthropic-hosted web_search_*/web_fetch_* server tools; see §9.8 Claude Hosted Web Search and Fetch)
  • apiProxy.hostedWeb.codex → (config-only; maps to AWF_CODEX_HOSTED_WEB_POLICY — AWF-owned domain policy for OpenAI Responses web_search and Codex /v1/alpha/search; see §9.9 Codex/OpenAI Hosted Web Policy)
  • apiProxy.maxEffectiveTokens → (config-only; no CLI equivalent)
  • apiProxy.maxAiCredits → (config-only; maps to AWF_MAX_AI_CREDITS)
  • apiProxy.defaultAiCreditsPricing → (config-only; maps to AWF_DEFAULT_AI_CREDITS_PRICING)
  • apiProxy.providers → (config-only; maps to AWF_API_PROXY_PROVIDERS)
  • apiProxy.modelMultipliers → --max-model-multiplier <model:multiplier,...>
  • apiProxy.defaultModelMultiplier → (config-only; maps to AWF_EFFECTIVE_TOKEN_DEFAULT_MODEL_MULTIPLIER)
  • apiProxy.maxTurns → (config-only; no CLI equivalent)
  • apiProxy.maxRuns → (deprecated alias for maxTurns; maps to AWF_MAX_RUNS)
  • apiProxy.maxModelMultiplierCap → --max-model-multiplier-cap <number>
  • apiProxy.maxCacheMisses → --max-cache-misses <number>
  • apiProxy.maxPermissionDenied → --max-permission-denied <number>
  • apiProxy.requestedModel → (config-only; maps to AWF_REQUESTED_MODEL for pre-startup validation)
  • apiProxy.modelFallback → (config-only; model fallback strategy)
  • apiProxy.fallbackModels → (config-only; maps to AWF_FALLBACK_MODELS — ordered model IDs retried on 5xx, timeout, or model-not-supported failures)
  • experimental.modelRouting → (config-only; experimental opt-in required for apiProxy.routing; defaults to off)
  • apiProxy.routing → (config-only; requires experimental.modelRouting: true; task-level routing objective and task conversation input)
  • apiProxy.routing.candidateModels → (optional glob patterns that limit router/classifier choices; intersected with the model policy; defaults to apiProxy.allowedModels)
  • apiProxy.modelRouter.providerType → (config-only; maps to COPILOT_PROVIDER_TYPE)
  • apiProxy.modelRouter.baseUrl → (config-only; maps to COPILOT_PROVIDER_BASE_URL)
  • apiProxy.allowedModels → (config-only; maps to AWF_ALLOWED_MODELS — JSON array of glob patterns; only matching models are permitted)
  • apiProxy.disallowedModels → (config-only; maps to AWF_DISALLOWED_MODELS — JSON array of glob patterns; matching models are rejected with HTTP 403)
  • apiProxy.models → (config-only; model alias rewriting)
  • apiProxy.logging.debugTokens → (config-only; maps to AWF_DEBUG_TOKENS)
  • apiProxy.logging.tokenLogDir → (config-only; maps to AWF_TOKEN_LOG_DIR)
  • apiProxy.diagnostics.captureBlockedRequests → (config-only; maps to AWF_CAPTURE_BLOCKED_LLM_REQUESTS)
  • apiProxy.diagnostics.maxCapturedBytes → (config-only; maps to AWF_MAX_BLOCKED_CAPTURE_BYTES)
  • apiProxy.auth.type → (config-only; maps to AWF_AUTH_TYPE)
  • apiProxy.auth.provider → (config-only; maps to AWF_AUTH_PROVIDER)
  • apiProxy.auth.oidcAudience → (config-only; maps to AWF_AUTH_OIDC_AUDIENCE)
  • apiProxy.auth.azureTenantId → (config-only; maps to AWF_AUTH_AZURE_TENANT_ID)
  • apiProxy.auth.azureClientId → (config-only; maps to AWF_AUTH_AZURE_CLIENT_ID)
  • apiProxy.auth.azureScope → (config-only; maps to AWF_AUTH_AZURE_SCOPE)
  • apiProxy.auth.azureCloud → (config-only; maps to AWF_AUTH_AZURE_CLOUD)
  • apiProxy.auth.awsRoleArn → (config-only; maps to AWF_AUTH_AWS_ROLE_ARN)
  • apiProxy.auth.awsRegion → (config-only; maps to AWF_AUTH_AWS_REGION)
  • apiProxy.auth.awsRoleSessionName → (config-only; maps to AWF_AUTH_AWS_ROLE_SESSION_NAME)
  • apiProxy.auth.gcpWorkloadIdentityProvider → (config-only; maps to AWF_AUTH_GCP_WORKLOAD_IDENTITY_PROVIDER)
  • apiProxy.auth.gcpServiceAccount → (config-only; maps to AWF_AUTH_GCP_SERVICE_ACCOUNT)
  • apiProxy.auth.gcpScope → (config-only; maps to AWF_AUTH_GCP_SCOPE)
  • apiProxy.auth.anthropicFederationRuleId → (config-only; maps to AWF_AUTH_ANTHROPIC_FEDERATION_RULE_ID)
  • apiProxy.auth.anthropicOrganizationId → (config-only; maps to AWF_AUTH_ANTHROPIC_ORGANIZATION_ID)
  • apiProxy.auth.anthropicServiceAccountId → (config-only; maps to AWF_AUTH_ANTHROPIC_SERVICE_ACCOUNT_ID)
  • apiProxy.auth.anthropicWorkspaceId → (config-only; maps to AWF_AUTH_ANTHROPIC_WORKSPACE_ID)
  • apiProxy.auth.anthropicTokenUrl → (config-only; maps to AWF_AUTH_ANTHROPIC_TOKEN_URL)
  • apiProxy.targets.<provider>.host → --<provider>-api-target (except antigravity.host, which maps to the Gemini flag below). Accepts a bare hostname or a full URL; a bare hostname or explicit https:// URL both dial the target over HTTPS on port 443 (the default), while an explicit http:// URL dials it in cleartext on port 80 — matching the runner-side allowlist, which already adds an http://-scoped allowlist entry for such targets. Custom ports are not supported: http://host always dials port 80 and https://host always dials port 443.)
  • apiProxy.targets.antigravity.host → --gemini-api-target
  • apiProxy.targets.copilot.extraHeaders → (config-only; non-sensitive supplemental BYOK headers, maps to AWF_BYOK_EXTRA_HEADERS)
  • apiProxy.targets.copilot.extraBodyFields → (config-only; non-sensitive supplemental BYOK body fields, maps to AWF_BYOK_EXTRA_BODY_FIELDS)
  • apiProxy.targets.copilot.sessionId → (config-only; opt-in x-session-id header / session_id body field for Copilot BYOK requests, maps to AWF_PROVIDER_SESSION_ID. Never auto-derived from GITHUB_RUN_ID.)
  • apiProxy.targets.openai.basePath → --openai-api-base-path
  • apiProxy.targets.openai.authHeader → --openai-api-auth-header
  • apiProxy.targets.openai.baseUrlEnv → --openai-base-url-env (names a runner environment variable holding a secret OpenAI-compatible base URL; see §9.7 Secret-Backed OpenAI Target)
  • apiProxy.targets.anthropic.basePath → --anthropic-api-base-path
  • apiProxy.targets.anthropic.authHeader → --anthropic-api-auth-header
  • apiProxy.targets.gemini.basePath → --gemini-api-base-path
  • apiProxy.targets.antigravity.basePath → --gemini-api-base-path
  • When both apiProxy.targets.antigravity and apiProxy.targets.gemini are set, antigravity takes precedence per field.
  • apiProxy.targets.vertex.host → --vertex-api-target
  • apiProxy.targets.vertex.basePath → --vertex-api-base-path
  • security.legacySecurity → --legacy-security
  • security.securityMode → --security-mode <strict|compat> ([DEPRECATED] Use security.legacySecurity instead)
  • security.sslBump → --ssl-bump
  • security.enableDlp → --enable-dlp
  • security.enableHostAccess → --enable-host-access
  • security.allowHostPorts → --allow-host-ports
  • security.allowHostServicePorts → --allow-host-service-ports
  • security.difcProxy.host → --difc-proxy-host
  • security.difcProxy.caCert → --difc-proxy-ca-cert
  • container.memoryLimit → --memory-limit
  • container.pidsLimit → --pids-limit
  • container.agentTimeout → --agent-timeout
  • container.enableDind → --enable-dind
  • container.workDir → --work-dir
  • container.containerWorkDir → --container-workdir
  • container.images → (config-only; a closed compiler-authorized manifest of literal, registry-qualified tag@sha256:<digest> OCI references. Supported keys are squid, agent, apiProxy, router, cliProxy, buildTools, dohProxy, enclaveScript, enclaveAgent, enclaveMcpServer, and dindStaging. Every image AWF runs — including consumers outside Docker Compose such as DinD staging, awf predownload --config, and rootless artifact repair — resolves through this manifest, and the effective per-role references are recorded in image-manifest.json. AWF rejects missing enabled roles and never falls back to the official registry. It cannot be combined with controls that would select a different image: container.imageRegistry, container.imageTag, container.agentImage, container.buildLocal, security.sslBump (requires a locally built Squid image), runner.sysrootImage, dind.stagingImage, or per-enclave image overrides. Registry credentials are intentionally not configured by AWF; use a pre-authenticated Docker daemon.)
  • container.imageRegistry → --image-registry
  • container.imageTag → --image-tag
  • container.skipPull → --skip-pull
  • container.buildLocal → --build-local
  • container.agentImage → --agent-image
  • container.tty → --tty
  • container.dockerHost → --docker-host
  • container.dockerHostPathPrefix → --docker-host-path-prefix
  • container.runnerToolCachePath → (config-only; checked first for optional read-only runner tool cache mount, before RUNNER_TOOL_CACHE and /home/runner/work/_tool auto-detection)
  • container.mounts[] → -v, --mount (repeatable; each array entry maps to one Docker volume mount in /host_path:/container_path[:ro|rw] format (both paths must be absolute; host path must exist); in chroot mode, container paths are automatically prefixed with /host)
  • container.containerRuntime → --container-runtime (user-facing runtime name: "gvisor" for an OCI runtime in Compose, "sbx" for a Docker sbx microVM, "cloud-hypervisor" for the explicit Cloud Hypervisor v53.0 workload preview (GitHub-hosted Ubuntu x86_64 KVM runners only; see §4.2), or "nvx" for the explicit NVX/OpenVMM one-shot workload preview (Linux x86_64 KVM-only; see docs/nvx-security-design.md). gVisor translates to "runsc" and injects extra_hosts for its DNS workaround. For sbx, Cloud Hypervisor, and NVX, infrastructure stays in Compose while the primary agent runs in a microVM.)
  • filesystem.allowWrite[] → (config-only; no CLI equivalent; narrows existing writable host binds to the listed guest-visible absolute paths, see §4.1)
  • cloudHypervisor.previewEnabled → --cloud-hypervisor-preview (requires container.containerRuntime: "cloud-hypervisor" or at least one enclaves[].runtime: "cloud-hypervisor" entry, plus a GitHub-hosted Ubuntu x86_64 KVM runner to execute a workload)
  • cloudHypervisor.mountPolicy → --cloud-hypervisor-mount-policy (workspace-only by default; use workspace-and-tool-cache only when the workload needs the runner tool cache)
  • cloudHypervisor.cloudHypervisorBinary → --cloud-hypervisor-binary
  • cloudHypervisor.kernelPath → --cloud-hypervisor-kernel
  • cloudHypervisor.rootfsPath → --cloud-hypervisor-rootfs
  • cloudHypervisor.supervisorPath → --cloud-hypervisor-supervisor
  • cloudHypervisor.artifactManifestPath → --cloud-hypervisor-artifact-manifest
  • cloudHypervisor.artifactManifestBundlePath → --cloud-hypervisor-artifact-manifest-bundle
  • cloudHypervisor.artifactReleaseTag → --cloud-hypervisor-artifact-release-tag
  • cloudHypervisor.developmentAllowUnattestedArtifacts → --cloud-hypervisor-development-allow-unattested-artifacts (development only; also requires the matching environment variable and complete legacy hashes)
  • cloudHypervisor.vcpuCount → --cloud-hypervisor-vcpus
  • cloudHypervisor.memoryMib → --cloud-hypervisor-memory-mib
  • cloudHypervisor.apiTimeoutMs → --cloud-hypervisor-api-timeout-ms
  • cloudHypervisor.sha256.cloudHypervisor → --cloud-hypervisor-binary-sha256 (development bypass only)
  • cloudHypervisor.sha256.virtiofsd → --cloud-hypervisor-virtiofsd-sha256
  • cloudHypervisor.sha256.kernel → --cloud-hypervisor-kernel-sha256
  • cloudHypervisor.sha256.rootfs → --cloud-hypervisor-rootfs-sha256
  • cloudHypervisor.sha256.supervisor → --cloud-hypervisor-supervisor-sha256
  • nvx.previewEnabled → --nvx-preview (requires container.containerRuntime: "nvx"; execution never falls back to Docker, Cloud Hypervisor, or another runtime on failure)
  • nvx.layerPath → --nvx-layer (guest distro layer; required)
  • nvx.artifactManifestPath → --nvx-artifact-manifest (required)
  • nvx.artifactManifestBundlePath → --nvx-artifact-manifest-bundle (required)
  • nvx.signerWorkflow → --nvx-signer-workflow
  • nvx.openvmmPath → --nvx-openvmm (required)
  • nvx.kernelPath → --nvx-kernel (required)
  • nvx.initramfsPath → --nvx-initramfs (required)
  • nvx.memoryMib → --nvx-memory-mib (default 512)
  • nvx.memoryMaxBytes → --nvx-memory-max-bytes (default 512 MiB)
  • nvx.pidsMax → --nvx-pids-max (default 128)
  • nvx.scratchBytes → --nvx-scratch-bytes
  • nvx.mountPolicy → --nvx-mount-policy ("workspace-only" (default) exports $GITHUB_WORKSPACE into the guest at /workspace read-write; "workspace-and-tool-cache" additionally exports the runner tool cache read-only. filesystem.allowWrite narrows the workspace export, and container.containerWorkDir must resolve inside /workspace.)
  • chroot.binariesSourcePath → (config-only; mounts a runner-side binaries directory at /tmp/awf-runner-bin inside chroot mode and prepends it to PATH)
  • chroot.identity.home → (config-only; forwarded as AWF_CHROOT_IDENTITY_HOME and applied after chroot pivot)
  • chroot.identity.user → (config-only; forwarded as AWF_CHROOT_IDENTITY_USER and applied to USER/LOGNAME after chroot pivot)
  • chroot.identity.uid → (config-only; forwarded as AWF_CHROOT_IDENTITY_UID for chroot user mapping)
  • chroot.identity.gid → (config-only; forwarded as AWF_CHROOT_IDENTITY_GID for chroot user mapping)
  • dind.preStageDirs → (config-only; enables daemon-side pre-staging of the DinD work directory tree before compose startup)
  • dind.workDir → (config-only; daemon-visible staging root, default /tmp/gh-aw)
  • dind.stagingImage → (config-only; image used for short-lived DinD staging containers)
  • dind.stageEngineBinary.path → (config-only; runner-side engine binary source path for DinD staging)
  • dind.stageEngineBinary.targetPath → (config-only; daemon-side destination path for staged engine binary)
  • environment.envFile → --env-file
  • environment.envAll → --env-all
  • environment.excludeEnv[] → --exclude-env (repeatable)
  • logging.logLevel → --log-level
  • logging.diagnosticLogs → --diagnostic-logs
  • logging.auditDir → --audit-dir
  • logging.proxyLogsDir → --proxy-logs-dir
  • logging.sessionStateDir → --session-state-dir
  • rateLimiting.enabled: false → --no-rate-limit
  • rateLimiting.requestsPerMinute → --rate-limit-rpm
  • rateLimiting.requestsPerHour → --rate-limit-rph
  • rateLimiting.bytesPerMinute → --rate-limit-bytes-pm
  • rateLimiting.maxGithubApiPointsRest → --max-github-api-points-rest (requires security.difcProxy.host)
  • rateLimiting.maxGithubApiPointsGraphql → --max-github-api-points-graphql (requires security.difcProxy.host)
  • rateLimiting.maxNumToolCalls → --max-num-tool-calls (requires enclaves, see §14.2a)

GitHub API point budgets apply to the entire AWF run and are enforced only by the protected gh CLI proxy enabled through security.difcProxy.host. REST read requests consume one point, REST mutations consume five points, and GraphQL requests consume one point, matching GitHub's secondary-rate-limit point accounting. A command that would exceed its REST or GraphQL budget is rejected before it is executed; the CLI proxy writes a github_api_points_limited structured audit record with the API kind, used points, configured limit, and remaining budget.

  • (no config equivalent) → --reflect (CLI-only; starts AWF, queries the API proxy /reflect endpoint, and prints its JSON response instead of running a command; mutually exclusive with a command argument)
  • platform.type → (config-only; maps to AWF_PLATFORM_TYPE)
  • runner.topology → (config-only; sets runner deployment model — standard or arc-dind; when arc-dind, enables sysroot staging and emits RUNNER_TOOL_CACHE warnings)
  • runner.sysrootImage → (config-only; sysroot init-container image for arc-dind topology; defaults to <container.imageRegistry>/build-tools:<container.imageTag>, where container.imageRegistry defaults to ghcr.io/github/gh-aw-firewall)
  • enclaves[] → (config-only; no CLI equivalent, see §14)
  • enclaves[].repos[] → (config-only; no CLI equivalent, see §14)
  • enclaves[].timeout → (config-only; no CLI equivalent, see §14)
  • enclaves[].runtime → (config-only; no CLI equivalent, see §14)
  • enclaves[].image → (config-only; no CLI equivalent, see §14)
  • enclaves[].memoryLimit → (config-only; no CLI equivalent, see §14)
  • enclaves[].cpuLimit → (config-only; no CLI equivalent, see §14)
  • enclaves[].pidsLimit → (config-only; no CLI equivalent, see §14)
  • enclaves[].tmpfsLimit → (config-only; no CLI equivalent, see §14)
  • enclaves[].maxOutputBytes → (config-only; no CLI equivalent, see §14)
  • enclaves[].maxInvocations → (config-only; no CLI equivalent, see §14)
  • enclaves[].script → (config-only; no CLI equivalent, see §14)
  • enclaves[].script.maxScriptBytes → (config-only; no CLI equivalent, see §14)
  • enclaves[].agent → (config-only; no CLI equivalent, see §14)
  • enclaves[].agent.engine → (config-only; no CLI equivalent, see §14)
  • enclaves[].agent.profile → (config-only; no CLI equivalent, see §14)
  • enclaves[].agent.model → (config-only; no CLI equivalent, see §14)
  • enclaves[].agent.maxTaskBytes → (config-only; no CLI equivalent, see §14)
  • enclaves[].agent.maxModelRequests → (config-only; no CLI equivalent, see §14)
  • enclaves[].agent.maxModelTokens → (config-only; no CLI equivalent, see §14)
  • enclaves[].agent.github.cli → (config-only; deprecated closed legacy profile, see §14.3)
  • enclaves[].agent.tools.github → (config-only; closed GitHub MCP tool contract, see §14.3)
  • enclaves[].dynamic → (config-only; no CLI equivalent; requires compiler handoff; see §14.1a)

When container.dockerHostPathPrefix points at a daemon-visible shared /tmp path, the implementation stages the invoking CLI binary together with /etc/passwd, /etc/group, and the generated chroot /etc/hosts under that shared path so chroot mode can bootstrap on split-filesystem ARC/DinD hosts.

When DinD is detected, AWF preserves the detected DOCKER_HOST value for the agent environment (including MCP servers) so DinD-aware tooling can reach the correct daemon without manual workflow env overrides.

security.allowHostPorts (--allow-host-ports) is accepted together with security.enableHostAccess (--enable-host-access) in strict security mode (the default, without --legacy-security), but it does not provide a direct route to raw-protocol GitHub Actions services: containers. Strict topology intentionally omits the agent's host.docker.internal mapping and host-access iptables bypass; use legacy security or a separately verified tunnel for direct service clients. security.allowHostServicePorts (--allow-host-service-ports), which relies on host iptables, remains suppressed in strict mode.

The following CLI flag has no config-file equivalent by design:

  • -e, --env <KEY=VALUE> — inject a single environment variable into the agent container (repeatable; CLI-only)

6. Standard Input Mode

A conforming implementation MUST accept --config - to read configuration from standard input, enabling programmatic and pipeline scenarios.

7. Error Reporting

On parse or validation failure, a conforming implementation MUST:

  1. exit with a non-zero status code;
  2. emit a diagnostic message identifying the location and nature of the error; and
  3. refrain from partial execution of the agent command.

8. Environment Merge Semantics

This section is normative.

The agent container's environment is constructed by merging variables from multiple sources. This section defines the merge order and exclusion rules.

Note: For usage guidance, examples, and troubleshooting, see docs/environment.md.

8.1 Merge Precedence

Variables from the following sources are merged in order of increasing precedence. A value set at a higher level MUST override the same-named value from any lower level.

Level Source Description
1 (lowest) AWF-reserved Proxy routing, DNS, container paths
2 --env-all Inherited host environment (when enabled)
3 --env-file Variables read from a file
4 (highest) -e / --env Explicit CLI key-value pairs

8.2 AWF-Reserved Variables

A conforming implementation MUST set the following variables in the agent container regardless of user configuration. Values from --env-all and --env-file MUST NOT override these variables.

Variable Value Purpose
HTTP_PROXY http://<squid-ip>:3128 Squid forward proxy for HTTP
HTTPS_PROXY http://<squid-ip>:3128 Squid forward proxy for HTTPS
https_proxy http://<squid-ip>:3128 Lowercase alias (Yarn 4, undici, Corepack)
NO_PROXY localhost,127.0.0.1,::1,... Loopback and container IPs bypassing Squid
SQUID_PROXY_HOST squid-proxy Proxy hostname (for tools requiring host separately)
SQUID_PROXY_PORT 3128 Proxy port
PATH (container default) MUST use the container's PATH, not the host's
HOME (host user's home) Derived via sudo-aware detection

Note: Lowercase http_proxy is intentionally NOT set. Certain curl builds on Ubuntu 22.04 ignore uppercase HTTP_PROXY for HTTP URLs (httpoxy mitigation), causing HTTP traffic to fall through to iptables DNAT interception — the intended defense-in-depth behavior.

8.3 Excluded Variables

The following variables MUST be excluded from --env-all and --env-file passthrough. A conforming implementation MUST NOT inherit them from the host:

Category Variables
System PATH, PWD, OLDPWD, SHLVL, _, SUDO_COMMAND, SUDO_USER, SUDO_UID, SUDO_GID
Proxy HTTP_PROXY, HTTPS_PROXY, http_proxy, https_proxy, NO_PROXY, no_proxy, ALL_PROXY, all_proxy, FTP_PROXY, ftp_proxy
Actions runtime credentials ACTIONS_RUNTIME_TOKEN, ACTIONS_RESULTS_URL, ACTIONS_ID_TOKEN_REQUEST_URL, ACTIONS_ID_TOKEN_REQUEST_TOKEN
AWF internal controls AWF_PREFLIGHT_BINARY, AWF_ENSURE_USR_LOCAL_BIN, AWF_GEMINI_ENABLED

Note: Host proxy variables are read for upstream proxy auto-detection (see --upstream-proxy) but MUST NOT propagate into the agent container. AWF sets its own proxy variables pointing to Squid.

8.4 Selectively Forwarded Variables

When --env-all is NOT active, a conforming implementation SHOULD forward the following host variables into the agent container:

Category Variables
GitHub authentication GITHUB_TOKEN, GH_TOKEN, GITHUB_PERSONAL_ACCESS_TOKEN
GitHub enterprise GITHUB_SERVER_URL, GITHUB_API_URL
Docker client DOCKER_HOST, DOCKER_TLS, DOCKER_TLS_VERIFY, DOCKER_CERT_PATH, DOCKER_CONFIG, DOCKER_CONTEXT, DOCKER_API_VERSION, DOCKER_DEFAULT_PLATFORM
User environment USER, XDG_CONFIG_HOME

When --env-all IS active, all host variables not in the excluded set (§8.3) SHALL be forwarded, subject to credential isolation rules (§9).

Actions OIDC request variables MUST be forwarded directly to the api-proxy sidecar when apiProxy.auth.type is github-oidc and MUST NOT be forwarded to the agent through any environment input path.

8.5 Explicit Overrides

Variables passed via -e / --env MUST override values from --env-all and --env-file.

Reserved proxy routing variables MAY be overridden only via -e / --env. Other AWF-reserved variables and source credentials protected by credential isolation (§9) MUST NOT be overridden.

Note: There is no config-file equivalent for -e / --env. Individual environment variable injection is a runtime concern, not a static configuration concern.

9. Credential Isolation Semantics

This section is normative.

AWF implements defense-in-depth credential isolation for LLM API keys. Behavior is governed by the value of apiProxy.enabled.

Note: For architectural diagrams and protocol-level details, see docs/authentication-architecture.md.

9.1 Source Credentials

A conforming implementation MUST recognize the following environment variables as source credentials — real API keys read from the host:

Variable Provider
OPENAI_API_KEY OpenAI
ANTHROPIC_API_KEY Anthropic (Claude)
COPILOT_GITHUB_TOKEN GitHub Copilot — enables sidecar routing to api.githubcopilot.com (CAPI BYOK / offline mode)
COPILOT_PROVIDER_API_KEY GitHub Copilot BYOK provider key (e.g., Azure OpenAI / OpenRouter API key); independently enables sidecar routing — typically combined with COPILOT_PROVIDER_BASE_URL to point at an arbitrary upstream
GEMINI_API_KEY Google Gemini
GOOGLE_API_KEY Google Vertex AI

The following secondary aliases SHOULD also be recognized: OPENAI_KEY, CODEX_API_KEY, CLAUDE_API_KEY.

9.2 API Proxy Enabled (apiProxy.enabled = true)

When the API proxy sidecar is enabled, the following rules apply:

  1. Source credentials (§9.1) MUST NOT be exposed in the agent container's environment. They SHALL be passed exclusively to the API proxy sidecar.

  2. The --env-all flag MUST NOT reintroduce excluded credentials into the agent environment.

  3. A conforming implementation MAY inject placeholder values into the agent container for tool compatibility (e.g., OPENAI_API_KEY=sk-placeholder-for-api-proxy). Placeholder values are not secrets and MUST NOT be treated as credentials.

  4. A conforming implementation MUST inject proxy-routing variables so that agent tools reach the sidecar rather than upstream APIs:

    Agent variable Value Purpose
    OPENAI_BASE_URL http://172.30.0.30:10000 Routes OpenAI calls to sidecar
    ANTHROPIC_BASE_URL http://172.30.0.30:10001 Routes Anthropic calls to sidecar
    COPILOT_API_URL http://172.30.0.30:10002 Routes Copilot calls to sidecar
    GOOGLE_GEMINI_BASE_URL http://172.30.0.30:10003 Routes Gemini calls to sidecar
    GEMINI_API_BASE_URL http://172.30.0.30:10003 Alias for compatibility
    GOOGLE_VERTEX_BASE_URL http://172.30.0.30:10004 Routes Vertex AI calls to sidecar
  5. The API proxy sidecar SHALL inject the real credentials into upstream requests. Sidecar port assignments: 10000 (OpenAI), 10001 (Anthropic), 10002 (Copilot), 10003 (Gemini), 10004 (Vertex AI).

  6. A conforming implementation MUST forward the following OpenTelemetry variables from the host into the api-proxy sidecar container so that the sidecar can participate in the distributed trace established by the workflow:

    Variable Description
    GH_AW_OTLP_ENDPOINTS JSON array of {url, headers} objects for fan-out export to multiple OTLP collectors. Takes priority over OTEL_EXPORTER_OTLP_ENDPOINT.
    OTEL_EXPORTER_OTLP_ENDPOINT OTLP/HTTP collector URL. Single-endpoint fallback when GH_AW_OTLP_ENDPOINTS is absent.
    OTEL_EXPORTER_OTLP_HEADERS Comma-separated key=value auth headers for the OTLP endpoint. Only used with OTEL_EXPORTER_OTLP_ENDPOINT.
    OTEL_SERVICE_NAME Service name tag. Defaults to awf-api-proxy when not set.
    GITHUB_AW_OTEL_TRACE_ID W3C trace-id of the parent workflow trace.
    GITHUB_AW_OTEL_PARENT_SPAN_ID W3C span-id of the parent workflow span.

    These variables are NOT forwarded to the agent container via this mechanism; the agent receives OTEL variables through the standard OTEL_* prefix forwarding described in §8.4.

    The sidecar selects its exporter using the following priority order:

    1. GH_AW_OTLP_ENDPOINTS (JSON array) — spans are exported concurrently to all listed endpoints (fan-out mode); partial failures on individual endpoints do not block others.
    2. OTEL_EXPORTER_OTLP_ENDPOINT (single URL) — legacy single-endpoint mode.
    3. Neither set — the sidecar writes span NDJSON to /var/log/api-proxy/otel.jsonl as a local fallback.

    When GITHUB_AW_OTEL_TRACE_ID / GITHUB_AW_OTEL_PARENT_SPAN_ID are present and valid hex, each sidecar span is created as a child of the specified parent span, enabling end-to-end distributed tracing from the GitHub Actions workflow through the api-proxy to the LLM provider.

9.3 apiProxy.enabled Deprecated

apiProxy.enabled is deprecated and ignored. The API proxy sidecar is always started; there is no disabled mode. Setting apiProxy.enabled: false in a config file is silently ignored for backward compatibility. The CLI flag --enable-api-proxy is similarly ignored; --no-enable-api-proxy is rejected at runtime with an error.

Credential isolation described in §9.2 therefore always applies: source credentials are never forwarded directly to the agent container.

9.4 Credential Exclusion

This constraint is normative for tools generating AWF configurations.

Because the API proxy sidecar is always active, source credentials (§9.1) are always excluded from the agent environment and held in the sidecar. A conforming implementation MUST NOT rely on environment.excludeEnv to suppress API keys — the sidecar handles exclusion automatically.

Tools that compile AWF configurations (e.g., gh-aw) MUST ensure that when an LLM agent requires an API key (OpenAI, Anthropic, Gemini, etc.), the real key is held by the sidecar and a placeholder is injected for tool compatibility.

9.4 One-Shot Token Protection

Real credentials forwarded to the agent — GitHub tokens (GITHUB_TOKEN, GH_TOKEN) and any non-LLM credentials — MUST be protected by the one-shot-token mechanism. Protected tokens are cached on first access and removed from /proc/self/environ to prevent environment variable inspection.

The default protected token list is:

COPILOT_GITHUB_TOKEN, GITHUB_TOKEN, GH_TOKEN, GITHUB_API_TOKEN,
GITHUB_PAT, GH_ACCESS_TOKEN, OPENAI_API_KEY, OPENAI_KEY,
ANTHROPIC_API_KEY, ANTHROPIC_AUTH_TOKEN, CLAUDE_API_KEY,
CODEX_API_KEY, COPILOT_PROVIDER_API_KEY

Placeholder compatibility values (§9.2 item 3) are not secrets. However, provider credential variable names such as ANTHROPIC_AUTH_TOKEN MAY remain on the protection list as defense-in-depth so unexpectedly forwarded real credentials are still scrubbed on first read.

9.5 OIDC Authentication

When apiProxy.auth.type is set to github-oidc, the API proxy sidecar exchanges a GitHub Actions OIDC token for a provider-specific access token. The apiProxy.auth.provider field (default: azure) selects the token exchange protocol. A conforming implementation MUST:

  1. Forward the common OIDC configuration to the sidecar via the following environment variables:

    Config path Environment variable Required Default
    apiProxy.auth.type AWF_AUTH_TYPE ✅ —
    apiProxy.auth.provider AWF_AUTH_PROVIDER No azure
    apiProxy.auth.oidcAudience AWF_AUTH_OIDC_AUDIENCE No (provider-specific)
  2. Forward the GitHub Actions OIDC runtime tokens (ACTIONS_ID_TOKEN_REQUEST_URL, ACTIONS_ID_TOKEN_REQUEST_TOKEN) to the sidecar when AWF_AUTH_TYPE=github-oidc. These are injected automatically by the Actions runner when the workflow declares permissions: id-token: write.

    If OIDC is requested for a provider but these runtime variables are not present in the sidecar environment, the provider adapter MUST fail closed and return an explicit configuration error (rather than falling back to static-key mode).

  3. NOT expose the exchanged provider token in the agent container environment. The sidecar SHALL inject it into upstream request headers.

9.5.1 Azure Provider (provider: azure)

Exchanges the GitHub OIDC JWT for an Azure AD / Microsoft Entra access token via workload identity federation. The sidecar injects the resulting token as a Bearer Authorization header on upstream requests.

Config path Environment variable Required Default
apiProxy.auth.azureTenantId AWF_AUTH_AZURE_TENANT_ID ✅ —
apiProxy.auth.azureClientId AWF_AUTH_AZURE_CLIENT_ID ✅ —
apiProxy.auth.azureScope AWF_AUTH_AZURE_SCOPE No https://cognitiveservices.azure.com/.default
apiProxy.auth.azureCloud AWF_AUTH_AZURE_CLOUD No public

Default OIDC audience: api://AzureADTokenExchange

Note: azureTenantId and azureClientId are required for Azure AD federated credential exchange but MAY be omitted when using managed identity. See docs/api-proxy-sidecar.md for protocol-level details.

9.5.2 AWS Provider (provider: aws)

Exchanges the GitHub OIDC JWT for temporary AWS credentials via sts.amazonaws.com AssumeRoleWithWebIdentity. The sidecar uses these credentials to sign upstream requests to AWS Bedrock using SigV4.

Config path Environment variable Required Default
apiProxy.auth.awsRoleArn AWF_AUTH_AWS_ROLE_ARN ✅ —
apiProxy.auth.awsRegion AWF_AUTH_AWS_REGION ✅ —
apiProxy.auth.awsRoleSessionName AWF_AUTH_AWS_ROLE_SESSION_NAME No awf-oidc-session

Default OIDC audience: sts.amazonaws.com

Note: AWS Bedrock uses IAM/SigV4 request signing rather than Bearer tokens. This means the sidecar MUST sign the complete request (method, path, headers, body hash) with the temporary credentials — it is not sufficient to inject a single Authorization header.

9.5.3 GCP Provider (provider: gcp)

Exchanges the GitHub OIDC JWT for a GCP access token via the Security Token Service (sts.googleapis.com), optionally followed by service account impersonation via iamcredentials.googleapis.com. The sidecar injects the resulting token as a Bearer Authorization header.

Config path Environment variable Required Default
apiProxy.auth.gcpWorkloadIdentityProvider AWF_AUTH_GCP_WORKLOAD_IDENTITY_PROVIDER ✅ —
apiProxy.auth.gcpServiceAccount AWF_AUTH_GCP_SERVICE_ACCOUNT No —
apiProxy.auth.gcpScope AWF_AUTH_GCP_SCOPE No https://www.googleapis.com/auth/cloud-platform

Default OIDC audience: the gcpWorkloadIdentityProvider value

When gcpServiceAccount is provided, the sidecar performs a two-step exchange:

  1. Exchange GitHub OIDC JWT for a federated access token via GCP STS
  2. Impersonate the service account to obtain a short-lived OAuth2 token

When gcpServiceAccount is omitted, only step 1 is performed and the federated token is used directly. This requires that the federated principal has direct access grants on the target resource.

9.5.4 Anthropic Provider (provider: anthropic)

Exchanges the GitHub OIDC JWT for an Anthropic Workload Identity Federation token via Anthropic OAuth token endpoint (default: https://api.anthropic.com/v1/oauth/token). The sidecar injects the resulting token as an Authorization header on upstream requests.

Config path Environment variable Required Default
apiProxy.auth.anthropicFederationRuleId AWF_AUTH_ANTHROPIC_FEDERATION_RULE_ID ✅ —
apiProxy.auth.anthropicOrganizationId AWF_AUTH_ANTHROPIC_ORGANIZATION_ID ✅ —
apiProxy.auth.anthropicServiceAccountId AWF_AUTH_ANTHROPIC_SERVICE_ACCOUNT_ID ✅ —
apiProxy.auth.anthropicWorkspaceId AWF_AUTH_ANTHROPIC_WORKSPACE_ID Conditional¹ —
apiProxy.auth.anthropicTokenUrl AWF_AUTH_ANTHROPIC_TOKEN_URL ❌ https://api.anthropic.com/v1/oauth/token

¹ AWF_AUTH_ANTHROPIC_WORKSPACE_ID is required when the federation rule covers multiple workspaces. When the rule is scoped to a single workspace, it may be omitted.

anthropicTokenUrl is non-sensitive and SHOULD be supplied via AWF config (including stdin config via --config -); env var support exists for compatibility.

Default OIDC audience: https://api.anthropic.com

9.6 DIFC Proxy Credential Isolation

When security.difcProxy.host is set, GITHUB_TOKEN and GH_TOKEN MUST be excluded from the agent environment. These tokens SHALL be held exclusively by the external DIFC proxy.

Under --network-isolation, the credential-bearing cli-proxy sidecar remains on awf-net only. When security.difcProxy.host is classified as external (host.docker.internal, an IP address outside awf-net's subnet, or a dotted DNS name), AWF creates a separate, credential-free cli-proxy-egress relay that is the only CLI-proxy component dual-homed onto the external bridge; it forwards solely to the configured DIFC host and port and never receives GITHUB_TOKEN or GH_TOKEN. When the DIFC proxy is an attached sibling container already reachable on awf-net, no relay is created.

9.7 Secret-Backed OpenAI Target

apiProxy.targets.openai.baseUrlEnv names a runner environment variable (typically bound to ${{ secrets.* }} by the gh-aw compiler) whose value is the base URL of a private OpenAI-compatible endpoint. This allows engine: codex workflows to route through a sensitive endpoint without writing the URL into workflow source or generated lockfiles.

{
  "apiProxy": {
    "targets": {
      "openai": {
        "baseUrlEnv": "CODEX_LB_BASE_URL"
      }
    }
  }
}

A conforming implementation:

  1. MUST read the named variable only in runner-side configuration code, before any container starts.
  2. MUST require an absolute https:// URL and MUST reject embedded credentials (user:pass@), query strings, fragments, malformed hosts, unsupported schemes, and non-default ports (the sidecar connects on port 443).
  3. MUST derive the host, host:port, and optional base path from the URL.
  4. MUST add the derived destination to the effective Squid policy without persisting it in repository configuration, and MUST keep it out of allowedDomains (it is carried in the sensitive allowlist instead).
  5. MUST configure the OpenAI api-proxy adapter with the derived host (OPENAI_API_TARGET) and base path (OPENAI_API_BASE_PATH); the derived values take precedence over apiProxy.targets.openai.host/basePath.
  6. MUST exclude the named variable from the primary agent environment, including under --env-all.
  7. MUST redact the URL, host, and host:port forms from logs, diagnostics, and uploaded audit artifacts (squid.conf, docker-compose.redacted.yml).
  8. MUST fail before agent startup, with an error message that does not contain the value, when the variable is unset or invalid.

The same rules apply to agent and detection phases, which share this configuration path.

9.8 Claude Hosted Web Search and Fetch

This section is normative.

Anthropic's hosted server tools — web_search_YYYYMMDD and web_fetch_YYYYMMDD — are executed by Anthropic infrastructure, not inside the agent container. Squid therefore only ever observes the Anthropic API endpoint; the searched or fetched destination is invisible to the domain ACL and to Squid access logs, and the retrieved content returns inside an already-allowed API response. Without additional enforcement this is a semantic egress gap: an indirect prompt-injection and blind-exfiltration channel driven by model-generated queries and URLs.

Because the api-proxy sidecar is trusted and already inspects Anthropic Messages bodies, AWF closes the gap there. apiProxy.hostedWeb.claude is config-only: there is no CLI flag and no environment alias. AWF serializes the validated policy into the sidecar as AWF_CLAUDE_HOSTED_WEB_POLICY; that variable is an implementation detail, is generated only from validated configuration, and MUST NOT be accepted from an agent-controlled source.

apiProxy:
  hostedWeb:
    claude:
      enabled: true
      allowedDomains:
        - docs.github.com
        - nodejs.org
      maxUses: 5

Blocklist mode:

apiProxy:
  hostedWeb:
    claude:
      enabled: true
      blockedDomains:
        - untrusted.example
        - ads.example

Prohibit Claude hosted web tools entirely:

apiProxy:
  hostedWeb:
    claude:
      enabled: false

9.8.1 Configuration Contract

apiProxy.hostedWeb and apiProxy.hostedWeb.claude are closed objects (additionalProperties: false).

Property Type Required Semantics
enabled boolean yes false rejects every matching Claude hosted search/fetch tool. true enables enforcement using exactly one configured domain mode.
allowedDomains non-empty array of unique domains conditional Provider allowlist upper bound. Mutually exclusive with blockedDomains.
blockedDomains non-empty array of unique domains conditional Provider blocklist lower bound. Mutually exclusive with allowedDomains.
maxUses positive integer no Maximum uses AWF permits per matching hosted tool. Request values may only reduce it.

There is deliberately no implicit unrestricted enabled: true state: when enabled is true, exactly one of allowedDomains or blockedDomains MUST be present and non-empty. This prevents a success-shaped configuration that appears protected while forwarding unrestricted hosted web access. Domains MUST NOT be combined with enabled: false.

Domain values use one canonical syntax: lowercase DNS hostnames with at least two labels and no URL scheme, path, query, fragment, credentials or port, and no raw IP address, CIDR, wildcard, localhost-style single-label name, Docker service alias, or empty string. Duplicates are removed after normalization. A parent domain also covers its subdomains, matching Anthropic's semantics.

Absence of apiProxy.hostedWeb.claude preserves the pre-existing pass-through behaviour. Warning: omitting the object does not constrain Claude-hosted egress. A compiler such as gh-aw can opt into enforcement by always emitting the object when workflow frontmatter declares a network policy.

The hosted-web list is intentionally explicit rather than derived from network.allowDomains. The network list may contain provider API endpoints, internal services, or ecosystem expansions that are inappropriate to disclose to or authorize for provider-hosted retrieval. Sensitive, secret-derived allowlist entries (see §9.7) MUST NOT be disclosed to the provider.

The behaviour is identical for a JSON file, a YAML file, JSON on stdin, and YAML on stdin. Validation errors from --config - identify stdin as the source and reject the run before any container starts. A compiler therefore only needs to emit a stable config document:

{
  "network": {
    "allowDomains": ["api.github.com", "docs.github.com"]
  },
  "apiProxy": {
    "hostedWeb": {
      "claude": {
        "enabled": true,
        "allowedDomains": ["docs.github.com"]
      }
    }
  }
}

and pipe it to awf --config - -- <command>. Compilers do not need to understand Anthropic request JSON, tool-version names, or api-proxy environment variables; AWF owns the provider-specific translation and enforcement.

9.8.2 Configuration-Source Precedence

The effective configured policy is selected before request processing:

  1. an explicit CLI option, if one is added in a later revision;
  2. the AWF configuration document, including JSON or YAML supplied via --config -;
  3. AWF's internal default (no enforcement; current compatibility behaviour).

9.8.3 Request-Level Precedence

The configured policy is an immutable security boundary. Request-provided tool settings MAY narrow it but MUST NEVER broaden, remove, replace, or disable it.

Configured policy Request tool fields Effective result
enabled: false any Reject (claude_hosted_web_disabled)
allowlist none Inject configured allowed_domains
allowlist allowed_domains Intersection of configured and requested
allowlist allowed_domains with empty intersection Reject (claude_hosted_web_empty_intersection)
allowlist blocked_domains Reject (claude_hosted_web_policy_conflict)
blocklist none Inject configured blocked_domains
blocklist blocked_domains Union of configured and requested
blocklist allowed_domains Reject (claude_hosted_web_policy_conflict)
any enabled mode both allowed_domains and blocked_domains Reject (claude_hosted_web_policy_conflict)
maxUses configured max_uses omitted Inject configured value
maxUses configured max_uses present min(configured, requested)
any enabled mode malformed max_uses (zero, negative, non-integer) Reject (claude_hosted_web_max_uses_invalid)

Cross-mode request restrictions are rejected rather than silently dropped because Anthropic cannot represent allowed_domains and blocked_domains on the same tool definition; rejecting is safer and more diagnosable than pretending both policies were enforced. AWF-generated values replace the corresponding request fields only after the effective policy has been calculated: a hosted tool is never silently stripped, and an unprotected hosted tool is never silently forwarded.

9.8.4 API-Proxy Behaviour

A conforming implementation:

  1. MUST parse and validate the serialized policy once at sidecar startup and MUST fail startup — not the first request — on an invalid internal policy.
  2. MUST apply enforcement only to Anthropic Messages request bodies that contain a matching hosted tool. A configured policy MUST NOT mutate an ordinary Messages request, including requests to custom Anthropic-compatible endpoints that implement no hosted server tools.
  3. MUST recognize versioned tool types with a validated web_search_YYYYMMDD / web_fetch_YYYYMMDD matcher — including versions published after this revision — rather than a fixed list of exact names.
  4. MUST fail closed on a value that claims to be a hosted web tool but does not match the versioned form (claude_hosted_web_tool_unrecognized), so a future or malformed shape can never be forwarded without policy.
  5. MUST apply enforcement after generic body parsing and after the other Anthropic transforms (model rewriting, prompt-cache optimizations, tool drop, custom transform file) so that no other transform can re-expand or remove the enforced policy.
  6. MUST enforce every matching tool in a request, not only the first.
  7. MUST return structured API-proxy errors with the stable error.code values listed above (HTTP 403 for policy rejections, HTTP 400 for malformed tool definitions, domains, or max_uses), and MUST NOT include prompts, queries, URLs, or request bodies in those errors.

9.9 Codex/OpenAI Hosted Web Policy

This section is normative.

OpenAI-hosted retrieval executes beyond the AWF network boundary. Squid sees the permitted OpenAI endpoint, not the searched or fetched destination. AWF therefore enforces apiProxy.hostedWeb.codex inside the trusted API proxy on:

  1. Responses requests containing a web_search or dated web_search_YYYY_MM_DD tool; and
  2. every request to /v1/alpha/search.

The config object is closed and references the same schema, normalization, and source-precedence contract as apiProxy.hostedWeb.claude: enabled is required; enabled: true requires exactly one non-empty allowedDomains or blockedDomains; enabled: false permits neither; and maxUses, when present, is a positive integer. Domains are normalized and validated identically for both providers. The two policies may coexist in one config.

apiProxy:
  hostedWeb:
    claude:
      enabled: true
      blockedDomains:
        - untrusted.example
    codex:
      enabled: true
      allowedDomains:
        - docs.github.com
        - nodejs.org
      maxUses: 5

Configuration-source precedence is:

  1. an explicit CLI option, if one is added in the future;
  2. the validated AWF document, including JSON/YAML via --config -; and
  3. compatibility behavior.

The initial implementation is config-only. AWF_CODEX_HOSTED_WEB_POLICY is an internal transport generated from validated config and cannot independently override it. Omission preserves pass-through behavior and does not constrain Codex-hosted egress.

9.9.1 Request Precedence

The configured policy is immutable. A request can only narrow it. OpenAI can represent allowed and blocked filters together, so the canonical cross-mode rule is to preserve the request's opposite-mode restriction while applying the configured mode; no restriction is silently removed.

Config mode Request filters Effective filters
disabled Any Responses hosted tool or standalone request Reject with codex_hosted_web_disabled
allow none Inject configured allowed_domains
allow allowed_domains Intersect with configured allowlist; reject an empty result
allow blocked_domains Inject configured allowlist and preserve request blocklist
block none Inject configured blocked_domains
block blocked_domains Union with configured blocklist
block allowed_domains Preserve request allowlist and inject configured blocklist

Both request filter arrays, when present, must be non-empty valid domain lists. On the standalone route, the effective settings.filters is applied again to each commands.search_query[].domains and commands.image_query[].domains list. An allowlist is intersected; blocked domains are removed; and an empty effective query scope rejects with codex_hosted_web_empty_query_scope rather than falling back to the request-level scope.

Literal HTTP(S) URLs in commands.open[].ref_id, commands.find[].ref_id, and commands.screenshot[].ref_id are host-checked against both effective lists. Non-URL provider reference IDs remain valid. A malformed URL/host rejects; a disallowed host rejects with codex_hosted_web_url_disallowed.

The most restrictive applicable configured, request, per-query, and URL scope governs each retrieval. No scope can relocate or widen authority granted by another. network.allowDomains and network.sensitiveAllowDomains are never copied into provider policy or request bodies.

9.9.2 Access Modes, Limits, and Unknown Shapes

Responses external_web_access and indexed_web_access values must be booleans. Standalone settings.external_web_access accepts booleans or the current cached, indexed, and live modes. Unknown future values fail closed with codex_hosted_web_access_invalid; they are never interpreted as permissive.

For Responses tools, configured maxUses is injected as max_uses, or resolved as min(configured, requested). Malformed, zero, negative, or non-integer request values reject. /v1/alpha/search exposes no equivalent per-request cap, so a configured maxUses rejects that route with codex_hosted_web_max_uses_unsupported rather than being silently ignored.

Recognized standalone command families are search_query, image_query, open, click, find, and screenshot. Unknown command shapes and malformed dated hosted-tool names fail closed. Enforcement runs after existing OpenAI body/model transforms and immediately before dispatch, so no later transform can remove or expand the effective policy.

9.9.3 Stable Errors and Compiler Contract

Policy failures use invalid_request_error envelopes with stable codex_hosted_web_* codes:

Code Meaning
codex_hosted_web_disabled Hosted retrieval is prohibited
codex_hosted_web_filter_invalid / codex_hosted_web_domain_invalid Malformed filters or domains
codex_hosted_web_empty_intersection / codex_hosted_web_empty_query_scope Narrowing produced no permitted scope
codex_hosted_web_url_invalid / codex_hosted_web_url_disallowed Literal URL is malformed or outside policy
codex_hosted_web_access_invalid Unknown hosted-access mode
codex_hosted_web_max_uses_invalid / codex_hosted_web_max_uses_unsupported Invalid cap or route cannot enforce it
codex_hosted_web_tool_unrecognized / codex_hosted_web_command_unrecognized / codex_hosted_web_shape_invalid Unknown future or malformed hosted-search shape

Errors do not include prompts, queries, literal URLs, or request bodies. Serialized internal policy is validated during sidecar startup. Ordinary OpenAI-compatible requests without a hosted-web surface are unchanged.

Compiler integrations only emit the stable config and pipe it to awf --config - -- <command>; they do not need to understand provider routes, body fields, tool versions, or internal environment variables.

10. Effective Token Budget Enforcement

This section is normative.

When apiProxy.maxEffectiveTokens is configured, the API proxy MUST enforce a cumulative effective-token budget across all LLM API requests in a single run. The budget limits total weighted token consumption, not raw token counts.

10.1 Token Weighting

Each upstream response's usage object is decomposed into four categories, each with a fixed weight:

Category Weight Usage field
Input 1.0 input_tokens / prompt_tokens
Cache read 0.1 cache_read_input_tokens / prompt_tokens_details.cached_tokens
Output 4.0 output_tokens / completion_tokens
Reasoning 4.0 reasoning_tokens / completion_tokens_details.reasoning_tokens

The base weighted tokens for a single response are:

base = (1.0 × input) + (0.1 × cache_read) + (4.0 × output) + (4.0 × reasoning)

10.2 Model Multipliers

When apiProxy.modelMultipliers is configured, each model name MAY have an associated positive multiplier. The effective tokens for a response are:

effective_tokens = model_multiplier × base_weighted_tokens

If no exact multiplier is configured, AWF MUST attempt to match apiProxy.modelMultipliers keys against the request model using a hyphen-suffix prefix match so family keys like claude-opus-4.7 apply to concrete model IDs like claude-opus-4.7-20260501.

If no exact or prefix match is found, and apiProxy.defaultModelMultiplier is configured, that default multiplier MUST be used.

Otherwise, if no exact or prefix match is found, the multiplier MUST default to the highest configured model multiplier. If no model multipliers are configured at all, the multiplier defaults to 1.

When AWF falls back to the default multiplier because no configured model key matched, it MUST emit a warning log entry.

10.3 Enforcement Behavior

The API proxy MUST enforce the budget as follows:

  1. Accumulation: After each successful upstream response, the proxy extracts the usage object, computes effective tokens, and adds them to a running total for the session.

  2. Pre-request check: Before forwarding each subsequent request to the upstream provider, the proxy checks whether the cumulative total has reached or exceeded maxEffectiveTokens.

  3. Rejection: When the budget is reached or exceeded, the proxy MUST reject the request with:

    • HTTP status: 403 Forbidden
    • Content-Type: application/json
    • Response body:
      {
        "error": {
          "type": "effective_tokens_limit_exceeded",
          "message": "Maximum effective tokens exceeded (1234.56 / 1000).",
          "total_effective_tokens": 1234.56,
          "max_effective_tokens": 1000
        }
      }
  4. WebSocket rejection: For WebSocket upgrade requests, the proxy MUST reject with HTTP/1.1 403 Forbidden and include the same JSON error body before destroying the socket.

  5. Finality: Once the budget is reached or exceeded, all subsequent requests in the same run MUST be rejected. The budget is not recoverable.

10.4 Threshold Tracking

The proxy MUST track when cumulative effective tokens cross the following percentage thresholds of maxEffectiveTokens:

Threshold
80%
90%
95%
99%

Each threshold MUST be recorded at most once per run.

10.5 Token Steering

Token steering is opt-in. It is active only when apiProxy.enableTokenSteering is true (CLI: --enable-token-steering), which sets AWF_ENABLE_TOKEN_STEERING=true in the api-proxy sidecar. When disabled (the default), thresholds are still tracked (for introspection) but no warning messages are injected. Setting the field to false, or omitting it, opts a workflow out; the env var is only emitted when the value is true.

When token steering is enabled and a threshold is first crossed, the proxy MUST inject a budget-warning system message into the body of the very next eligible request sent by the agent, then discard the pending message so that it is injected at most once per threshold per run.

The injected message has the format:

[AWF TOKEN WARNING] <threshold-specific text>
Threshold Injected text
80% You have used 80% of your effective token budget. Begin planning to wrap up your current work.
90% You have used 90% of your effective token budget. Complete your current task and prepare final output.
95% You have used 95% of your effective token budget. Finalize and submit your work now.
99% You have used 99% of your effective token budget. You are about to be cut off. Submit immediately.

If multiple thresholds are crossed simultaneously (e.g. a single large response crosses both 80% and 90%), the proxy MUST inject only the highest crossed threshold on the next request and queue the remaining thresholds for subsequent requests (one per request).

Provider-specific injection rules:

  • OpenAI / Copilot — the proxy inserts a { "role": "system", "content": "<message>" } entry into the messages array immediately after any pre-existing system messages.
  • Anthropic — the proxy appends the warning to the system field: if system is a string it is concatenated (separated by \n\n); if system is an array of content blocks a { "type": "text", "text": "<message>" } block is appended; if system is absent it is created as the warning string.
  • Gemini — the proxy appends a { "text": "<message>" } part to systemInstruction.parts; if systemInstruction is absent it is created.

If the request body cannot be parsed as JSON, or if the body format does not match the expected structure, the proxy MUST silently skip injection for that request and NOT re-queue the message.

When token steering is enabled and container.agentTimeout is configured, the proxy MUST also inject runtime warnings at 80/90/95/99% of elapsed run time using the same queueing behavior (highest crossed threshold first, then one pending warning per subsequent request):

[AWF TIME WARNING] <threshold-specific text>

10.6 Introspection

The API proxy exposes a GET /reflect endpoint on every provider port (10000–10004). Each port returns the same aggregate reflection payload, whose endpoints array lists all provider adapters. Only the management port (10000, OpenAI) serves /metrics and the aggregate /health; non-management ports still serve provider-local /health responses.

10.7 Max AI Credits Configuration

maxAiCredits is a positive number. It is supplied via the AWF config file (including stdin config via --config -) and maps to the AWF_MAX_AI_CREDITS environment variable injected into the api-proxy container.

When configured, the proxy MUST enforce this budget in addition to any configured maxEffectiveTokens budget. Once cumulative AI credits reach or exceed maxAiCredits, subsequent requests MUST be rejected with HTTP 403 and error type ai_credits_limit_exceeded.

Regardless of maxAiCredits configuration, AWF also enforces a non-overridable hard cap of 10,000 AI credits. When cumulative AI credits reach this hard cap, subsequent requests MUST be rejected with HTTP 403 and error type ai_credits_limit_exceeded, and the error/log payload MUST include hard_cap: true.

If both limits are present, the effective enforcement threshold is the lower of:

  • configured maxAiCredits
  • the fixed hard cap (10,000)

Setting maxAiCredits above 10,000 MUST NOT raise the effective limit.

10.7.1 Model Name Resolution for Pricing

The AI credits guard resolves model names using this lookup order:

  1. Operator provider overlay — model prices configured under apiProxy.providers.
  2. Runtime provider metadata — authoritative token prices discovered from the configured provider. Copilot supports this today.
  3. Curated pricing table — a built-in table of known models with exact pricing.
  4. Bundled models.dev catalog — a bundled snapshot of the models.dev catalog used as a fallback when the model is not found in the curated table.

Model names are canonicalized before lookup: provider prefixes (e.g. copilot/) are stripped, and separators (., _, -) are treated as interchangeable. For example, copilot/claude-sonnet-4.6, claude_sonnet_4_6, and claude-sonnet-4-6 all resolve to the same pricing entry.

If none of these sources resolves the model, the defaultAiCreditsPricing fallback (if configured) is used. If that is also absent, the request is rejected. Models whose catalog entry carries zero-cost pricing are recognized as known models with zero AI credit impact, so they are never rejected as "unknown".

Runtime tiered pricing uses the provider's default-tier prompt threshold. When the total input exceeds that threshold, all token categories use the long-context tier. Pricing source, API version, observation time, selected tier, and any provider-advertised promotion are retained in provenance. Promotions are informational only because provider discovery does not prove that a discount applies to a specific request; they never reduce accounting. Failed or empty discovery responses do not replace the last successful runtime snapshot.

Provider overlays use the models.dev provider structure and per-token dollar rates:

apiProxy:
  providers:
    github-copilot:
      models:
        custom-model:
          cost:
            input: "3e-06"
            output: "1.5e-05"
            cache_read: "3e-07"
            cache_write: "3.75e-06"

The overlay is passed to both normal and threat-detection API proxy instances through AWF_API_PROXY_PROVIDERS. Provider aliases github-copilot and copilot resolve to the Copilot proxy.

10.7.2 Default AI Credits Pricing (Fallback)

defaultAiCreditsPricing is an optional object with input and output fields (both required, in $/1M tokens), plus optional cachedInput and cacheWrite fields.

It is supplied via the AWF config file and maps to the AWF_DEFAULT_AI_CREDITS_PRICING environment variable (JSON string) injected into the api-proxy container.

When configured, any model not found in the curated built-in pricing table or the bundled models.dev catalog uses these rates as a fallback for AI credits calculation.

10.7.3 Unknown Model Rejection

When maxAiCredits is active and the proxy encounters a request whose model cannot be resolved from the curated built-in pricing table or the bundled models.dev catalog:

  1. If defaultAiCreditsPricing is configured: the fallback rates are used and the request proceeds normally.

  2. If defaultAiCreditsPricing is NOT configured: the proxy MUST reject the request with HTTP 400 and error type unknown_model_ai_credits. The error payload includes:

    • model: the unresolved model name
    • message: human-readable instructions to configure apiProxy.defaultAiCreditsPricing

    This fail-closed behavior prevents unaccounted spending from models whose pricing is unknown to the proxy.

Note: Requests without a model field in the body (e.g. non-chat endpoints) are not subject to this check.

Recognized dynamic selectors. The Copilot auto model (copilot:auto) is not subject to unknown_model_ai_credits rejection. Because its concrete runtime model is not known at request time, the proxy accounts it using a conservative fallback ceiling (the maximum per-token rate across the curated pricing catalog) rather than rejecting the request or silently under-counting spend. Token-usage records for these requests set pricing_source and accounting_policy to dynamic_selector_fallback, pricing_tier to conservative, fallback_pricing_used to true, and dynamic_selector to copilot:auto. This dynamic-selector accounting path applies only to the recognized Copilot auto selector; other unresolved models still follow the fallback/rejection behavior above.

10.7.4 Token Usage JSONL Schema Extensions

When AI credits and/or effective tokens are computed, the token-usage.jsonl records include additional optional fields:

Field Type Description
effective_tokens_this_response number Weighted tokens for this request
effective_tokens_total number Running total of effective tokens
model_multiplier number Cost multiplier applied for this model
ai_credits_this_response number AI credits consumed by this request
ai_credits_total number Running total of AI credits

These fields are only present when the respective guard is active.

11. Max-Runs Enforcement

This section is normative.

When apiProxy.maxTurns is configured, the API proxy MUST enforce an absolute maximum number of LLM invocations per run.

11.1 Counting Invocations

An invocation is counted each time the proxy receives a successful (2xx) HTTP response from an upstream LLM provider. Each response increments a per-run counter by one, regardless of the number of tokens consumed.

11.2 Enforcement Behavior

The API proxy MUST enforce the max-runs limit as follows:

  1. Pre-request check: Before forwarding each request to the upstream provider, the proxy checks whether the invocation count has reached or exceeded maxTurns.

  2. Rejection: When the limit is reached or exceeded, the proxy MUST reject the request with:

    • HTTP status: 403 Forbidden
    • Content-Type: application/json
    • Response body:
      {
        "error": {
          "type": "max_runs_exceeded",
          "message": "Maximum LLM invocations exceeded (5 / 5).",
          "invocation_count": 5,
          "max_runs": 5
        }
      }
  3. WebSocket rejection: For WebSocket upgrade requests, the proxy MUST reject with HTTP/1.1 403 Forbidden and include the same JSON error body before destroying the socket.

  4. Finality: Once the limit is reached, all subsequent requests in the same run MUST be rejected. The counter is not recoverable.

11.3 Introspection

The /reflect endpoint (available on all provider ports 10000–10004; see §10.6) MUST include the current max-runs state:

{
  "runs": {
    "enabled": true,
    "max_runs": 5,
    "invocation_count": 3,
    "remaining_runs": 2
  }
}

When maxTurns is not configured, the enabled field MUST be false and max_runs and remaining_runs MUST be null.

11a. Permission-Denied Guard

This section is normative.

When apiProxy.maxPermissionDenied is configured, the API proxy MUST halt further LLM requests after the upstream returns a configurable number of 401 or 403 responses, preventing token waste when API credentials are misconfigured or expired.

11a.1 Counting Permission Errors

A permission error is counted each time the proxy receives an HTTP 401 or 403 response from an upstream LLM provider. Each such response increments a per-run counter by one.

11a.2 Enforcement Behavior

The API proxy MUST enforce the permission-denied limit as follows:

  1. Post-response counting: After receiving a 401 or 403 from upstream, the proxy increments the denied count.

  2. Pre-request check: Before forwarding each subsequent request to the upstream provider, the proxy checks whether the denied count has reached or exceeded maxPermissionDenied.

  3. Rejection: When the limit is reached or exceeded, the proxy MUST reject the request with:

    • HTTP status: 403 Forbidden
    • Content-Type: application/json
    • Response body:
      {
        "error": {
          "type": "permission_denied_limit_exceeded",
          "message": "Permission denied limit exceeded (3 / 3). The run has been stopped due to repeated permission errors — check that all API keys and tokens are correctly configured.",
          "denied_count": 3,
          "max_permission_denied": 3
        }
      }
  4. Finality: Once the limit is reached, all subsequent requests in the same run MUST be rejected until the configured limit changes (changing AWF_MAX_PERMISSION_DENIED resets the counter).

11a.3 Introspection

The /reflect endpoint (available on all provider ports 10000–10004; see §10.6) MUST include the current permission-denied guard state:

{
  "permission_denied": {
    "enabled": true,
    "max_permission_denied": 3,
    "denied_count": 1
  }
}

When maxPermissionDenied is not configured, the enabled field MUST be false, max_permission_denied MUST be null, and denied_count MUST be 0.

11a.4 Configuration

maxPermissionDenied is a positive integer. It is supplied via the AWF config file (stdin config) or the --max-permission-denied CLI flag, and maps to the AWF_MAX_PERMISSION_DENIED environment variable injected into the api-proxy container.

Example:

apiProxy:
  maxPermissionDenied: 3   # stop run after 3 upstream 401/403 responses

11b. Cache-Miss Guard

This section is normative.

When apiProxy.maxCacheMisses is configured, the API proxy MUST halt further LLM requests after the configured number of consecutive responses that had no prompt-cache hits, preventing runaway token spend caused by a broken or expired cache (e.g., mismatched cache keys, context window overflow, or prompt drift).

11b.1 Counting Cache Misses

A cache miss is counted for a response when all of the following are true:

  • The response is a successful upstream completion (not a proxy-level error).
  • input_tokens > 0 (zero-input responses such as empty tool calls are excluded so they do not inflate the streak counter).
  • cache_read_tokens === 0 (no prompt-cache hit occurred).

A cache hit (cache_read_tokens > 0) resets the consecutive miss streak to zero.

11b.2 Enforcement Behavior

The API proxy MUST enforce the cache-miss limit as follows:

  1. Post-response counting: After receiving each successful upstream response, the proxy inspects the normalized token usage and increments or resets the miss streak counter.

  2. Pre-request check: Before forwarding each subsequent request to the upstream provider, the proxy checks whether the miss streak has reached or exceeded maxCacheMisses.

  3. Rejection: When the limit is reached or exceeded, the proxy MUST reject the request with:

    • HTTP status: 403 Forbidden
    • Content-Type: application/json
    • Response body:
      {
        "error": {
          "type": "max_cache_misses_exceeded",
          "message": "Maximum consecutive cache misses exceeded (3 / 3).",
          "consecutive_cache_misses": 3,
          "max_cache_misses": 3
        }
      }
  4. WebSocket rejection: For WebSocket upgrade requests, the proxy MUST reject with HTTP/1.1 403 Forbidden and include the same JSON error body before destroying the socket.

  5. Finality: Once the streak limit is reached, all subsequent requests in the same run MUST be rejected. Changing AWF_MAX_CACHE_MISSES resets the streak counter.

11b.3 Introspection

The /reflect endpoint (available on all provider ports 10000–10004; see §10.6) MUST include the current cache-miss guard state:

{
  "cache_misses": {
    "enabled": true,
    "max_cache_misses": 3,
    "consecutive_cache_misses": 1,
    "remaining_cache_misses": 2
  }
}

When maxCacheMisses is not configured, the enabled field MUST be false, max_cache_misses MUST be null, consecutive_cache_misses MUST be 0, and remaining_cache_misses MUST be null.

11b.4 Configuration

maxCacheMisses is a positive integer. It is supplied via the AWF config file (stdin config) or the --max-cache-misses CLI flag, and maps to the AWF_MAX_CACHE_MISSES environment variable injected into the api-proxy container.

Example:

apiProxy:
  maxCacheMisses: 3   # stop run after 3 consecutive cache misses

12. Model Multiplier Cap

This section is normative.

When apiProxy.maxModelMultiplierCap is configured, the API proxy MUST reject any request whose resolved model multiplier exceeds the cap before forwarding the request to the upstream provider.

12.1 Multiplier Resolution

The proxy resolves the effective multiplier for the requested model using the same algorithm as the effective-token guard:

  1. Exact match: if apiProxy.modelMultipliers contains the exact model name, use its multiplier.
  2. Longest-prefix match: if any configured model name is a prefix of the requested model name (followed by -), use the multiplier of the longest-matching prefix.
  3. Default: use apiProxy.defaultModelMultiplier if configured, otherwise default to 1.

12.2 Enforcement Behavior

Before forwarding each POST/PUT/PATCH request to an upstream LLM provider, the proxy MUST:

  1. Extract the model field from the request body.

  2. Resolve the model's effective multiplier (§12.1).

  3. If the multiplier exceeds maxModelMultiplierCap, reject the request with:

    • HTTP status: 400 Bad Request
    • Content-Type: application/json
    • Response body:
      {
        "error": {
          "type": "model_multiplier_cap_exceeded",
          "message": "Model multiplier cap exceeded: model \"claude-opus-4.7\" has multiplier 27 which exceeds the configured maximum of 5.",
          "model": "claude-opus-4.7",
          "model_multiplier": 27,
          "max_model_multiplier": 5
        }
      }
  4. If the model field is absent or the multiplier is within the cap, the request MUST be forwarded normally.

12.3 Configuration

maxModelMultiplierCap is a positive number. It is supplied via the AWF config file (stdin config) and maps to the AWF_MAX_MODEL_MULTIPLIER environment variable injected into the api-proxy container. The CLI flag --max-model-multiplier-cap <number> may also be used.

Example:

apiProxy:
  maxModelMultiplierCap: 5       # reject any model with multiplier > 5
  modelMultipliers:
    claude-opus-4.7: 27
    gpt-4o: 2

13. Model Fallback

This section is normative.

When apiProxy.modelFallback is configured, the API proxy provides automatic model selection when a requested model is unavailable. The fallback mechanism ensures requests complete gracefully without requiring explicit agent-side handling.

12.1 Configuration

Model fallback is controlled via apiProxy.modelFallback:

{
  "apiProxy": {
    "modelFallback": {
      "enabled": true,
      "strategy": "middle_power"
    }
  }
}
Field Type Default Description
enabled boolean true Enable/disable the fallback mechanism
strategy string middle_power Selection strategy (middle_power is currently the only strategy)
excludeEngines string[] [] Engines for which middle-power fallback is suppressed (e.g. ["openai"]). Excluded engines receive native model-unavailable errors instead of silent rewrites.

12.2 Middle-Power Strategy

When strategy is middle_power, the proxy selects the median capability-tier model from the available models for the current provider.

Capability tiers are assigned based on model family and version:

Provider Tier 5 Tier 4 Tier 3 Tier 1
Anthropic claude-opus* claude-sonnet* claude-haiku* (other)
OpenAI / Copilot gpt-5* gpt-4*, gpt-4o* gpt-3.5* (other)
Gemini (reserved) (reserved) (reserved) (all)

Selection algorithm:

  1. Sort available models by capability tier (highest first), then lexicographically
  2. Select the median model from the sorted list
  3. Log the selection with the reason and full candidate list

Example:

Available: ['gpt-3.5-turbo', 'gpt-5.2', 'gpt-4.1']
Sorted:   ['gpt-5.2' (tier 5), 'gpt-4.1' (tier 4), 'gpt-3.5-turbo' (tier 3)]
Median:   gpt-4.1 (index 1 of 0-2)

12.3 Activation Conditions

The fallback is activated when:

  1. Direct match fails: The requested model is not found in the available models list for the provider.
  2. Family version fallback doesn't apply: For gpt-5.* models on OpenAI, if a lower gpt-5.* version is available, use that before triggering middle-power fallback.
  3. Alias has no candidates: An alias pattern matched but produced no resolvable models on the current provider.

The fallback is NOT activated when:

  • A direct model match is found (return it immediately)
  • A family version fallback is available (for gpt-5.* only)
  • The fallback is disabled (enabled: false)
  • An alias has fallback: false (see §12.4)
  • The provider is in the excludeEngines list
  • Copilot engine in standard mode (no BYOK env vars): the Copilot CLI is authoritative for its own model catalogue, so retired/restricted model names should fail fast with a clear upstream error rather than being silently rewritten to a middle-power fallback
  • Copilot BYOK that still targets a GitHub Copilot catalog host (for example api.githubcopilot.com): the catalog remains authoritative, so fallback is still suppressed
  • Copilot is configured for a BYOK non-githubcopilot target (for example Azure OpenAI deployment endpoints), where deployment names are provider-local and must not be rewritten to catalog model IDs

12.4 Extended Alias Syntax

Model aliases now support an extended syntax that permits per-alias fallback control:

Legacy syntax (string array):

{
  "models": {
    "sonnet": ["copilot/*sonnet*", "openai/*sonnet*"]
  }
}

Fallback is enabled by default for legacy syntax.

Extended syntax (object with patterns):

{
  "models": {
    "sonnet": {
      "patterns": ["copilot/*sonnet*"],
      "fallback": false
    }
  }
}
Field Type Default Description
patterns string[] — Glob patterns to match against available models
fallback boolean true Enable fallback for this alias if no candidates are found

When fallback: false, if the alias patterns produce no candidates, the resolution returns null instead of activating middle-power fallback.

12.5 Introspection

The health endpoint (GET /health) includes a model_fallback field in the response:

{
  "status": "healthy",
  "service": "awf-api-proxy",
  "model_fallback": {
    "enabled": true,
    "strategy": "middle_power"
  }
}

The /reflect endpoint does not include fallback state by design (it is static per run).

12.6 Pre-Startup Model Validation

When apiProxy.requestedModel is configured, the API proxy validates at startup that the specified model is available in at least one provider's model catalogue.

Configuration:

{
  "apiProxy": {
    "requestedModel": "gpt-4o"
  }
}

Mapping: apiProxy.requestedModel → AWF_REQUESTED_MODEL (config-only; set by AWF CLI)

Behavior:

  1. After fetchStartupModels() completes, the proxy checks AWF_REQUESTED_MODEL against all cached provider model lists.
  2. If the model is found directly or resolves via model aliases, a confirmation model_validation log is emitted.
  3. If the model is NOT found, a model_unavailable_at_startup error log is emitted listing available models as a diagnostic aid.
  4. Validation is non-blocking — the proxy continues serving requests regardless of the outcome, so agents that ignore the model hint are not affected.

This enables workflow authors to get clear, early feedback when a retired or misspelled model is specified, rather than waiting for the first API request to fail with an opaque error.

12.7 Ordered Fallback Models

When apiProxy.fallbackModels is configured, the API proxy retries a failed request with the next model in an ordered list. The middle-power fallback above picks a model when the request is resolved. This chain applies after the upstream has rejected the model.

{
  "apiProxy": {
    "fallbackModels": ["gpt-5.4", "claude-sonnet-4.6"]
  }
}

Mapping: apiProxy.fallbackModels → AWF_FALLBACK_MODELS (JSON array; a comma-separated list is also accepted when the variable is set directly)

Behavior:

  1. Fallback triggers only on model-specific failures:
    • any upstream 5xx response (504 is reported as upstream_timeout)
    • a connection error or timeout before any upstream response (upstream_connection_error)
    • a 400/404 whose body says the model is unsupported, not found, or not accessible, such as model_not_supported, model_not_found, or not accessible via the … endpoint (model_not_supported)
  2. 401, 403, and 429 responses never trigger a fallback. A generic 400 validation error, such as a bad tool schema or a too-long context, is returned to the client unchanged.
  3. The proxy rewrites the request body's model field (OpenAI, Anthropic, Copilot). For Gemini, where the body has no model, it rewrites the /models/<model>:<method> path segment instead. The proxy strips a redundant <provider>/ prefix on fallback entries.
  4. Entries are tried in order. The proxy skips models already attempted for the request and models rejected by the model-policy, retired-model, multiplier-cap, or budget guards. When the chain is exhausted, the proxy returns the last upstream error to the client.
  5. Each switch emits a model_fallback warning log with from_model, to_model, requested_model, attempt, reason, and status. The token-usage record for the successful response holds the model that served the request in model, plus a model_fallback object with requested_model, model, attempt, reason, and status. gh-aw can report that model in GH_AW_INFO_MODEL and telemetry.
  6. Copilot's existing transient model not supported retries and the alias endpoint-blocked candidate retry still run first. The ordered chain applies only after those have been exhausted.
  7. WebSocket (Responses API) upgrades and AWF-internal routing classifier requests are not covered.

12.1 Alias Candidates Are Restricted to Configured Providers

Alias resolution MUST only consider provider slots that are actually configured for the run. Before an alias is expanded (and before aliases are advertised via /reflect and models.json), the cached model lists of providers that report configured: false are treated as empty. A provider-scoped pattern such as copilot/*sonnet* therefore yields no candidate when no Copilot credential is present, even if a model list was cached earlier in the run.

Configuration is determined from each provider's reflected configured slot, not from request readiness. A configured OIDC provider remains eligible while its token is being minted. When configured providers have no model catalogue yet, aliases scoped only to unconfigured providers are still omitted, while aliases that can target a configured provider remain advertised until model data is available.

Without this filter, a Copilot-first alias group would steer every request to a slot that answers provider_not_configured, producing a 100% call-failure rate and, for retry-happy clients, a non-terminating retry loop.

A provider_not_configured response is a terminal run-level misconfiguration: it is returned with HTTP 403 and "retryable": false so clients fail fast. Only transient OIDC readiness states, such as a token that has not been minted yet, use HTTP 503 and "retryable": true.

13a. Task-Level Model Routing

Task-level model routing is experimental and opt-in. To enable it, set root-level experimental.modelRouting: true alongside the compiler-authored apiProxy.routing request. An existing config containing apiProxy.routing without this gate must add the experimental block shown below; otherwise validation fails rather than silently ignoring the request. The gate alone, without apiProxy.routing, is valid but does not activate routing. Omitting the gate (or setting it to false) without a routing request preserves normal operation without routing infrastructure or environment.

As of this release, both the proxy-side and host-side halves of task-level routing are wired and shipped on main. The API proxy's routing controller is wired into the running server (PR #8966): when AWF_ROUTING_CONFIG is present, a routing session starts after key validation and model discovery and selects one model/effort for the run. The host workflow stages and validates apiProxy.routing input before the proxy starts (PR #8985): it writes the task conversation into a private per-run routing directory, rejects unsupported configurations (non-Linux, non-runc, disabled API proxy, --keep-containers, DinD/split filesystems, Docker-socket exposure, or an unpinned router image), and waits for selection.json before starting the agent.

The selection is advisory, not admitted-only. The agent is seeded with the selected model, effort, and endpoint, but the proxy does not pin requests to it: an agent (or a sub-agent that declares its own model:) MAY send any model that model policy permits, and such a request completes normally. Model choice is not a containment boundary — allowedModels / disallowedModels (AWF_ALLOWED_MODELS / AWF_DISALLOWED_MODELS) remain the enforcement surface that bounds cost and policy, independently of routing, and a request for an excluded model is still rejected by that policy guard. WebSocket upgrades are proxied normally while a routing session exists. For each inference request the proxy logs a model_routing event with stage: "request", recording the requested and selected provider, model, effort, and endpoint, routed: "as_selected" or "deviated", and the list of deviations, so routing quality stays measurable without enforcement.

A genuine routing failure still surfaces as host exit code 78 instead of the run silently continuing: no selection could be produced (no_route, router unreachable, contract or configuration errors), or an upstream failure on a request that used the selected provider and model (a native provider error code, an SSE error event, or a prematurely closed response). A request that uses another model is not a routing failure, and neither is its upstream error or a model-policy rejection of it.

The agent learns the selected model, effort, and endpoint from the API proxy's GET /reflect routing field (see api-proxy-sidecar.md); the private selection.json is not visible to the agent.

experimental:
  modelRouting: true
apiProxy:
  routing:
    provider: copilot
    objective:
      goal: cost
      mode: balanced
    task:
      conversationFile: /tmp/gh-aw/routing-conversation.json
Field Allowed values Description
provider copilot (default), openai, anthropic Restricts routing to one configured native API-proxy provider; AWF does not switch credentials or translate across providers.
candidateModels non-empty array of model glob patterns Limits router/classifier choices without widening the request policy; defaults to apiProxy.allowedModels.
objective.goal cost, cost-speed Optimization goal used by the router
objective.mode economy, balanced, robust, auto Fixed routing profile, or auto classification
task.conversationFile non-empty string Host path to the task conversation whose description the router classifies

The task conversation must be written by the workflow host before AWF starts. Callers are responsible for preparing task-relevant conversation content before AWF starts. AWF classifies the supplied conversation as-is; it does not parse workflow prompt markup or remove injected system instructions. If a rendered prompt contains a system block, the caller should provide the task conversation without that unrelated block. The conversation is a JSON array in the router's conversation format, for example [{"role":"user","parts":[{"text":"Fix the failing unit test."}]}], with at least one non-blank user message and at most 1 MiB. Its user messages form the task description, which the router classifies once per run. The resulting classification (task type, scope, complexity, and, for mode: auto, the routing profile) selects the one model and effort the agent is seeded with for the whole run. The router is not invoked again per request or per sub-agent.

The routing object is closed: objective and task are required, provider and candidateModels are optional, and unknown properties are rejected. Omitting provider preserves Copilot routing. OpenAI and Anthropic routing require the matching provider to be configured for the agent and currently require the native api.openai.com or api.anthropic.com target; AWF does not route custom gateways or translate or forward requests across provider boundaries. A supported routed run also requires a complete container.images manifest containing digest-pinned references for router and every other enabled image role. The legacy latest router default is kept only for resolver compatibility and is not a supported tag-only routed configuration.

The candidate pool uses models discovered for the selected native provider. The /reflect model_api_mapping includes maintained routing metadata for selected model families where provider /models endpoints do not publish context limits or reasoning-effort support. The initial maintained set covers OpenAI gpt-5.4 and gpt-5.4-2026-03-05, plus Anthropic claude-opus-5-5, claude-opus-5, claude-fable-5-1, claude-fable-5, claude-mythos-5-1, claude-mythos-5, claude-mythos-preview*, claude-sonnet-5-5, claude-sonnet-5, claude-opus-4-6/4-7/4-8, and claude-sonnet-4-6. This is deliberately not exhaustive: discovered IDs are eligible only when endpoint and effort support are known from runtime metadata or an exact maintained mapping. For example, o3, gpt-5-nano, and gpt-5.4-mini do not inherit metadata from the GPT-5.4 base model and can yield no_route. Positive context limits are also maintained in /reflect; without one, a model is excluded from classifier preflight but may still remain available to the router.

An explicitly empty effort list (or explicit lack of reasoning-effort support) allows one effortless choice. OpenAI uses its mapped Responses or Chat Completions protocol; Anthropic uses Messages. Unsupported effort values are discarded, and a model with no remaining advertised effort is excluded rather than converted into an effortless choice.

Request guards and alias resolution use the provider-aware allowedModels / disallowedModels policy. Candidate filtering additionally uses apiProxy.routing.candidateModels when supplied; those patterns only narrow the router/classifier pool and never widen the request policy. When omitted, candidates continue to be derived from allowedModels. Native patterns such as gpt-* match the native model name; qualified patterns such as github-copilot/gpt-* match that provider only. Copilot recognizes the existing copilot, github-copilot, and github provider aliases. Matching remains case-insensitive with * wildcards, and deny rules take precedence. A provider-prefixed auto remains subject to dynamic-model verification: under a denylist it must also match an explicit allow rule.

13. Model Alias Logging

The API proxy emits structured logging events during model alias resolution. These events are critical for debugging model routing decisions in production.

13.1 Always-On Events (stdout)

The following events are emitted as JSON lines to the API proxy's stdout (captured by Docker logging). They are always active when model aliases are configured (apiProxy.models):

Event Trigger Key fields
model_resolution Every request where a model alias resolves requested_model, resolved_model, provider, resolution_log[]
model_rewrite Every request where the model field is rewritten original_model, rewritten_model, provider
model_fallback_activated Fallback strategy selected a replacement reason, selected, candidates[]
model_fallback_skipped Fallback was available but explicitly suppressed reason, requested_model
model_fallback_candidates Informational: available fallback models candidates[], strategy

These events are written by logRequest() in containers/api-proxy/logging.js.

13.2 Diagnostic Events (token-diag.jsonl)

When apiProxy.logging.debugTokens is true (or AWF_DEBUG_TOKENS=1), additional diagnostic events are written to token-diag.jsonl in the directory specified by apiProxy.logging.tokenLogDir (default: /var/log/api-proxy):

Event Description
model_alias_resolution_step Each step in the alias resolution chain (input → pattern match → candidate)
model_alias_rewrite Final rewrite decision with before/after model names and matched pattern

Each diagnostic record follows the token-diag/v<version> schema:

{
  "_schema": "token-diag/v0.25.40",
  "timestamp": "2025-01-15T10:30:00.000Z",
  "event": "model_alias_resolution_step",
  "data": {
    "alias": "sonnet",
    "pattern": "anthropic/*sonnet*",
    "candidate": "claude-sonnet-4-5",
    "provider": "anthropic"
  }
}

13.3 Configuration

apiProxy:
  models:
    sonnet: ["copilot/*sonnet*", "anthropic/*sonnet*"]
  logging:
    debugTokens: true
    tokenLogDir: "/var/log/api-proxy"
  diagnostics:
    captureBlockedRequests: summary  # false | summary | redacted | full
    maxCapturedBytes: 250000
Property Type Default Env var Description
apiProxy.logging.debugTokens boolean false AWF_DEBUG_TOKENS Enable diagnostic token/model-alias logging to file
apiProxy.logging.tokenLogDir string /var/log/api-proxy AWF_TOKEN_LOG_DIR Directory for token-usage.jsonl and token-diag.jsonl
apiProxy.diagnostics.captureBlockedRequests string | boolean false AWF_CAPTURE_BLOCKED_LLM_REQUESTS Capture body-shape info for guard-blocked requests (false/true/summary/redacted/full; true is an alias for summary)
apiProxy.diagnostics.maxCapturedBytes integer 250000 AWF_MAX_BLOCKED_CAPTURE_BYTES Max bytes per record in full capture mode

13.4 Log File Inventory

AWF produces the following structured and unstructured log files at runtime. All JSONL files use the .jsonl extension.

All AWF JSONL records MUST include the following top-level fields:

  • timestamp (string, required): ISO 8601 UTC with milliseconds (YYYY-MM-DDTHH:mm:ss.SSSZ).
  • event (string, required): Stable snake_case record discriminator.
  • _schema (string, required): Schema identifier in the form <record-type>/v<version>.

Squid Proxy Logs

Directory: configured by logging.proxyLogsDir (default: <workDir>/squid-logs/)

File Format Description Always written
access.log Custom text (firewall_detailed logformat) L7 HTTP/HTTPS traffic decisions with timestamps, client IP, domain, status, and decision codes Yes
audit.jsonl JSONL (audit/v<version> schema) Structured version of access log; preferred for programmatic consumption Yes
cache.log Squid native text Squid internal diagnostics (startup, shutdown, errors) Yes

API Proxy Logs

Directory: configured by apiProxy.logging.tokenLogDir / AWF_TOKEN_LOG_DIR (default: /var/log/api-proxy/; must be /var/log/api-proxy or a subdirectory to be preserved by AWF's default bind mount)

On the runner, these files are preserved under <logging.proxyLogsDir>/api-proxy-logs/ (or /tmp/api-proxy-logs-<ts>/ when proxyLogsDir is not set). After cleanup AWF logs Token usage log available at: <path> and, when $GITHUB_ENV is set, exports AWF_TOKEN_USAGE_LOG=<path> so later workflow steps can locate token-usage.jsonl without hardcoding a path (see ARC + DinD).

File Format Description Always written
token-usage.jsonl JSONL (token-usage/v<version> schema) Per-API-call token usage and cost records Yes (when API proxy is active)
token-diag.jsonl JSONL (token-diag/v<version> schema) Diagnostic events: model resolution steps, alias rewrites, token budget decisions Only when apiProxy.logging.debugTokens: true
blocked-request-diag.jsonl JSONL (blocked-request-diag/v<version> schema) Body-shape diagnostics for guard-blocked requests (effective tokens, AI credits, etc.) Only when apiProxy.diagnostics.captureBlockedRequests is set
otel.jsonl JSONL (OpenTelemetry spans) Distributed tracing spans; written as local fallback when no OTLP collector is configured Only when OTEL is active and no collector endpoint set

CLI Proxy Logs

Directory: /var/log/cli-proxy/ (or AWF_CLI_PROXY_LOG_DIR)

File Format Description Always written
access.jsonl JSONL CLI proxy request audit records (gh CLI invocations routed through DIFC proxy) Yes (when CLI proxy is active)

API Proxy stdout (Docker logs)

The API proxy also emits JSON lines to stdout (captured by docker logs). These are always active and include model resolution events (model_resolution, model_rewrite, model_fallback_*). Use docker logs awf-api-proxy or the AWF diagnostic log collection to access them.

13.5 Availability

Model alias logging was introduced in v0.25.40 (PR #2329). The diagnostic file mechanism (token-persistence.js) was refactored into a dedicated module in v0.25.50 but the logging events and their format have been stable since initial release.

13.6 Blocked Request Diagnostics (blocked-request-diag.jsonl)

When a guard hard-rails a request (e.g. effective_tokens_limit_exceeded, ai_credits_limit_exceeded, max_runs_exceeded), the api-proxy can write a structured diagnostic record to blocked-request-diag.jsonl. This is opt-in and disabled by default.

Enabling

Set the environment variable or config key before starting the container:

# Minimal (body-shape only, no content):
AWF_CAPTURE_BLOCKED_LLM_REQUESTS=summary

# Include first 200 chars of each message (for debugging over-large tool results):
AWF_CAPTURE_BLOCKED_LLM_REQUESTS=redacted

# Full body up to AWF_MAX_BLOCKED_CAPTURE_BYTES (default 250 000 bytes):
AWF_CAPTURE_BLOCKED_LLM_REQUESTS=full
AWF_MAX_BLOCKED_CAPTURE_BYTES=250000

Or via config YAML:

apiProxy:
  diagnostics:
    captureBlockedRequests: summary   # false | summary | redacted | full
    maxCapturedBytes: 250000

Capture modes

Mode Content Use case
false (default) Nothing written Production default
summary Counts, sizes, hashes — no content Safe for normal debugging; identify which message/tool-result was large
redacted Summary + first 200 chars per message Debug prompt growth without full disclosure
full Full body up to maxCapturedBytes Local/private runs only; explicitly document and review

Record format

Each record follows the blocked-request-diag/v<version> schema:

{
  "_schema": "blocked-request-diag/v0.26.0",
  "timestamp": "2025-01-15T10:30:00.000Z",
  "event": "blocked_request_diag",
  "capture_mode": "summary",
  "request_id": "bc446626-a67b-4a78-a8c3-7293a2bc7306",
  "provider": "anthropic",
  "path": "/v1/messages",
  "guard_type": "effective_tokens_limit_exceeded",
  "guard_totals": {
    "total_effective_tokens": 27198679,
    "max_effective_tokens": 25000000
  },
  "body_transformed": true,
  "inbound_bytes": 184320,
  "body_bytes": 185040,
  "body_sha256": "a3f2b1c8d9e0f1a2",
  "model": "claude-opus-4.7",
  "streaming": true,
  "message_count": 52,
  "tool_result_count": 14,
  "message_sizes": [
    { "role": "user",      "content_type": "text",        "chars": 312,   "bytes": 312,   "estimated_tokens": 78 },
    { "role": "assistant", "content_type": "text",        "chars": 1840,  "bytes": 1840,  "estimated_tokens": 460 },
    { "role": "user",      "content_type": "tool_result", "chars": 94321, "bytes": 94321, "estimated_tokens": 23580, "tool_blocks": 3 }
  ]
}

Security considerations

  • summary mode captures no message content and is safe for shared/public workflow runs.
  • redacted mode includes short previews; review before attaching to public issues.
  • full mode captures potentially sensitive prompt and tool-result content. Use only for private runs and rotate or delete the artifact promptly.
  • The file is written to AWF_TOKEN_LOG_DIR alongside token-usage.jsonl and is governed by the same artifact-retention policy.

14. Unified Enclaves

The optional top-level enclaves array defines AWF's sole supported private-repository execution surface. It is structurally identical to the gh-aw compiler's enclave frontmatter: every entry declares exactly one script or agent executor, its own non-empty repos list, and entry-level shared controls including timeout, runtime, image, resource limits, and disclosure limits. AWF stages immutable repository seeds on the host, starts one AWF-owned enclave-mcp-server, maintains one shared per-repository ledger for the run, and exposes configured executors only through compiler-launched gh-aw-mcpg.

Dynamic repository-policy entries (dynamic in place of repos on an agent entry, per docs/adr/0001-agent-enclaves.md) select one canonical repository at runtime, receive one short-lived github-repository-read-v1 identity, and read that repository through GitHub MCP without cloning or mounting a seed. They run only when the gh-aw compiler has started mcpg's github-repository-delegation-v1 controller and handed AWF its private control endpoint and capability; otherwise AWF rejects the run. See §14.1a for the envelope, the handoff, and the identity lifecycle.

14.1 Executors and shared configuration

enclaves:
  - script: {}
    repos:
      - repo: octo-org/private-service
        sensitivity: confidential
    timeout: 45
  - agent:
      model: gpt-5
      maxModelRequests: 3
      maxModelTokens: 10000
      tools:
        github:
          allowed:
            - list_issues
            - issue_read
          allowedRepos:
            - octo-org/private-service
          minIntegrity: none
    runtime: gvisor
    memoryLimit: 256m
    maxOutputBytes: 2048
    maxInvocations: 3
    repos:
      - repo: octo-org/private-service
        sensitivity: confidential
    timeout: 180
  • Script executor — an entry keyed by script; launches a no-network, read-only, single-use Python sandbox. An empty script: {} object is valid and selects AWF's pinned defaults.
  • Agent executor — an entry keyed by agent; launches a bounded single-use Copilot enclave. agent.model is REQUIRED. Optional agent.tools.github (or the deprecated legacy agent.github.cli: issues-read-v1 marker; the two are mutually exclusive) adds compiler-owned shared mcpg as the only additional peer.
  • Entry-level controls — runtime, image, memoryLimit, cpuLimit, pidsLimit, tmpfsLimit, maxOutputBytes, and maxInvocations apply to the entry's selected executor. script.maxScriptBytes and agent maxTaskBytes, maxModelRequests, and maxModelTokens remain executor-specific. Network and interpreter are AWF-owned invariants, not input fields.

At most one entry MAY exist per executor kind, and each entry MUST declare exactly one executor key. Every entry's repos list is merged into one trusted repository catalog: a repository shared by both entries MUST declare the same sensitivity, because sensitivity fixes one shared per-run information budget that both executors debit.

timeout is a per-invocation wall-clock bound in seconds. It defaults to 30 for script entries and 120 for agent entries, and values above 4740 are rejected. Responses use fixed timing buckets at 100 ms, 1 second, 10 seconds, 60 seconds, 120 seconds, 180 seconds, 240 seconds, 300 seconds, 600 seconds, 1200 seconds, 2400 seconds, and 4800 seconds, followed by a cryptographically random, secret-independent delay from 0 through 1000 ms. The canonical enclave MCP tools use a fixed toolTimeout of 4860 seconds, covering the maximum bucket, response jitter, and a bounded transport allowance.

gvisor requires an exactly registered runsc runtime and never falls back. sbx remains fail-closed for both executors until the audited capability proof lands.

cloud-hypervisor is a reserved enclave runtime value governed by ADR 0002. AWF preserves the selection through parsing and validates the shared top-level cloudHypervisor preview, host, and attested-artifact configuration, but currently fails closed before launching an enclave. Execution remains disabled until the host executor, dedicated script and agent rootfs artifacts, workload-specific networking, resource parity, and durable recovery gates land. The initial scope is static script and static agent entries only; dynamic entries and custom image overrides are rejected, and no configuration falls back to another runtime.

An enclave-only cloud-hypervisor selection requires top-level cloudHypervisor configuration but does not select Cloud Hypervisor for the primary agent or alter its mounts, TTY, or container runtime.

The agent executor additionally requires enableApiProxy, a configured provider route for its fixed engine/profile, a configured model, and the absence of enableDind. AWF validates those requirements before repository staging.

14.1a Dynamic repository admission (agent-only)

enclaves:
  - agent:
      model: gpt-5
    dynamic:
      allowedOwners:
        - octo-org
      allowedRepositories: []
      sensitivity: confidential
      executor: agent
      githubPolicy:
        version: github-repository-read-v1
        tools:
          - list_issues
          - issue_read
      maxRepositories: 4
      limits:
        timeoutSeconds: 180
        memoryLimit: 256m
        cpuLimit: "1"
        pidsLimit: 128
        tmpfsLimit: 256m
        maxOutputBytes: 2048
        maxTaskBytes: 4096
        maxModelRequests: 8
        maxModelTokens: 4096
      quotas:
        maxInvocations: 100
        maxOutputBytes: 100000
        maxExecutionSeconds: 3600
      auditLabels:
        - awf-enclave-dynamic
      expiresAt: "2030-01-01T00:00:00Z"

The dynamic object is byte-for-byte the envelope the gh-aw compiler emits. An agent entry MUST declare exactly one of repos or dynamic, never both, and a script entry MUST NOT declare dynamic. The object is closed (unknown fields are rejected) and every field is REQUIRED:

  • allowedOwners / allowedRepositories — exact canonical lowercase ASCII owner scopes (owner) and/or owner/repo selectors. Either list MAY be empty, but at least one selector MUST be reachable for the envelope to admit anything. AWF performs no trimming, case folding, Unicode normalization, or URL decoding before matching; a selector that is not already in this exact canonical form is rejected.
  • sensitivity — one of public, trusted, internal, confidential, sealed; fixes the shared per-repository information budget an admitted repository debits (§14.4). An admitted repository opens its balance in the same run-wide ledger the static executors debit, so re-admitting a repository can never refill a budget it has already spent.
  • executor — fixed to agent; any other value is rejected.
  • githubPolicy — fixed to { version: "github-repository-read-v1", tools: ["list_issues", "issue_read"] }. Any other version, tool set, or additional tool is rejected; this is the sole supported dynamic GitHub policy.
  • maxRepositories — integer 1..1000: distinct repositories this envelope may admit for the run.
  • limits — per-invocation trusted bounds, all REQUIRED: timeoutSeconds (1..4740, whole seconds), memoryLimit, cpuLimit, pidsLimit (1..4096), tmpfsLimit, maxOutputBytes (1..8192), maxTaskBytes (1..65536), maxModelRequests (1..64), maxModelTokens (1..32768). These are the same resource and response controls a static agent entry declares at entry level.
  • quotas — run-wide totals debited across every admission under this envelope, all REQUIRED: maxInvocations (1..10000), maxOutputBytes (1..1048576), maxExecutionSeconds (1..86400).
  • auditLabels — a non-empty, unique array of at most 32 opaque labels matching ^[A-Za-z0-9][A-Za-z0-9_.:-]{0,127}$. Labels are what AWF and mcpg reconcile dynamic state against at shutdown; they are never repository names or credentials.
  • expiresAt — an absolute ISO-8601 timestamp, never later than the workflow job lifetime; admission at or after this time is denied.

Runtime prerequisites

A dynamic entry runs only when the gh-aw compiler has started mcpg's github-repository-delegation-v1 controller and handed AWF both private values:

  • AWF_ENCLAVE_GITHUB_DELEGATION_CONTROL_ENDPOINT — the loopback-only control endpoint, published by Docker on the runner's own 127.0.0.1 (http://127.0.0.1:<port>/internal/awf-enclave-mcp-control/github-repository-delegation-v1). AWF accepts only the literal loopback hosts 127.0.0.1 and [::1]; a resolver-dependent name such as localhost, a routable or wildcard address, embedded credentials, a query string, a fragment, or any other path is rejected. An omitted or explicit :80 is accepted as port 80 (the WHATWG URL parser normalizes both to the same value).
  • AWF_ENCLAVE_GITHUB_DELEGATION_CONTROL_CAPABILITY — the AWF-only 256-bit hex control capability.

A missing, partial, or malformed handoff is a hard failure. AWF never falls back to a static seed catalog, a job-lifetime identity, or a broader policy.

Because the control listener is published on host loopback, only the AWF host process can reach it through the published port, which is why the control client runs there rather than in a container. That is a property of the publication rather than a general routing guarantee: under network isolation the in-container listener binds 0.0.0.0, and a peer co-attached to a Docker network with mcpg addresses the container IP directly without traversing the published port. The control plane is therefore protected by authentication — every request must carry the AWF-only capability, and mcpg rejects anything else with 403 delegation_access_denied.

AWF takes custody of both values before any inherited environment is assembled, stages them into the 0700 enclave private root with exclusive 0600 files, and never mounts either one into the enclave MCP broker, the single-use executor, the model sidecar, the general MCP route, or the delegated data plane. Both variables are also in the primary agent's environment exclusion set.

The broker routes each enclave_run_agent through AWF's host-side admission authority over a 0700 request/response directory that is bind-mounted only into the broker. That channel carries the caller's selector, the exact finite output-schema hash, and the invocation's settlement; it never carries the control endpoint, the control capability, the identity handle, the compiler envelope, mcpg's state path, or its policy generation.

Dynamic-only runs stage nothing

A dynamic-only entry needs no GH_TOKEN/GITHUB_TOKEN, clones no repository, writes no seed catalog (not even an empty one), and mounts neither /awf/seed nor a seed map. A run that also declares a separate static entry keeps the full static staging path unchanged.

Admission semantics (enforced by the AWF registry)

Admission runs in src/enclave/dynamic-registry.ts before any repository content is exposed and before any control call is made:

  • Admission is idempotent by (run, enclave entry, invocation id, canonical repository). A retried request with the same key returns the previously recorded outcome, and joins an in-flight admission rather than reserving capacity a second time. A different repository under an already-bound invocation id is rejected rather than rebinding.
  • maxRepositories and all three quotas are reserved synchronously before any asynchronous lookup, so concurrent admissions can never both observe capacity and both commit. maxOutputBytes and maxExecutionSeconds are reserved at their per-invocation worst case (limits.maxOutputBytes, limits.timeoutSeconds) and replaced by the actual reported usage once the invocation settles. Charges committed after admission stay committed even if the invocation later fails.
  • Every outcome — malformed selector, policy denial, resolution failure, or success — is delayed to the same fixed timing bucket (§14.3's bucket list) plus secret-independent jitter, so elapsed wall-clock time cannot distinguish failure causes. Every failure returns one non-disclosing canonical denial.

Identity lifecycle

For each admitted invocation AWF calls mcpg's control API, authenticated with the capability in Authorization, using bounded request/response bodies, explicit timeouts, and strict JSON validation:

Operation Path
create or confirm /internal/awf-enclave-mcp-control/create-or-confirm
status /internal/awf-enclave-mcp-control/status
reconcile /internal/awf-enclave-mcp-control/reconcile
revoke /internal/awf-enclave-mcp-control/revoke
revoke by labels /internal/awf-enclave-mcp-control/revoke-by-labels

Operation paths are siblings of the controller name in the exported endpoint, not children of it. requested_ttl is a positive whole number of seconds, matching the user-configured limits.timeoutSeconds and mcpg's max_identity_ttl unit; timestamps are RFC 3339.

Every create-or-confirm response is verified before it is trusted: non-empty handle and executor bearer, an exact repository match against the admitted selector, tool_policy equal to github-repository-read-v1, tools exactly list_issues and issue_read, an admitted SHA that matches when one was requested, and an expiry no later than the requested TTL or the invocation deadline. Only the executor bearer leaves the AWF host process; the handle stays in AWF-private state.

At startup AWF calls status, revokes any stale labelled identity, and only then calls the transactional reconcile. New admissions stay blocked until that sequence succeeds. Every terminal path — success, agent failure, schema failure, timeout, cancellation, broker error — settles the reserved quota and revokes the identity, and revoke-by-labels sweeps the run at teardown. An unresolved revocation re-blocks admissions and downgrades the invocation to the canonical error rather than returning a success-shaped result.

Live versus pinned reads

admitted_default_branch_sha is optional in both the ADR and mcpg's contract. AWF has no already-authorized, repository-confined path to resolve a default-branch SHA before the delegated identity exists, and github-repository-read-v1 grants only list_issues and issue_read afterwards. AWF therefore omits the field, never widens a token or tool to obtain one, and audits every repository read as a live read. A read is marked pinned only when the control binding actually carries a resolved SHA.

14.2 MCP-only tool surface

The primary agent reaches private-repository execution only through these MCP tools:

enclave_run_script({
  privateRepo: "owner/repo",
  schema: <finite disclosure schema>,
  script: <bounded UTF-8 Python source>
})

enclave_run_agent({
  privateRepo: "owner/repo",
  schema: <finite disclosure schema>,
  prompt: <bounded UTF-8 task prompt>
})

Both tool schemas are closed (additionalProperties: false). A call can never provide or override images, runtimes, models, engines, profiles, mounts, network settings, credentials, repository catalogs, budgets, timeouts, or any other trusted control.

The primary agent MUST NOT receive a broker socket, wrapper binary, direct server URL, capability token, repository seed, ledger state, or alternate enclave transport.

14.2a Optional tool-call cap (cost control)

rateLimiting.maxNumToolCalls (CLI: --max-num-tool-calls <n>) caps the number of enclave tool calls the primary agent may make in one AWF run. It is OPTIONAL and defaults to unlimited: when omitted, enclave-mcp-server keeps no counter, persists no state, and publishes no advisory text. Setting it without any configured enclave is a startup error.

When configured:

  • Scope — one run-wide budget shared by enclave_run_script and enclave_run_agent, independent of each entry's maxInvocations and of the per-repository disclosure ledger. Dynamic repository admission is performed by the broker inside an enclave_run_agent call and does not consume additional units; tool calls made inside an enclave agent are bounded by maxModelRequests/maxModelTokens instead.
  • Counting — every attempted, well-formed tools/call to a published enclave tool consumes one unit before any other admission decision, including calls later rejected as busy, oversized, invalid, or failed. Malformed JSON-RPC requests that never name a published tool do not count.
  • Advisory — each published tool description, and the initialize result's instructions, state: "You are allowed to make at most N enclave tool calls in this run. After that, the system will deny any further enclave tool calls."
  • Denial — once exhausted, calls never reach an executor and return an in-band isError result whose text is "Max tool call count reached, no more tool calls are allowed. Make a decision based on what you already have in context." The decision depends only on the caller's own call count and trusted configuration, so it discloses no repository information. On the first denied call the broker logs one warning, Max tool call count reached. {"toolName":…,"agentName":<executor kind>,"sessionID":<run id>,"maxToolCalls":N}, so repeated retries cannot flood the logs.
  • Persistence — the count is persisted (mode 0600) in the broker's private control directory keyed by the AWF run id, so a restarted broker for the same run resumes the count rather than resetting it.
  • Exemptions — the enclave server publishes no final-answer or structured-result submission tool; the agent's own result/safe-output tools are served elsewhere and are never counted or denied by this cap.

14.3 Topology, gateway contract, and readiness

enclave-mcp-server joins only the private awf-enclave-mcp-control network. The compiler launches gh-aw-mcpg, labels it for the run, and passes AWF the private gateway endpoint plus a run-unique capability/identity handoff. The server is reachable only through that gateway.

When the agent executor is enabled, each invocation joins only the dedicated internal awf-enclave-agent network. Its mandatory peer is the dedicated enclave API proxy. When GitHub access is configured (agent.tools.github or the deprecated legacy agent.github.cli: issues-read-v1 marker), the only additional peer is compiler-owned shared mcpg. AWF attaches that existing container directly at 172.31.0.40 under the fixed alias awf-enclave-github-mcp; the enclave uses /mcp/github on port 8080. Squid, the primary agent, general proxies, safe outputs, and the enclave MCP server itself remain excluded.

The base compiler handoff from github/gh-aw#50920 and late backend rediscovery from github/gh-aw-mcpg#10784 are present in mcpg v0.4.15, which reports MCP Gateway spec 1.16.0. The base floor remains spec 1.15.0 and a post-v0.4.8 mcpg release. That floor covers static entries only; dynamic repository admission requires mcpg v0.4.18 or newer for the whole-second delegation duration encoding (see §14 and docs/enclaves-architecture.md).

GitHub access additionally requires compiler support for mcpg multi-agent identities and policies, tracked by github/gh-aw#57787. The compiler MUST gate or pin the first supporting AWF release. Older AWF versions reject the closed github/tools.github fields; AWF has no compatibility fallback.

While the backend is still starting, mcpg may return retryable HTTP 503 backend_unavailable. AWF retries initialize with bounded backoff until AWF_ENCLAVE_MCP_READINESS_TIMEOUT_MS expires, then fails closed before the primary agent starts.

The gateway has two separate authorization hops. AWF's upstream Authorization header authenticates mcpg to the AWF-owned enclave server with AWF_ENCLAVE_MCP_CAPABILITY. mcpg then generates the client-facing gateway Authorization header from its gateway agent ID/API key; this is the header present in mcpg's rewritten gateway output consumed by engine config adapters. That downstream credential is distinct from the AWF capability.

Adapters consuming mcpg's rewritten output MUST keep the client-facing Authorization value runtime-only: they must not resolve it while generating configuration or persist the resolved credential under GITHUB_WORKSPACE (or any other agent-readable path). The AWF upstream contract cannot enforce this requirement on the downstream output/converter path.

When GitHub access is configured, the compiler supplies the shared gateway contract (AWF_ENCLAVE_MCP_GATEWAY_CONTAINER, AWF_ENCLAVE_MCP_GATEWAY_ENDPOINT, and AWF_ENCLAVE_MCP_GATEWAY_IDENTITY) plus a distinct AWF_ENCLAVE_GITHUB_MCP_AGENT_ID. The compiler configures that identity in mcpg gateway.agentIds and restricts it with gateway.agentPolicies to the github server, agent.tools.github.allowed (or the legacy marker's fixed list_issues/issue_read pair), and the trusted enclave repository catalog named in agent.tools.github.allowedRepos:

  • list_issues
  • issue_read with method get
  • issue_read with method get_comments

AWF itself only validates the closed agent.tools.github contract — that allowed is a non-empty subset of the two supported tools, that allowedRepos is a non-empty list of exact owner/repository slugs each present in this same agent entry's own repos list (not the run-wide merged catalog shared with the script executor), and that minIntegrity (when set) is one of none, unapproved, approved, or merged — and wires the shared gateway connection. The legacy agent.github.cli: issues-read-v1 marker has no allowedRepos field of its own; its repository scope comes entirely from the compiler-defined trusted catalog baked into that fixed marker, not from anything AWF validates. Repository and integrity enforcement live entirely in the compiler-created, enclave-specific mcpg identity; AWF never broadens or replaces that policy.

AWF stores the enclave identity in a mode-0600 private file, removes it from the host environment, and gives each invocation a read-only private copy. The enclave sends the identity directly as Authorization to http://172.31.0.40:8080/mcp/github. It contains no gh executable and fails preflight if one is present. Before primary-agent work begins, AWF initializes the endpoint and fails closed unless tools/list advertises exactly the configured allowed tools — no more, no fewer.

This mcpg identity lasts for the job rather than one invocation. Its repository policy therefore covers the union of trusted repositories configured for the enclave agent; it is not independently expired, revoked, or narrowed to the repository assigned to a particular invocation. This is an explicit tradeoff of direct shared-mcpg connectivity. AWF's per-invocation process, seed, admission, shared-ledger debit, finite output schema, and timing controls remain in force.

14.4 Shared ledger and disclosure

Script and agent calls debit the same live per-repository balance and share one AWF-owned admission lane. Switching executor kinds never resets or forks the ledger.

Repository sensitivity selects the response schema and per-run disclosure policy:

Sensitivity Per-run budget Response schema
trusted Unmetered Structured schemas, including free-form string nodes
public Unmetered Finite schemas only
internal 64 bits Finite schemas only
confidential 8 bits Finite schemas only
sealed 0 bits No invocation can be admitted

The trusted class is intended only for repositories whose content may be returned to the primary agent without confidentiality accounting. It permits an exact { "type": "string" } schema node at any otherwise valid schema position. Strings remain bounded by maxOutputBytes and the global 8192-byte result ceiling. The schema remains strict and structured: floats, optional fields, extra properties, $ref, recursion, regex schemas, and untagged unions are unsupported. Every other sensitivity rejects a schema containing a free-form string node before launching an enclave.

enclave_run_agent necessarily sends repository-derived content to the configured model provider through the dedicated API proxy. The ledger bounds what the calling agent learns; it does not bound what the provider sees.

14.5 Validation coverage

Legacy bounded smoke and runtime-matrix workflow assets have been removed from the owned surface. Until a unified gh-aw enclave smoke workflow exists, local coverage remains unit-focused:

  • src/services/enclave-mcp-service.test.ts
  • src/services/enclave-agent-service.test.ts
  • src/enclave/script-runner-spec.test.ts
  • src/enclave/agent-runner-spec.test.ts
  • src/enclave/manager.test.ts
  • src/enclave/mcp-server.test.ts
  • src/enclave/agent-mcp-server.test.ts

See Unified Enclave Architecture for the operator-facing summary.

Normative References

  • RFC 2119 — Key words for use in RFCs to Indicate Requirement Levels
  • docs/awf-config.schema.json — Machine-readable JSON Schema for configuration documents (normative)

Runtime JSONL Schemas

AWF emits structured JSONL artifact files at runtime. Most record types have a corresponding JSON Schema in the schemas/ directory; opt-in diagnostic formats are documented inline in this spec instead:

Schema JSONL file Description
schemas/audit.schema.json audit.jsonl L7 HTTP/HTTPS traffic decisions (allowed/denied) from the Squid proxy
schemas/token-usage.schema.json token-usage.jsonl Per-API-call token usage records from the api-proxy sidecar
schemas/otel-span.schema.json otel.jsonl OpenTelemetry span records emitted by the local file exporter
schemas/cli-proxy-access.schema.json access.jsonl (cli-proxy) CLI proxy request audit records
(inline, see §13.2) token-diag.jsonl Model alias resolution steps and diagnostic events (opt-in via apiProxy.logging.debugTokens)
(inline, see §13.6) blocked-request-diag.jsonl Body-shape diagnostics for guard-blocked requests (opt-in via apiProxy.diagnostics.captureBlockedRequests)

Versioning

Schema files do not carry an independent version. The repository release tag serves as the version:

  • The $id field in each schema resolves to a stable release download URL.
  • Each JSONL record includes a _schema wire-format field encoding the record type and AWF version (e.g., "_schema": "audit/v0.26.0").
  • Consumers SHOULD use a prefix match (_schema.startsWith("audit/")) rather than an exact match to handle future versions gracefully.

Published locations

Versioned (release assets):

https://github.com/github/gh-aw-firewall/releases/download/<tag>/awf-config.schema.json
https://github.com/github/gh-aw-firewall/releases/download/<tag>/audit.schema.json
https://github.com/github/gh-aw-firewall/releases/download/<tag>/token-usage.schema.json
https://github.com/github/gh-aw-firewall/releases/download/<tag>/otel-span.schema.json
https://github.com/github/gh-aw-firewall/releases/download/<tag>/cli-proxy-access.schema.json

Latest (main branch):

https://github.com/raw/github/gh-aw-firewall/main/docs/awf-config.schema.json
https://github.com/raw/github/gh-aw-firewall/main/schemas/audit.schema.json
https://github.com/raw/github/gh-aw-firewall/main/schemas/token-usage.schema.json
https://github.com/raw/github/gh-aw-firewall/main/schemas/otel-span.schema.json
https://github.com/raw/github/gh-aw-firewall/main/schemas/cli-proxy-access.schema.json

Informative References