Every entry here is a trade we made on purpose: what it costs, and what it buys. Changing one means re-arguing the trade, not just editing the code. The README states what the stack does; this file states why it does nothing more.
- TLS 1.3 only, one profile, nothing negotiated. Cost: no interop with TLS 1.2-only peers. Gain: no downgrade or agility surface. The client offers exactly one of everything; the server takes it or the handshake fails closed.
- No 0-RTT, no compression, no renegotiation-era features. Cost: none we accept (the IETF IoT profile forbids 0-RTT anyway). Gain: replay and compression-oracle bug classes are structurally absent.
- The MUSTs stay despite minimalism: HelloRetryRequest with transcript restart, KeyUpdate both directions, NewSessionTicket with resumption, RFC 9257 binder discipline. Cost: real complexity. Gain: a conforming client, not a toy that works until it meets a strict server.
- The stack tracks RFC 9846, the 2026 revision of TLS 1.3 (no wire changes; same version number). The audit against its tightened requirements: fresh KeyShare per connection and no legacy version negotiation were already true by design; the sender-side KeyUpdate epoch cap and reading through user_canceled to the close_notify were implemented; the receive-side epoch cap is deliberately not enforced, exactly as the RFC requires of receivers. Section citations follow 9846's numbering.
- Strict parsing. Trailing bytes in an extension, duplicate extension types, and malformed CCS are fatal; streams that make no progress hit hard caps. Cost: no tolerance for sloppy peers. Gain: RFC decode errors actually fail, and a hostile stream cannot pin the client forever.
-
ChaCha20-Poly1305 only as a cipher suite; AES exists for QUIC's public-key packets alone. Cost: the IoT profile's mandatory AES-CCM suite and AES-only servers. Gain: constant time by construction on any core for every secret this tree holds — no lookup table is ever indexed with a key from the TLS key schedule, so there is no timing story to defend. The one AES in the tree protects QUIC Initial packets and checks the Retry tag, where RFC 9001 §5 states the keys are public, and entry 38 and INV-26 state how the build keeps it there. An AES-CCM build flag is the most likely future concession, and it would not reuse that AES. Entries 45, 50, 58 and 68 later admit AES under traffic keys, in a
SUITE=aesgcmbuild that takes AES instructions or an AES peripheral whose timing the build vouches for; the default build is still ChaCha20 alone. -
x25519 in 16-bit limbs (the TweetNaCl scheme). Cost: a scalar multiplication takes about 78 ms on the mips32r2 reference target, or 57 ms in a build that asserts
CH_NATIVE_WIDEMUL(bench/results-insn.csv), and wider limbs would be faster. Gain: a machine-checked overflow lemma and citable prior formal work on the same scheme. Provability over speed; revisit if the workload becomes many short connections. Entry 52 adds wider limbs for 64-bit hosts asX25519=wide, and the 16-bit field stays the default. -
One pinned signature algorithm per build: RSA-PSS by default, P-256 behind
make TRUST=raw-ecdsa, never both. Cost: switching means rebuilding. Gain: no signature-algorithm negotiation surface and a smaller binary. RSA won the default on measurement — its verify is 4x faster and 3.5 kB smaller in flash than P-256's — and RSA is what stock endpoints hold. -
RSA is verify-only: exponent fixed at 65537, moduli of 256 to 384 bytes, the pin is the raw modulus, no PKCS#1 v1.5, deliberately variable time. Cost: exotic keys are unsupported. Gain: a tiny fixed-shape verifier with no ASN.1 in it. DER handling lives in
x509_der.calone, and only CA-mode builds package that file;rsa.cstays ASN.1-free in every build. Variable time is safe because every input to verification is public. -
No Ed25519. Cost: none today — no real server-certificate population uses it, and PSK already covers endpoints we control. Gain: no second hash function (Ed25519 needs SHA-512) and no third pin mode.
TRUST=webpkidoes carry SHA-512, because a public chain's signatures use SHA-384; the gain stays whole for the device modes, and docs/webpki.md says why Ed25519 stays out of the public chain too. -
Rejection sampling reads a fixed 1536-byte XOF budget per polynomial. FIPS 203's SampleNTT reads an unbounded SHAKE128 stream;
mlk_sample_nttstops afterMLK_SAMPLE_GROUPS(512) three-byte groups. Cost: a seed that needed more than 1536 bytes would leave the polynomial's tail holding the caller's values — and no reachable seed does: 704 bytes already put the probability of needing more below 2^-128 (C2SP/CCTV's bound), and the adversarial unluckysample vector, built to need over 575 bytes, passes. Gain: every loop in the ML-KEM module has a static bound, so the CBMC harness proves memory safety with unwinding assertions on rather than assuming an unproven loop bound. -
Hybrid key exchange is a build, not a negotiation.
make KEX=pqoffers X25519MLKEM768 alone; the share carries the ML-KEM-768 bytes first, as RFC 10024 orders them, despite the name. Classic builds offer x25519 alone. No build offers both groups. Cost: a pq build cannot talk to a classic-only server, and a classic build cannot talk to a pq-only server; each pairing fails the handshake closed. Gain: no group negotiation, the same rule the rest of the profile follows (entry 1).Entry 39 narrows this to the device modes.
KEX=pq TRUST=webpkioffers both groups, for the reason that mode already offers several signature schemes and several application protocols.KEX=x25519 TRUST=webpkistill offers x25519 alone and carries no ML-KEM. Entry 53 later changed the webpki half: everyTRUST=webpkibuild carries ML-KEM and offers both groups, andmakerefuses aKEXvalue beside that trust mode. Entry 54 gives every server role both groups and a preference for the hybrid; this entry describes clients. Entry 63 adds secp256r1 to the webpki offer, listed last with no share, and to every server, taken last.Offering both groups and taking whichever the server picks was considered and rejected. It fails where it would matter most: the threat is harvest-now-decrypt-later, so a client that offers both and meets a server without ML-KEM completes a classically protected session, the recording stays decryptable later, and no part of the API reports which exchange ran. Fail-closed answers instead that a completed handshake was post-quantum. This is not a downgrade attack — the transcript hash and the server's signature or binder authenticate the group choice — it is the honest server that has no ML-KEM. The cost is also paid whether or not the hybrid is used: measured with gcc -Os on x86-64, the library sources are 25.0 kB of .text plus .rodata classically and 32.3 kB in the hybrid build, and the session struct 1,056 bytes against 2,328, so a negotiating build carries ML-KEM on every connection including the ones that never run it.
Two fields close, in every build, the one gap the argument above names — that no part of the API reports which exchange ran.
ch_tls.groupreports the NamedGroup the ServerHello's key_share selected,CH_GROUP_X25519orCH_GROUP_X25519MLKEM768, andch_cfg.require_pqfails the handshake when that group is notCH_GROUP_X25519MLKEM768. UnderKEX=pqthe flag asserts a build-time property at run time: the build offers the hybrid alone, so the check reads the field the parser wrote and never the constant the build offered. A classic build refuses the flag atch_connectwithCH_EINVAL. The decision itself does not move for the raw and ca builds: each offers one group.
-
Raw-pin builds hash certificates into the transcript, never parse them. No X.509, no chains, no names, no expiry, no revocation, no trusted clock. Cost: no PKI; the operator provisions a key. Gain: the DER-parser vulnerability class does not exist in those builds. CA pinning was first declined on three objections: it readmits the parser class, needs a trusted clock for validity, and a pinned CA without name checking turns every certificate that CA ever issued into a skeleton key. The CA-mode build (entry 16) later answered each objection on its own terms. The profile removes the clock objection by design: the device reads no validity values, and freshness moves to reissuance policy. It contains the parser objection by proof: a fixed-grammar canonical-DER parser with CBMC memory-safety proofs, a Lean differential oracle, and a fuzz harness. And it accepts the skeleton-key objection and scopes it: the pinned key must belong to a CA dedicated to the fleet, every server certificate carries
extendedKeyUsageexactly serverAuth, and docs/ca.md makes exclusivity of the pinned key the operator's contract. -
Key rotation is a second pin slot. Cost: 16 bytes of config and an out-of-band recovery path for devices that miss both pushes. Gain: rotation without a fleet flag day, inside the trust model we already have. CA builds layer CA indirection on top of the same two slots: the slots hold CA keys, routine server-key rotation becomes reissuance and never touches devices, and the slot pair rotates the CA key itself. See docs/rotation.md.
-
Tickets make both auth modes cheap. Reconnects resume over PSK, so pinned mode pays its signature verification once per ticket lifetime; the recurring cost of any handshake is the key exchange alone — two x25519 operations, plus the ML-KEM keygen and decaps in a
KEX=pqbuild (entry 12). -
CA trust is a build, not a negotiation.
make TRUST=ca-rsapins a CA public key in the pin slots and verifies the server's chain — a server certificate alone, or that plus one intermediate — against it with a profiled parser: canonical DER, the build's one signature algorithm throughout, a fixed extension profile, signatures and shape only. Cost: the parser's flash and stack, one or two extra signature verifies on each full handshake, a receive-buffer floor derived from the certificate cap, and no device-side revocation — a stolen server key keeps authenticating until the CA key rotates, because the device reads no dates. Gain: server keys rotate by reissuance with zero device touches, and one pin covers a fleet of servers. Freshness is issuance policy, not device state: short certificate lifetimes and a monitored reissuance pipeline do the work that expiry checking would. docs/ca.md is the operational contract that makes the small device-side check sufficient. -
Public CAs stay a non-goal for the device modes. Cost: an operator of a device fleet runs a dedicated CA or contracts a dedicated intermediate. Gain: a raw or ca device keeps needing no clock and no name matching — a public CA's trust model requires both, and a public CA's signature algorithms sit outside the device profile besides. A public-CA-fronted server still works from a device through raw-pin mode with a stable server key. docs/ca.md records the argument and the workable arrangements with external CAs.
A host-side client has the clock, the memory and the hostname a public chain needs, and it may have no other way to reach the endpoint it was written for. That is what
TRUST=webpkiis, and entry 36 records it as a separate mode rather than as a change to this one: raw and ca do not move.
-
Zero heap. One caller-allocated session struct plus one caller-provided receive buffer. Cost: the caller sizes memory up front. Gain: no allocator, no out-of-memory paths, and bounds the proofs can state exactly.
-
record_size_limitis the receive buffer's size. Cost: strictness toward peers that ignore RFC 8449 — an oversized record is a protocol error, not a resize. Gain: a peer can never send what the buffer cannot hold. -
Single task, single connection. The reference random generator has global state, so every session in an image draws from one stream. Cost: no multi-session generator isolation. Gain:
ch_rand_bytesstays a clean import that firmware replaces; see docs/entropy.md.Which generator an image uses is declared, never defaulted (#41).
RAND=externleavesch_rand_bytesundefined, so an image that never wired a generator fails to link;RAND=drbgpackagesdrbg.cand exportsch_drbg_seed, which also puts the generator'sCH_ASSERT(g_seeded)into a shipped object for the first time. Naming neither reaches an#errorincfg.hrather than one in the Makefile, because a firmware tree compiles these sources with its own build system and a Makefile-only check would miss exactly the integrator this targets. Cost: every consumer build writes one more line, and the break reaches every existing consumer at once. Gain: "I supply my own generator" is a statement somebody made rather than a step nobody took. Rejected: a weakch_rand_bytesdefault, which would convert the link error into a build that succeeds — wolfSSL'sUSE_TEST_GENSEEDshape, where a skipped decision looks like a made one.The hook returns
void, and giving it ach_errreturn was considered and deferred (#41). A return code is the honest answer for a transient hardware fault, but it breaks the one function every consumer firmware tree implements, for a caserand.halready covers by contract: a generator that cannot produce bytes blocks or faults rather than returning short. What thevoidreturn does not catch is a hook that returns without writing, and that is why every draw inhandshake.cis followed by an all-zero check againstrand.h's contract. Neither the check nor any signature can tell a weak generator from a strong one; nothing in a library can. -
Every operational error fails closed: alert, wipe keys, dead session, caller reconnects. Cost: no graceful recovery. Gain: the entire resumable-error state space is removed from the code and the proofs. Corollary:
CH_EINVAL(invalid configuration, nothing sent) is distinct fromCH_ECAP(runtime capacity), so provisioning corruption never reads as an attack on the wire. -
One TX staging array, sized per build. The ClientHello builder and the sealed-record path share the session's TX array; their lifetimes never overlap.
CH_TX_STAGEis whichever is larger, per build, and the hello wins in both: 617 bytes classic, 1801 with a hybrid key share, against the 529 a sealed record needs. The classic figure used to be the sealed record's, which left a maximum ticket identity plus a maximum retry cookie failing closed withCH_ECAPmid-handshake (#46); covering the hello costs that build 88 bytes and lets the same compile-time assert run everywhere. A second array would cost every build the hello's bytes, and the handshake proof keeps one array to model. Streaming the hello stays rejected: the PSK binder is an HMAC over the contiguous truncated hello, so a streaming builder would buffer the message anyway. Entry 71 lets a TCP build raise the sealed record's plaintext withTX_RECORD, and where the record then outgrows the hello, the record wins. -
The receive-buffer floor is a build constant.
ch_connectchecksbuf_lenagainstCH_MIN_RXBUFbefore anything is sent. A feature that needs more room raises the constant, so a too-small buffer fails at setup withCH_EINVAL, not mid-handshake withCH_ECAP, where it would read like an attack. The floor covers only what the build can know: a raw-pin server's chain is the server's choice, so raw-pin deployments size above the floor. A CA build does know its worst case — the certificate cap bounds the chain — so it derives the floor from the cap instead of picking a number:2*(CH_X509_MAX+5)+16. -
The ML-KEM key pair lives as a 64-byte seed. The handshake keeps the (d, z) seed in
handshake_state, which lives onch_handshake's frame for the length of the handshake, instead of the expanded key pair.mlkem_keygen_dkre-expands it into a stack buffer at ClientHello build and again at decapsulation: two keygen runs per handshake, three when a HelloRetryRequest makes the client rebuild its hello.mlkem_keygen_dkwrites the encapsulation key in place at dk + 1152, FIPS 203's own dk layout, so one dk-sized buffer serves both sites. The expansion is deterministic, so that retry sends the identical share without storing it. Cost: one keygen run more than a stored key pair would need, two more on a retry. Gain: 64 bytes of session state instead of the 2,400 a stored dk costs, and the dk lives only for the length of one call rather than for the length of the session.
- Bounded model checking, layered, with a published ledger. Leaf modules prove concrete; upper layers prove against contract-checking stubs of the proven layer below. Where a formula will not converge, the harness pins a representative bound and documents it. Cost: this is not functional verification. Gain: proofs that finish, and claims nobody has to take on faith — docs/verification.md states what is proved, at what bound, and what is only tested.
- The Lean spec is written from the RFCs, never from the C, is partial exactly where the RFCs are partial, and carries theorems about itself. Cost: everything is implemented twice. Gain: a shared misreading of an RFC cannot make both sides agree, and each C-versus-spec agreement transfers a proven property, not just a matching answer.
- The RFC 8448 replay stops at secrets and MACs. The traces protect records with AES-128-GCM, which this stack excludes, and sign with an RSA-1024 key, below the verifier's floor — so the tests check the floor holds, then verify the trace's CertificateVerify one layer down through the raw modexp. Cost: the replay never opens a record. Gain: third-party byte-exact checks of the transcript and key schedule without weakening the profile.
-
Four exported symbols. The library packages as one relocatable object; partial linking plus symbol localization does the namespacing, so sources keep natural names and applications cannot collide with internals. The four are the calls of the first build. Each axis now sets its own list of calls (entries 38, 41 and 43), and entry 56 adds one data symbol,
ch_build, to every object. -
CI compiles with gcc on purpose while development machines run clang: consumers are firmware trees whose vendor SDKs ship gcc cross-compilers, so gcc-only diagnostics belong in CI. Between the two, both major compiler families stay covered without a second CI leg.
-
Tool versions pin to the development machine's. When the local toolchain upgrades, the CI pins bump in the same commit. Code never adapts to an older checker.
-
Two proof solvers, each where its memory profile fits. kissat runs the fast tier (measured: verdicts in seconds to minutes where the built-in solver ran for hours); CI's slow tier keeps the built-in solver, because the external-solver path materializes the whole formula and exhausts a 16 GB runner. Verdicts are solver-independent. A content-keyed cache re-proves only what changed.
-
Third-party audit is the optimization target. Every review-facing trade — spelled-out names, the complexity-15 gate, pure predicates with state changes on their own lines, pinned byte constants instead of decode-and-judge — pays a little compactness for a lot of reviewability. Cost: more lines and more named helpers than the terse form. Gain: a security library earns trust through reviewers who did not write it, and every clever compression taxes each of them.
-
Assertions live at proof time; runtime keeps contract-point guards. TigerStyle asserts the negative space at runtime, two per function, on in production. chapulin moves that space into the CBMC layer — 193 proof assertions checked over every input at the bound — because on an unattended device an abort on hostile input is the denial of service. Runtime keeps only contract-point guards that no input can trigger (state-enum validity at the API entries), which also cover corrupted memory. Cost: hardware faults mid-connection surface as failed handshakes, not named aborts. Gain: exhaustive checking where inputs are hostile, and no abort path an attacker can reach.
-
Revocation travels in the server certificate's notBefore, on a restricted set of dates. A clockless device cannot check expiry or fetch a CRL, so reissuance alone never revokes a stolen server key. The epoch turns notBefore into a counter the CA already signs: dates restricted to UTCTime, YY 00..49, DD 01..28, midnight, compared as the exact index
YY*336 + (MM-1)*28 + (DD-1). Alternatives declined: a new certificate extension (every CA would have to learn it, and the device would parse more), a serial-number counter (issuance tools own serials), and GeneralizedTime dates from 2000 on (RFC 5280 §4.1.2.5 mandates UTCTime through 2049, so those certificates are unissuable). Cost: the CA must write an absolute notBefore, which Vault, step-ca, and AD CS will not do; advancing the epoch requires reissuing every server before any device sees it; and a device isolated from the fleet never learns of a it. Gain: a device-side revocation check with no clock, no CRL, no OCSP, and four bytes of device state. The epoch revokes certificates; revoking a stolen key means reissuing that server on a fresh key pair and advancing the epoch, so the old certificate falls below the stored epoch. Advancing is what makes key rotation stick. -
The stored epoch moves only after the server authenticates. A CA-signed certificate is public, so presenting one proves nothing about the presenter: an attacker can replay a genuine higher-epoch certificate harvested from any real server. Rejecting on the chain verdict alone is safe — it only fails the handshake closed — but raising the stored epoch outlives the session, so it waits until CertificateVerify and Finished have proven a real server is there. Cost: the rule splits across two call sites instead of one. Gain: an unauthenticated peer cannot move device state that outlives the session.
-
TRUST=webpkiis a third trust mode, not a change to the other two. It verifies a public chain against caller-supplied anchors, with hostnames and validity dates, so a host-side client can reach a public endpoint. Cost: a clock the caller supplies, a receive buffer measured in kilobytes, every signature family a public chain uses in one object, and a mode that refused PSK and resumption because nothing bound a ticket to a hostname, until entry 47 bound one. Gain: the raw and ca objects do not change — their sources, their defines and their SRAM rows stay where they are, andmake lint-trust-separationholds the partition. docs/webpki.md states the profile, the measured bounds, and what the mode does not check.Extending the CA modes with optional dates and names was considered and rejected: it would put clock and name logic inside the object a device links. Verifying the chain in the caller instead of here was considered and rejected: it duplicates a certificate parser in a second language, outside this tree's proofs.
-
A
TRUST=webpkicaller offers a list of application protocols and the server picks one. ALPN (RFC 7301) is the mode's second exception to the rule that the client offers exactly one of everything, after the signature schemes. Cost: a negotiation surface. The ClientHello carries up toCH_ALPN_MAXnames, the server chooses among them, and the outcome differs per connection, so the caller readsch_tls.alpn_selectedand branches on it — including onCH_ALPN_NONE, which says the server selected no protocol. Gain: one handshake instead of a failed one and a reconnect. An HTTP client that could offer one name would have to guessh2, and a server that speakshttp/1.1would cost it a second full handshake.Offering one protocol per build, the way
PINandKEXfix one algorithm, was considered and rejected: a build cannot know what a given endpoint speaks, and the fallback costs a whole connection. Failing the handshake when the server sends no ALPN extension was considered and rejected too: RFC 7301 §3.2 lets a server that does not implement ALPN leave it out, so refusing there would refuse every server that speakshttp/1.1by convention. docs/webpki.md states what the caller branches on and what the client refuses. -
QUIC is a transport axis, and chapulin owns packet protection on it, AES included. RFC 9001 §4.1.3 removes the record layer: QUIC carries bare handshake messages in CRYPTO frames and protects packets itself. A
TRANSPORT=quic-nonblockingbuild takes handshake bytes in, hands handshake bytes out, and seals and opens every packet at every level, so no traffic secret leaves the object. colibri, the HTTP/3 caller, owns everything that is not cryptography: packet numbers, ACKs, loss recovery, congestion control, flow control, streams, connection IDs, Retry and version negotiation logic, and path validation. Cost, and most of it lands here: a fifth value inLIB_VARIANTbeside PIN, TRUST, KEX and RAND; fifteen exported transport calls against entry 28's four, on aPUBLICwhose first term the axis selects —$(PUBLIC_TRANSPORT) $(PUBLIC_RAND) $(PUBLIC_CA), wherePUBLIC_TRANSPORTdrops the four TLS names rather than adding to them and the other two terms keep their meaning, so a CA-mode or RAND=drbg QUIC object exports sixteen calls and one with both exports seventeen; five library files replaced —record.c,io.c,session.c,handshake.candtls.c, 990 lines — and five more given a second arm under#ifdef; entry 19'srecord_size_limitdropped, which leavescfg.buf_lenas the only cap on what a peer can send; the KeyUpdate entry 3 keeps turned into a connection error (RFC 9001 §6); two exceptions to entry 21 inside this tree, because RFC 9001 §5.5 discards a packet it cannot authenticate instead of closing and because a caller-order mistake leaves the session live; entry 37's ALPN lifted out ofTRUST=webpkiinto every trust mode, because RFC 9001 §8.1 requires it; and AES-128-GCM and AES-128-ECB entering the codebase against entry 6. The largest piece is none of those. The driver blocks on the wire at fivehsr_next_msgcall sites, two of them inhandshake_auth.c, so the caller-driven interface runs one step per whole handshake message, moves the handshake frame — 448 bytes raw, 856 under TRUST=ca-rsa, 976 under TRUST=webpki, 512 under KEX=pq — into the session, moves 233 ofhandshake.c's 395 lines into ahandshake_flight.cboth transports compile, and changeshandshake.oin every tcp-blocking build.handshake_pskandhandshake_pin, two ofproof/run.sh's 83 launch lines, are re-measured because that file moves under them. Gain: no second crypto stack. Every traffic secret a QUIC connection uses is derived, held, used and wiped inside one object this tree proves, and RFC 9001 §9.5's requirement that header protection removal, packet number recovery and packet protection removal happen together without timing side channels is met inside this tree's constant-time rules, measured by lint-wide-multiply on the 1-RTT path and argued from public keys on the Initial one, instead of re-argued in a second repository. Measured at3432a5d: the driver touches the record layer at 15 call sites in three files and the key schedule at none, and unchanged lines are 1,335 of 3,460 in a raw-mode object and 4,392 of 6,517 in a TRUST=webpki one.Entry 6 does not fall; it gains one exception with a checkable boundary. The AES ban exists to keep a secret key out of table-driven code, and every key AES touches in QUIC is public: the Initial keys come from the client's Destination Connection ID and a salt RFC 9001 §5.2 prints, the Retry key and nonce are printed in §5.8, and §5 states that neither packet type is considered to have confidentiality or integrity protection. So AES may exist only where the key is public — Initial packet protection (§5.2), Initial header protection (§5.4.3) and the Retry integrity tag (§5.8) — and never under a key from the TLS key schedule. A key type,
aes_public_key, that onlyaes.c,quic_initial.candquic_retry.ccan build is the first guard, and the compiler runs it:quic.hstores no key, only the Destination Connection ID the keys come from, so the type is incomplete everywhere but the three sources that includeaes_public_key.h, and a fourth file that declares one gets an error. The calls are the part a rule reads, so the new invariant is Semgrep-tripwire there, the grade this tree gives an identifier ban: it permits exactly two callers and exactly three key sources, and no other source may call a symbol whose name beginsaes_orgcm_. A reintroduction under another name is what a tripwire does not catch, and the reviewer reading the diff is what does. A Semgrep rule beside the thirteen in.semgrep/invariants.ymlfails the build on a third caller,lint-codegen-partitionholds both files inWIDEMUL_PUBLICwhere a maintainer has to move them in plain view to admit a secret,lib-checkkeeps their symbols out ofPUBLIC, and a.violationfile proves the first gate works. TheCLAUDE.mdsentence and this file's entry 6 change in the commit that lands the first AES source; docs/quic.md carries both replacement texts. Be suspicious of this: the safety sits in the call graph, not in the code, and a constrained primitive tends to grow callers.Handing per-level secrets out and letting the caller protect packets was considered and rejected. A 32-byte secret is not a packet protection layer: the caller would still write HKDF-Expand-Label, HMAC-SHA-256, ChaCha20-Poly1305 and the §5.4.4 mask before it needed the AES, so that split moves seven primitives outside this tree's proofs to keep two out, and five of the seven touch real traffic secrets — the cost entry 36 refused to pay for the certificate parser. One secret leaves the object today, the resumption PSK in
ch_ticket.psk, throughon_ticket(handshake_post.c:53-54); it is a key for a future connection, not a live traffic secret, and a per-level traffic secret would be the first live key to leave. A second TLS stack for HTTP/3 was considered and rejected too: it gives one product two trust models and two failure disciplines, and entry 12's fail-closed post-quantum property would hold over TCP and not over QUIC. Entry 12 stands in a QUIC build: one key-exchange group per build, on the same PIN, TRUST, KEX and RAND axes.The driver landed in
33978f6, which moved the client's flight handlers intohandshake_flight.cand putquic_step.con top of them. docs/quic.md states the interface, the suspendable driver step by step with the state each step leaves behind, the measured reuse per build, the bounds that still need measuring, and the verification owed. -
A
TRUST=webpkicaller offers both key exchange groups and the server picks one. This is the mode's third exception to the rule that the client offers exactly one of everything, after the signature schemes (entry 36) and the application protocols (entry 37), and it has the same cause: a host-side client cannot know what the endpoint it dialled supports. Entry 12's one-group-per-build rule stands unchanged for the raw and CA modes, where the device already pins the key of the endpoint it will talk to and therefore knows what that endpoint speaks.The ClientHello lists X25519MLKEM768 and x25519 in
supported_groupsand carries the X25519MLKEM768key_share. A server that wants x25519 answers with a HelloRetryRequest naming it, and the second hello carries an x25519 share over the x25519 half of the key pair the hybrid share already held. This entry first said the client already handled that retry; it did not. The retry path it had took a cookie and refused every retry that named a group, and the change that built this offer built that path with it (CH_KEX_TWO_GROUPS,hsf_read_server_hello). This applies underKEX=pq:KEX=x25519 TRUST=webpkihas no hybrid to offer and lists x25519 alone. Cost: one extra round trip against a classic-only server. Gain: no round trip on the post-quantum path, which is the path worth making fast, and one ML-KEM key generation per handshake rather than one for every hello whether or not it is used.Carrying both shares was considered and rejected. It never costs a round trip, but it puts an ML-KEM-768 share — 1216 octets — in every hello a webpki client sends, generates a key pair that most handshakes discard, and grows
CH_HELLO_MAXfor a build that already carries the largest hello in the tree. HelloRetryRequest is an RFC 9846 MUST this client implements and proves, so the fallback costs a round trip on a path that is becoming rare rather than bytes on every path.The consequence worth stating plainly: entry 12's fail-closed property does not survive negotiation. A webpki client that offers both groups will complete a handshake with a classic-only server rather than refusing one, which is the point of offering both. A caller that wants the old guarantee sets
ch_cfg.require_pq, which already refuses a handshake whose selected group is not the hybrid, so the property becomes the caller's to ask for rather than the build's to enforce. The flag also drops x25519 from the hello, so that caller sends the one-group hello a rawKEX=pqbuild sends: a server without the hybrid finds no common group and fails the handshake, rather than asking for x25519 and being refused one round trip later.Entry 53 replaces the key shares this entry describes, and the paragraph above that rejected carrying both. Every webpki build now sends a share for each group, so a server without the hybrid selects x25519 in one round trip, and the HelloRetryRequest path described here is gone. The rest stands: the mode offers two groups, and
ch_cfg.require_pqgives the fail-closed property back. -
The pinned algorithm is half of a
TRUSTvalue, not an axis.PINchose RSA-PSS or P-256 for the key a raw or ca build pins. It selected nothing in the other two builds: aTRUST=webpkiobject carries every verifier because a public chain's links are signed by different algorithm families, and aROLE=serverobject carries both becausech_srv_checkverifies both provisioned identities at boot. An axis that names nothing in two of the builds that read it is a suffix.TRUSTnow spells it:raw-rsa(the default),raw-ecdsa,ca-rsa,ca-ecdsaandwebpki.What this buys is not one fewer flag. It is that
make TRUST=webpki PIN=ecdsaasked for one verifier and got every one, and the Makefile answered by quietly emptyingPIN_FILTERbehind the caller's back; that build can no longer be written. TheROLE=serverblock carried two refusals for the same reason and now carries one, andTRUST=ca-ecdsagained alint-trust-separationrow, which it never had while it was a combination rather than a value.A server names
TRUST=none, a sixth value, and is refused without it. Letting it take the client default described the object with a value naming one algorithm where it holds both, and left the trust value unwritten at every server call site:make check TRUST=raw-ecdsadied because one recursion had not named one and inherited it.TRUST=noneis refused for a client, whose whole job is to judge a peer certificate, so neither role can build under the other's value. The object it produces is the same 34 sources as before.The C stays as it was:
CH_PIN_ECDSAandCH_TRUST_CAare unchanged, so firmware that compiles the sources directly sees nothing move. A stalePIN=on a build line is refused by name rather than ignored, because ignoring it would hand back an object built around the other verifier.TRUST=rawandTRUST=caare refused the same way, each naming the two values that replaced it. -
ROLE=bothis a host-side value, and one-role-per-object stays for devices.ROLE=clientandROLE=servereach carry one role because a device carries one for the life of the deployment, so the flash the other costs buys it nothing. That is a firmware argument, and it does not reach a host library: colibri serves HTTP/2 and HTTP/3 and also fetches over them, and stompy will do both in one process.Two objects are not a substitute, which is the fact that decided this. Each carries the shared half, so linking a client object and a server object into one program makes
ldreportch_read,ch_write,ch_closeandch_drbg_seeddefined twice — measured, four duplicate symbols. colibri avoids it today only by attaching one object per module, so the two never meet in one binary.The combined object needs no dispatch and no second name.
srv.halready states why:ch_read,ch_writeandch_closeare the same functions over the samech_tls, "because record.[ch] names no side". So the roles differ in one call each way, andROLE=bothexports seven calls where the halves export five and six. It is also smaller than what it replaces: 132,960 bytes against 183,980 for the two tcp-blocking objects, and 148,520 against 213,804 for the two QUIC ones.TRUST=noneis refused here, because the client half judges a peer. -
A tcp-nonblocking server pushes its flight; only the client pulls.
TRANSPORT=tcp-nonblockingexists because a blocking callback cannot sit under a completion-based event loop: colibri drives rotor, whose loop is single-threaded with no fibers, so acfg.recvthat waits stalls every connection the loop holds.ch_srv_acceptblocks by contract (cfg.h:371), which is whyROLE=server TRANSPORT=tcp-nonblockingwas refused until the driver existed.The client's shape does not carry over.
ch_record_outhands a staged record to the caller, and that works because a client's messages fitch_tls.tx. A server's do not: one Certificate message is larger thanCH_TX_STAGE, andsrv_flight.cstages a protected message on the handler's own stack frame and streams the chain throughsrv_frag. There is nothing to pull from. A pull would need a resume point insidesrv_out_sealed's record loop, andtcp_nonblocking_step.hrules that out: a step runs on a whole message and waits nowhere inside it.srv_quic.hreached the same place for the same reason, soch_srv_cfg.on_record_outison_crypto_outwithout the level.That still solves the problem: the callback copies each record into a buffer the caller owns and returns, so nothing waits on a socket. INV-28 states the claim and
bin/srv_tcp_nonblocking_testmeasures it with asendand arecvthat fail the run if the driver calls them. -
The exporter is a build axis, and it widens one cap rather than adding a second serializer.
EXPORTER=oncompilesch_export(RFC 9846 §7.5) and addsexp_mastertoch_tls;EXPORTER=off, the default, compiles neither.ch_tlsmeasures 1144 bytes off and 1176 on, so a device that exports nothing pays nothing and docs/performance.md's SRAM figures are the default build's, unchanged. colibri asked for the call for h2 (docs/chapulin.mdin that tree), and a host is the only caller.The label is the caller's, and RFC 9266's is 24 bytes against the 12 TLS 1.3 itself writes, so the axis sets
HKDF_LABEL_MAXto 32.hkdf.hmakes that cap a build parameter with a floor of 12 instead of growing a secondhkdf_expand_labelfor long labels: the only thing the cap sizes is one stack buffer, and two serializers of oneHkdfLabelwould be two places for its layout to drift.HKDF_INFO_MAXis derived from the cap for the same reason; it was a literal 64 while the cap was fixed, and the default build's buffer shrinks by ten bytes as a result.tls.casserts that the publicCH_EXPORT_LABEL_MAXand hkdf's cap are one number.ch_exportrefuses rather than asserts. A label is data a caller may compute, so an over-long one is an operational error and returnsCH_EINVAL, andhkdf_expand_label'sCH_ASSERTon the length, whichks_exporterreaches, is then unreachable from the public call. It refuses every state butCH_ST_CONNECTED, because the secret does not exist until the peer's Finished verifies and a closed session has wiped it with the rest.EXPORTER=onwithTRANSPORT=quic-nonblockingis refused by name, in the Makefile and again inkeysched.hfor a tree with its own build system. The call sits intls.c, whichQUIC_REPLACEDdrops, so that object would listch_exportand never define it — which is whatlib-checkcaught when the pair was first tried. RFC 9001 keys QUIC from the handshake secrets and uses no TLS exporter, and the h2 caller runs over records, so a QUIC exporter is a separate change with an entry of its own inquic.hif anyone asks for one.No published vector exists: RFC 9846 prints none and RFC 8448's trace stops short of it.
bin/exporter_test's four vectors were produced by this code and confirmed byte for byte against an implementation written from §7.5's text in Python, which catches a misreading of the spec and not a shared one. docs/verification.md says cross-checked, not published. -
The key log is an axis, a link-time hook, and refused for a device client. colibri's interop endpoint must write an NSS key log in both roles and over QUIC (its design §9), and colibri holds no secret to write, so chapulin hands each traffic secret out as it derives it.
KEYLOG=oncompiles fourch_keylogcalls at the two placesks_handshakeandks_masterrun in each role;KEYLOG=off, the default, compiles none.It is a hook the image defines, the way
ch_rand_bytesandch_aes_blockare, rather than ach_cfgfield. AKEYLOG=onobject importsch_keylog, so an image that turned the axis on and wired nothing fails to link rather than logging into nowhere; ach_cfgfield would have needed a function pointer in every session and two linescfg.h, at its 500-line cap, does not have. The hook getscfg.io, so one hook tells connections apart.The refusal covers a client in a raw or ca trust mode, which is what a pinned firmware image is, and admits
TRUST=webpki, the server'sTRUST=noneandROLE=both. The first draft admitted only webpki andROLE=both; that would have refused colibri's own server, which buildsROLE=server TRUST=none. A firmware server can therefore carry the axis. The refusal is a guard against building it by accident, not a security boundary: anyone compiling these sources can pass the define.Four labels and no more: the two handshake and the two
_0application secrets. The format has no label for a KeyUpdate's next generation, which a reader derives itself, and none of this build's handshakes has early data.EXPORTER_SECRETis left out because no reader in colibri's matrix asks for it.The first build logged 32 zero bytes as the client's random. The client wipes
h->randomwhen the key exchange finishes, before the handshake secrets it logs, and the secrets still matched across the two ends, so a test comparing secrets alone would have passed.bin/tcp_nonblocking_loop_testcompares the random too, and INV-29 records the rule with a mutant that restores the bug. -
A
SUITE=aesgcm TRUST=webpkiclient offers both cipher suites and the server picks one. This is the mode's fourth exception to the rule that the client offers exactly one of everything, after the signature schemes (entry 36), the application protocols (37) and the key exchange groups (39), and it has their cause: a host-side client cannot know which suites the endpoint it dialled accepts, and RFC 9846 §9.1 makesTLS_AES_128_GCM_SHA256the one a conformant server must implement. The ClientHello listsTLS_CHACHA20_POLY1305_SHA256first andTLS_AES_128_GCM_SHA256after it, the ordersrv_selectprefers for the reason it states: ChaCha20 is constant time by construction, and AES is constant time because the build asserted it (CH_NATIVE_AES). The client keys every record direction with the suite the ServerHello selected, a ServerHello after a retry must repeat the retry's suite, andch_tls.suitereports the one that ran. Entry 80 puts the AES-GCM suites first in a build onAES=hwwithCH_NATIVE_AES, and lets a caller name the client's order.Cost: a negotiation surface, the AES sources in the object, and the build's statement about its hardware.
ct.hrefuses the suite withoutAES=hwandCH_NATIVE_AES, so the offer exists only on a host whose AES instructions the builder vouches for, andquic_packet.crefused it over QUIC, where packet protection ran ChaCha20 alone, until entry 58. Gain: the client completes a handshake with a server that accepts AES-128-GCM alone, which the e2e suite checks against OpenSSL.A raw or ca client refuses
SUITE=aesgcm, the way it refusesKEYLOG=on(entry 44). It pins the endpoint it talks to, so it knows that endpoint's suite, and it offers ChaCha20 alone; the Makefile andhandshake_message.cstop the define there. A build with a server role takes it whatever its trust mode, because its server selects AES from a client that offers nothing else, and the client beside it in a raw or caROLE=bothbuild still offers ChaCha20 alone.Offering AES-128-GCM alone under
SUITE=aesgcmwas considered and rejected: the build would then fail against every server that accepts ChaCha20 and not AES, and the reason to offer AES is to reach more servers, not different ones. -
A tcp-nonblocking
ch_readreturnsCH_RECORD_AGAINwhen no record has arrived, and the session stays connected. Entry 21 makes every operational error fatal, and an emptyrecvinTRANSPORT=tcp-nonblockingis not an error. The caller owns the socket and hands over whole records as they arrive, so between records it has nothing to hand over. Before this result existed,ch_readturned that emptyrecvintoCH_EIO. A caller that received a record with no application data, such as a NewSessionTicket, had to hold it back until a data record arrived, or lose the session.TRANSPORT=tcp-blockingdoes not change: itsrecvblocks, and a 0 there is the end of the stream.Cost: a second live result from
ch_read, and one field that lives across calls,ch_tls.post_fill, which counts the bytes of a post-handshake message split across records. INV-13 states the terms. Gain: the caller passes each record toch_readas it arrives and keeps no queue of its own.CH_QUIET_CAPstill bounds the records onech_readcall handles without application data. A peer that sends such records without end now costs the caller one call per record, so the loop that repeats is the caller's, andch_readdoes not spin.Treating a 0 inside a record as
CH_RECORD_AGAINtoo was considered and rejected. The partial header would have to be kept across calls, andtcp_nonblocking.halready asks the caller for whole records. -
A
TRUST=webpkiclient resumes a ticket bound to the hostname and anchors that received it. A resumed handshake checks no certificate, so entry 36 refused resumption in this mode. RFC 8310 §9 makes resumption a MUST for a DNS-over-TLS client, and RFC 9846 §4.7.1 lets a client resume only under aserver_namevalid for the original certificate. Each ticket now carries a binding: HMAC-SHA256 keyed by its PSK over a hash of the lowercased hostname and the anchor array.ch_connectrecomputes it and refuses a mismatch withCH_EINVALbefore it sends a byte. docs/webpki.md, "Resumption", states the rules.Cost: 32 bytes in
ch_ticket, 40 inch_tlsasbench/sram.shmeasures it (a 32-byte field and its alignment), one field inch_cfg, and a second path through the handshake for this mode, the one the raw and ca modes already take after a ticket. Gain: a reconnect to a public endpoint skips the chain walk and its signature checks, and a ticket stored under the wrong name fails at configuration, not after a session with a server the caller did not name.Keying the binding by the PSK, not hashing the configuration alone, ties it to one ticket: a caller cannot pair one ticket's PSK with another's binding by mistake. Storing the hostname in the ticket and comparing it was considered and rejected: the caller supplies both sides of that comparison, so it checks nothing a storage mistake would break.
Offering the certificate path beside the ticket, so a server that declines the ticket can still authenticate by chain, was considered and rejected for now. It would put
signature_algorithmsin the resumed hello and give the client two ways through one handshake, where every mode here has one per hello. The cost of failing closed is one reconnect after a declined ticket. Entry 55 reversed this: dns.google resumes no hello that lackssignature_algorithms, and declines about one ticket in three. -
A QUIC server's Retry token is an HMAC chapulin computes under a key the caller holds, bound to the client's address, with the caller's clock. colibri runs the QUIC Interop Runner's
retrycase as a server over oneROLE=bothobject and holds no key, soquic_token.[ch]mints and checks the token.docs/quic_server.md, "The Retry token", states the format and what the caller still owns. Cost: two exported calls, so aROLE=server TRANSPORT=quic-nonblockingobject exports eighteen calls, and one HMAC-SHA-256 per mint and per check. Gain: a stateless server gets both connection IDs back for its transport parameters, and a key stays on chapulin's side of the linedocs/quic_server.mddraws.The token is authenticated and not encrypted. RFC 9000 §8.1.4 asks integrity of a Retry token and nothing more, and its fields are ones the path saw in the clear. Sealing it with the ChaCha20-Poly1305 the object already carries would need a fresh nonce per token, and a nonce needs randomness or a stored counter: the first breaks the seeded replay colibri needs, and the second breaks the statelessness a Retry exists for. The §8.1.4 alternative of a random value the server remembers breaks the same two things.
The instant is the caller's, in seconds, because chapulin reads no clock and
ch_cfg.now_secondsalready counts seconds. The check tells the caller which RFC answer applies:CH_EPROTOfor a token that is not a Retry token, which §8.1.3 treats as no token, andCH_EAUTHfor a Retry token that fails, which §8.1.2 answers with INVALID_TOKEN. The first byte decides, because §8.1.1 requires the two kinds to be told apart. A check that accepted each token once was considered and left to the caller: it needs state, and the window and the address binding already limit replay as §8.1.4 requires. -
A
TRUST=webpkiclient takes SPKI pins, and with them RFC 7250 raw public keys. RFC 8310 §9 makes RFC 7250 a MUST for a DNS-over-TLS client and lets it offer raw keys only with an SPKI pin set, so the host-side mode that caller builds takes pins. Pins alone are a whole configuration, the "SPKI + IP" profile for a server with no public certificate. With anchors too, a chain must pass the walk, the clock and the hostname, and a pin must name a key on the path the walk verified, as RFC 8310 §6.4 and RFC 7858 §4.2 ask. docs/webpki.md, "Raw public keys and SPKI pins", states the rules. Entry 65 widens pins alone to a certificate chain whose leaf key a pin names.Cost: a fifth thing the mode offers more than one of, the certificate types, and a second way for a Certificate message to authenticate a server, which is why the device modes stay without it. Gain: the caller reaches a pinned server whether it presents a raw key or a chain, and a pin change is a configuration change, not a CA.
A new trust mode for pins alone was considered and rejected: a configuration with both a name and pins needs the chain code anyway, and a second host-side mode would split the tickets, the ALPN offer and the tests between two objects. Matching a pin on the leaf alone was considered and rejected: RFC 7858 pins the validated chain, and an operator who pins an intermediate would be locked out on the next leaf rotation.
-
CH_NATIVE_AEScovers the carry-less multiply as well as the AES instructions. UnderAES=hw, GHASH multiplies on PMULL or PCLMULQDQ inghash_hw.c, becausegcm.c's portable multiply was 98% of anAES=hwseal and the instruction runs GHASH 60 to 65 times faster from 1200 bytes up (docs/quic.md, "What the AES axis costs in time, measured"). AES-GCM needs both instructions under one key: the AES rounds produce the keystream and the hash subkey, and the carry-less multiply multiplies by that subkey. SoCH_NATIVE_AES, the build's statement that this part's AES instructions run in constant time, now also states that its carry-less multiply does.ct.hwrites the terms, and aSUITE=aesgcmbuild still names one define. A build whose keys are the public QUIC Initial and Retry keys needs no statement, as before (INV-26).Cost: one define now asserts two things, so the vendor statement behind it has to cover both instructions. A part whose AES rounds are constant time and whose carry-less multiply is not cannot carry
SUITE=aesgcmhonestly, and nothing here detects that part. Gain: one statement per AEAD, written once in the build files by someone who can answer for the part. On Arm the two are one feature already: the Arm C Language Extensions put the 64-bit PMULL in the AES extension, and__ARM_FEATURE_AESnames both.A second macro,
CH_NATIVE_CLMUL, was considered and rejected. No build here runs one instruction without the other:AES=hwcompilesaes_hw.candghash_hw.ctogether,AES=runtimecompiles both and runs both on one answer about the part (entry 81), and every otherAESvalue compiles neither. The second define would be required exactly when the first is, so it would add a line to every suite build and a refusal toct.hwithout separating any build that exists.CH_NATIVE_WIDEMULstays apart because it covers a different instruction in different files, and a build asserts it without any AES at all. -
A server issues one self-sealed ticket per connection and resumes it under
psk_dhe_ke, on the caller's clock. colibri needs the QUIC Interop Runner'sresumptioncase in both roles, and Camilo answereddocs/server.md's open question five yes. After the client Finished verifies, the server seals the ticket's PSK, the suite, the ALPN protocol and an issue instant under a ChaCha20-Poly1305 key the caller supplies (srv_ticket.[ch]), and sends it in one NewSessionTicket with noearly_data. A later ClientHello that carries it resumes with a fresh key exchange and no Certificate (srv_resume.[ch]).docs/server.md, "Resumption", states the rules.Cost: one caller-held key as valuable as the signing keys, one more
ch_rand_bytessite (INV-4), one more AEAD caller (INV-1), a caller clock the server did not read before, and 104 bytes of ticket per connection. Gain: a reconnect skips the signature and the certificate on both ends, and two chapulin endpoints resume each other, which is the deployment the client was written for.One ticket, because a client that resumes one connection after another gets a fresh ticket on each; a client racing parallel connections wants more, and none asks. The instant is the last full handshake's, carried forward through every resumed one, so a chain of resumptions ends one lifetime after the certificate last signed. A server with no clock issues no ticket and accepts none. A ticket binds its ALPN protocol, because application state may follow a ticket across connections (RFC 9001 §4.5), and it does not bind the server name or the signing identity: RFC 9846 §4.3.11 tells a server it need not bind the name, and a deployment that retires an identity rotates the ticket key.
A database of tickets was considered and rejected: it needs storage that outlives a connection, which the zero-heap rule forbids. An HMAC over a readable ticket, the cookie's shape, was rejected because the ticket carries the PSK, which must stay secret. Deriving a key per ticket from a random salt, so no nonce can repeat, was considered and left aside: a random 96-bit nonce is safe for 2^32 tickets under one key, and
srv_ticket.htells the operator to rotate before then. -
The X25519 field is a build axis,
X25519=portableorX25519=wide, and the wide field asserts its own multiply.X25519=portable, the default, isx25519.c's 16 limbs of 16 bits: 256 products of 32 by 32 bits per field multiply, whichct.hcan build from 16x16 pieces on any core.X25519=wideaddsx25519_wide.c, five limbs of 51 bits: 25 products of 64 by 64 bits into 128, which a 64-bit core computes as MUL and UMULH on arm64 and as one MUL or MULX on x86-64. On an Apple M1 Pro the wide field takes about 34 µs per scalar multiplication, against 428 µs for the 16-limb field on the native multiply and 953 µs on the 16x16 decomposition the packaged object ships. The x25519 pair was 74% of an RSA-3072 client handshake there, and the wide field takes that client side from 2.57 ms to 0.77 ms (bench/notes-primitives.md).It is an axis rather than a replacement because a device cannot run the wide field. A 32-bit core has no 64x64->128 multiply, so its compiler would build each product from a runtime routine that branches on its operands;
ct.hrefuses the build instead, when the compiler has nounsigned __int128. The 16-limb field stays the default and the device path, unchanged, with its ten proofs, INV-24 and the 32-bit codegen specs. The axis follows the AES one: the Makefile variable picks, one field per object, and the compiler's predefined macros are the whole detection. Nothing probes a CPU at run time, for the reasonsaes_hw.cgives.The values name what each field needs from the target, because that is what the person choosing has to know.
portableruns on every core this tree builds for;wideneeds the wide multiply and a statement about its timing. "fast" would name a property of one machine, and "radix51" names the representation, which says nothing about which targets can build it.The wide field has its own timing assertion,
CH_NATIVE_MUL128, rather than readingCH_NATIVE_WIDEMUL, for two reasons. The two macros name different instructions:CH_NATIVE_WIDEMULis about the 32x32->64 multiply, and a part can promise one and not the other. And the Makefile setsCH_NATIVE_WIDEMULfor every host test binary, so a field keyed on it would move every host test off the 16-limb field, which would lose its native-multiply unit, Wycheproof and differential runs.ct.hwrites the terms: Arm's FEAT_DIT list and Intel's DOIT list both name the instructions, and each holds only in the mode its vendor names.Cost: two fields to keep correct instead of one. The wide field has seven harnesses of its own (INV-34), an equivalence binary that compares it with the 16-limb field on every
make check, a Wycheproof leg, a unit leg, a timing leg and a differential leg. Its constant-time claim rests on a vendor statement this tree cannot check, and on the code the pinned clang emits for arm64 and x86-64, whichlint-wide-multiplyholds; no gcc spec measures it, because no CI lane runs a 64-bit gcc through that gate. Gain: host builds, which open many connections, stop paying for a representation chosen for a core with no wide multiply. The Lean model needs no second copy:spec/lean/Spec/X25519.leancomputes over natural numbers with a reduction mod p after every operation, so it states no limb layout, andbin/diff_x25519_wideruns the x25519 rows against it with the wide field. -
A
TRUST=webpkiclient sends a key share for both groups, andKEXselects nothing for it. Every webpki build carries ML-KEM-768 and lists X25519MLKEM768 and then x25519, the offer entry 39 made underKEX=pq. Its ClientHello now carries a key share for each: the hybrid share, then an x25519 share. The x25519 share repeats the x25519 half of the hybrid share. RFC 9846 §4.3.8 asks for the key_exchange of each KeyShareEntry to be generated independently (rfc9846.txt:2182-2184), and RFC 9954 §3.2 relaxes that rule for a value of the same algorithm reused across the entries of one ClientHello. So the second share needs no second key generation. A server selects either group in one round trip. When it selects x25519, the client wipes the ML-KEM seed,handshake_state.dz, before it runs the exchange.Cost: 36 bytes in every webpki hello, the second KeyShareEntry's group, length and 32-byte value, so
CH_HELLO_MAXgoes from 2,335 to 2,371 against theKEX=pq TRUST=webpkibuild. Against the classic webpki build, which this entry removes, the hello grows by 1,222 bytes,ch_tlsfrom 1,800 to 3,016 bytes on arm64, and thech_connectstack peak from 7,152 to 16,416 bytes, because every webpki object now carries ML-KEM (bench/results-sram.csv). A host pays those bytes easily, and the device modes do not pay them.Gain: no round trip to a server without the hybrid, and no key exchange choice for a host client to get wrong.
make TRUST=webpki KEX=x25519built a client that could never run the hybrid, andmake TRUST=webpki KEX=pqone that paid a HelloRetryRequest round trip to every classic server. Neither build exists now; the one webpki build completes with either server in one round trip.The change also deletes code. Both groups carry a share, so a HelloRetryRequest that names either one names a group the hello already shared, and RFC 9846 §4.3.8 makes that an illegal_parameter abort (rfc9846.txt:2205-2212). A retry can ask this client for a cookie and nothing else.
server_hello_info.retry_group,handshake_state.share_group,take_retryand the x25519-only retry hello are gone, with the mutants that guarded them, and the retry hello's record version depends on the cookie alone again. A cookie retry still works: the retry hello resends both shares and echoes the cookie. Entry 63 lists secp256r1 after the two shared groups with no share, so a retry may name that group again, and the retry hello's record version depends on its position rather than on the cookie.ch_cfg.require_pqkeeps its meaning. It drops x25519 fromsupported_groupsand fromkey_share, so the hello is the one-group hybrid hello a raw or caKEX=pqbuild sends, and a ServerHello that selects x25519 fails with illegal_parameter.KEXnow chooses the group of a raw or ca device client and nothing else, and the Makefile refuses both values everywhere else, for entry 40's reason: a variable must not let a build ask for something it will not get. BesideTRUST=webpki,KEX=x25519asks for an x25519-only hello, andKEX=pqasks for the one-group hybrid hello, whichrequire_pqgives at run time. A server role's key exchange is not a build choice either: the server offers x25519 until its hybrid half lands, and then carries ML-KEM in every build, soROLE=serverandROLE=bothrefuseKEXtoo.$(origin KEX)tells the default from a value on the command line or in the environment. The object directory names a webpki build's key exchangeboth.cfg.hrefuses-DCH_KEX_PQbeside-DCH_TRUST_WEBPKIfor a client-only tree with its own build system.A
ROLE=bothwebpki object, the one colibri links, carries the two-group client beside a server that still offers x25519 alone. The client's defines,CH_KEX_TWO_GROUPSandCH_KEX_HYBRID, come fromCH_TRUST_WEBPKIand not fromCH_KEX_PQ, so the server half keepsCH_KEX_GROUPat x25519 and never meetssrv_flight.h's refusal ofCH_KEX_PQ. The server's hybrid half, when it lands, replaces that refusal and those x25519 constants. Entry 54 landed it: the refusal and the server's use ofCH_KEX_GROUPare gone, and that object's server selects the hybrid its client offers.Entry 39 rejected carrying both shares for three costs: 1,216 octets of ML-KEM share in every hello, a key pair most handshakes discard, and a larger
CH_HELLO_MAX. The build that entry made already paid the first two: its hello carries the hybrid share, and so draws the ML-KEM key pair, whether or not the server takes it. What both shares add is the 36-byte x25519 entry, and that entry's key pair is the x25519 half the hybrid share already carries.A second, independent x25519 key pair for the x25519 share was considered and rejected. RFC 9954 permits the reuse, the server selects one group so only one share enters a key exchange, and the second pair would cost one more scalar multiplication per handshake.
-
A server holds both groups and prefers X25519MLKEM768, and asks for it with a HelloRetryRequest when the client did not share it. Camilo answered
docs/server.md's open question ten on 2026-09-24: every server build,ROLE=serverandROLE=bothover all three transports, carries ML-KEM-768 and selects the hybrid for any client that lists it.srv_kex.[ch]holds the choice, the server's key share and the shared secret, andch_tls.groupreports the group on the server as it does on the client.The order is the server's, and it reads
supported_groups: X25519MLKEM768 when the client lists it, and x25519 otherwise. A hello that carries the hybrid share gets the hybrid in one round trip, even when an x25519 share comes beside it. A hello that lists the hybrid and shares x25519 alone gets a HelloRetryRequest that names the hybrid, through the cookie the retry path already had, and its second hello must carry that share. A hello that lists x25519 alone gets x25519. RFC 9846 §4.3.8 describes this shape for a server that respects preferences: select fromsupported_groupsfirst, then send a ServerHello or a HelloRetryRequest from whatkey_sharecarries (rfc9846.txt:2172-2177). Here the preference is the server's own.The trade is one round trip, paid only by a client that lists the hybrid and shares x25519 alone. Taking that x25519 share would save the round trip and complete a classic key exchange with a client that offered post-quantum protection, which is the recording entry 12 names: harvested now and decrypted later. OpenSSL with x25519 first in its group list pays the round trip. A
TRUST=webpkiclient, aKEX=pqclient and OpenSSL's default list share the hybrid and pay nothing.The hybrid follows RFC 10024. The server encapsulates to the ML-KEM encapsulation key at the front of the client's share and answers with the ciphertext and then its x25519 value, and the shared secret is the ML-KEM secret and then the x25519 one, the order
handshake_flight.c'shybrid_secretalready reads. An encapsulation key that fails FIPS 203 §7.2's modulus check, and a share of any length but 1,216 bytes, end the handshake with illegal_parameter, as RFC 10024 asks; the x25519 half keeps the all-zero check (INV-3). The 32 bytes of encapsulation randomness are a newch_rand_bytessite, drawn only when the server selects the hybrid (INV-4), and the ML-KEM secret lives inhandshake_state.mlkem_ssfrom the ServerHello to the key schedule and no longer (INV-17).Cost: every server object carries ML-KEM-768 and SHA-3. The server's
ch_tlsgrows from 1,368 to 1,968 bytes on arm64, because the ServerHello it stages in the TX array carries a 1,120-byte share, andch_srv_accept's stack peak goes from 5,248 to 10,304 bytes through the encapsulation (bench/results-sram.csv), which puts every server build on the hybrid's 6,656-byte frame budget (INV-19). There is no classic-only server: a device that cannot spare the stack cannot build one. Gain: every client that can run the hybrid gets it from a chapulin server, colibri's QUIC server included, with no build choice to get wrong.Three alternatives were considered and rejected. Selecting whichever group the client shared saves the round trip and gives a classic key exchange to exactly the clients that shared x25519 first. A
KEXaxis for servers would be a build choice a server does not need, and entry 53 already refusesKEXbeside a server role. Drawing the encapsulation randomness insrv_beginbeside the x25519 scalar would draw 32 bytes a classic handshake never uses and keep them in the handshake state until the ServerHello.Entry 63 adds secp256r1 as a third group, taken after x25519, only from a client that lists neither of the other two; the order above stands for them.
-
A
TRUST=webpkiclient offers the certificate path beside a ticket, and a declined ticket becomes a full handshake in the same connection. cocuyo, a DNS-over-TLS client that links the webpki tcp-nonblocking object, measured dns.google (8.8.8.8:853) on 2026-09-24. With the hello entry 47 wrote, which offered the ticket and no signature scheme, it resumed 0 of 10 connections: dns.google answered every resuming hello with a handshake_failure alert, a good ticket included. It appears to choose its certificate and signature scheme before it decides whether to resume, so a hello with no scheme fails there. RFC 9846 §4.3.3 names missing_extension for that case (rfc9846.txt:1813-1816). With signature_algorithms added beside the ticket and no fallback, cocuyo resumed 5 of 6; the sixth was an ordinary decline, and the client failed it closed withCH_EAUTH. OpenSSL 3.6.4's s_client resumed 6 of 9 and completed the other 3 as full handshakes in the same connection. cloudflare-dns.com and dns.quad9.net resumed with either hello.So the change has two halves, and each one is in the RFC. RFC 9846 §4.3.3 requires signature_algorithms of a client that wants a server to authenticate with a certificate (rfc9846.txt:1811-1813), and §9.2 lets a hello leave it out only when the hello offers a PSK (rfc9846.txt:4595-4597). §2.2 asks a client that offers a PSK to send a key share "to allow the server to decline resumption and fall back to a full handshake" (rfc9846.txt:699-702), and this client always sends one.
- The hello. Every webpki ClientHello carries signature_algorithms with the five schemes, and server_certificate_type when SPKI pins are set, whether or not it presents a ticket. pre_shared_key stays the last extension, because the binder covers every byte before the binders list (rfc9846.txt:2564-2565). A retry hello after a HelloRetryRequest carries the same extensions and a binder computed again over the transcript the retry replaced.
- A selected ticket. The handshake resumes as before: no
Certificate, and
ch_tls.psk_selectedis 1. - A declined ticket. A ServerHello with no pre_shared_key makes
hsf_accept_server_hellowipe the PSK's early secret and binder key and derive the early secret of no PSK, HKDF-Extract over 32 zero bytes (rfc9846.txt:4182-4185). The handshake then reads EncryptedExtensions, Certificate, CertificateVerify and Finished, and checks the chain, the clock, the hostname, the anchors and any SPKI pins exactly as a handshake with no ticket does.ch_tls.psk_selectedis 0. A pre_shared_key that names any identity but 0 is no decline, and fails with illegal_parameter (rfc9846.txt:2551-2557).
The three drivers decide whether a Certificate comes next from
ch_tls.psk_selected, whichhsf_accept_server_hellowrites, and no longer fromcfg.psk. In the raw and ca modes the two agree on every path, because those modes still fail a decline.Cost: 16 bytes in every resuming webpki hello, 23 with SPKI pins beside anchors.
CH_HELLO_MAXandCH_TX_STAGEgrow by 23 bytes in the webpki builds, from 2,371 to 2,394 over TCP and from 2,625 to 2,648 over QUIC, andch_tlsgrows from 3,016 to 3,040 bytes on arm64 (bench/results-sram.csv). A webpki hello now offers two ways for a server to authenticate, the sixth thing the mode offers more than one of, and one handshake has two paths through it. The device modes pay nothing.Gain: cocuyo resumes with dns.google, and a server that declines a ticket costs one handshake instead of a failed connection and a reconnect. That matches what OpenSSL does and what §2.2 describes.
The raw and ca modes stay as they were: their resuming hello offers the ticket alone, and a decline fails with handshake_failure (
CH_EAUTH). Their server is the one endpoint a device pins, usually a chapulin server that holds its own ticket key, so a decline is rare and costs one reconnect. Offering the pinned scheme beside the ticket would put a second path through the device driver for a case nothing measured asks for, and in the ca modes a declined ticket would also have to run the epoch check that a resumed handshake skips (INV-21).Two alternatives were considered and rejected. Adding the schemes and failing a decline closed, the build cocuyo measured, resumes with dns.google and still turns about one connection in three into a reconnect. Reconnecting inside
ch_connectwithout the ticket needs a second connection, and the caller owns the socket. -
Every packaged object exports a build record,
ch_build, and a consumer compares it with its own headers. A consumer linksbin/chapulin.oand compiles against the headers under defines it writes itself. cocuyo reads them through Zig's@cImportunderCH_TRUST_WEBPKI,CH_TRANSPORT_TCP_NONBLOCKINGandCH_RAND_EXTERN, and colibri and stompy write their own lists. Nothing checked that those defines were the object's, and a mismatch links and runs. A consumer that forgetsCH_TRUST_WEBPKIbeside a webpki tcp-nonblocking object passes a 152-bytech_cfgto a call that reads 232 bytes, and declares a 1,624-bytech_recordthat the object writes 4,120 bytes of (arm64). Sobuild.cdefines one constch_build_infoin every object: the record format, one bit per define that changes a public layout or bound, the sizes ofch_cfg,ch_tls,ch_ticket,ch_record,ch_quicandch_rsa_priv, and four bounds:CH_TX_STAGE,CH_MIN_RXBUF,CH_X509_MAXin a CA mode andCH_TRANSPORT_PARAMS_MAXin a QUIC build.build.hcomputes the same values from the consumer's defines asCH_BUILD_macros, andch_build_matches(&ch_build)compares them all.The axes are the ten defines that change a size or a bound the record holds, or, for
CH_PIN_ECDSAin a raw mode, the length a pinned key has:CH_TRUST_CA,CH_TRUST_WEBPKI,CH_PIN_ECDSA,CH_KEX_PQ,CH_TRANSPORT_QUIC_NONBLOCKING,CH_TRANSPORT_TCP_NONBLOCKING,CH_SUITE_AES_GCM,CH_ROLE_SERVER,CH_EXPORTERandCH_KEYLOG.build.hsays what each one changes. Left out, each measured with and without the define under all three transports:- The RAND pattern. It changes no size or bound:
rand.hdeclares the same call in every build, anddrbg.hadds only the declaration ofch_drbg_seed, underCH_RAND_DRBG. A mismatch still reports itself. An extern-pattern program that links a drbg object never has itsch_rand_bytescalled, and the object's generator stops atCH_ASSERTon its first draw, unseeded, before a handshake sends a byte. A drbg-pattern program that links an extern object fails to link. CH_ROLE_BOTH. AROLE=bothobject has the layouts and bounds of theROLE=serverobject with the same trust defines. The define declaresch_connect, and a program that calls it against a server object fails to link.- The AES implementation,
X25519=wideandWIDEMUL. They change code and no size or bound.AES=externimportsch_aes_block, and a program that does not define it fails to link. CH_NATIVE_AESandCH_NATIVE_MUL128, statements about the part thatct.hreads to refuse a build.CH_KEX_TWO_GROUPSandCH_KEX_HYBRID.cfg.hcomputes both fromCH_TRUST_WEBPKIandCH_KEX_PQ, so their bits would repeat those two.
The record is data, and not a check inside
ch_connectand the init calls, for three reasons:- No exported name changes and no call gains a parameter or a result code, so every consumer links as it did.
- A consumer in another language reads it. The struct is twelve
uint32_tfields with no padding, and the expected values are object-like macros, which Zig's translate-c turns into constants. A Zig 0.16 program that@cImportsbuild.hunder cocuyo's three defines readsc.ch_build, callsc.ch_build_matches, and gets 1 against the webpki tcp-nonblocking object and 0 withCH_TRANSPORT_TCP_NONBLOCKINGleft out. - A consumer is asked, not forced. No library source reads
ch_build, so a firmware tree that compiles the sources into its own build, and so cannot disagree with itself, carries 48 bytes of read-only data and runs nothing.
Cost: one more exported symbol in every object, and the first that is data rather than a call, so every export list names it and every count of exported symbols grows by one, while the counts of calls stay as they were. 48 bytes of read-only data per object. A consumer that never compares gets nothing from it.
lib-checkbuilds a consumer twice for every object it checks, about 0.2 s a leg, andmake checkgained one leg,TRUST=raw-ecdsa KEX=pq, 2.6 s cold and 1.0 s warm, so the record is read in aKEX=pqobject.Gain: a disagreement between an object and a consumer's defines is one comparison at startup instead of a struct written past its end.
Two alternatives were considered and rejected. A check inside each init call, against a size the caller passes, changes every call's signature or wraps each one in a macro that a Zig consumer cannot use, and a size alone misses a bound such as
CH_X509_MAX. A symbol name per build, so a mismatch fails to link, would encode every axis and every overridable bound in the name and change every consumer's link line with each axis. The record compares layouts and bounds, not behavior, and defines, not revisions: headers from another commit are caught only where a size, a bound or a bit moved (INV-35).Entry 61 changed one thing here: the record's symbol name now carries the object's transport, and
build.hmapsch_buildto it, because two objects that both definedch_builddid not link into one image.Entry 77 changed one thing here:
CH_RAND_SESSIONtakes bit0x400ofaxes, because it adds two fields toch_cfg.RAND=externandRAND=drbgstill take no bit. - The RAND pattern. It changes no size or bound:
-
A failed QUIC session keeps its write keys for one CONNECTION_CLOSE per level, and a call of its own seals it. RFC 9001 §4.8 turns a TLS alert into a CONNECTION_CLOSE frame whose error code is 0x0100 plus the alert (rfc9001.txt:883-887). Until this entry a failure wiped every key at once, as entry 21 has every error do, so colibri could seal no close and the peer waited out its idle timeout. h3spec's TLS cases, a KeyUpdate at the Handshake level, no_application_protocol and missing_extension, check for that close (colibri#59).
- What a failure keeps.
quic_failwipes every read key,hsand the traffic secrets, and clears every read bit ofch_quic.levels_ready, as before. It keeps the write keys of each level whose write bit is set:initial_dcidat the Initial level, which the seal derives the Initial send key from,handshake_txandhandshake_hp_tx, andapp_txandapp_hp_tx. Both drivers fail through it, so the rule is the same for the client and the server. - What the caller does. It reads
ch_quic_error_codeand callsch_quic_seal_closeonce at each level it can send, with one CONNECTION_CLOSE frame as the plaintext. The call seals the packet asch_quic_sealdoes, then wipes that level's write keys and clears its write bit, so a second call there returnsCH_EINVAL.ch_quic_closewipes whatever keys remain. docs/quic.md, "When a session fails", lists the steps. - Its own call, not
ch_quic_seal.ch_quic_sealkeeps its rule that a dead session seals nothing. So the caller's ordinary send path cannot seal a queued ACK or STREAM frame in place of a level's one close, and a reader of a call site knows which of the two it is. - A bound on the packet.
CH_QUIC_CLOSE_MAXis 1200 bytes, header and tag included: RFC 9000 §14's smallest maximum datagram size (rfc9000.txt:4598-4599). A close needs a few dozen bytes, so the bound refuses nothing a close needs. chapulin checks no frame content, because that means parsing QUIC frames, which colibri owns; the bound caps what a caller that passes other bytes can send under a kept key at one such packet per level.
Initial keys are public: RFC 9001 §5.2 derives them from a printed salt and a connection ID that travels in the clear. Handshake and 1-RTT keys are secret, and this entry lets a secret write key outlive the failure, until its one seal or until
ch_quic_close. Three facts make that acceptable. RFC 9000 §10.2.3 asks for the close at the Handshake level, and at the Initial level from a server, before the handshake is confirmed, because the peer may not yet read a higher level (rfc9000.txt:3306-3308, rfc9000.txt:3316-3320); the kept key is what that packet needs. The key protects only the one packet the caller builds, and the read keys, which would open the peer's packets, die with the failure. And the key is wiped right after that packet, so it lives as long as the caller takes to build one packet. When the failure is the peer's authentication, the Handshake key is shared with a peer this endpoint did not authenticate, and the close tells that peer the error code and nothing else.Cost: one exported call, so a
TRANSPORT=quic-nonblockingclient object exports sixteen calls and aROLE=server TRANSPORT=quic-nonblockingobject nineteen. INV-17 gains an exception, stated there. A caller that neither seals nor closes keeps a failed session's secret write keys for the life of the session struct. Gain: the peer learns why the connection ended, at each level it can read, and stops waiting on its idle timeout.Three alternatives were considered and rejected. Letting
ch_quic_sealrun once per level on a failed session changes the meaning of every existing call site, and a queued packet could take the close's place. Keeping every key and letting the caller seal any number of packets leaves nothing bounding what a failed session encrypts under a secret key. Building the frame inside chapulin would put a QUIC frame encoder on chapulin's side of the line entry 38 draws. - What a failure keeps.
-
A
SUITE=aesgcmbuild holdsTLS_AES_128_GCM_SHA256andTLS_AES_256_GCM_SHA384beside ChaCha20, over every transport, on the AES instructions alone. Entry 45 gave the build one AES suite, and refused it over QUIC. colibri serves QUIC clients it does not control, and such a client may offer AES-GCM alone: h3spec's ClientHello listsTLS_AES_256_GCM_SHA384,TLS_AES_128_GCM_SHA256andTLS_AES_128_CCM_SHA256, and no ChaCha20. RFC 9846 §9.1 makes the AES-128 suite a MUST and the AES-256 one a SHOULD (rfc9846.txt:4540-4543).- AES-256 runs on the instructions, never on the S-box.
aes_traffic_key_initexpands a 16-byte or a 32-byte key, andaes_encrypt_scheduleruns ten or fourteen rounds by the round count the schedule records. A traffic key is secret, soct.h's refusal of-DCH_SUITE_AES_GCMwithoutAES=hwandCH_NATIVE_AEScovers both key sizes.quic_aes_soft.cholds a software AES-256 that only-DCH_AES_256_TESTcompiles: it is the referencebin/aes_equiv_testand theaes256proof hold the instructions to, andlint-trust-separationrefuses that define in every packaged object. - The key schedule takes the hash length first.
hmac,hkdf_extract,hkdf_expand,hkdf_expand_label,hkdf_derive_secretand everyks_call take a leadingsize_t hash_len:SHA256_LEN, orSHA384_LENin a-DCH_SUITE_AES_GCMbuild. HMAC-SHA-384 is a body of its own behind a two-arm dispatcher, the shapedocs/server.mdmeasured, because one body over both hashes exceeds the complexity limit. Every secret array isHKDF_HASH_MAXwide, which isSHA256_LENin every build without the suite, so those builds keep their sizes. - The transcript runs both hashes in a suite build. A client hashes
its ClientHello before the ServerHello names a suite, so
ch_transcriptfeeds SHA-256 and SHA-384 the same bytes, and each read names its hash by length. A HelloRetryRequest writes the synthetic message at the retry suite's hash. The client derives the early secret of no PSK when the ServerHello arrives, at the suite's hash. - A PSK carries its hash in its length. A ticket's PSK is as long
as its suite's hash (
rfc9846.txt:3298-3301), soch_ticketgainspsk_len, a client presents it asch_cfg.psk_len, and the binder and the early secret take the hash that length names. A client aborts with illegal_parameter when the server selects the PSK under a suite of another hash (rfc9846.txt:2551-2556). A server passes a ticket over unless its suite has the selected suite's hash (rfc9846.txt:3219-3220), and the handshake goes on with a certificate. - The record layer gains AES-256.
rec_dirrecords its suite, the key and IV derive at the suite's hash, and a KeyUpdate runs "traffic upd" at that hash and keeps the suite. - QUIC protects Handshake and 1-RTT packets with the suite. Each
quic_keysandquic_hp_keyrecords its suite, packet protection runs AES-GCM or ChaCha20-Poly1305 by it (RFC 9001 §5.3), and header protection runs AES-ECB or ChaCha20 by it, at the suite's key length (§5.4.3, §5.4.4). The Initial level keeps AES-128-GCM under its public keys, unchanged.quic_packet.cbuilds anaes_traffic_keyon its frame for each packet and each mask and wipes it there. Each AES-GCM key set counts what it seals and refuses the packet that would reach §6.6's confidentiality limit of 2^23 (rfc9001.txt:1812-1813). The count lives in the key set, so a key update starts it again andch_quicgains no field; initiating that update before the limit is colibri's (rfc9001.txt:1803-1805). The integrity limit stays ChaCha20's 2^36, which is stricter than AES-GCM's 2^52 (rfc9001.txt:1829-1831). - The order. The server selects ChaCha20, then AES-128-GCM, then
AES-256-GCM, the first of them the client listed, and the webpki
client offers them in that order. ChaCha20 comes first for entry
45's reason. AES-128-GCM comes before AES-256-GCM because its
schedule is SHA-256, the hash every handshake proof covers; the
SHA-384 schedule is proved in its own harnesses alone. So h3spec's
offer selects AES-128-GCM, which
bin/webpki_loop_aesfeeds the server through the real parser.ch_srv_cfg.cipher_suitesreplaces the order: a host whose AES instructions outrun its ChaCha20, or one that wants AES-256's margin, names the order it wants, and a list may leave a suite out. Every server init refuses a code point the build does not hold. Entry 80 orders both roles AES-256-GCM, then AES-128-GCM, then ChaCha20 in a build onAES=hwwithCH_NATIVE_AES, where h3spec's offer then selects AES-256-GCM. - The key log takes the secret's length.
ch_keyloggainssecret_len, because a SHA-384 secret is 48 bytes and the NSS format writes all of them.
Cost, measured by
bench/sram.shon arm64:ch_tlsgrows from 1,984 to 2,256 bytes in aROLE=server SUITE=aesgcmbuild and from 3,064 to 3,336 in aTRUST=webpki SUITE=aesgcmone, against the same probes run on the tree before this change.ch_quicin theROLE=both TRUST=webpkiQUIC object colibri links is 5,432 bytes with the suite and 4,944 without it. The suite build'sch_srv_acceptpeaks at 10,448 bytes of stack where it peaked at 10,304, and its webpkich_connectat 16,544 where it peaked at 16,432.ch_readpeaks at 1,728 bytes where it peaked at 1,696, in every build, on the KeyUpdate path through the hash-agile HKDF. Every other build keeps its session size. A suite build hashes every handshake byte twice.ch_ticketgains 8 bytes, forpsk_len, in every build, and the key log hook changes signature in everyKEYLOG=onbuild, colibri's included. A server ticket in a suite build is 120 bytes where it was 104, because its body holds a 48-byte PSK. Gain: the server meets §9.1 for AES-only clients over every transport, and a webpki client reaches an AES-only endpoint.No publication prints a QUIC packet protected under an AES-256-GCM key or a SHA-384 schedule: RFC 9001 Appendix A protects its Initial packets with AES-128-GCM and its 1-RTT packet with ChaCha20.
test/quic_suite_test.cholds both AES suites to an independent Python computation instead, one that reproduces Appendix A.5 from its printed secret first.TLS_AES_128_CCM_SHA256stays out: no client this tree serves needs it, and it would add a second AES mode with its own tag construction.Adding the AES-256 rows to the webpki loop found a defect entry 45 shipped. The server read
cipher_suitesonce per suite it held, and each read walked the whole list, so aSUITE=aesgcmserver read the compression bytes as suites and refused every real ClientHello. The parser now reads the list once. - AES-256 runs on the instructions, never on the S-box.
-
A server compares the extensions of a retried ClientHello as a set: the frozen digest takes them in ascending type order. colibri found the need through the QUIC Interop Runner on 2026-09-24. Since entry 54 (4d897ed) a server asks for X25519MLKEM768 with a HelloRetryRequest when a client lists it and shares x25519 alone, and ngtcp2's interop client does exactly that. Its second ClientHello carries the same extensions in another order: supported_versions moves from after key_share to the front. The server keeps no state across the retry. It compares a SHA-256 digest, carried in the cookie, over every field and extension §4.2.2 does not let the client change, and it computed that digest in the order the extensions sat on the wire. The reorder changed the digest,
srv_check_retry_helloanswered illegal_parameter, and every ngtcp2 handshake that needed a retry failed.test/srv_quic_retry_vectors.hholds the two hellos colibri recorded.§4.2.2 has the client send "the same ClientHello without modification" apart from five listed changes (rfc9846.txt:1191-1213), and §4.3 lets extensions "appear in any order" (rfc9846.txt:1669-1670). Camilo decided on 2026-09-24 that a reorder is not a modification the server refuses. A second hello that carries exactly the covered extensions of the first, each with the same type and bytes, is accepted in any order. One that adds, drops or changes a covered extension, or changes a head field from legacy_version through legacy_compression_methods, is refused with illegal_parameter as before. The five extensions §4.2.2 lets change, key_share, early_data, cookie, pre_shared_key and padding, stay out of the digest.
The construction. The digest is SHA-256 over the head, then each covered extension whole, its type, length and body, in ascending type order (
add_frozen_extensions,srv_parser.c). The parser already refuses a second extension of any type, unknown types included, with illegal_parameter (srv_ext_duplicate, rfc9846.txt:1673-1674), before it computes the digest, so the types in a block it hashes are distinct. With distinct types, one set of covered extensions gives one byte string: the sort order is fixed, and each extension carries its own type and length, so the string reads back into exactly one set. Two different sets therefore give two different strings, and a digest that matches across them is a SHA-256 collision. That is the argument the wire-order digest rested on, with a list replaced by a set. The cookie format does not change: 32 bytes of digest, opened and compared withct_memeqas before. A cookie minted before this change carries a wire-order digest and fails the comparison for any hello whose extensions were not already in ascending order, and a cookie lives for one retry, so nothing supports the old form.The cost. The walk keeps no list: each pass reads the whole block for the covered extension with the smallest type above the last one it added, and the pass after the last one finds none. A block of n extensions costs at most (n + 1) * n extension headers read, beside the n * (n - 1) / 2 that
srv_ext_duplicatealready read. The construction does not bound n. The block is at most 65,535 bytes and never longer thancfg.buf_len, and an extension is at least 4 bytes, so n is at most 16,383, and a device's buffer holds far fewer. A ClientHello of n empty extensions of distinct unknown types, parsed on an arm64 M1 Pro (clang -O2) by the parser before this entry and after its construction, took:Extensions Message Duplicate check Whole parse, before Whole parse, after 100 443 B under 1 ms under 1 ms under 1 ms 1,000 4,043 B 5 ms 5 ms 15 ms 4,086 16,387 B 75 ms 74 ms 252 ms 16,383 65,575 B 1.2 s 1.2 s 4.1 s A browser or ngtcp2 hello carries 10 to 20 extensions, a few hundred header reads. The last row is a host that accepts a 64 KiB ClientHello: one hello costs it 4.1 s of processor time where it cost 1.2 s before, all of it before any key exists.
The bound on the extension count. Camilo decided on 2026-09-24 to bound n. The parser refuses a ClientHello whose extension block holds more than
SRV_CLIENT_HELLO_EXT_MAXextensions, 128, with illegal_parameter (srv_parser.h).srv_ext_over_maxcounts them in one walk that stops at the 129th, so it reads at most 129 headers however long the block is.srv_parse_client_hellocalls it after the block's length check and beforesrv_ext_duplicate, so neither walk whose cost is n squared ever runs over more than 128 extensions. The count is not taken inparse_extension's walk, becausesrv_ext_duplicateruns before that walk: a count there would bound the frozen digest's walk and leave the duplicate check unbounded. An unknown type counts the same as a recognized one. Every server path reads its hellos throughsrv_parse_client_hello: the blocking server (srv_handshake.c), the tcp-nonblocking server (srv_tcp_nonblocking.c) and the QUIC server (srv_quic.c), for a first hello and a retried one alike, so the refusal is the same on each.Why 128. Each count below was published, or read from the library's source, as of 2026-09-24.
- Browsers. JA4 fingerprints count a hello's extensions and leave GREASE out (https://github.com/FoxIO-LLC/ja4). Chrome's carry 16 to 18, to which it adds two GREASE extensions, and one more, pre_shared_key, when it resumes; a capture of Chrome 146 to 153 is bogdanfinn/tls-client#281. Firefox's carry 17, Safari's 11 to 14, and Chromium's over QUIC 12.
- Libraries. A client sends at most the extensions its source can write: OpenSSL defines 32, BoringSSL 31 plus padding, pre_shared_key and two GREASE extensions, rustls 23 and Go's crypto/tls 20. curl over OpenSSL sends 12.
- QUIC. ngtcp2's two recorded hellos carry 10 and 11.
- The registry. IANA's TLS ExtensionType Values registry (https://www.iana.org/assignments/tls-extensiontype-values) assigns 64 values, 43 of them allowed in a TLS 1.3 ClientHello, and RFC 8701 reserves 16 GREASE values for extensions. A client that sent every assigned type and every GREASE value once would send 80, and the parser refuses a type sent twice.
So 128 is six times what Chrome sends and 48 more than a client can send without inventing types. The bound departs from one RFC 9846 rule, and this record says so: a server must ignore an unrecognized extension (
rfc9846.txt:1299), and §9.3 restates that as an invariant (rfc9846.txt:4636-4637). A hello of 129 extensions, most of them unknown, is refused here where those lines would have it negotiate. No client above sends one. tlsfuzzer'stest-tls13-large-number-of-extensions.py(https://github.com/tlsfuzzer/tlsfuzzer) does: it sends thousands of empty unknown extensions and expects a handshake, so those conversations fail against this server by design. None of the libraries above bounds the count. Each refuses only duplicates, and BoringSSL, Go and rustls find them by sorting the types or with a map or a set, each held in memory that grows with n; a parser with no heap has no such memory.Why illegal_parameter. RFC 9846 bounds the block's bytes,
Extension extensions<7..2^16-1>(rfc9846.txt:1237), and not its count, so a block of 129 extensions parses under the syntax. §6 sends decode_error for a message that "cannot be parsed according to the syntax" and illegal_parameter for one that is "syntactically correct but semantically invalid" (rfc9846.txt:3784-3791). The alert definitions draw the same line (rfc9846.txt:3947-3950,rfc9846.txt:3961-3966), and add that decode_error "should never be observed in communication between proper implementations", which would point an operator at a corrupted message that is not there. So the refusal is illegal_parameter, this design's choice under §6, the rule the refusal of a second extension of one type follows too. It is checked before every rule on the extensions themselves, so a hello past the bound is illegal_parameter whatever else its extensions break.The cost with the bound, measured the same way on the same machine, best of 2,000 runs where a run takes under a millisecond. The clock counts whole microseconds, so a refusal shows as its smallest step:
Extensions Message Whole parse, without the bound Whole parse, with it 128 555 B 0.25 ms 0.25 ms 128, with 507-byte bodies 65,451 B 0.58 ms 0.57 ms 129 559 B 0.25 ms 1 µs or less, refused 4,086 16,387 B 262 ms 1 µs or less, refused 16,383 65,575 B 4.2 s 1 µs or less, refused The most a hello can now cost the parser is the second row: the 128 extensions the bound admits, with bodies that fill a 64 KiB block. What it adds to the row above is SHA-256 over those bodies, the only work the parser does per byte of an unknown extension. Before this entry, 16,383 empty extensions in the same 64 KiB cost 1.2 s.
Three alternatives were considered and rejected. An XOR of one digest per extension is order-independent, but two copies of one extension cancel and a client could add a pair unseen. A sum of per-extension digests modulo 2^256 does not cancel, but its collision resistance is a generalized birthday bound this record cannot state plainly. Sorting into a buffer costs n log n and needs memory in proportion to n. With the bound, the buffer would fit in 256 bytes of stack, but the walks it would replace cost a quarter of a millisecond at the bound, so it would add a buffer and a sort to audit and save nothing a caller could measure.
bin/srv_quic_testreplays the recorded hellos: the first draws a HelloRetryRequest that matches colibri's server's byte for byte up to the digest, and the second completes the handshake through the client Finished, with this server's cookie and the test's own key share written over the recorded ones. The same binary holds the boundary pairs: the first hello's order and ngtcp2's accepted; one covered byte changed, one covered extension dropped, one added, one head byte changed and one extension sent twice refused.bin/srv_testholds the digest to a vector over the covered extensions in ascending order. Two CBMC harnesses in the slow tier cover the parser.srv_parser_frozenproves the construction above over every extension block up to 24 bytes: the duplicate check answers exactly when two types match, and the walk hands SHA-256 each covered extension once, whole, in strictly ascending type order.srv_parser_walkproves the whole ClientHello walk memory safe up to 64 bytes with the readers stubbed; it had no launch line before this entry, because harness.h's SHA-256 stub made the formula too large, and it now keeps a stub of its own. OpenSSL'ss_clientoffers no option that reorders a retried hello, sotest/e2e.shkeeps its TCP retry leg, which sends the same order twice and still passes.The bound has tests of its own.
bin/srv_testfills the golden hello out with unknown extensions: at exactly 128 it parses, and its frozen digest matches a vector over every covered extension, the filler included; at 129 it is illegal_parameter. A supported_versions with a trailing byte is decode_error at 128 and illegal_parameter at 129, and so is a block that ends in half a header, so the bound is checked first.bin/srv_tcp_nonblocking_testsends this tree's own hello, filled out to 128 and to 129, through the tcp-nonblocking server: the first draws the flight and the second illegal_parameter.bin/srv_quic_testfills ngtcp2's two recorded hellos out with the same unknown extensions, so the frozen digest matches and only the count can refuse. A first hello of 127 and its retried hello of 128 draw the flight; a first hello of 128 draws a HelloRetryRequest and its retried hello of 129 is illegal_parameter; a first hello of 129 is refused before any HelloRetryRequest. The ngtcp2 replay above still passes.Two CBMC harnesses hold the bound at smaller values, which
srv_parser.hadmits for a harness.srv_parser_countproves, over every block up to 24 bytes with the bound at 3, thatsrv_ext_over_maxanswers 1 exactly when the block begins with more than the bound's whole extensions, and that it never reads a header past the one that passes the bound. It runs in the fast tier, in 1 s.srv_parser_walktakes the bound at 4 and holds the loops of both walks to four extensions, one fewer than its 64-byte message holds, so an unwinding assertion fails if either walk ever runs over a fifth. The bound let that formula grow from 60 bytes to 64: before it, 64 bytes returned no verdict in 18 minutes, and with it the formula takes 489 s and 1.34 GB, still the slow tier. Five violations guard the bound. srv-parser-ext-max-off-by-one, srv-parser-ext-max-refuses-bound and srv-parser-ext-max-after-walk requirebin/srv_testto fail, srv-parser-ext-max-removed requiresbin/srv_quic_testto fail, and srv-parser-ext-max-after-duplicate, which moves the count below the duplicate check, requires thesrv_parser_walkproof to fail. Both checks answer illegal_parameter, so no test can see that order, and the proof's loop bound is what catches it. -
The peer's close_notify closes the peer's direction alone, and
ch_closecloses this side's. RFC 9846 §6 says a close_notify closes one direction of the connection (rfc9846.txt:3767-3768), and §6.1 says sending one has no effect on the sender's read side and drops TLS 1.2's rule of answering one at once with a close_notify of one's own (rfc9846.txt:3857-3864). Until this entrych_readkept the TLS 1.2 rule: on the peer's close_notify it calledch_close, which sent this side's close_notify and wiped both directions. The caller could not send what it still owed, becausech_writerefused the closed session. UnderTRANSPORT=tcp-nonblockingthe reply went throughcfg.sendfrom insidech_read, which colibri's adapter does not expect, so the alert was lost, and the caller's ownch_closesent nothing because the keys were gone.- What the read does. The
ch_readthat reads the peer's close_notify returns 0 and sends nothing. It wipesrd,rd_secretandres_master, the secrets only a read uses (INV-17), and setsch_tls.read_closed. Every laterch_readreturns 0 before it callscfg.recv. §6.1 says data after a closure alert MUST be ignored (rfc9846.txt:3837-3839), and a record that is never read is never decrypted or acted on. - What stays.
wr,wr_secretandexp_masterstay, andch_tls.statestaysCH_ST_CONNECTED.ch_writesends as before, andch_closesends this side's close_notify under the write key and wipes the rest, as it always did. - How a caller learns of it. From
ch_read's 0, which already meant the peer closed, and fromread_closed. A new state value was considered and rejected. Callers testCH_ST_CONNECTEDbefore they write, and writing is still allowed, so every such test would be wrong until the caller learned the new value;CH_ASSERT(t->state <= CH_ST_FAILED)andtcp_nonblocking_session_deadwould each need to judge a fifth value too.CH_ST_CLOSEDkeeps its one meaning: this side calledch_closeorch_record_close, and no key is left. The field sits in the padding beforesend_epochs, sosizeof(ch_tls)did not change in any build, host or rv32, andch_buildrecords the same sizes. - Every driver at once. The blocking client and server and the
tcp-nonblocking client and server all read through the one
ch_readintls.c. A QUIC object compiles notls.c: QUIC carries no close_notify, and RFC 9001 §4.8 treats every TLS alert as fatal (rfc9001.txt:888-893), so nothing there changes.
Cost: one public field. A caller whose
ch_readreturned 0 must still callch_close, or the peer never receives this side's close_notify, and the write key lives until it does. The examples already called it on that path, andtest/tls_client.cnow does. Gain: this side can finish sending after the peer is done, as RFC 9846 intends, andch_readsends nothing when the peer closes.A call that sends this side's close_notify and keeps reading, the other half of a half close, was not added. No caller has asked for it, and
ch_closestays the one call that ends a session. §6.1 lets a party close its read side without waiting for the peer's close_notify (rfc9846.txt:3864-3867), whichch_closedoes. - What the read does. The
-
An image links one packaged object of each of two transports: the three exports every transport carries take the transport into their symbol names, and the two pairs that share calls are refused at the link. cocuyo wants DNS over TLS, DNS over QUIC and DNS over HTTP/3 in one binary, so it links a
TRUST=webpki TRANSPORT=tcp-nonblockingobject for the first and colibri'sTRUST=webpki TRANSPORT=quic-nonblocking ROLE=bothobject for the other two. The link failed onch_build, which both objects defined (entry 56). The two objects' headers disagree aboutch_cfgandch_tls, so a program calls each object from a translation unit compiled under that object's defines. A name both objects define fails the link, and a linker that kept one definition would hand one of those units the other object's.- The build record. Its symbol is
ch_build_info_tcp_blocking,ch_build_info_tcp_nonblockingorch_build_info_quic_nonblocking, andbuild.hdefinesch_buildas an object-like macro for the one the defines in force select.ch_build_matches(&ch_build)compiles unchanged in C and throughchapulin.hpp, reads the record of the object whose headers the unit compiles against, and a unit compiled for another transport than its object's fails to link. - Zig. translate-c turns the macro into
pub const ch_build = ch_build_info_tcp_nonblocking;, and Zig 0.16 refuses to evaluate that constant, because its initializer is an extern variable (checked 2026-09-25). An asm label on the declaration is dropped by translate-c, and a macro that dereferences the record's address meets the same refusal. So a Zig program writes the record's own name,c.ch_build_matches(&c.ch_build_info_tcp_nonblocking), and an@cImportmissingCH_TRANSPORT_TCP_NONBLOCKINGdeclares no such name, so that mistake stops the compile. A function name maps without this: Zig evaluatespub const ch_srv_check = ch_srv_check_quic_nonblocking;, because a function is known at compile time.
The audit. Twenty-two objects were built on 2026-09-25 and each pair of different transports was compared by the names
nmlists as defined: the fourTRUSTclient modes andRAND=drbgover each transport,ROLE=serverandROLE=bothover each,EXPORTER=on, andSUITE=aesgcm AES=hwover tcp-nonblocking and QUIC. No object defines a common or weak symbol. Six names were shared:Name Objects that export it Resolution ch_buildevery object named per transport, mapped in build.hch_srv_checkevery ROLE=serverandROLE=bothobjectnamed per transport, mapped in srv.hch_pubkey_from_pemevery TRUST=ca-rsaandTRUST=ca-ecdsaobjectnamed per transport, mapped in x509_ca.hch_drbg_seed, andch_rand_bytesfrom entry 67 onevery RAND=drbgobjectrefused: two RAND=drbgobjects do not linkch_read,ch_write,ch_closeevery TRANSPORT=tcp-blockingandTRANSPORT=tcp-nonblockingobjectrefused: a tcp-blocking object and a tcp-nonblocking object do not link ch_exportEXPORTER=onobjects, the two TCP transports onlyrefused with the pair above ch_srv_checkandch_pubkey_from_pemfollow the record for two reasons. A server pair is the plain case, an HTTP/2 server beside an HTTP/3 one, and two definitions ofch_srv_checkare not interchangeable, since each reads its own transport'sch_cfg. A CA pair is rarer, but the rule is then one sentence: every export that objects of more than one transport carry takes the transport into its name. The Makefile applies it in one list,TRANSPORT_NAMED, soPUBLICandlib-checkread symbol names.The two refusals, and why neither is renamed:
ch_drbg_seed. EachRAND=drbgobject carries its own generator, with its own state, local to the object. Two objects are two generators and two seeds, and an image that hands one seed to both draws the same bytes in each: the same key share in a record session and a QUIC session. Renaming the call would make that image link. ARAND=drbgobject beside aRAND=externone links, because the generator'sch_rand_bytesis local, and it still keeps a second generator the image's hook does not feed, sodocs/porting.mdrefuses that pair in words. Entry 67 exportsch_rand_bytesfrom aRAND=drbgobject, and that pair now has one generator.ch_read,ch_writeandch_close. The tcp-nonblocking transport keeps the connected session's calls under the blocking transport's names (tcp_nonblocking.h), and a tcp-nonblocking object does everything a blocking one does with the caller driving the socket. Renaming them would move the three calls every TLS program links against, for an image that gains nothing by carrying both transports.
Camilo decided on 2026-09-24 that an image defines
ch_rand_bytesandch_assert_failonce, for every chapulin object and every user of chapulin it links, and that no session takes a randomness callback of its own. One entropy source per image is simpler to audit, INV-4 keeps a single hook, and two users already share one object per transport: cocuyo, and a consumer that reaches chapulin through colibri. Soch_rand_bytesmust be safe to call from several threads at once, because a thread-per-core image runs sessions on every core.docs/porting.mdstates that and lists the calls that draw from it.What holds it:
test/lib-pair-check.sh, inmake check, links four pairs and runs each half: webpki tcp-nonblocking client beside the webpki QUICROLE=bothobject (cocuyo's), tcp-nonblocking server beside QUIC server, raw-rsa tcp-blocking beside raw-rsa QUIC, and ca-rsa tcp-blocking beside ca-rsa QUIC. Each half reads its own record, runsch_srv_checkorch_pubkey_from_pemwhere its object has one, and starts a session. The drbg pair and the tcp-blocking and tcp-nonblocking pair must fail to link, and the linker must name each shared name. It reuses the objects thelib-checklegs build and builds two more: 5.2 s with every object built, 8 s with the two to build.lib-check's consumer compiled with the transport moved now must fail to link, naming the other transport's record. It used to read a record that differed, which can no longer happen, so a consumer withCH_PIN_ECDSAmoved reads the difference instead. Every header admits that define, and it changes no symbol name and no hook.inv35-build-record-shared-namegives the QUIC transport's build record the tcp-nonblocking transport's name, andtest/lib-pair-check.shcatches it.
One more change came with it.
test/violations.pyeditsPUBLIC, links, restores it and links again within one second, and make 3.81 compares mtimes to the second, so the second link was skipped and the object kept the edited export list:inv35-build-record-not-exportedfollowed byinv35-build-record-omits-axisfailed the second one's baseline on a correct tree. The link stamp now forces the link when its line differs from the build's, whatever the clocks say.Cost: three symbol names change. A C or C++ consumer changes nothing; a Zig consumer writes
ch_build_info_tcp_nonblockingorch_build_info_quic_nonblockingwhere it wrotech_build, which cocuyo does on three lines and colibri's tests on four.lint-quic-partitionlistsx509_ca.handx509_ca.cbesidebuild.handbuild.c, because a QUIC build compiles the provisioning call under its own name.bench/stack.pyreportsch_pubkey_from_pem_tcp_blocking.make checkgains the pair test and one more consumer build perlib-checkleg.Gain: one image links the objects of two transports, and each unit of it reads the record of its own object.
Rejected: a weak
ch_buildin every object, because the linker keeps one, and the other transport's unit reads it; renaming at the link withobjcopy --redefine-sym, because Apple's toolchain has noobjcopyand the consumer's header must name the symbol anyway; and one object carrying two transports, becauseTRANSPORTis one per object and the two layouts ofch_tlscannot share one name. Entry 56 rejected a symbol name per build because it would encode every axis in the name. The transport is one axis, and it already changes the link line, since each transport exports its own calls.Entry 77 changed one thing here: a
RAND=sessionobject takes a randomness callback per session in itsch_cfgand imports noch_rand_bytes, so an image whose objects are allRAND=sessiondefines none. The hook stays one per image forRAND=externandRAND=drbg. - The build record. Its symbol is
-
Each
TRANSPORTvalue names what TLS runs over and who does the I/O:tcp-blocking,tcp-nonblockingandquic-nonblocking. The axis tooktls,recordandquic. TLS is not a transport: it runs over TCP or inside QUIC.recordnamed the unit the caller passes, a TLS record, and not what sets the mode apart, which is that the caller's code does the I/O. It also gave "record" two meanings, the build record and a transport, soch_build_record(entry 61) read as the struct itself. The values now say both halves:tcp-blocking, the default: TLS records over a byte stream, and chapulin calls the blockingcfg.sendandcfg.recv.tcp-nonblocking: the same records, and the caller passes bytes in and takes bytes out (tcp_nonblocking.h).quic-nonblocking: TLS handshake messages inside QUIC, always driven by the caller (quic.h).
Every value names its I/O style, the QUIC one included, which has no blocking twin, so no reader takes an unmarked name to block. The defines follow the values (
CH_TRANSPORT_TCP_NONBLOCKING,CH_TRANSPORT_QUIC_NONBLOCKING), and so do the symbol names entry 61 gave the transport:ch_build_info_tcp_blocking, named for the typech_build_info, andch_srv_check_tcp_blocking. The old values stop the build with a message naming the new ones, so no consumer gets a different transport without noticing.Gain: a reader learns from the name alone what the object runs over and whether a call can wait on the network.
Rejected: marking only one I/O style (
tlsbesidetls-async, ortls-blockingbesidetls), because the unmarked names then read as the other style;https, because chapulin carries any protocol over TLS, and HTTP is one; and renaming the API, becausech_record_instill takes TLS records. "TCP" names the byte stream every consumer uses; any reliable byte stream works under either TCP value. -
A
TRUST=webpkiclient lists secp256r1 after the two groups it shares and answers a HelloRetryRequest for it, and every server takes secp256r1 last. colibri's webpki client could not connect to nghttpd 1.52.0 on OpenSSL 3.0, which accepts secp256r1 and no other group: the hello listed X25519MLKEM768 and x25519, and the server answered handshake_failure. RFC 9846 §9.1 makes key exchange with secp256r1 a MUST and X25519 a SHOULD (rfc9846.txt:4548-4550).p256_ecdh.[ch]already held a constant-time P-256 key exchange that no library object called.- The client's offer.
supported_groupslists X25519MLKEM768, x25519 and secp256r1, in that order, andkey_sharecarries the hybrid and x25519 shares entry 53 set and no secp256r1 share. §4.3.8 lets the shares be a subset of the list (rfc9846.txt:2161-2165). A server that wants secp256r1 answers with a HelloRetryRequest naming it, which is legal because the hello listed secp256r1 and sent no share for it (rfc9846.txt:2205-2215). The retry hello replaceskey_sharewith one secp256r1 entry, the 65-byte uncompressed point of §4.3.8.2 (rfc9846.txt:1194-1196, 2261-2275), and keeps the rest of the first hello, a cookie beside it when the retry sent one. The ServerHello must then select secp256r1 (rfc9846.txt:2233-2238). A retry naming the hybrid or x25519 is still illegal_parameter, as entry 53 made it, and so is a ServerHello that selects secp256r1 when no retry named it.ch_cfg.require_pqkeeps its meaning: the hello lists and shares the hybrid alone, so a retry naming secp256r1 names a group the hello never listed and is refused. - The client's key.
handshake_groups.cdraws the P-256 scalar throughch_rand_byteswhen a retry names secp256r1 and at no other time, a new INV-4 site. A draw outside [1, n-1] happens with probability below 2^-32 and is drawn again, up to four draws (P256_ECDH_DRAWS); four refusals in a row from a working generator happen with probability below 2^-128, so CH_ASSERT treats them as a hook that wrote nothing, as every draw site treats an all-zero draw. The retry wipes the first hello's x25519 and ML-KEM key pairs, which no later message uses. The scalar and the point live inhandshake_stateuntil the ServerHello; the call that computes the secret wipes the scalar on both exits (INV-17). - The checks and the secret.
p256_ecdhrefuses a server point whose form byte is not 0x04, whose coordinates are not below p, or which is not on the curve, the three steps §4.3.8.2 lists (rfc9846.txt:2277-2286); the point at infinity has no 65-byte encoding and fails the curve equation. The parser refuses any length but 65. Each refusal is illegal_parameter. The shared secret is the 32-byte X coordinate with no leading zero dropped (§7.4.2, rfc9846.txt:4266-4276), and the key schedule extracts from it as it extracts from an x25519 secret.ch_tls.groupreportsCH_GROUP_SECP256R1. - The server. Every server role,
ROLE=serverandROLE=bothover all three transports, holds secp256r1 as its third group and takes it last: X25519MLKEM768 when the client lists it, then x25519, then secp256r1. A client that lists secp256r1 alone gets it, in one round trip when it shared it and after a HelloRetryRequest when it did not. A client that lists x25519 or the hybrid never gets P-256.srv_kex_sharechecks the client's point before it draws the server's P-256 key, a new INV-4 site insrv_kex.c, and a refused point is illegal_parameter before any ServerHello goes out.srv_kex_secretwipes the scalar on both exits. The PQ-first rule of entry 54 stands unchanged.
Cost: the webpki hello grows by the 2 bytes of the third NamedGroup, so
CH_HELLO_MAXand the webpkiCH_TX_STAGEgo from 2,394 to 2,396, and the QUIC one from 2,648 to 2,650. The retry hello to secp256r1 needs no term: its one 69-byteKeyShareEntryreplaces the 1,256 bytes of the first hello's two, so it is 1,187 bytes shorter than a cookie retry.bench/sram.shmeasures the webpkich_tlsat 3,048 bytes on arm64, against 3,040, andch_connect's stack peak at 16,528 bytes, against 16,416, because the handshake state holds the retry group, the 65-byte point and the 32-byte scalar;ch_srv_acceptpeaks at 10,336 bytes, against 10,304, because it holds the scalar. Every webpki object now packagesp256_ecdh.c,p256_point.c,p256_scalar.candp256_field.c, and every server object addsp256_ecdh.cto the three it already carried for its ECDSA signer. A server that holds secp256r1 alone costs this client one round trip, and one P-256 scalar multiplication takes 1,228 microseconds against x25519's 953, measured on an M1 Pro (bench/notes-primitives.md). Gain: the client reaches a server that holds only the group §9.1 requires, and the server serves a client that offers only that group.Two alternatives were considered and rejected. Sending a secp256r1 share in the first hello saves that round trip, and puts 69 bytes and a P-256 key generation in every hello, a key pair nearly every handshake discards, which is the trade entry 39 declined for the hybrid. Preferring secp256r1 to x25519 on the server gives the slower group to every client that lists both, OpenSSL's default list and every webpki client among them.
The earlier text overclaimed.
docs/server.md's list of §9.1 residuals, written on 2026-09-18, said the server held both secp256r1 and X25519, while its profile table andsrv_parser.hsaid it held X25519MLKEM768 and x25519 alone, and the code agreed with the table. The server held no secp256r1 until this entry, and that sentence is true from this entry on. - The client's offer.
-
A
TRUST=webpkiQUIC client takes SPKI pins with the meaning they have over TCP. cocuyo, a DNS resolver, runs DNS over QUIC (RFC 9250) through colibri with chapulin's QUIC object as its TLS. RFC 9250 §5.1 gives a DNS-over-QUIC client the authentication requirements RFC 7858 and RFC 8310 give a DNS-over-TLS one. RFC 8310 §6.3 lists an SPKI pin set and an address as one way to authenticate a server, and §6.4 has a client configured with a name and pins require both (rfc8310.txt:805-814). cocuyo's DNS-over-TLS path already uses all three configurations entry 49 allows: a hostname alone, pins alone, and both.ch_quic_initrefused pins and required a hostname, so its QUIC path could use the first alone.- The rules.
quic_config.ccallswebpki_cfg_ok, the functionch_connectandch_record_initcall, where it kept its own copy of the anchor, hostname and clock rules. So a configuration is valid over QUIC exactly when it is valid over TCP, and RFC 9001's three rules stay on top: transport parameters,on_level_ready, and an ALPN offer that names a protocol.CH_SPKI_PIN_MAXbounds the pins on both transports. - The handshake. Nothing else changes. The ClientHello builder,
the EncryptedExtensions parser and
webpki_server_keywere already shared with TCP, and the configuration rule alone kept pins out. The hello offersserver_certificate_typeas over TCP, and the Certificate is judged as over TCP: a raw key by the pins alone, a chain beside anchors by the walk, the name and a pin on the path the walk verified, and a chain answering pins alone refused with unsupported_certificate. - The ticket binding.
webpki_ticket_config_hashhashes the pins beside the hostname and the anchors, andch_quic_inittakes that hash when the session starts, asch_connectdoes. So every ticket a pinned QUIC session receives is bound to its pins, andch_quic_initrefuses the ticket under another pin set or none. A resumed handshake sends no certificate, so no pin is checked in it; the binding is what holds it to the pins that judged the first session's key. - What is tested, and what is not.
bin/quic_loop_webpkiruns the three configurations against this tree's QUIC server, which sends a chain and never a raw public key. A hostname alone and a hostname with pins pass end to end. Pins alone are tested for their offer and for the refusal of that server's chain. A raw public key accepted over QUIC is not tested: no QUIC server this tree runs sends one, and OpenSSL 3.6.4'ss_serverhas-enable_server_rpkand no QUIC. The raw-key rule is the TCP one, tested end to end over TCP.
Cost: the 7-byte
server_certificate_typeoffer now goes out over QUIC.CH_HELLO_MAXalready counted it there, because the builder is shared, so the QUIC webpkiCH_TX_STAGEstays 2,650 bytes, andch_tlsandch_quickeep their sizes, read from the build record of each object on arm64 macOS: 3,200 and 4,808 bytes in the client object, 3,408 and 5,056 underROLE=both.quic_config.cdrops its second copy of the ALPN name rules in this build, becausewebpki_cfg_okruns them, and its text falls from 452 to 144 bytes. Gain: a DNS-over-QUIC client uses the configurations its DNS-over-TLS path uses, and a pin means one thing on every transport.A QUIC copy of the pin rules was considered and rejected: two copies of one rule can drift apart, and a copy has no reason to exist when the handshake code that reads the configuration is shared.
- The rules.
-
SPKI pins alone accept a certificate chain whose leaf key a pin names. Entry 49 made pins alone RFC 8310's "SPKI + IP" profile (
rfc8310.txt:683-685) and read it as a raw public key alone: an X.509 answer was refused with unsupported_certificate. cocuyo's DNS-over-QUIC targets, AdGuard and NextDNS, send ordinary ECDSA chains that end at USERTrust ECC, and cocuyo reaches them by address and pin. RFC 7858 §4.2 has the client hash the keys of the validated server chain, or the raw key the server sent, and match a pin (rfc7858.txt:434-440). With no anchor, no clock and no hostname there is no validated chain, so this entry fixes what pins alone check. Pins alone now mean that the server proves it holds a pinned leaf key, sent raw or inside a certificate.- The offer. Pins alone offer RawPublicKey, then X509. With anchors the offer already listed both.
- The rule. A chain under pins alone passes when one pin is the SHA-256 of its leaf's SubjectPublicKeyInfo, and CertificateVerify then verifies under the leaf's key. The certificate is public; the signature is what proves the server holds the key.
- What is not read. The chain above the leaf, the dates and the
names, because there is no anchor, clock or hostname to check them
against. Every entry is framed by the walk's own framing, but only
the leaf is kept, so pins alone take any number of entries the
message holds, where the walk takes
CH_WEBPKI_FLIGHT_ENTRIES. For the same reason every entry, the leaf included, may takeCH_WEBPKI_LEAF_PIN_CERT_MAXbytes, 16375, the largest certificate one entry carries in the 0x4000-byte message body every handshake reader admits, where the walk holds each certificate toCH_WEBPKI_CERT_MAX, 3072 bytes. The walk parses up to three certificates and verifies their signatures; pins alone parse only the leaf, and only as far as its key. This entry first kept both of the walk's caps, which refused the QUIC Interop Runner's amplificationlimit chain, a leaf of 5,514 bytes under eight intermediates, though no entry past the leaf is read and the twenty 250-byte subjectAltName entries that make the leaf that large are not read either. The webpki_cert_key proof covers the reader to one byte past the new cap. The receive buffer bounds the message as well, so a caller whose server sends a Certificate message larger thanCH_TRUST_MIN_RXBUFsizescfg.buf_lento hold all of it. The leaf is read only as far as its key, bywebpki_cert.c's own field readers. The fields after the key are skipped as whole TLVs, so each container still ends where its fields end (INV-25), and their content is not read. INV-20's containment holds: the reader has one caller, that caller has one caller inhandshake_auth.c, and neither reads a clock. - Other keys do not count. A pin that names only an intermediate or a root key is refused with bad_certificate, the alert a pin miss gets. With no name to check, a pin on a CA key would accept every certificate that CA issued, to anyone.
- Unchanged. A raw public key, pins beside anchors, and the ticket binding, which holds a resumed session to the pins that judged the first one.
Cost: a pinned leaf key breaks when the operator rotates it, and a leaf rotates more often than a CA key, so a caller pins a backup key too, as RFC 7858 §4.2 asks. Under pins alone,
webpki_cert.c's readers now parse peer input, where onlywebpki_spki.cparsed it before. The pins-alone hello grows by 1 byte, the X509 entry, inside theCH_HELLO_MAXevery webpki build already counted, soCH_TX_STAGE,ch_tlsandch_quickeep their sizes. Gain: pins alone reach a server whose leaf key the caller knows, whether it sends the key raw or in a certificate, over both TCP transports and QUIC.What changes in entry 49: its "SPKI + IP" profile no longer means a server with no public certificate alone. Its rejection of matching a pin on the leaf alone stands where anchors are set, because there a pin on any key of the validated path counts. Under pins alone the leaf is the one key the signature proves, so it is the one key a pin may name.
Two alternatives were considered and rejected. Accepting a pin on any key of the chain and checking the signatures from the leaf up to it, with the pinned certificate as an anchor, accepts every leaf that CA issued, to anyone, when no name is checked; a caller who pins a CA sets anchors and a hostname instead. Refusing X.509 under pins alone, as entry 49 did, leaves a DNS client unable to reach a public resolver by address and pin.
-
ch_drbg_seedtakes a seed of any length from 32 bytes up and hashes it into the generator key (#164). docs/entropy.md told aRAND=drbgintegrator to concatenate several entropy sources and hash them withsha256_ofinto the 32 bytesch_drbg_seedtook. The packaged object exportsch_drbg_seedand keepssha256_oflocal, so that recipe did not link, and nothing compiled it.- The call.
ch_drbg_seed(const uint8_t *seed, size_t seed_len). The generator key is the SHA-256 of allseed_lenbytes. The caller concatenates its sources into one buffer and passes the whole buffer, and needs no hash of its own. - The floor. A seed shorter than
CH_DRBG_SEED_MIN, 32 bytes, is a programmer error, andCH_ASSERTfires. 32 is the key length. A seed shorter than the key cannot carry a full key of entropy. The floor does not measure entropy: a 32-byte counter passes it. - Reseeding. A second call replaces the state, as before. The reseed recipe concatenates fresh bytes with output the generator just drew.
- Wipes. The digest goes straight into the generator key, so no stack copy of it exists, and the SHA-256 context is wiped before the call returns. The caller wipes its own buffer.
- The check.
test/entropy_recipe.cis docs/entropy.md's boot-seed recipe in C.make lib-check RAND=drbglinks it against the packaged object and runs it, so a recipe that calls a function the object keeps local fails there.
Cost: an API break. Every caller changes its call, colibri's test endpoints among them, and the known answers in
bin/drbg_testchanged with the key.sha256.cwas already in every object, so the object gains no module;drbg.c's text grows from 240 to 313 bytes andch_drbg_seed's frame from 8 to 128 bytes on Cortex-M3 (Arm GNU gcc 15.3,-Os), measured with-fstack-usage. The call runs at boot, beside no handshake, so no stack peak bench/sram.sh reports moves. Under the two riscv32 gcc specs,drbg.c's branch ceiling rises from 9 to 10: the test ofseed_lenand a stack-protector canary replace the copy loop's back edge, and neither reads a seed byte. Gain: the documented recipe links and runs, and a part with no hash of its own seeds from several sources in one call.What this entry does not change: rewriting the seed file and both reseed recipes read generator output, which takes a
ch_rand_bytescall. The packaged object keptch_rand_byteslocal, so an image that linked the object could not follow those three steps. Entry 67 exports it.Two alternatives were considered and rejected. Exporting a
ch_sha256underRAND=drbgadds a sixth public call, and adds it because one recipe needed it, not because the API calls for a hash. Fixing only the documentation leaves a part with no SHA-256 of its own with nothing to hash with. - The call.
-
A
RAND=drbgobject exportsch_rand_bytes. docs/entropy.md has the image read generator output three times: it rewrites the seed file at boot, it reseeds when entropy arrives later, and a part with a TRNG reseeds on a schedule. The packaged object keptch_rand_byteslocal, so an image that linked it had no call that returned generator output, and entry 66 left those steps as a known gap. Camilo chose on 2026-09-26 to export the call.- The export.
PUBLIC_RANDisch_drbg_seed ch_rand_bytes, so aRAND=drbgobject exports six calls, and seven under a ca mode. - An image that also defines
ch_rand_bytes. Before, the object used its own generator and the image's definition was never called. Now the link fails with a duplicate symbol, which names the mistake at build time. - Pairs of objects (entry 61). Two
RAND=drbgobjects still do not link, and the linker now namesch_rand_bytesbesidech_drbg_seed. ARAND=drbgobject beside aRAND=externone links when the image defines noch_rand_bytes, and theRAND=externobject then draws from the other's generator. The image has one generator, as entry 61's rule of one entropy source per image asks. That generator is single-task (drbg.h), so an image that runs sessions on several threads keepsRAND=externeverywhere. - The check.
test/entropy_recipe.cnow rewrites its seed file and reseeds withch_rand_bytesbefore its handshake, andlib-check RAND=drbglinks it against the object, so ach_rand_bytesmade local again fails there.
Cost: one more public call, and a
RAND=drbgimage's hook has a name it must not also define. Gain: every step docs/entropy.md gives can be followed with the packaged object alone.The alternative, rewriting the three steps to need no generator output, was rejected: it drops the seed-file rewrite, the step that keeps two devices imaged from one flash from replaying one stream.
- The export.
-
A
SUITE=aesgcmbuild runs both AES-GCM suites onAES=extern, and the build states the peripheral's timing withCH_AES_EXTERN_CONSTANT_TIME(#177). Entry 58 held the suites toAES=hw, so a part with an AES peripheral and no AES instructions could not offer them.AES=externalready left QUIC's public keys to the image'sch_aes_block. Camilo decided on 2026-09-26 to let that hook take traffic keys too, over both TCP transports and QUIC.- The hook takes a key length. It is
ch_aes_block(key, key_len, in, out), withkey_lenAES_128_KEYorAES_256_KEY.aes_extern.cgainsaes_expand_round_keys_256andaes_cipher_block_256: the first stores the 32 key bytes where the round keys go and zeros the rest, and the second passes them to the hook with length 32, as the AES-128 pair does with 16. SoSUITE=aesgcmholdsTLS_AES_128_GCM_SHA256andTLS_AES_256_GCM_SHA384under everyAESvalue it takes. - A separate timing flag.
ct.hadmits-DCH_SUITE_AES_GCMonAES=hwwithCH_NATIVE_AES, onAES=runtimewith the same statement since entry 81, or onAES=externwithCH_AES_EXTERN_CONSTANT_TIME, and refuses every other pairing,AES=softincluded.CH_AES_EXTERN_CONSTANT_TIMEis the firmware author's statement, from the vendor, that the peripheral behindch_aes_blockruns in constant time for 16-byte and 32-byte keys. The Makefile never writes it, as it never writesCH_NATIVE_AES, andlint-trust-separationbans both from every suite object's defines. The test binaries state it on their own lines. The build record leaves it out for entry 56's reason: no public layout or bound reads it. - What the flag covers. The AES blocks, and nothing else. Under
AES=extern, GHASH runs ongcm.c's portable multiply: 128 masked steps per block, with no table, no multiply instruction and no branch on a subkey bit, andlint-wide-multiplyholds its branch count. So the flag claims nothing about GHASH. It claims nothing about what the hook or the peripheral keeps after a call either, such as a key register or a cached expansion;aes_block.hleaves that to the image, and this tree wipes its own copies, the stored key among them, where it wiped the round keys before. And no mechanism in this tree can observe a peripheral's timing. The statement is the whole of the claim, asCH_NATIVE_AESis for the instructions. - The rename.
quic_aes_extern.cis nowaes_extern.c, because a suite build compiles it over TCP, so it is no longer QUIC-only (INV-27).quic_aes_soft.ckeeps its prefix: its S-box is indexed with the key, so a suite build refuses it, and the one suite build that holds it, anAES=runtimeQUIC object, runs QUIC's public keys alone on it (entry 81). - The tests. Every
AES=externtest binary linkstest/aes_extern_hook.cas the hook:quic_aes_soft.c's cipher under other names, for both key lengths, which aborts on any other length. Over it run FIPS 197, SP 800-38D and RFC 9001 Appendix A, RFC 8448's record, the QUIC suite computation, both loop tests, the Wycheproof AES-GCM suite, the AES rows of the Lean differential, and e2e's client and server against OpenSSL under each suite. None needs an AES instruction, so every host runs them. Theaes_externproof holds the four entries to the hook's contract over a stub of it.
Cost:
- The hook's signature changes. An image that defined the old three-argument hook still links, because C checks no signature at link time, and that hook then takes the key length for its input pointer. No known image defines the hook; colibri and cocuyo do not.
- A part whose peripheral has no AES-256 cannot build the suite with
AES=extern, and keeps ChaCha20. - A second timing statement a firmware author must make and answer
for. It is as weak as
CH_NATIVE_AES: this tree cannot check it. - A traffic key now leaves code this tree compiles. Before this entry, every suite build ran its traffic keys on instructions the compiler emitted from this tree's sources. Now a suite build can hand them to a function the image supplies, whose code nobody here reads.
make checkbuilds and runs five more binaries, a third Wycheproof leg and one morelib-checkobject;make diffruns one more binary, andtest/e2e.shten more legs.
Gain: a part with an AES peripheral offers the suite RFC 9846 §9.1 makes mandatory, and the SHOULD one beside it, over every transport, and the suite build is no longer host-only.
Three alternatives were considered and rejected.
- AES-128 alone under
AES=extern. It keeps the hook's signature and serves the mandatory suite. It was rejected becauseSUITE=aesgcmwould then name two suites underAES=hwand one underAES=extern, so which suites a build holds would depend on a second axis, and ach_srv_cfg.cipher_suitesorder that namesTLS_AES_256_GCM_SHA384would be valid in one build and refused in the other. - A second hook for AES-256.
ch_aes_block_256beside the 16-byte hook would keep an old definition linking. It was rejected because it doubles what the image implements and what the INV-26 rule matches, while a part with AES-256 serves both lengths from one peripheral driver; and no known image defines the old hook, so its signature protects nobody. - Reusing
CH_NATIVE_AES. One define for both values. It was rejected becauseCH_NATIVE_AESstates the timing of the AES instructions and of the carry-less multiply, and anAES=externobject runs neither: its blocks are the peripheral's and its GHASH is the portable multiply. One define would let a statement about one piece of silicon be read as a statement about another.test/quic-builds.shrefuses each flag on the other's value.
- The hook takes a key length. It is
-
A Zig project depends on chapulin as a package:
build.zigbuilds the objectmake libbuilds, andmake lint-zig-buildholds the two builds to each other. colibri linked abin/*.ofrom a checkout its caller had built with make. Camilo decided on 2026-09-26 that the Makefile stays the source of truth, that a Zig build beside it must produce the same object, and that a check inmake checkcompares the two. The Zig version is colibri's, 0.16.0.- The options.
b.dependency("chapulin", .{ ... })takes the Makefile's variables under their own names and values:TRANSPORT,ROLE,TRUST,SUITE,AES,RAND,EXPORTER,KEYLOG,KEX,X25519andWIDEMUL. The three hardware statements a builder adds to make'sCFLAGS,CH_NATIVE_AES,CH_AES_EXTERN_CONSTANT_TIMEandCH_NATIVE_MUL128, are options of their own that default off, so a build that needs one and lacks it stops atct.h, as make's does.build.zigrepeats each axis block of the Makefile as one function, and refuses the same combinations with the same words.AES=hwadds the target's AES and carry-less multiply features, asAES_HW_CFLAGSadds flags for cc. - What a dependent gets. The named lazy path
chapulin.o, the localized object, andinclude, the directory of the headers. The dependent compiles the headers under the object's defines, as a C program does, and callsch_build_matchesonce (entry 56). There is no static library: the object links with oneaddObjectFilecall, and Zig 0.16's archiver leaves an odd-sized last member unpadded, which Apple's nm and llvm-ar refuse to read. Entry 70 adds the modulechapulin, so a Zig dependent no longer writes the defines. - One object.
addObjectover every source partially links them with Zig's own linker, for ELF and Mach-O, asld -rdoes for make. - The localizer.
tools/localize_symbols.zigdoes to that object whatobjcopy -Gandnmedit -sdo to make's: every defined symbol but the public names becomes local. In ELF32 and ELF64 of either byte order it sets STB_LOCAL, moves the locals before the globals, rewrites the symbol table's sh_info, and renumbers the symbol index in every relocation, section group and SHT_SYMTAB_SHNDX entry. It clears the sh_link of SHT_LLVM_ADDRSIG, which marks that table stale the wayld -rdoes, and lld then ignores it. In 64-bit Mach-O it clears N_EXT and N_PEXT, reorders the table into LC_DYSYMTAB's three ranges, rewrites the ranges, renumbers every external relocation and indirect symbol entry, and rewrites a localized variable's N_GSYM debugging entry as nmedit does. No section moves and no size changes. - What it refuses. Anything it cannot rewrite in full, rather than a guess: another format, a section or load command that may hold symbol indices it does not know, a common symbol, a name to keep that the object does not define, and a MIPS GOT16 or CALL16 relocation against a symbol it would make local. The MIPS ABI reads those two differently against a local symbol, so localizing position independent MIPS code changes what it computes; objcopy does that without a word, and lld only warns. A MIPS object compiled without PIC has neither relocation.
- Flags. The defines,
-std=c11and-O2are make's, and so are the warnings. Zig adds-DNDEBUG, which nothing here reads,-fPICand a kept frame pointer. It turns on no stack protector, which Apple's clang and Ubuntu's gcc turn on by default, and on Linux no_FORTIFY_SOURCE, which Ubuntu's gcc defines by default. The comment abovecflagsinbuild.ziglists each and why it does not change what the object computes or exports. - The check (INV-36).
test/zig-build-check.shcopies exactly the filesbuild.zig.zon's.pathsnames, which is what a dependent receives, after requiring that list to name every root source and header git tracks. It builds the default object and colibri's four both ways, requires the same sources, defines and exports, linkstest/build_test.cagainst each Zig object under make's defines, and links the tcp-nonblockingROLE=bothand QUIC objects into one image and runs it. check-slow repeats the comparison over everylib-checkleg's configuration.test/localize-check.shcompares the localizer withllvm-objcopy -Gon nine ELF targets, big-endian mips32r2 among them, and withnmedit -sas well on two Mach-O targets, and links every result. Five mutants breakbuild.zigor the localizer. - The pin.
tools/toolchain.envpinsZIG_VERSIONand the hash of the x86_64 Linux tarball,.github/actions/install-zigchecks the download against it, the check job and the nightly's violation job install it, andlint-toolchainchecks the version on every machine.
Cost:
- The Makefile's axis logic exists twice, and a change to an axis is made in both files. The check catches a change made in one of them in each configuration it builds; a combination it does not build is caught when a dependent builds it.
make checkneeds zig.lint-zig-buildtakes 8 s with every object built and 90 s with none, and check-slow adds 37 s.- About 1,300 lines of Zig,
build.zigand the localizer, that this tree reads and tests as it does its C.
Gain: a Zig project builds chapulin with
zig build, for any target Zig compiles C for, with no make, no binutils and no Xcode tools, and links an object that exports whatlib-checkholds make's to.Two alternatives were considered and rejected.
- Host
ld -randobjcopyornmedit, run frombuild.zig. It would repeat the Makefile's recipe, and it would tie a Zig build to the host's tools: macOS ships no objcopy,nmeditreads no ELF, and a host linker partially links only its own target's objects, so a cross build would need a binutils per target.zig objcopyin 0.16 has no option that keeps some globals and localizes the rest. - A plain static library whose internal symbols stay global. It
is what Zig builds with no tool of ours. But every internal name is
then global, so two objects of different transports define the same
names and one image cannot link both (entry 61), an application's own
name can collide with one, and the export list
lib-checkholds means nothing for that object.
- The options.
-
The Zig package exports a module of the object's API, which translate-c makes from the public headers under the defines
build.zigcompiled that object with. The public headers change shape with the object's defines: fields ofch_cfgandch_tlsappear and disappear, and array sizes change. A Zig program used to write the define list for its own@cImportby hand. A wrong list compiles, links and corrupts memory at run time, and onlych_build_matchesat startup catches it. colibri's list was already wrong: it left outHKDF_LABEL_MAX=32, whichEXPORTER=onadds. colibri also links objects of two transports into one image, and one@cImportcannot includetls.htwice under two sets of defines. Camilo decided on 2026-09-26 that the package exports the module, and thatch_build_matchesstays as the run-time check.-
The module.
build.zigwriteschapulin.h, which includes the headers that declare what the object exports and imports. translate-c translates it for the object's target under every-Din the flag list the sources compile with. Both take their defines from that one list, so the object and the module cannot differ. Each dependency gives a module of its own, so colibri's image of two transports imports two modules and links two objects:const h2 = b.dependency("chapulin", .{ .target = target, .RAND = .@"extern", .TRANSPORT = .@"tcp-nonblocking", .ROLE = .both, .TRUST = .webpki, .EXPORTER = .on, }); const quic = b.dependency("chapulin", .{ .target = target, .RAND = .@"extern", .TRANSPORT = .@"quic-nonblocking", .ROLE = .both, .TRUST = .webpki, .SUITE = .aesgcm, .AES = .hw, .KEYLOG = .on, .CH_NATIVE_AES = true, }); module.addImport("chapulin_h2", h2.module("chapulin")); module.addImport("chapulin_quic", quic.module("chapulin")); module.addObjectFile(h2.namedLazyPath("chapulin.o")); module.addObjectFile(quic.namedLazyPath("chapulin.o"));
-
The name. The module is
chapulin, the package's name. A Zig package names its main module after itself, and this module is the object's API for Zig, aschapulin.hppis its API for C++. -
The headers. The
declarationstable inbuild.zignames the header that declares each name an object can export and each hook it can import (docs/porting.md).chapulin.hincludes the headers of the names one object exports and imports, and a name the table lacks stops the build. So a QUIC client's module declares noch_connect, which that object does not define, and a CA mode's module declaresch_pubkey_from_pem. -
What translate-c handles. Zig 0.16.0's translate-c translates every public struct, union, enum constant, function pointer field and call. It turns
_Static_assertinto a comptime check. It turns the static inlinech_build_matchesinto a Zig function whosesizeofterms become@sizeOfof the translated types, so the comparison checks the layout the Zig program itself uses. It makes the incomplete AES key types (aes_public_key,aes_traffic_keyandaes_key_schedule) opaque, so a Zig program cannot build one, as a C file cannot (INV-26). No public header has a bit field or a flexible array member. The macros that give a call its transport's symbol name,ch_srv_checkandch_pubkey_from_pem, become constants that name an extern function, and a Zig program callsc.ch_srv_checkas a C program does. Two of chapulin's macros do not work.CH_ASSERTuses__FILE__, which translate-c does not translate, and no caller needs it.ch_buildbecomes a constant whose value is an extern variable, and Zig refuses to evaluate it, so a program names the transport's record,&c.ch_build_info_quic_nonblocking. The module declares only its own transport's record, so a wrong name stops the compile. -
The check (INV-36).
test/zig-build-check.shcopiestest/zig-consumer, a Zig project that depends on the staged package as colibri does. For each configuration it buildsmatches.zig, which imports the module and links the object. The program compiles only when the module declares every name the object exports, and it requiresch_build_matchesto return 1. For colibri's two objects it buildspair.zig, which imports both modules, links both objects and starts a client on each.inv36-zig-module-drops-definetranslates the module without-DCH_EXPORTER, andinv36-zig-module-misses-headernames the wrong header for the tcp-nonblocking server's calls. The check catches both. -
The lengths the headers name. A module that declares every export can still lack a constant a public header sizes a field by. colibri's QUIC server could not name the ticket key's length, which
srv_cfg.hcited and onlysrv_ticket.hdefined, until64e2f25moved it intosrv_cfg.h.tools/public-constants.pynow preprocesses each object's headers under its defines, comments kept, and lists every name shaped like a length or a cap in a comment of a public header. It fails when the consumer cannot see one, andmatches.zigdeclares and evaluates each. Across the 21 configurations it found five more comments naming four internal constants, and those comments now give the number instead.
Cost:
- It serves Zig alone. A C program, firmware among them, still writes
its own defines, and only
ch_build_matchescatches a mistake. - The module is translate-c's output, and translate-c changes between Zig releases: 0.16.0's is built on Aro. Zig is pinned at 0.16.0, so translate-c changes only when the pin moves, and the pull request that moves it runs the check against the new output.
- From nothing built,
lint-zig-buildtakes about 25 s longer: 66 s against 39 s before, and 103 s against 81 s, in two pairs of runs on an M-series Mac with a load average above 18. With everything built,test/zig-build-check.shtakes 3.5 s where it took 3.0 s.
Gain: a Zig program names no define. Its types come from the defines the object compiled with, so no program keeps a list that can go out of date as colibri's did, and each object in an image of two transports has its own module.
Two alternatives were considered and set aside.
- A generated configuration header every consumer includes.
make libandbuild.zigwould write a header that defines the object's defines, and every public header would include it first. It would serve C firmware as well as Zig. But it changes the public headers and the Makefile, so it waits until a C consumer asks for it. - The define list as a file the package exports. A Zig program
would read it and pass each define to its own
@cImportor translate-c step. The program still does that work, and every consumer writes it again.
-
-
A TCP build sets how much plaintext one outgoing record carries,
TX_RECORD, from 512 to 16384 bytes, and the default stays 512. stompy uploads 5 MiB log segments to S3 over aTRUST=webpki TRANSPORT=tcp-nonblocking ROLE=bothobject. At 512 bytes a record, one segment takes 10,240 records and 10,240 send callbacks, and the 22 bytes each record adds come to 4.3% of the data. At 16384 bytes it takes 320 records, and the overhead is 0.13%. Camilo approved the option on 2026-09-26 at stompy's request. This entry amends entry 22.- The value.
CH_TX_PT(cfg.h) is the most plaintext one outgoing record carries. It sits under#ifndef, and the Makefile'sTX_RECORD=Nwrites-DCH_TX_PT=Ninto the object, andbuild.zig'sTX_RECORDoption does the same.ch_writeand a server's sealed flight still put the smaller ofCH_TX_PTand the peer'srecord_size_limit(RFC 8449) into each record, so a peer that asks for smaller records gets them. - The range. 512 is the floor: nothing needs smaller records, and
every test and proof in this tree runs at 512 or above. 16384 is
the ceiling, the most plaintext RFC 9846 §5.1 lets one record carry
(
rfc9846.txt:3514-3516). The Makefile andbuild.zigtake a decimal integer with no leading zero, because C reads a leading zero as octal.cfg.hrefuses a value outside the range for a firmware tree that builds these sources its own way. - The staging array.
session.hnames each build's hello literalCH_TX_HELLO, andCH_TX_STAGEis the larger of it and one sealed record,CH_TX_PT + 1 + AEAD_TAG. At the default the hello still wins in every build, so every defaultch_tlskeeps its size byte for byte.handshake.candquic.ccheckCH_HELLO_MAXagainstCH_TX_HELLO, notCH_TX_STAGE, because a raisedCH_TX_PTcan make the array larger than any hello, and a stale literal would then pass. The 2^14 assertion moves toCH_TX_HELLOtoo: the hello ships as one plaintext record, while a sealed record's body is ciphertext, which §5.2 caps at 2^14 + 256 bytes (rfc9846.txt:3595-3596). - The Certificate. A server streams its Certificate through
srv_frag, which lives onsrv_send_certificate's stack frame. Its buffer staysSRV_FRAG_MAX, 512 bytes, whateverCH_TX_PTis, so a Certificate still goes out in records of at most 512 bytes. A buffer of 16,384 bytes would be four times the 4,096-byte frame budget of aTRUST=webpkiobject and more than six times the 2,560 bytes of every other (INV-19). A Certificate goes out once per full handshake, so larger fragments would save a few records per connection. - The receive side. The option changes what this endpoint sends
and nothing else. What it receives is bounded by its own
cfg.buf_len, which it advertises asrecord_size_limit, as before. A peer of aTX_RECORD=16384object receives records of 16,384 bytes only when its own buffer holds 16,406 bytes: the record header, the plaintext, the inner content type and the tag. - The check.
bin/webpki_loop_tx_recordistest/webpki_loop_test.catTX_RECORD=16384: a write ofCH_TX_PTbytes goes out as one record andCH_TX_PT + 1bytes as two, a client whose buffer advertises a smaller limit gets records of that limit, and 32,868 bytes move each way between the object's two drivers. The handshakes it shares with the default loop count the Certificate's three records, which holdssrv_fragatSRV_FRAG_MAX.test/tx-record-builds.shcompiles each edge of the range in the headers, runs it through make andbuild.zig, and checks whereCH_TX_STAGEturns from the hello to the sealed record.make checkruns both, withlib-checkandlint-stack, on theTX_RECORD=16384object, and check-slow's Zig roster builds that object both ways. Five mutants intest/violations/break the rules:inv19-srv-frag-sized-by-tx-record, which the loop catches, andinv38-tx-record-past-2-14,inv38-tx-record-quic-accepted,inv38-tx-record-makefile-quic-acceptedandinv36-zig-build-tx-record-past-2-14, which the script catches.
Cost:
ch_tlsgrows byCH_TX_PT + 17bytes less the build's hello literal, rounded to the struct's alignment. Measured with asizeofprobe under Apple clang 21 on arm64, stompy's object,TRUST=webpki TRANSPORT=tcp-nonblocking ROLE=both, goes from 3,264 bytes to 17,264 atTX_RECORD=16384, and itsch_recordfrom 4,480 to 18,480. The formula gives 16,401 - 2,396 = 14,005 bytes and the struct grew by 14,000, becausetxis its last field and 5 of those bytes fill the padding the old struct already carried at its end. The default classic client would go from 1,144 bytes to 16,928.- A device cannot spare that, which is why the default stays 512 and the figures in docs/performance.md stay the default build's.
- The build record holds
CH_TX_STAGEand notCH_TX_PT, soch_build_matchestells an object and a consumer apart only where the two values give different arrays. A value whose sealed record stays under the hello, such as 1,024 in aKEX=pqbuild, changes no layout, and the record cannot see it.
A
TRANSPORT=quic-nonblockingobject refuses the option, in the Makefile, inbuild.zigand insession.h. RFC 9001 §4.1.3 removes the record layer (rfc9001.txt:462-464), so a QUIC build seals no record and stages its hello alone, and the one place it readsCH_TX_PTissrv_out_limit, which onlysrv_fragreads there andSRV_FRAG_MAXcaps anyway. A value there would change nothing, so the build refuses it rather than accept a setting it ignores.Two alternatives were considered and set aside.
- A transmit buffer the caller supplies in
ch_cfg, the waycfg.bufholds received records. One object could then send at either size. But every build'sch_cfgwould carry a pointer and a length, every init call a second buffer rule, and no consumer needs one object at two sizes. - Sealing straight from the caller's bytes, a scatter-gather
rec_sealthat writes no staged copy. It would costch_tlsnothing at any size, and it is a second sealing entry point, which INV-1 exists to prevent.
- The value.
-
Where a record ends, how much plaintext a send buffer holds, and a ticket's age are computed in C, not by the caller. Camilo approved a Zig API that forwards to C and holds no TLS logic, and on 2026-09-26 he decided that three computations colibri makes in Zig move into C first:
whole_record_len,sealable_lenandowed_len_maxin its record adapter, and the obfuscated ticket age in its resumption code. Each one is a TLS rule, so a caller that computes it keeps a copy of the rule that no proof and no test here checks.ch_record_whole_len(p, n)(tcp_nonblocking.h,tcp_nonblocking_frame.c, everyTRANSPORT=tcp-nonblockingobject in either role). It answers the length of the record at the front ofp, header included, or 0 whilepholds less than that record. A caller passes that many bytes toch_read, whoserecvmust hand over whole records. A length field above 2^14 + 256, which RFC 9846 §5.2 forbids a peer (rfc9846.txt:3595-3596), answersREC_HDR, the header alone:ch_readreads those five bytes and refuses the record, where an answer of 0 would leave the caller waiting for a body that must not come.ch_writable_len(t, cap)(tls.h,tls.c, both TCP transports). It answers the most plaintext onech_writeseals intocapbytes of records.ch_writeand this call read the record limit, the smaller ofpeer_limitandCH_TX_PT(INV-38), through one helper,record_plaintext_max, so the two cannot disagree about it.CH_ALERT_RECORD_LENandCH_KEY_UPDATE_RECORD_LEN(tls.h),REC_OVERHEAD + 2andREC_OVERHEAD + 5, whichtls.casserts are 24 and 27 bytes.ch_closesends one alert record.ch_readsends one KeyUpdate record for each KeyUpdate that asks for an answer, and an alert record when it fails. A record carries at most one KeyUpdate, as its last message, because RFC 9846 §5.1 lets no handshake message span the key change it makes (INV-39), so a record gets at most one answer;bin/unitchecks that. This entry first said one record could carry two KeyUpdates and get two answers, which is the record INV-39 refuses.ch_ticket_obfuscated_age(ticket, age_ms)(ticket.h,handshake_post.c, every object with a client). It answersage_ms + ticket->age_addmodulo 2^32 (RFC 9846 §4.3.11.1,rfc9846.txt:2574-2578) and reads no other field, so a copy of the ticket kept afteron_ticketreturned serves.age_msis 64 bits wide, like the configuration field below, so a caller passes one value to both and the reduction happens here.ch_cfg.ticket_age_msandch_cfg.ticket_lifetime_s. RFC 9846 §4.3.11.1 says a client MUST NOT use a ticket older than itsticket_lifetime(rfc9846.txt:2572-2574), and §4.6.1 that it MUST NOT use one more than 7 days after issuance whatever the lifetime (rfc9846.txt:3259-3261). Withresumptionset,ch_connect,ch_record_initandch_quic_initreturnCH_EINVALbefore a byte is sent when the age is above the lifetime or aboveCH_TICKET_LIFETIME_MAX, 604,800 seconds. An age equal to either is still offered. The rule is one predicate,hspost_ticket_age_ok(handshake_post.h), whichtls.c's twotlsi_config_okdefinitions andquic_config_okeach call.- What 0 means. An age of 0 is a ticket that arrived this
millisecond. A lifetime of 0 is one the caller did not give, and the
age is then held to seven days alone. That reading is sound because
handle_ticketnow drops a NewSessionTicket whoseticket_lifetimeis 0, which §4.6.1 says to discard at once (rfc9846.txt:3258-3259), so no ticket with that lifetime is handed toon_ticket. A configuration that sets neither field, as every one written before them does, is refused nothing. A lifetime above seven days, which a server must not send, keeps the seven days. - Symbol names. An image links one object of each of two
transports (entry 61), and every client object exports
ch_ticket_obfuscated_age, so it joinsTRANSPORT_NAMED: its symbol isch_ticket_obfuscated_age_tcp_blocking,_tcp_nonblockingor_quic_nonblocking, andticket.hmaps the name a caller writes, assrv.hmapsch_srv_check.build.zignames the same symbols.ch_writable_lenkeeps its name: the two TCP transports export it besidech_write, and entry 61 refuses an image of those two objects forch_read,ch_writeandch_closealready.test/lib-pair-check.sh's refused pair now names it. Each client half that script links, and each half oftest/zig-consumer/pair.zig, calls its own object's ticket call. - The header.
ch_ticketandCH_TICKET_ID_MAXmoved fromcfg.htoticket.h, whichcfg.hincludes.cfg.hwas at the 500-line limit, and the ticket, the call that reads it and the seven-day cap are one concern. - The build record. The two fields grow
ch_cfg, and with itch_tls,ch_recordandch_quic: 16 bytes on arm64, 1,144 to 1,160 for the defaultch_tls, and 8 on rv32, 1,072 to 1,080 (bench/results-sram.csv).ch_build_matchescomparessizeof(ch_cfg), which grew in every build measured: 152 to 168 bytes on arm64, and by 8 or 12 on rv32 and Cortex-M3. A consumer compiled under these headers and linked against an object built at02515f2readsizeof_ch_cfg152 against its own 168 and exited 1, in the default, the webpki tcp-nonblockingROLE=both, the webpki QUICROLE=bothand the tcp-nonblocking server configurations. - The proofs.
record_whole_lenproves the whole contract over everynup to 2^20, in a heap object exactlynbytes long.writable_lenproves the call safe over anypeer_limitandcap, and runs its answer through the realch_writefor everycapup to 1,603 bytes and everypeer_limitfrom 63: whatch_writesends fitscap, and one byte more does not. The second claim is an equality over a division, and at 16 bits ofcapit returned no verdict in 600 seconds, sobin/unitchecks the answer atSIZE_MAX.quic_config_webpkiproves the age rule over any age and lifetime, andhandshake_postthat no ticket with a lifetime of 0 is handed over. Thirteen mutants intest/violations/break the rules, and each is caught. - The spec. Unchanged.
spec/lean/Spec/Record.leanmodels one record's protection, and itsseal_sizetheorem states the 22 bytes a record adds, which is the sizewritable_len'srec_sealstub asserts. It models no byte stream cut into records, no write loop and no configuration refusal, andSpec/Handshake.leanmodels a NewSessionTicket as a message in the handshake's order and reads none of its fields. The three additions are length arithmetic and one configuration rule, which CBMC proves on the C itself, so a spec model would add a second statement of the same arithmetic and a differential that compares two copies of one formula.
Cost:
ch_cfggrows by 16 bytes on arm64 and by 8 or 12 on rv32, and every session struct with it: the defaultch_tlsby 16 and 8.- Every client object exports one more call and every TCP object one
more; every tcp-nonblocking object exports
ch_record_whole_len. ch_writable_lendivides once, sotls.ctakes a ceiling of 1 inlint-wide-multiplybesidesha3.c's, andRV_ALLOWEDrecords its__mulsi3and__udivsi3on rv32ic. Both operands are public: the caller's buffer length and a record's length. It counts the overhead per record rather than taking the remainder of the division, because gcc turned that remainder into a second division on riscv32.ch_cfgcarries both the age andobfuscated_age, which the caller computes from it. C cannot check that the two agree, becausech_cfgcarries noage_add.
Gain: the Zig API and colibri keep no TLS rule of their own, and the ticket lifetime that RFC 9846 puts on a client is checked where every client entry checks its configuration.
Rejected:
ch_record_readandch_record_write, bytes in and bytes out, the design's option (b). A tcp-nonblocking session would then need nosendorrecvafter the handshake, but a connected session would have two ways to read and two to write, each with a contract and tests of its own, where the two calls above read a header and a limit and change nothing.- Leaving the three in the caller, the design's option (c). It keeps a record rule outside C, which is what the API was approved to avoid.
- A 32-bit age. RFC 9846 says 32 bits hold any plausible age, and an age kept in 32 bits wraps after about 49.7 days, where a ticket reads as young again. The field is 64 bits wide and the call reduces its sum itself.
- Refusing a lifetime of 0 with
resumptionset. Every configuration written before the field existed sets 0, so every resumption in them would be refused.
-
The Zig package's module
chapulinis a Zig API that forwards to the C calls and carries the object, and the translated headers are itschapulin.c. colibri and cocuyo each wrote an adapter over the translated headers: callbacks,ch_cfgbuilding, record framing and error codes. Camilo approved an API that replaces them on 2026-09-26, colibri and cocuyo reviewed its design, and entry 72 moved the three computations it would otherwise have held into C. This entry amends entry 69, whose dependents added the object themselves, and entry 70, whose module becomeschapulin.c. docs/zig.md is the API's reference.- What it adds. Values a session is configured from and the
ch_cfgeach builds (toCfg), one Zig error per result code, the callbacks C calls, which copy bytes between the caller's slices and the session, and storage: the receive buffer, the latest ticket and the peer's transport parameters. It is chapulin.hpp's kind of wrapper. A record's length, a write's size and a ticket's age come fromch_record_whole_len,ch_writable_lenandch_ticket_obfuscated_age, so the API keeps no TLS rule. - The module.
build.zigmakes the translated headers a private module, imported by the API aschapulin_cand declared aschapulin.c, and roots the modulechapulinatchapulin.zig. Zig 0.16.0 refuses one file in two modules of one program, and colibri imports the modules of two objects, sobuild.zigcopies the three files into a directory of each configuration's own, withdefines.txt, the object's define list, beside them. - The object. The module carries it with
addObjectFile. A program links it once however many of its modules import the module, and a program that also addschapulin.odefines every public name twice and fails to link. Sotest/zig-consumeradds no object, and neither does the example in docs/building.md. - Per configuration. Each file declares every call, and a call
whose C name the object lacks is a
@compileErrornaming the option that adds it. The headers declared the client'sch_record_initandch_quic_initin aROLE=serverobject too, which defines neither, so the client sessions first required the client's ticket call as well. Each header now declares a call only in the objects that define it, which matches.zig holds (INV-36), and the sessions ask for the call itself. A TCP object'sErrorlacksDiscardandAeadLimit, whose codes only a QUIC object's headers declare. - Error sets. Each call's set is the part of
chapulin.Errorits C call returns, which a reading of every return path confirmed for the 22 calls that return a code.fromCodemaps a code, and panics on one the call's header says it never returns. It is public, for a program that calls a function the API leaves out. - Reading and writing.
readpasses at most one whole record toch_read, and itsconsumedis 0, a record, or the 5-byte header of a record no peer may send, whichch_readthen refuses.ch_readanswers a KeyUpdate that asks for one, and a record carries at most one KeyUpdate (INV-39), soreplytakes at most one record ofkey_update_record_lenbytes, and one ofalert_record_lenwhen the read fails.writeis all or nothing: it seals nothing whenptis longer thanch_writable_lenallows.closeisch_close, which sends the close_notify and wipes the keys. - Tickets. A Ticket holds the
ch_ticketby value and the identity bytes, soch_ticket_obfuscated_agetakes it as it is andsizeof_ch_ticketin the build record covers it.takeTicket,recordCloseand the QUICclosezero the slot.Ticket.fromFieldsrebuilds one from stored fields, andTicket.fromOnTicketcopies whaton_tickethands over. - The reserved
keyUpdate. No C call starts a record-mode KeyUpdate, so a record session'skeyUpdateis a@compileError. When C gains the call, RFC 9846 §4.7.3's cap on updates sent gets a result of its own. - The design's open questions. Entry 72 settled the first two: the
age is 64 bits wide in
ch_cfgand inch_ticket_obfuscated_age, so the Zig passes it whole and truncates nothing, and the call readsage_addalone, so a ticket whose identity pointer is null serves. Camilo decided the third: stompy's object,TX_RECORD=16384, runs incheck. - The check (INV-36).
test/zig-build-check.shnow builds six configurations, stompy's among them, and for each runsunit.zig, the API's unit tests, and for eachROLE=bothobjectloop.zig, a client and a server of that object against each other through the API alone, in record mode and over QUIC. The loops take the r2 chain, its anchor, its clock and its leaf key fromtest/webpki_corpus.h, the fixture the C loop tests use, through a translate-c step. Seven mutants break the API or the module, and the script catches each.
Cost:
- About 1,150 lines of Zig in the three files, each under 500, and
about 1,170 in
test/zig-consumer. A change to a public C call changes its Zig call in the same commit, aschapulin.hppis kept. lint-zig-buildbuilds one more configuration and runs more programs. With every object built,test/zig-build-check.shtook 4.2 and 5.1 s where it took 3.65 s; with no Zig build, 41.1 and 41.7 s where it took 35.2 s, on an M-series Mac with a load average between 7 and 14.- A program's two objects give two sets of types, so a program that serves both converts its own values once per object.
Gain: colibri, cocuyo and stompy run chapulin through Zig values and errors, with no callback and no
ch_cfgof their own, and the code that did that work in three programs is tested here against both roles of the objects they link.Rejected:
- Keeping
chapulinas the translated module and exporting the API as a second module. colibri's two imports would stay as they are, but the package's own name would keep naming the headers rather than the API, and a program would still add the object. - The API's unit tests as
testblocks in its own files. Zig runs a module's tests only when that module is the test's root, and the API's files cannot define the hooks the object imports, so the tests would need a second copy of the module. They sit intest/zig-consumer/unit.zig, over the public calls, and run against each configuration's module.
Entry 78 changed one thing here:
ch_writesends a KeyUpdate by itself as the last record an AES-GCM write key seals, no call that a caller starts one through is planned, andkeyUpdatestays reserved. - What it adds. Values a session is configured from and the
-
A server checks every provisioned identity against what its flight needs before a session starts, and a refusal inside the flight is
CH_EAUTH. The Zig API's error table (entry 73) found a server flight that returnedCH_EINVALafter its ServerHello went out:srv_auth.creturned it when a slot's key lengths did not match its scheme or its signer refused the key, andsrv_flight.cwhen a certificate did not fit its cert_data field.CH_EINVALsays nothing was sent (cfg.h), and a caller that read it that way would retry a dead session. Each of those conditions is a fact about the configuration, so a check at entry can find it.- The rules.
srv_identities_usable(srv_auth.h) asks of every provisioned slot: key lengths srv_cfg.h states, with an RSApub_lenof at mostSRV_SIG_MAX, the bytes the flight signs into; a private key the scheme's signer takes; and a chain a Certificate message can carry (srv_certificate_fits, srv_message.h).srv_config_ok, whichch_srv_accept,ch_srv_record_initandch_srv_quic_initrun, andch_srv_checkask it. Each rule reads lengths, pointers, two bytes of a public modulus and one private scalar, and signs nothing. - The signers own their key tests.
p256_sign_key_okandrsa_pss_sign_key_okare the testsp256_signandrsa_pss_signalready ran, now calls of their own that each signer runs first. They read the private key, andsrv_auth.creads no byte behindch_identity.priv, so each sits in its signer's module.rsa_pss_sign_key_okis inline in rsa_sign.h, so rsa_sign.c compiles one copy of the test, as it did before, andlint-wide-multiply, which counts the branches in rsa_sign.c, reads the same count. Measured before and after: identical assembly under the three clang specs and under mips gcc at -Os, and the same branch and multiply counts under mips gcc at -O2 and under Ubuntu's arm and riscv gcc at the m3 and rv32 specs' flags. - The chain rule closes a second gap. A chain whose Certificate
message runs past the handshake header's 3-byte length field was
never refused:
srv_build_certificate_headerwrote the length cut to its low 24 bits. The same rule refuses it, and a certificate of no bytes, which cert_data<1..2^24-1> forbids. - The code inside the flight. Once the check has passed, a
refusal inside the flight comes from ECDSA's own signing, no nonce
candidate in range or an r or s of zero, each below 2^-127
(p256_sign.h), from key or chain bytes the caller changed after
init, or from a fault. The ServerHello has gone out by then. The
code is
CH_EAUTH, because this side's authentication failed, and the alert stays internal_error, which tells the peer the fault is local.
Rejected:
CH_ASSERTinside the flight. CLAUDE.md keeps it for programmer error. The nonce refusal is not one, and a fault or a caller's later write can cause the others; an assertion would stop the whole device for one connection's failure.- A new result code for a local failure. Every wrapper would map
it, chapulin.hpp and the Zig API among them, for events the check
now excludes, where
CH_EAUTHalready says what failed. CH_ECAP. The flight answers a message that does not fit with it, and a signer's refusal is not that.
Cost: each server init runs the rules, one pass over each chain's lengths and one scalar range test per ECDSA slot. And a configuration with one broken slot beside a sound one, which served every client that selected the sound one, now serves none: init refuses it whole, as
ch_srv_checkalready did.Gain:
CH_EINVALfrom a server means nothing was sent, on every path, and a broken identity is found at init rather than at the first handshake that selects it. - The rules.
-
Every object reports what ended a session through two calls, and a session that reads the peer's fatal alert answers it with nothing. colibri could not tell its own failure from its peer's: the alert a tcp-nonblocking session chose after the handshake was sent by
ch_readand reported nowhere, becausech_record_alertnames only the handshake's, and a peer's alert was answered with unexpected_message, which RFC 9846 §6.2 forbids: on a fatal alert both sides close the connection at once (rfc9846.txt:3890-3893).- The calls.
ch_alert_sentreadsch_tls.alert_sent, which each failure funnel writes:tlsi_fail,tcp_nonblocking_failbesidech_record_alert's field, andquic_failbesidech_quic_alert's.ch_alert_receivedreadsch_tls.alert_received, whichhsr_refuse_alertwrites for every TCP reader, and answers 0 in a QUIC object, which carries no alert record and declares no such field. alert.h declares both, and tls.h and quic.h include it: tls.h is the TCP transports' header alone (INV-27), and a QUIC object exports the two calls as well. Every object of every transport exports them, so their symbol names carry the transport, asch_ticket_obfuscated_age's do (entries 61 and 72). Both fields sit in padding, soch_tls,ch_recordandch_quickeep their sizes in every build bench/sram.sh measures. - The peer's fatal alert. Any 2-byte alert record whose
description is not close_notify or user_canceled is one, whatever
its level byte (§6, rfc9846.txt:3779-3782). The call that reads it
returns
CH_EPROTO, the code a peer's alert already gave, wipes, fails the session and sends nothing, and the funnels write noalert_sentoncealert_receivedis set. A record of the alert type of any other length is not an alert (§5.1, rfc9846.txt:3475-3478), and every reader now answers it with decode_error, where the tcp-nonblocking handshake and the record layer answered unexpected_message. close_notify and user_canceled keep what they did: after the handshake the first closes the peer's direction and the second is read past, and in the handshake either ends it with the alert that reader gives any record it cannot use. - An alert in the clear during the handshake. Both handshake readers take one whether or not this side has installed its read key. A peer protects an alert under its own write key (§6, rfc9846.txt:3759-3760), and a client that could not use the ServerHello has none, so it answers in the clear to a server that already reads protected records. The tcp-nonblocking reader used to decrypt such a record and answer bad_record_mac.
- Where the post-handshake reader fails.
hspost_readnow callstlsi_failitself, with bad_record_mac, decode_error or unexpected_message, instead of returning a code fortls.cto map. A wrapper intls.cholding the alert as a local grewch_read's peak stack by 48 bytes; with the call inhspost_readit is 16 bytes lower than before.
Rejected:
- Both declarations in tls.h, with a second pair in quic.h. Two contracts to keep in step, a quic.h already at the 500-line cap, and build.zig's header table names one header per call.
- A new result code for the peer's alert. Every wrapper would map
it, chapulin.hpp and the Zig API among them, and
CH_EPROTOwithch_alert_receivedalready says what happened. - Reading an alert in the clear only before the read key. The server would decrypt the alert of a client that failed on the ServerHello, fail it with bad_record_mac and answer an alert that closed the connection.
Cost: two more exported calls in every object, and a peer's alert in the clear is read during the whole handshake, where nothing authenticates its description; it ends a handshake an on-path attacker could end anyway by dropping bytes.
ch_read's peak stack is 1,712 bytes, from 1,728, and no struct grew.Gain: a caller logs what ended each session through one pair of calls on every transport, and neither side answers an alert that closed the connection.
Entry 76 changed one thing here: a tcp-nonblocking driver sends the alert of its own handshake failure, and
ch_record_alertand its field are gone. - The calls.
-
A tcp-nonblocking driver sends the alert of its own handshake failure, sealed once it has a write key, and
ch_record_alertis gone. A failure inch_record_inorch_srv_record_inwiped every key and named its alert inch_record_alertfor the caller to send. The caller holds no key, so it could send the alert only in the clear. RFC 9846 §6 encrypts an alert under the current connection state (rfc9846.txt:3758-3760), and a strict peer that already reads protected records takes a record in the clear for a bad one. The blocking drivers did this right:tlsi_failseals underch_tls.wrwhench_tls.keysis set.- The seal.
tcp_nonblocking_failwrites the alert as one record before it wipes: sealed underch_tls.wrwhench_tls.keysis set, and in the clear before. A client setskeysright after the ServerHello, and a server right after it has sent its own. The wipe then clears every key, so none outlives the failure. The record isCH_ALERT_RECORD_LENbytes sealed andREC_HDR + 2in the clear. - The client stages it. The record goes into
ch_tls.tx, where the client stages every record it writes.ch_record_outhands it over on the failed session, in as many calls as the caller's buffer needs, and answersCH_EINVALafter the last byte, the answer a failed session gave before. - The server pushes it. The record leaves through
cfg.srv.on_record_outinside the failingch_srv_record_in, after the records of the flight that went out first. The push is best effort, astlsi_fail's send is. A failedrecordInin the Zig API returns noProgress, so a server'soutputLen()returns what the lastrecordInwrote intooutput, a failed one included. The QUIC server'scryptoInneeds no such call: the caller'sOutgoingkeeps its own counts, and a QUIC failure pushes no bytes, because the caller seals the CONNECTION_CLOSE (entry 57). - Nothing after the peer's alert. A failure on the peer's fatal alert stages and pushes nothing (§6.2, entry 75).
ch_record_alertand its field go. The call would read 0 after every failure now, andch_alert_sentnames the alert in every case. The field sat in padding, soch_recordkeeps its size in every build.
Rejected:
- Keeping the write key for the caller, as
quic_failkeeps a level's write keys forch_quic_seal_close(entry 57). The caller would need a call to seal with, and the key would outlive the failure until the caller made that call. A QUIC caller builds the CONNECTION_CLOSE packet around its seal; a record-mode alert is one fixed record the driver can finish itself. - A pull for the server's alert. A server pushes every record it
writes and has no
ch_srv_record_out(docs/server.md). One record to pull would give the server a second output path. - Keeping
ch_record_alertfor the callers that read it. It would read 0 after every failure, and a caller that sent what it named would send nothing.
Cost:
ch_record_outreturns bytes on a failed session once, a caller ofch_record_alertno longer links, and a server whose sink refused a record pushes the alert into the same sink. A record sealed after a refused one is one sequence number ahead of what the peer read, so the peer cannot open it unless the refused bytes reached it; the blocking drivers have the same limit, and the last item below says why it stays.Gain: a caller sends the bytes the API hands it and nothing else, and a strict peer reads every handshake alert: in the clear before the failing side's write key, and protected after. No key survives a failure.
Three cases the change left open are closed the same way in every driver:
- A refused record names internal_error. A record the caller's
transport refused, through
cfg.send,cfg.srv.on_record_outorcfg.srv.on_crypto_out, kept the handler's default alert, decode_error, which says the peer sent a malformed message. RFC 9846 §6.2 names internal_error for an error unrelated to the peer (rfc9846.txt:3979-3981), andch_writealready named it for a failed send. Only the send call knows that the failure is this side's, so each send names it there:srv_out.cfor the server,send_stagedinhandshake.cfor the blocking client's ClientHello and Finished, andtake_key_updateinhandshake_post.cfor the KeyUpdate replych_readsends. A failed receive keeps its reader's alert, because a peer that closes the connection fails a receive the same way. A tcp-nonblocking or QUIC client sends nothing during its handshake: its caller collects the bytes. - The ticket push. A NewSessionTicket that the sink refused after
the client Finished failed the tcp-nonblocking and the QUIC server
with no alert recorded: each wiped its handshake state,
hs.alertamong it, beforetcp_nonblocking_failorquic_failread the alert. Each now returns first.tcp_nonblocking_failthen records internal_error and seals the alert thatfailpushes, andquic_failrecords it and keeps the close keys, and each wipes the state after. The QUIC server reported 0x0100, which names close_notify. The blocking server'ssrv_handshakereads the alert before its wipe, so it always recorded one. - The sequence number after a refused record. The refused record
used its sequence number, and the driver does not take it back. The
sink may have written the record whole or in part before it
reported the failure, so the peer may hold those bytes. Sealing the
alert under the same number would put two plaintexts under one
nonce, which breaks ChaCha20-Poly1305 and AES-GCM alike. So the
alert goes out one number ahead, and it opens only for a peer that
took the refused bytes, as
test/tcp_nonblocking_failure_alert_tests.hshows. A QUIC server has no such limit: its caller seals the CONNECTION_CLOSE under a packet number of its own.
- The seal.
-
A
RAND=sessionobject draws every random byte from a source each session'sch_cfgnames, and packages no generator. colibri's owner ruled in its decision 94 that cocuyo, its sibling project, replays a connection from a seeded stream. With onech_rand_bytesper image, a replay had to know which chapulin calls draw and in what order, and it broke with no error when a draw moved from one call to another. Sessions on different threads also shared that one hook, which then needed a lock or state held per thread. Camilo chose on 2026-09-27 a third value of theRANDaxis, and firmware keepsRAND=extern.- The fields. Under
-DCH_RAND_SESSION,ch_cfggainsrand_bytes, which the library calls asrand_bytes(rand_io, p, n), andrand_io, which it hands over unread and which may be NULL. The contract isch_rand_bytes's (rand.h): the source writes n bytes and returns nothing, so a source that cannot fill must not return. It is a CSPRNG in production, because a deterministic stream gives predictable keys to anyone who knows its seed. A seeded stream that replays a connection is for tests. - One draw path. Every library draw calls
rand_draw(rand_draw.h), which calls the session'srand_bytesunderRAND=sessionandch_rand_bytesunder the other two patterns, so those builds draw as they did. The ten sites stay ten. One moved:rsa_pss_signtakes its salt as an argument, andsrv_auth.c'ssign_rsa_pssdraws it from the configuration the signature serves, asmlkem_encaps_derandandp256_ecdh_keygentake their random bytes from their callers.rsa_pss_signkeeps the assertion against an all-zero salt.inv-4-one-draw-pathrefuses ach_rand_bytescall outsiderand_draw's body, and the files and calls rules now countrand_drawcalls. - The refusal.
ch_connect,ch_record_init,ch_quic_init,ch_srv_accept,ch_srv_record_init,ch_srv_quic_initandch_srv_checkreturnCH_EINVALfor a NULLrand_bytes, before they draw or send anything.ch_srv_checksigns with the RSA identity, and that signature draws a salt, so it refuses whichever identities the configuration holds, as the init calls do. - No fallback. The object neither defines nor imports
ch_rand_bytesand packages nodrbg.c, so an image that links only such objects defines no entropy hook.rand.hdeclares the hook only for the other two patterns.lib-checkholds the object to name neitherch_rand_bytesnorch_drbg_seed, the counterpart of theRAND=externimport check, and linkstest/build_test.c, which defines no hook under the define. - The build record.
CH_BUILD_RAND_SESSIONis bit0x400ofaxes: the define changesch_cfg, and with itch_tls,ch_recordandch_quic.RAND=externandRAND=drbgstill take no bit, because neither changes a layout (entry 56). - Threads. A session draws from its own source alone, so the library holds no entropy state that sessions share, and sessions on different threads need no lock around the library's draws. A lock is the source's business, and only where two sessions share one.
- The Zig API.
Client.randomandServer.randomare?std.Random, null by default. A session's init stores the value in the session and pointsrand_ioat that copy, and an adapter fills each draw from it. A null value leavesrand_bytesNULL, and C refuses it withCH_EINVAL, which init returns aserror.Invalid.Server.checkholds the value for the one call.
Rejected:
- A per-session callback beside the hook, which falls back to
ch_rand_byteswhen NULL. A forgotten field would draw from the image's generator with no error, which is the silent break this entry exists to remove, and every object would still import the hook. rsa_pss_signtaking a fill function and its context. The signer would call through a pointer into the session's source. With the salt as an argument it draws nothing, and a signature is a function of the key, the digest and the salt.- A required Zig field. A Zig field's default cannot depend on the build, so a field that is absent from the other builds and required in this one needs a second copy of each value struct. A null value that C refuses keeps the rule in C, where every other configuration rule is.
Cost: in a
RAND=sessionbuildch_cfg,ch_tls,ch_recordandch_quiceach grow by two pointers, 16 bytes on arm64 and 8 on rv32ic, measured. No other build's layout moves, and bench/sram.sh reads the same numbers as before.rsa_pss_signtakes one more argument.make checkgains one packaged-object leg and three loop binaries, and check-slow's Zig roster one configuration.Gain: a replay needs only each session's seed, whichever calls draw and in whatever order, and two sessions never draw from each other's stream, on one thread or several.
- The fields. Under
-
An AES-GCM write key seals at most 2^24 records, and
ch_writesends the KeyUpdate that retires it by itself. RFC 9846 §5.5 has a sender close the connection or send a KeyUpdate while a key is still below its AEAD's usage limit (rfc9846.txt:3743-3744), and gives AES-GCM's as up to 2^24.5 full-size records under one set of keys (rfc9846.txt:3750-3753). ASUITE=aesgcmbuild's record layer held only the 64-bit sequence wrap, so a long connection sealed past that limit under one key. docs/server.md saidrec_sealhad a per-suite ceiling and that the session rekeyed withrec_dir_updatebefore it. Neither existed, and a rekey with no KeyUpdate would have left the peer reading under the old key. colibri needed the rekey and asked for a call that starts a KeyUpdate. Camilo decided on 2026-09-27 that chapulin rekeys by itself and adds no public call.- The ceiling.
REC_AES_GCM_RECORDS_MAX(record.h) is 2^24, the largest power of two at or below 2^24.5. It counts every record under the key, whatever its size, the reading that can never let a key pass the RFC's figure. An AES-GCM write key seals at sequence numbers 0 to 2^24 - 1 and at none above:rec_sealrefuses a record at or past the ceiling, beside its wrap guard, so no path seals past it.rec_openkeeps no count, because §5.5 says a receiver SHOULD NOT enforce the limit (rfc9846.txt:3747-3748). - The KeyUpdate. Before each record it seals,
ch_writereads the write key's suite and sequence number. When an AES-GCM key's next record would take its last sequence number,ch_writesends a KeyUpdate there, under the old key, then moves the write direction to the next key and seals the record at sequence number 0 of it. So the KeyUpdate is the last record under the old key. Its request_update is 0: the key at its limit is this side's write key, and the peer's write key keeps a count of its own.hspost_send_key_update(handshake_post.c) seals, sends and rekeys, for this and for the answer to a peer that asked for one.ch_writeserves both roles and both TCP transports. - The sender cap. §4.7.3 lets a sender send 2^48 - 1 KeyUpdates
(
rfc9846.txt:3400-3402),HSPOST_SEND_EPOCHS_MAX, and says the limits of §5.5 may then end the connection (rfc9846.txt:3408-3410). A key at its ceiling after that many fails the write:tlsi_failseals internal_error at the key's last sequence number and wipes the keys, andch_writereturnsCH_ECAP. That isch_write's code for a record it cannot seal, which it already returned at the sequence wrap, and it leaves the session dead, as INV-13 and tls.h say of everych_writeerror butCH_EPROTO. internal_error, because the cause is this side's count and not the peer's input. - Other sends. A server's NewSessionTicket goes out at sequence
number 0 of the key its handshake installed, the answer to a peer's
KeyUpdate rekeys right after its one record, and an alert is the
last record before the keys are wiped.
ch_writeleaves every AES-GCM key with its last sequence number free, so each of the three takes at most that one, and none checks. ch_writable_len. It counts the KeyUpdate record when the write it sizes would cross the ceiling, so the Zig API's all-or-nothingwrite(entry 73) never handsch_writean output too short for what it sends. It counts one KeyUpdate and no second: it answers at most the plaintext of the records the key has left and 2^24 - 1 records under the next key, over a gigabyte at the smallest record a peer may ask for. Counting any number of KeyUpdates would divide cap by the length of 2^24 - 1 records and a KeyUpdate, which passes 2^32 at records of 235 bytes, where a 32-bitsize_tcannot hold it. The crossing path divides by the record length a second time, in aSUITE=aesgcmbuild alone, over public lengths.- The file.
ch_writeandch_writable_lenmoved fromtls.ctotls_write.c, whichtls.hdeclares, because the change tooktls.cpast 500 lines.lint-wide-multiply's division ceiling andRV_ALLOWED's runtime calls moved with them, andtls.cjoinedtools/proof-cover.py'sAUDITED, becausewritable_lennow includestls_write.c. - QUIC. Unchanged. RFC 9001 §6 forbids the TLS KeyUpdate message
there (
rfc9001.txt:1566-1568), and each AES-GCM key set counts its packets against §6.6's confidentiality limit (docs/quic.md). - The Zig API.
writeis unchanged, becausewritableLenisch_writable_len. The reservedkeyUpdatestays a@compileError, and entry 73's result for a caller-started update at the cap goes with the call. - The check.
test/key_limit_cases.hruns inbin/webpki_loop_aesandbin/webpki_loop_aes_externover tcp-nonblocking, and inbin/tcp_blocking_key_limitover tcp-blocking: a client and a server of one object each write across the ceiling under both AES-GCM suites, with both ends' sequence numbers moved near it rather than 2^24 records sealed.bin/aes_suite_testholdsrec_seal's refusal to its exact boundary.writable_len_suiteprovesch_writable_len's answer against the realch_writein the suite build, for every cap up to three records of the session's limit, a KeyUpdate record and a byte, and that no data record takes an AES-GCM key's last sequence number;writable_len_suite_anyproves the call safe over any input;record_suiteproves thatrec_sealrefuses the wrap and the ceiling and nothing else. Seven mutants intest/violations/break the rules, and each is caught. - Two repairs the checks needed.
bin/webpki_loop_aes_extern's rule sat aboveWEBPKI_LOOP_SRCS, and make expands a rule's prerequisites where it reads the rule, so the list held none of the library sources and an edit to one never rebuilt the binary. Two of the mutants here passed on that stale binary until the rule moved below the list. And alint-tidypass now readstls_write.candrecord.cunder the suite define, which readsct.h's suite guard for the first time; its#if defined(CH_AES_HW)is#ifdefnow, as readability-use-concise-preprocessor-directives asks. - The spec. Unchanged.
spec/lean/Spec/Record.leanmodels one record's protection under ChaCha20-Poly1305, and itssealsays it does not modelrec_seal's refusal;Spec/Handshake.leanmodels a KeyUpdate as a message a connected session accepts. Neither models the records a session sends or when it moves to a new key.
Rejected:
- A KeyUpdate call the caller starts, entry 73's reserved
keyUpdate. colibri needed it for this limit alone, and a caller that forgot to call it would seal past the limit. - Closing the connection at the ceiling, which §5.5 also allows. A host moving large uploads would lose its connection every 2^24 records, 8 GiB at 512 bytes a record.
- Rekeying without a KeyUpdate, what docs/server.md described. The peer's read direction would stay on the old key, and the next record would fail its tag.
- Counting full-size records alone, the unit the RFC's figure is
in.
ch_writewould have to judge each record's size, and the key could seal more records than the count here allows. - request_update 1. The peer's write key keeps its own count, so its answer would rekey a direction that is at no limit.
Cost: one comparison before each record a
SUITE=aesgcmbuild seals throughch_write, and one 27-byte KeyUpdate record per 2^24 records under an AES-GCM key. No struct grew, and bench/sram.sh reads the same numbers. A caller that sizes its output for one message addsCH_KEY_UPDATE_RECORD_LENbytes underSUITE=aesgcm, or the write that crosses the ceiling returnserror.Capthrough the Zig API with nothing sealed.make checkruns one more binary and one more tidy pass, and the fast proof tier two more harnesses, 49 seconds and under one.Gain: a
SUITE=aesgcmsession keeps RFC 9846 §5.5's AES-GCM limit for as long as it runs, with nothing for the caller to call, and every write sized withch_writable_lenstill fits. - The ceiling.
-
A QUIC object derives packet keys for QUIC version 1 and version 2, and the caller names the version of each packet. RFC 9369's QUIC version 2 is version 1 with a new Version field, new long header type codes, a new Initial salt, new HKDF labels and a new Retry integrity key and nonce (
rfc9369.txt:130-188). colibri reads and writes the wire, so the header codes, Version Negotiation and RFC 9368's version_information transport parameter are its half, tracked in c4milo/colibri#54. The keys are chapulin's, and a QUIC object derived version 1's alone. Camilo decided on 2026-09-29 how the version enters the calls.- The caller names the version. Each packet call takes the version
beside the encryption level, and chapulin reads no header byte to
learn either. An Initial packet may carry either version, because a
server keeps its original version's Initial receive keys until it
processes a Handshake packet in the negotiated version
(
rfc9369.txt:250-254). A Handshake or 1-RTT packet in any version but the negotiated one is refused, the drop §4.1 requires (rfc9369.txt:256-259). A Retry uses the original version (rfc9369.txt:221-227). - The client. Its configuration names its original version, and a
configuration that names none is refused. It learns the negotiated
version from the first long header whose Version field differs from
the original (
rfc9369.txt:240-244), and switches once, deriving the new version's Initial keys from the same Destination Connection ID. The switch is refused once the first CRYPTO byte from the server has been delivered, because the server sends every CRYPTO frame in the negotiated version, and a CRYPTO frame in the original version makes the original the negotiated one (rfc9369.txt:236-244). - The server. A callback from the server's caller chooses the
negotiated version once, after the client's transport parameters
and before ticket selection, the HelloRetryRequest and the
ServerHello, because a server sends no CRYPTO frame before it has
processed those parameters (
rfc9369.txt:236-237).srv_selectran before the parameters were handed to the caller, and the HelloRetryRequest path never handed them out, so the parameters moved ahead of it. The caller answers from RFC 9368's version_information, which chapulin does not parse. The callback isch_srv_cfg.choose_version: NULL keeps the original version, and an answer the build does not derive fails the session withCH_EIOand internal_error, the code a refusedon_crypto_outreturns, because the failure is the caller's and not the peer's. - Tickets and tokens. A ticket belongs to the version of the
connection that issued it, the negotiated one after a switch
(
rfc9369.txt:268-284). A server ticket records that version, and the server passes over a ticket of another version for a full handshake. A TCP server's tickets record no version, so no ticket crosses transports. A client offers no ticket issued under a version other than its original one:ch_ticket.quic_versioncarries the version andch_cfg.ticket_quic_versionpresents it, andch_quic_initrefuses a mismatch. A raw or ca client refuses a declined ticket, so its resumption fails closed when the server switches. The Retry token binds the version, which §4.1 permits (rfc9369.txt:224-227): its tag covers the version the way it covers the client's address, so the layout andCH_QUIC_TOKEN_MAXstay as they were, and a token checked under the other version fails its tag. - INV-7 and INV-26. INV-7's one version per build is the TLS
version, 1.3, and it stays one; its claim will say that the QUIC
version is the caller's value, like the level, and that chapulin
chooses neither. Version 2's Initial keys come from a salt the RFC
prints and a connection ID sent in the clear, and its Retry key and
nonce are printed (
rfc9369.txt:158-188), so INV-26's public-key argument covers them as it covers version 1's.
Rejected:
- A
QUIC_VERSIONbuild axis. No field would grow and no signature would change, but a build could not negotiate: the interop runner's v2 case starts in version 1 and has the server answer in version 2. - The version as a pseudo-level. The packet calls would keep their signatures, but a level would then name a version, and chapulin could not drop a Handshake packet in the wrong version.
- A switch allowed until the Handshake keys. It would admit a switch after server CRYPTO bytes arrived in the original version, which §4.1 makes the negotiated version.
- An unset original version meaning version 1. Every caller names the version on every packet call anyway, and a default would hide a configuration that forgot it.
- Dropping a mismatched ticket at init. colibri asked on
2026-09-29 whether
ch_quic_initshould drop a ticket issued under another version and run a full handshake instead of refusing the configuration. RFC 9369 §5 forbids offering the ticket (rfc9369.txt:268-271) but does not say which to do. Camilo kept the refusal the same day: a silent drop would change what the caller asked for without telling it, and hide the mistake that passed the ticket.
Cost: a version argument on each packet call, fields for the original and the negotiated version, and a version in each ticket. The two version fields grow
ch_quicin colibri's object,ROLE=both TRUST=webpki TRANSPORT=quic-nonblocking, from 5,072 to 5,080 bytes on arm64, and from 5,544 to 5,560 underSUITE=aesgcm, measured by bench/sram.sh; the TCP sessions docs/performance.md measures do not change. The server'schoose_versionpointer grows it to 5,088 bytes, and to 5,584 underSUITE=aesgcm, measured the same way. The two ticket version fields sit in padding, so neitherch_ticketnorch_cfggrows. A server ticket grows by four bytes, to 108, and to 124 underSUITE=aesgcm, and its format version moves to 2, so a ticket sealed before the change fails to open and costs its client one full handshake.Gain: colibri can negotiate version 2 in both roles, as the interop runner's v2 case asks, and chapulin enforces every rule that §4.1 and §5 state for the keys, tickets and tokens it holds.
Status: this entry and the two RFCs landed first, then the interface:
ch_cfg.quic_original_version, the version argument on the packet calls and on the two Retry calls,ch_quic_switch_versionandch_quic_negotiated_version. Version 2's keys and the client's switch have landed since.quic_version.h'squic_version_derived, the one rule that says which versions a build derives, admits version 1 and version 2.aes.cholds version 2's Initial salt and Retry key beside version 1's,quic_retry.cits Retry nonce, andquic_version.hits four labels, each copied fromrfc9369.txt:158-188, and RFC 9369 Appendix A checks every one of them against the vendored text. A client that starts in either version may switch to the other once, before the server's first CRYPTO byte, and a whole handshake runs in version 2. A server's caller chooses the negotiated version throughch_srv_cfg.choose_versionat the first ClientHello, after the client's transport parameters and before the ticket selection, a HelloRetryRequest and the ServerHello. A ticket records the negotiated version of the connection that issued it in both roles, the server passes over a ticket of another version, the client offers none in another, and the Retry token binds the original version through its tag. QUIC version 2 is complete in both roles. - The caller names the version. Each packet call takes the version
beside the encryption level, and chapulin reads no header byte to
learn either. An Initial packet may carry either version, because a
server keeps its original version's Initial receive keys until it
processes a Handshake packet in the negotiated version
(
-
A build on
AES=hwwithCH_NATIVE_AESoffers and prefers AES-256-GCM, then AES-128-GCM, then ChaCha20, and a caller may set a client's order. Entries 45 and 58 put ChaCha20 first in both roles: aSUITE=aesgcm TRUST=webpkiclient offeredTLS_CHACHA20_POLY1305_SHA256, thenTLS_AES_128_GCM_SHA256, thenTLS_AES_256_GCM_SHA384, and a server with noch_srv_cfg.cipher_suitespreferred the same order. A server that follows the client's order then selects ChaCha20 when both ends have the AES instructions, and so does Go's, which reads the client's order only to choose between AES-GCM and ChaCha20 and then takes its own list, AES-128-GCM first (crypto/tls,handshake_server_tls13.go). stompy measured the cost through colibri's h11 client, over the TCP object builtSUITE=aesgcm AES=hwat 0adcf33, on one Apple M1 core in OrbStack with 5 MiB PUTs to MinIO: 1 ms of CPU per MB and 935 MB/s over 8 connections in cleartext, and 5 ms of CPU per MB and 150 MB/s, the same from 1 to 8 connections, over TLS. MinIO's Go server selected ChaCha20, 0x1303, because the client listed it before AES-GCM. The server role showed the same effect: against Go's client, on an AMD EPYC 7763 runner with the AES instructions, colibri's server over the object at 0adcf33 selected 0x1303. Camilo decided on 2026-09-29 (#180). This entry amends entry 45's offer and the order bullet of entry 58.- The default order, chosen at build time, the same in both
roles. A
SUITE=aesgcmbuild onAES=hwthat definesCH_NATIVE_AESoffers, as a client, and prefers, as a server,TLS_AES_256_GCM_SHA384, thenTLS_AES_128_GCM_SHA256, thenTLS_CHACHA20_POLY1305_SHA256. The build asserted that its AES instructions and its carry-less multiply run in constant time (entry 50), and AES-128-GCM on them outran ChaCha20-Poly1305 at every size docs/quic.md measured, on arm64 and on x86-64. Every other build keeps entry 58's order, ChaCha20 first,AES=externincluded: GHASH runs there ongcm.c's portable multiply, and nothing in this tree measures a peripheral.suite.hstates the order once, assuite_default_order, andSUITE_AES_FIRSTmarks the build that takes the AES-first one; the ClientHello andsrv_selectread that one definition.ch_srv_cfg.cipher_suitesstill overrides the server's default. - AES-256-GCM first. Camilo chose AES-256-GCM ahead of
AES-128-GCM to align chapulin with NSA's CNSA 2.0 suite, which
requires AES-256 and SHA-384. Entry 58 put AES-128-GCM first
because every handshake proof covers the SHA-256 schedule, and that
stays true: the SHA-384 schedule has only its own harnesses
(
transcript384,keysched384and thehkdf384ones). The AES-256 rows ofbin/webpki_loop_aes,bin/quic_loop_aesandtest/e2e.shrun it through whole handshakes. OnAES=hw, a default handshake with this tree's server, or with any server that follows the client's order and holds AES-256-GCM, now runs the SHA-384 schedule, and its tickets carry 48-byte PSKs. Go's server still selects AES-128-GCM, from its own list; a caller who wants AES-256-GCM from it leaves AES-128-GCM out ofch_cfg.cipher_suites. - The caller's order. A client that offers more than one suite
(
CH_CLIENT_AES_SUITES) takesch_cfg.cipher_suitesandch_cfg.cipher_suite_count, with the element type and the count convention ofch_srv_cfg.cipher_suites. NULL with a count of 0 offers the build's default.ch_connect,ch_record_initandch_quic_initreturnCH_EINVALfor a list that names a suite the build does not hold, repeats a suite, or is longer than theSUITE_HELD_COUNTsuites the build holds (webpki_cfg.h). The ClientHello offers exactly the caller's list, in its order. The parser still takes any suite the build holds, sohsf_read_server_hellorefuses a ServerHello or a HelloRetryRequest that names a suite the list left out, with illegal_parameter (rfc9846.txt:1373-1376,rfc9846.txt:1484-1485). The Zig API gainsvalues.Client.cipher_suites, which forwards to the C field and is gated with@hasField, as the server's is.chapulin.hppexposes no suite order for either role and gains none. - No CPU probe. Nothing in this tree asks a CPU what it has
(CLAUDE.md), and neither does colibri: its decision 97 picks the
chapulin object from the build target's features, so an
AES=hwobject exists only for a target with the AES instructions. colibri's callers probe the CPU when they want a choice at run time, and pass an order that colibri forwards throughvalues.Client.cipher_suitesandvalues.Server.cipher_suites. Without one, the build-time default holds. - The check.
bin/webpki_session_test,bin/webpki_session_aesandbin/webpki_session_aes_externhold the ClientHello's cipher_suites bytes to each build's default, and the last two hold the caller's list to each rule's edge and to the suites the client takes back.test/webpki_loop_order.hfeeds this tree's server a hello that lists the three suites in each of their six orders, in the three webpki loop builds, and requires the first suite of the server's default order;bin/webpki_loop_aes,bin/webpki_loop_aes_externandbin/quic_loop_aesrun the caller's list end to end.test/e2e.shhas OpenSSL'ss_server, which follows the client's order, select AES-256-GCM from theAES=hwclient, ChaCha20 from theAES=externone and AES-128-GCM from either underWEBPKI_SUITES=1301,1303, and hass_clientoffer AES-128-GCM first to this tree's server, which selects its own first suite. Thehello_build_suiteproof holds the builder toCH_HELLO_MAXover every list shape, andquic_config_webpki_suiteproves the list rule. Eight mutants intest/violations/break the new rules, and each is caught.
Rejected, both on 2026-09-29:
- A CPU probe in chapulin. An arm64 core cannot read its ID
registers from EL0 without the operating system's help, so a probe
needs per-OS code, and the bare-metal lanes have no OS to ask.
Whether a build has the AES instructions stays the compiler's
answer at build time, or under
AES=runtimethe caller's (entry 81). - A probe function the caller supplies. It would move no rule into chapulin that a caller's order does not already carry, and it would add a callback to every configuration. A caller runs its own probe and passes the order it chose.
Cost:
ch_cfggains a pointer and a count in aSUITE=aesgcm TRUST=webpkibuild, 16 bytes on arm64, and every struct that holds a copy grows with it. Measured bybench/sram.shand the samesizeofprobes run on the tree before this change:ch_tlsfrom 3,352 to 3,368 bytes, theROLE=both TRUST=webpkitcp-nonblockingch_recordfrom 4,896 to 4,912 and its QUICch_quicfrom 5,544 to 5,560. No stack peak moved. OnAES=hwa default handshake runs AES-256-GCM, fourteen rounds a block where AES-128-GCM runs ten, and the SHA-384 key schedule; this tree has not measured what that costs. The fast proof tier gains two formulas,hello_build_suite(659 properties, 131 s, 0.17 GB) andquic_config_webpki_suite(705 properties, 6 s, 0.14 GB), andmake checkone binary,bin/webpki_session_aes_extern.Gain: on a host whose build asserts the AES instructions, a chapulin client gets AES-GCM from a server that follows the client's order or, as Go's does, reads it to choose between AES-GCM and ChaCha20, and a chapulin server gives it to a client that offers it, with no call to make. A caller that learns at run time what its CPU has sets either order through the Zig API.
- The default order, chosen at build time, the same in both
roles. A
-
An
AES=runtimeobject holds the AES instructions and a fallback, and the caller's CPU probe picks between them for each session (#183). Every otherAESvalue fixes the choice when the object is built. AnAES=hwobject runs the AES instructions for every QUIC Initial packet and for AES-GCM traffic, so it stops with SIGILL on a CPU without them, and every other object holds ChaCha20 alone for traffic, so a caller whose probe finds the instructions cannot use them. colibri picks its chapulin object from the build target's features (its decision 97), and its callers probe the CPU when they want the choice at run time (entry 80). Camilo decided on 2026-09-29. This entry amends entries 50 and 68 where they state theAESaxis.- The build value.
AES=runtimecompilesaes_hw.candghash_hw.cwith no instruction flag. A target pragma at the top of each file puts the target attribute on each function in it,+aeson arm64, whose AES extension holds the 64-bit PMULL, andaesandpclmulon x86-64, so the instructions appear in those two files alone and the rest of the object runs on any CPU of its architecture. A QUIC object also holdsquic_aes_soft.c, whose two entries take the namesaes_soft_expand_round_keysandaes_soft_cipher_blockthere, for QUIC's public keys. A TCP object has no public key and holds no table. Both files refuse a target other than arm64 or x86-64 with an#error. - The answer.
ch_cfg.aes_instructions, declared only in anAES=runtimeobject, takesCH_AES_INSTRUCTIONS_PRESENTorCH_AES_INSTRUCTIONS_ABSENT: whether the CPU has the AES and carry-less multiply instructions, as the caller's probe found.ch_connect,ch_record_init,ch_quic_init,ch_srv_accept,ch_srv_record_initandch_srv_quic_initreturnCH_EINVALfor any other value, 0 included, before anything is sent, as they do for an unset QUIC version, so every caller states what its probe found. chapulin probes nothing. - Present behaves as
AES=hw. QUIC Initial packets and their header protection run on the instructions, and aSUITE=aesgcmbuild offers and prefers entry 80's order. - Absent runs neither instruction. QUIC Initial packets and their
header protection run on the table, and their GHASH on
gcm.c's portable multiply; both keys are public (INV-26). The session holds ChaCha20 alone: a client offers it alone and a server's default order is it alone. Init refuses ach_cfg.cipher_suitesor ach_srv_cfg.cipher_suitesthat names an AES-GCM suite rather than dropping the suite, and a server refuses a retry cookie that names one with illegal_parameter, because a server with the instructions may have minted it under the same cookie key. - Each schedule records its cipher.
aes_key_schedulegainsinstructionsin a QUIC object, andaes_encrypt_schedule,aes_encrypt_block_hpandgcm.crun a schedule on the cipher it records. The Initial constructor records the caller's answer, the Retry constructor the table, andaes_traffic_key_initthe instructions under either answer, so the table runs no traffic key. Each branch reads the caller's answer or a constant, not a key. - The Retry key runs on the table under either answer.
ch_srv_quic_retry_tagtakes no configuration to read an answer from, and RFC 9001 and RFC 9369 print the key, so the table leaks nothing. Under the present answer this departs fromAES=hwin which cipher computes the tag, and not in the tag. CH_NATIVE_AESkeeps its meaning. It is the builder's statement that the part's AES instructions and carry-less multiply run in constant time where the part has them, andct.hrefusesSUITE=aesgcmonAES=runtimewithout it. The caller's answer says whether the instructions exist; the build states their timing.- What the build refuses. The Makefile and
build.zigrefuseAES=runtimein a TCP object withoutSUITE=aesgcm, which carries no AES to choose, andcfg.hrefuses the same build for a tree with its own build system.aes_block.hrefusesCH_AES_RUNTIMEbesideCH_AES_HWorCH_AES_EXTERN. - The build record carries
CH_BUILD_AES_RUNTIME, because the field changesch_cfg's layout, which is entry 77's reason forRAND=session. - The wrappers.
chapulin.hppgainsAesInstructionsandConfig::aes_instructions(). The Zig API gainsAesInstructionsandaes_instructionsinClientandServervalues, declared where@hasFieldfinds the field; null leaves 0, which init refuses. - The checks.
bin/aes_runtime_testruns RFC 9001 and RFC 9369 Appendix A under both answers and the SP 800-38D and FIPS 197 vectors under traffic keys, and counts every call into the table, the instructions and the carry-less multiply: under the absent answer no call goes to the instructions, under the present one the table runs no Initial key, and the table runs no traffic key.test/aes-runtime-qemu.shbuilds that binary and both loop binaries for x86-64 and runs them underqemu-x86_64 -cpu max,-aes,-pclmulqdq: the absent answer passes the vectors and whole QUIC and TCP handshakes, and the present answer and anAES=hwbuild die of SIGILL. CI's mips job runs it on every push, with the qemu-user package that job installs, andtest/docker-aes-runtime-qemu.shruns it in a container elsewhere. QEMU's arm64 models all implement the AES extension, so on arm64 the claim rests on the counts and ontest/aes-runtime-disasm.sh, which CI's arm64 job runs: it disassembles the threeAES=runtimeobjectsmake checklinks and finds the AES and PMULL instructions inaes_hw.c's andghash_hw.c's functions alone.bin/webpki_session_aes_runtime,bin/webpki_loop_aes_runtime,bin/quic_loop_aes_runtime,bin/tcp_blocking_loop_aes_runtimeandbin/srv_flight_test_aes_runtimehold the field to 0, 1, 2 and 3 at every init call, each pair of answers to the suite both ends run, and each refused list and cookie. Theaes_runtime,srv_select_runtimeandquic_config_webpki_runtimeproofs hold the cipher each key takes, the default order and the answer rule over every byte an answer can be.lint-trust-separationadmitsaes_hw.c,ghash_hw.candquic_aes_soft.ctogether inAES=runtime's QUIC rows alone, andaes-two-implementations-in-one-object.violationstays caught. Thirteen mutants intest/violations/break the new rules, and each is caught.inv26-runtime-initial-seal-ignores-answerseals every QUIC Initial packet under the present answer, which no test on a CPU with the instructions can see: both ciphers compute the same packet, andbin/aes_runtime_testlinks noquic.c. The qemu run catches it.
Cost:
- One object holds two AES implementations, which CLAUDE.md forbade.
lint-trust-separationholds the exception to one value and one pair. - Every block in a QUIC object reads which cipher its schedule records, and every schedule carries one more byte.
ch_cfggrows 8 bytes on arm64 in anAES=runtimebuild: one byte, and the alignment of the pointer after it (docs/performance.md).- A present answer on a CPU without the instructions dies of SIGILL, and an absent one on a CPU with them runs ChaCha20. The answer is the caller's, and chapulin cannot check it.
make checkbuilds and runs eight more binaries and three morelib-checkobjects, andlint-zig-buildone more configuration. CI's mips job builds four x86-64 binaries and runs them under qemu, and its arm64 job builds the three objects and disassembles them.
Gain: one object per architecture serves CPUs with and without the AES instructions, and colibri's callers pass the answer their probe found instead of choosing an object at build time.
Rejected:
- An attribute macro on each function. The first draft put
AES_HW_TARGETbefore every function in the two files. Semgrep could not parse a function whose definition starts with an unknown macro, solint-invariantsread those files only in part. The pragma gives each function the same target attribute and leaves each definition asAES=hwwrites it. - Cutting a caller's list down to ChaCha20 under the absent answer. A caller who named AES-GCM asked for an offer the session cannot make, and a silent drop would hide that.
- The build value.
-
CHACHA=vectorcomputes ChaCha20 four blocks at a time in 128-bit vectors, NEON on arm64 and SSE2 on x86-64, andCHACHA=portablestays the default and the reference. At fc02391,chacha20.ccomputed one 64-byte block at a time and XORed one byte at a time. On an Apple M1 Pro under clang those two stages took 33 of the 66 µs thatrec_sealspent on a 16 KiB record, and every build carries ChaCha20-Poly1305 (docs/performance.md, "Where a record's time goes"). Camilo decided on 2026-09-29 (#181) to make the AEAD cheaper with a vector path. The path keeps entry 45's reason for ChaCha20, constant time by construction, gives the portable C's output, and leaves the portable C as the reference.- Four blocks per call, one per lane.
chacha20_vector.cholds each of the 16 state words of four consecutive blocks in one vector of four 32-bit lanes. It runschacha20.c's rounds on the 16 vectors with adds, exclusive-ors and fixed rotations, transposes the result into block order, and XORs it into the output 16 bytes at a time.chacha20_xorcalls it under-DCH_CHACHA_VECTOR.chacha20_block, which derives the Poly1305 key, stays the portable function in both builds. chacha20.h's contract, unchanged. The lanes hold the counters counter to counter + 3, and the adds wrap modulo 2^32, asstate[12]++does. A last group of 1 to 255 bytes computes four blocks into a buffer and XORs its bytes one at a time, aschacha20.cXORs its last block. Each 16 bytes are read before the 16 at the same offset are written, in ascending order, so the output may sit on the input or below it, whererec_openputs it.- NEON and SSE2, from the compiler's macros. Every AArch64 core
has NEON and every x86-64 core has SSE2, so a compiler for either
target defines
__ARM_NEONor__SSE2__. Nothing probes a CPU, and no caller passes a probe's answer: unlike the AES instructions, these come with the architecture.chacha20_vector.hstops a build for any other target, and for a big-endian one, whose lanes would store their bytes in the wrong order. SoCHACHA=vectornever falls back to the portable loop.test/chacha-builds.shchecks both refusals, and that a vector object calls the path. - An axis, as
X25519is (entry 52). A device core has neither instruction set, so the portable loop stays the default and the device path, with its proof. The values name what each path needs from the target. The vector path asks for no timing statement of its own, whereX25519=wideasks forCH_NATIVE_MUL128: it multiplies nothing, reads no table, and runs the operations the portable loop runs. The build record leaves the axis out, as entry 56 leavesX25519=wideout, because no public layout or bound reads it. - 128-bit vectors only. They are the widest that every core of both architectures has. No machine here runs x86-64, so nothing could show that AVX2 pays, and the ruling admits AVX2 only on a measurement. An eight-block NEON variant, two groups of four in one round loop, ran faster in a scratch timing loop under clang, which is no measurement this tree records. It is a second tradeoff, with more register pressure on SSE2's 16 registers, so it waits for a change of its own, measured alone.
- Poly1305 stays as it is. Its cost in the packaged object is the
widening multiply:
ct.hbuilds each 32x32->64 product from 16x16 pieces.WIDEMUL=native, the builder's statement that the multiply runs in constant time, already runs the same Poly1305 in about a third of the time. A vector Poly1305 needs a widening multiply too, so it would need the same statement. Entry 83 adds one under that statement.
Cost: a second ChaCha20 to keep equal to the first. The pair is 313 lines. On arm64 under clang at
-O2it adds 1,192 bytes of text to a host object, and takes the stack belowchacha20_xorfrom 320 to 768 bytes, whichlint-stackholds to the device budget in theCHACHA=vectorleg. Nothing proves the path, because CBMC cannot read an intrinsic.make checkholds it withbin/chacha20_equiv_test, 30,771 cases against the portable loop;bin/unit_chacha_vector, which runs RFC 8439's vectors and every record the unit suite seals; a Wycheproof leg; a packaged-object leg; and the codegen gate's two 64-bit specs, which hold its branches at 12. Seven violations break the path, and each is caught (docs/verification.md, "The CHACHA=vector path").Gain: the record bench's second run, a filter's figures, took a 16 KiB
rec_sealfrom 67.3 to 47.6 µs on macOS clang, from 61.6 to 41.6 µs on the Linux VM's clang and from 83.7 to 65.6 µs on gcc 13. WithWIDEMUL=nativeas well, it takes 26.1 to 34.2 µs, 39% to 42% of the packaged portable record (docs/performance.md). - Four blocks per call, one per lane.
-
CHACHA=vectorwithWIDEMUL=nativeruns Poly1305's block loop four blocks at a time in two vector lanes, NEON on arm64 and SSE2 on x86-64, andCH_NATIVE_WIDEMULstates the timing of every widening multiply the object runs, scalar or vector. Entry 82 left Poly1305 as it was. In that entry's record runs, with the vector ChaCha20 in place andWIDEMUL=native, Poly1305 took 11.2 to 12.8 µs of the 26.1 to 34.2 µs a 16 KiBrec_sealtook, most of what the vector ChaCha20 left. A vector Poly1305 multiplies, so it needs a statement about the multiply's timing. Camilo ruled on 2026-09-30 (#181) thatCH_NATIVE_WIDEMULis that statement for every widening multiply the object runs, scalar or vector, NEON's UMULL and UMLAL and SSE2's PMULUDQ among them, with no new flag, and that the vector Poly1305 builds only underCHACHA=vectorwithWIDEMUL=native.- Two lanes, four blocks a group.
poly1305_vector.cholds two accumulators of five 26-bit limbs, one per lane, aspoly1305.cholds one. Lane 0 takes the first and third block of each group of four, and lane 1 the second and fourth. For each group both lanes compute (h + the lane's first block) * r^4 + (the lane's second block) * r^2, two steps of Horner's rule over every other block, as five sums of 32x32->64 products, and carry the sums back into limbs. The last group multiplies lane 1 by r^3 and r instead, the powers its blocks are owed, and the two lanes' sums add up to the accumulator. Each call computes r^2, r^3 and r^4 fromp->rwith three scalar multiplies, so thepoly1305context gains no field. poly1305.cstays the reference and chooses in one place.whole_blocksinpoly1305.cis the one call site that picks between the loops, so a choice the caller makes at run time (#186) can go there alone. It hands the path every whole group of an update that holds at least 128 bytes of whole blocks, and its own loop takes the blocks after the last group. The buffered partial block and the final reduction stay inpoly1305.cin every build. The path hands back the accumulator's value modulo 2^130 - 5 with every limb at most 2^26, inside the boundspoly1305.c's loop keeps, so the loop andpoly1305_finalread it as they read their own.- A threshold of two groups. Below it, the powers of r cost more
than the lanes save. In a scratch timing loop on the M1 Pro the path
was slower than the portable loop over one group and faster over
two, under Apple clang 21, clang 18 and gcc 13. That is no
measurement this tree records, and no x86-64 machine has timed the
SSE2 arm, so
POLY1305_VECTOR_MINis 128 bytes until one does. - Two carry rounds where the portable loop has one chain. In that timing loop, clang moved each carry of a chain into the next sum's chain of multiply-adds, which the adds allow, and so each sum waited on the one below it. A round adds a shifted value to a masked one, which leaves no chain to move a carry into, and two rounds leave limbs below 2^26 + 2^10. gcc kept every array that a loop indexes in memory, so the file indexes limbs by constants alone.
- One statement for both multiplies. A builder who defines
CH_NATIVE_WIDEMULfor an object that is alsoCHACHA=vectorstates the timing of the part's vector widening multiplies, not only its scalar one;ct.hsays so. Without the define aCHACHA=vectorobject runspoly1305.c's loop and its 16x16 decomposition.poly1305_vector.hturns the path on only whereCH_CHACHA_VECTORandct.h'sCH_WIDEMUL_NATIVEmeet, soCH_CT_WIDEMULturns it off with the scalar multiply, and the Makefile andbuild.zigpackagepoly1305_vector.conly in an object that is both. - Each call wipes the powers of r. r and a tag seen on the wire
give the pad s, and r and s forge any message under that one-time
key, such as a lost QUIC packet with its bits flipped. Each power
gives r back, by a root modulo 2^130 - 5. So the call keeps r^2,
r^3, r^4 and the two multipliers built from them in one struct, and
wipes it through
ct_wipeonce when it ends: 208 bytes on NEON and 352 on SSE2. Registers and the spill slots the compiler picks stay out of reach, as they do for every wipe written in C.bin/poly1305_equiv_testcopies the stack below a call and requires none of the three powers there, in any layout the call holds them in. The wipe came after the scratch timing that set the threshold, which does not measure it, and after the paired record runs below. The record runs docs/performance.md holds now measure the code with it. - Entry 82's limits hold. 128-bit vectors only, no probe of the CPU, and nothing new in the build record: no public layout or bound reads the path.
Cost: a second Poly1305 block loop to keep equal to the first. The pair is 443 lines. On arm64 under Apple clang 21 at
-O2it adds 1,872 bytes of text to a host object, 1,824 inpoly1305_vector.cand 48 inpoly1305.c. It takes the stack belowpoly1305_updatefrom 80 to 400 bytes in aCHACHA=vector WIDEMUL=nativebuild, 336 of them the path's frame with its struct of powers, and the deepest path fromaead_sealthroughmacfrom 400 to 720. The peak belowaead_sealstays where the vector ChaCha20 put it, 848 bytes:aead_seal's 80 overchacha20_vector_xor's 768, asbench/stack.pyreports.lint-stackholds the path's frames to the device budget in thecheck-lib-chacha-vector-widemulleg. Nothing proves the path, because CBMC cannot read an intrinsic.make checkholds it withbin/poly1305_equiv_test, 43,282 cases against the portable loop and the check of the stack a call leaves;bin/unit_chacha_vector, which runs RFC 8439's A.3 and A.5 vectors on it; a Wycheproof leg; a packaged-object leg on each CI architecture; and the codegen gate's two 64-bit specs, which hold its branches at 4. Nine violations break the path, and each is caught (docs/verification.md, "The CHACHA=vector Poly1305").Gain: in paired runs of the record bench, a filter's figures, a
CHACHA=vector WIDEMUL=nativebuild took a 16 KiBrec_sealfrom 26.2 to 17.4 µs on macOS clang, from 25.5 to 16.8 µs on the Linux VM's clang and from 33.3 to 23.1 µs on gcc 13, and Poly1305 over its ciphertext from 11.6, 11.2 and 12.8 µs to 2.7, 2.6 and 2.7 µs. That record takes 26% to 28% of the packaged portable one's time (docs/performance.md, "Where a record's time goes"). These runs came before the wipe of the powers, onect_wipeof 208 or 352 bytes per call. The runs that performance.md held at d10b6e2, taken with the AES-GCM changes of #184, measure the code with the wipe: that build's Poly1305 takes 2.8 µs on macOS clang and gcc 13, and its 16 KiBrec_seal17.2 and 18.4 µs. On the VM's clang the Poly1305 takes 4.7 µs in that bench build, where the linker's placement of the function, not its code, causes the difference (docs/performance.md, the pitfalls). - Two lanes, four blocks a group.
-
Every client entry takes a PSK identity of 1 to
CH_TICKET_ID_MAXbytes, refuses a ClientHello it cannot stage before a byte goes out, and keeps the session's epoch report when it refuses. Two answers differed betweench_connectand the two non-blocking client entries,ch_record_initandch_quic_init, and one of them was reachable because the raw and ca modes put no bound on a PSK identity.- The hello.
ch_connectbuilt its first ClientHello inside the handshake, so a hello too long forch_tls.txfailed there: an internal_error alert went out in the clear, andch_connectreturnedCH_ECAP. The two non-blocking entries build that hello before they return, and they refused it withCH_EINVALand nothing staged.ch_handshakenow builds the first hello before its first send and refuses it the same way, so all three returnCH_EINVAL, send nothing and record no alert (INV-13). The retry hello a HelloRetryRequest asks for still fails the handshake withCH_ECAPand internal_error in every driver, because the first hello has gone out by then. - The identity bound.
hello_buildproves thatCH_HELLO_MAXholds every hello whose PSK identity is at mostCH_TICKET_ID_MAXbytes.webpki_cfg_okheld a ticket's identity to 1 toCH_TICKET_ID_MAXbytes, and the raw and ca configuration checks held an external identity to no length at all. A longer identity reached the first hello's refusal, and one the first hello still held could make the retry hello, which echoes a cookie, too long in the middle of the handshake. Camilo ruled on 2026-09-30 that every raw and ca client entry refuses an identity that is not 1 toCH_TICKET_ID_MAXbytes withCH_EINVAL, throughpsk_id_len_okintls.candquic_config.c. That check is what makes both branches unreachable for every configuration a client entry accepts, and each branch still fails closed as it did, with no test that reaches it. The empty identity those checks took goes too: RFC 9846 §4.3.11 gives an identity at least one byte (rfc9846.txt:2468-2471). - The epoch report.
ch_record_initandch_quic_initzeroed the session on every refusal. After refusing a ticket that the stored epoch retired, they leftCH_EPOCH_NONEinch_tls.epoch_status, wherech_connectleavesCH_EPOCH_REVOKED(docs/ca.md). Each now zeroes the session but forch_tls.epoch,epoch_seenandepoch_status, throughtcp_nonblocking_refuse_initandquic_refuse_init. The two server entries call the same functions, and a server's three fields are zero.
Rejected:
- Zeroing
ch_connect's epoch report instead. No rule asks for a zeroed session. The non-blocking entries zero theirs so that nothing stays staged and no key share outlives the call, and three fields that hold no secret change neither. Zeroing them would remove the one field that tells a retired ticket from every other refusal. CH_ECAPfrom the two non-blocking entries. A non-blocking call that returnsCH_ECAPleaves its session live (INV-13), and a peer that has read no byte is owed no alert.- Holding a longer identity. A larger
ch_tls.txwould carry more thanCH_TICKET_ID_MAXbytes of identity, at an SRAM cost in every build, for an identity no ticket this tree takes can have.
Cost: two functions of ten lines, one per non-blocking transport, and
ch_handshakecarries the first hello's length into the handshake. An external PSK identity longer than 320 bytes, which RFC 9846 allows, no longer connects in the raw and ca modes.Gain: whichever client entry a caller uses, a refusal returns one code and leaves one epoch report, and no configuration a client entry accepts makes a ClientHello its staging array cannot hold.
- The hello.
-
An AES-GCM open decrypts while it hashes, and wipes the plaintext it wrote when the tag does not match.
gcm_opencomputed the tag over the ciphertext, compared it and only then decrypted, so it read the ciphertext twice. On the AES instructions the seal runs counter mode and GHASH in one loop (gcm_hw.c), and the open could not, because it wrote no plaintext before the comparison. OpenSSL's open runs both in one loop and wipes its output on a mismatch. Camilo ruled on 2026-09-30 (#184) that chapulin's open does the same, and thatgcm.hpromises this instead: no plaintext is returned on a bad tag, and the open wipes its output before it returns.- The order.
open_scheduleingcm.cstarts GHASH over the associated data, hands the whole passes of eight blocks under a schedule the AES instructions run togcm_open_passes_hw, runs GHASH and then counter mode over the rest, and compares the tag throughct_memeq.gcm_open_passes_hwruns one pass ahead: iteration p hashes pass p and decrypts pass p - 1, so the GHASH of one pass runs beside the AES rounds of the one before. Every AES value takes this order, the table's included, because one body serves them all. - Aliasing.
gcm.hstill admitspt == ctandptbelowct, andquic_packet.c,quic_initial.candrecord.copen in place. A plaintext write can then overwrite ciphertext, so both loops hash each ciphertext byte before they write to its address. - The wipe. On a mismatch the open wipes the n bytes it wrote
through
ct_wipeand returns 0, and it writes no byte outside them. The branch on the comparison's verdict tells an observer nothing the return value does not tell the caller. The plaintext of a forged ciphertext is that ciphertext exclusive-ored with the keystream of the nonce, so leaving it would hand the sender that keystream. - What a caller finds after a discard.
ch_quic_openworks in place. A packet whose tag does not match leaves its header with the protection removed, which the call does before the AEAD runs, as it did before, zeros where the payload was, and the tag and every byte pastpkt_lenas they arrived. colibri compares a datagram's last 16 bytes with its Stateless Reset tokens (RFC 9000 §10.3.1,rfc9000.txt:3486-3497), and those bytes are the last packet's tag, which no open writes.quic.handquic_initial.hnow state this where they called those bytes unspecified. A failedrec_openends its session (INV-13), and its record holds zeros where the payload was. - ChaCha20-Poly1305 keeps its order.
aead_openstill verifies first and writes nothing on a mismatch. The promise above admits both orders, so moving it later changes no caller.
Rejected:
- Comparing first on the instructions. It reads the ciphertext twice, which is the time the gain below measures.
- Decrypting into a scratch buffer and copying on a match. That takes a second buffer as large as a record, 16 KiB of SRAM, and a second pass over the data.
- Leaving the plaintext for the caller to drop. The caller would hold unauthenticated plaintext and the keystream it gives away.
Cost: a forged record now costs counter mode and a wipe on top of GHASH.
ct_wipewrites one byte at a time. In a scratch timing loop on the M1 Pro under Apple clang, a failed open of 1 KiB took 845 ns where the comparison first took 375 ns, and of 16 KiB 8.0 µs where it took 1.3 µs; a genuine open of 16 KiB went from 3.0 to 2.5 µs in the same loop. A failed TLS record ends its session, so it happens once per connection. A QUIC endpoint discards a forged packet and goes on, and a packet that fits a 1,500-byte path costs about the 1 KiB figure.Gain: in paired runs of the record bench, a 16 KiB
aead_openwent from 2.99 to 2.57 µs under AES-128-GCM and from 3.39 to 2.99 µs under AES-256-GCM on macOS clang, from 3.00 to 2.86 and 3.44 to 3.27 µs on the Linux VM's clang 18, and from 3.23 to 2.84 and 3.65 to 3.28 µs on gcc 13. The seal did not move.Guards.
test/gcm_tests.hforges one tag bit in every SP 800-38D case and requires zeros after the call and the byte after them untouched, inbin/quic_testand the legs that run those vectors.bin/ghash_equiv_test, the Wycheproof AES-GCM suite andbin/diff_quic_testrequire zeros from every refused open, andbin/ghash_equiv_testopens in place and five bytes below the ciphertext on both paths.bin/quic_suite_testforges the tag of the first of two packets in one datagram and checks every byte of the datagram after the failed open.gcm-open-keeps-plaintextdrops the wipe, andbin/quic_testfails;gcm-hw-open-hashes-after-decrypthashes each pass after decrypting it, andbin/ghash_equiv_testfails;gcm-open-tail-hashed-after-decryptruns counter mode over the rest before GHASH, andbin/quic_testfails.gcm_refusalproves the zeros below n and nothing written past n for any tag, on the portable path; no harness readsgcm_hw.c. - The order.
-
A
CHACHA=vectorpass computes eight blocks on NEON, two groups of four side by side, and four on SSE2, and XORs its keystream into the data from the registers that computed it. Entry 82 computed one group of four blocks a call, stored it to memory and XORed it into the data in a second loop, and it left an eight-block NEON variant for a change of its own, measured alone. In the record bench at 9b72b67, ChaCha20 took 14.0 of the 17.3 µs a 16 KiBrec_sealtook on macOS in aCHACHA=vector WIDEMUL=nativebuild, where OpenSSL sealed the same record in 9.4 (docs/performance.md).- Where the time went. In a scratch timing loop on the M1 Pro under Apple clang 21, over 16 KiB in place, entry 82's path took 13.9 µs. The same group with its XOR from registers took 13.4; that group without its four transposes, 13.0; with the rotation by 8 as one table lookup (TBL) in place of a shift and an insert, 12.5. Two groups a pass took 8.3. The rounds of one group kept every word in a register, with no load or store, so they took most of the time, and they waited on themselves: each quarter round is a chain of 15 operations on NEON, each waiting on the one before, and one group runs four such chains at once.
- Two groups on NEON. A pass holds two groups' 32 words, which
fill arm64's 32 vector registers. The compilers keep a few of them
on the stack: per double round the loop loads or stores a vector 11
times under Apple clang 21, 23 times under clang 18 and 42 under
gcc 13, beside 240 vector operations. The pass's three loops over
its groups carry
#pragma GCC unroll 2, for the reasongcm_hw.c's loops carry theirs (docs/performance.md, the pitfalls): gcc keeps an array that a rolled loop indexes in memory. - One group on SSE2. SSE2 has 16 vector registers, which one group's 16 words fill. No x86-64 machine here can time a second group, and an emulated one says nothing about time, so SSE2 keeps one group a pass. It takes the XOR from registers and the last bytes below as NEON does.
- The XOR from registers. A pass XORs each row of 16 bytes from
the vector that holds its keystream, in ascending order, reading
each row before it writes it, as entry 82 did from memory, so the
output may still sit on the input or below it. The last pass tests
each row against the bytes left, and the last 1 to 15 bytes pass
through a buffer of 16 bytes that
ct_wipeclears. Entry 82 left its last 256 bytes of keystream on the stack. No wipe written in C clears a register or a spill slot the compiler picks, here as elsewhere. - Rejected: the rotation by 8 as a table lookup. It took 6% off Apple clang's two-group pass. Under gcc 13 it took from 5% off to 27% more, depending on the order of the rounds in the source: the lookup's index takes a register, and gcc then kept more of the state on the stack. gcc is the compiler CI runs, so the rotation stays a shift and an insert.
- Unchanged. Entry 82's constant-time argument and its limits:
128-bit vectors, no probe of the CPU, nothing in the build record.
chacha20_blockstays the portable function.
Cost: on arm64 under Apple clang 21 at
-O2,chacha20_vector.c's text grows from 1,744 to 3,108 bytes. Its conditional branches inmake lint-wide-multiplyrise from 12 to 40 on arm64 and 23 on x86-64: most test whether the last pass's limit covers a row, and each tests the byte count.bench/stack.pyputs the stack belowchacha20_xorin aCHACHA=vectorbuild at 544 bytes, from 768, because no call stores a group to memory; withWIDEMUL=native, the peak belowaead_sealfalls from 848 to 720, and its deepest path now runs throughmacand the vector Poly1305.bin/chacha20_equiv_testruns every length to 2,048 bytes, which crosses a NEON pass's edge four times, and the counter's last 17 values at every length to 20 blocks: 49,211 cases. The four violations that break the path's rows, counters and last bytes moved to the new code, and the equivalence test catches each, on NEON and, under emulation, on SSE2.Gain: in paired runs of the record bench, a filter's figures, a
CHACHA=vector WIDEMUL=nativebuild took a 16 KiBrec_sealfrom 17.8 to 12.1 µs on macOS clang, from 17.5 to 12.2 µs on the Linux VM's clang and from 18.4 to 12.7 µs on gcc 13, and ChaCha20 in place from 14.5, 14.0 and 15.0 µs to 8.8, 8.4 and 9.2 µs. The first pair agreed: 0.67, 0.64 and 0.67 of the time before, where the second gave 0.68, 0.70 and 0.69. OpenSSL seals the same record in 9.4, 10.2 and 10.2 µs on the same machine. -
A
WIDEMUL=runtimeobject holds both widening multiplies, and the caller's answer picks one for each session (#186).WIDEMUL=decomposed, the default, builds every widening product fromct.h's 16x16 pieces, andWIDEMUL=nativetakes the CPU's multiply on the builder's statement that it runs in constant time (#53, entry 83). Both fix the multiply when the object is built. On arm64 the multiply runs in constant time only on a core with FEAT_DIT and only while the thread has set PSTATE.DIT, and on x86-64 only on a part in Intel's DOIT list while the operating system has set DOITM. So a host program whose threads differ, or which runs on CPUs that differ, had no object that fit. Camilo decided on 2026-09-30.- The build value.
WIDEMUL=runtimedefinesCH_WIDEMUL_RUNTIMEand compiles twice each file built on the multiply that the object carries:poly1305.c,x25519.c,mlkem_poly.c,p256_field.c,p256_scalar.candrsa_sign.c(WIDEMUL_COPIEDin the Makefile,widemul_copiedinbuild.zig). UnderCHACHA=vectorit addspoly1305_vector_native.c, the vector Poly1305, as a native copy alone, because that path runs on the native multiply only. - How a file compiles twice. The file under its own names compiles
as a
WIDEMUL=decomposedobject compiles it:test/widemul-builds.shrequires the same assembly with the runtime define as without it, so the proofs and the recorded ceilings of that file hold for this copy unchanged. The native copy is<file>_native.c, two lines:#include "widemul_native.h", then#include "<file>.c".widemul_native.hdefinesCH_WIDEMUL_NATIVE_COPY, which makesct.htake the native multiply in that translation unit alone, and gives each of the 52 names the seven files define outside their unit a second name ending in_native. A name missing from that list is defined by both copies, and the link of the object refuses it. An auditor reads one source per file, one two-line wrapper and one list of renames. - The answer.
ch_cfg.widemul, declared only in aWIDEMUL=runtimeobject, takesCH_WIDEMUL_CONSTANT_TIME, which says the multiply runs in constant time on this CPU in the mode the session's thread runs in, orCH_WIDEMUL_NOT_STATED.ch_connect,ch_record_init,ch_quic_init,ch_srv_accept,ch_srv_record_init,ch_srv_quic_initandch_srv_checkreturnCH_EINVALfor any other value, 0 included, before they send anything, so every caller states its answer. - The mode is the caller's. Setting PSTATE.DIT on arm64, and the
DOITM policy on x86-64, belong to the caller and its operating
system. chapulin writes no CPU state and probes nothing (
cpu_cfg.h), for the reason entry 81 gives for the AES instructions: arm64 code that reads PSTATE.DIT on a core without FEAT_DIT takes SIGILL, and code in user mode cannot read DOITM. - One branch per operation.
widemul.hholds one dispatcher per entry built on the multiply that code outside the seven files calls:poly1305_update,poly1305_final,x25519,x25519_base,mlk_polyvec_compress,mlk_poly_compress,mlk_poly_tomsg,p256_fe_mul,p256_fe_sqr,p256_fe_to_mont,p256_fe_from_mont,p256_fe_inv,p256_scalar_mul,p256_scalar_inverse,rsa_pss_signandrsa_sp1. Each branches once on the answer, which the caller chose and which is not secret: once per Poly1305 update and final, once per X25519 scalar multiplication, once per ML-KEM compression, once per P-256 field or scalar multiply, once per RSA signature, and never once per product. No function pointer is involved.CH_WIDEMUL_CONSTANT_TIMEruns the native copy, and every other byte runs the file under its own names, so a direction whose answer was never written takes the decomposition. - Where the answer travels. Every operation built on the multiply
takes the answer as its first argument in every build, and an object
that holds one multiply passes
WIDEMUL_BUILD_ANSWER, which its dispatchers ignore. A session passes its configuration's answer. Each TCP init call writes it into both record directions once it accepts the configuration (rec_dir.widemul, in bytes the alignment of the sequence number left unused), and each QUIC packet call reads it from its session. - The vector Poly1305. Under
CHACHA=vectorthe constant-time answer also runs the vector Poly1305:poly1305_native.c's block loop callspoly1305_vector_blocks_native, andpoly1305.cunder its own names calls no vector path (entry 83). - What the build refuses.
ct.hrefusesCH_NATIVE_WIDEMULbesideCH_WIDEMUL_RUNTIME, which would give the files under their own names the native multiply, and a native copy outside aWIDEMUL=runtimeobject. The Makefile,build.zigandct.heach refuseX25519=widebeside it: that field'sCH_NATIVE_MUL128states the 64x64->128 multiply's timing when the object is built (entry 52), which is the statement this value moves to each session. A copy of the wide field per answer is left for a later decision. - The build record carries
CH_BUILD_WIDEMUL_RUNTIME, because the field changes the layout ofch_cfg, and with it ofch_tls,ch_recordandch_quic, which is entry 77's reason forRAND=session. - The wrappers.
chapulin.hppgainsWidemulandConfig::widemul(). The Zig API gainsWidemulandwidemulinClientandServervalues, declared where@hasFieldfinds the field; null leaves 0, which init refuses. - The checks.
bin/widemul_runtime_testcompiles the seven files again under counted names and runs the AEAD, X25519, ML-KEM, P-256, RSA signing and record operations under each answer: the native copies alone under the constant-time answer, the files under their own names alone and as many times under the other, the decomposition under 0, 3, 0x80 and 0xff, the same bytes under all of them, and the vector Poly1305 under the constant-time answer alone. The unit, ML-KEM, P-256 and RSA signing vectors and both Wycheproof legs run once per answer.bin/tcp_blocking_loop_widemul,bin/tcp_nonblocking_loop_widemul,bin/quic_loop_widemulandbin/webpki_session_widemulhold the field to 0, 1, 2 and 3 at every init call andch_srv_check, and run a whole handshake for each pair of answers with the same counts. Where both ends are this tree's sessions they count each end's calls around its own calls, so each end runs the copy its own answer names, the QUIC one through a Handshake and a 1-RTT packet each way.test/widemul-builds.shholdsct.h's three refusals, the assembly of each file under its own names, the vector calls, and the Makefile's andbuild.zig's lists and refusal.lint-trust-separationadmits the native copies inWIDEMUL=runtime's rows alone and requires each beside its file.lint-wide-multiplyrecords each native copy's own ceilings under every spec: the products it asks of the native multiply, which is what the copy is for, and its branches, read against the file under its own names, which keeps its ceilings.lint-runtime-symbolsadmits rv32ic's__muldi3in the native copies, whichsoftmul.csupplies in constant time. Fifty-three mutants intest/violations/break the new rules, and each is caught. - The proofs. No harness compiles
CH_WIDEMUL_RUNTIME, and no copy needs a run of its own.proof/run.shcompiles each file on the native multiply, which is the native copy's text under other names, andctwidemulcarries those verdicts to the decomposition, which is the file under its own names (docs/verification.md). Under twelve build configurations every harness preprocesses to the text it had before this change, one assertion's line number aside. The dispatchers have no harness; the counting test holds them. - CI. The check job runs all of it on x86-64. The arm64 and macOS
jobs run the
WIDEMUL=runtimebinaries under both answers insuite-checkand package aCHACHA=vector WIDEMUL=runtimeobject.
Cost:
- Flash. The default build on
WIDEMUL=runtimetakes 35.9 kB on mips32r2 at-Os, 5.7 kB more than the default (docs/performance.md). On arm64 at-O2, counted asbench/device-ram.shcounts host flash, the packaged object grows from 35,754 to 44,647 bytes for the raw client, from 97,264 to 139,882 for the server, and from 158,145 to 203,023 for theCHACHA=vectorQUIC object colibri links onAES=runtime. A native copy keeps the entries no dispatcher calls, such asp256_fe_add_native, and a link that drops unreferenced functions can remove them. ch_cfggrows 8 bytes on arm64, one byte and its alignment, so the default session struct is 1,168 bytes against 1,160. In colibri's QUIC object the byte sits besideaes_instructions, andch_quicdoes not grow.- One compare and branch per dispatched call. No bench measures a
WIDEMUL=runtimeobject. - A constant-time answer on a thread without DIT, or on a part outside the DOIT list, runs the native multiply where it is not constant time. The answer is the caller's, and chapulin cannot check it.
make checkbuilds and runs fifteen more binaries, two more Wycheproof legs, three morelib-checkobjects and one more script, andlint-zig-buildone more configuration.
Gain: one object serves threads and CPUs whose multiply runs in constant time and those whose multiply does not, and its caller states which at each init, as it states the AES instructions (entry 81).
Rejected:
- The renames in the build. The Makefile could compile each file a
second time with
-Drenames. The renames would then live in the Makefile, inbuild.zigand in every build a firmware tree writes, where a reader of the sources does not see them. - A template expanded twice. Each file could define its functions through a naming macro and be included twice with two suffixes. Every function of the seven files would change shape, and the proofs and lints would read macro-built names.
- A function pointer per operation, or a branch per product. A table of pointers would make every call an indirect call an auditor has to trace to its table, and a branch in each product loop would put the answer in every inner loop. One branch per operation keeps the files as they were.
- The build value.
-
A recipe or script that builds outside
make checktakes its sources from a list a rule check builds links, and check builds what a script still lists itself. Recipes outsidemake checknamed their sources by hand, so a change that adds a call from one file into another broke only them. It happened twice on 2026-09-30: the AES=hw build inbench/aead.shlackedgcm_hw.c, and bench.yml's aead job failed to link (9b72b67);san-check's build ofbin/san/chacha20_equiv_testlackedct.concechacha20_vector.ccalledct_wipe, and CI's san job failed to link (c798fb8).make checkbuilt neither, so both landed on main.- The lanes. A test that check builds keeps the sources it links
in a variable its rule reads, such as
RSA_TEST_SRCSorCHACHA20_EQUIV_TEST_SRCS, andsan-check,cross-check,m3-check,coverage, theCH_CT_WIDEMULbuilds and theWIDEMUL=runtimebuilds read the same variable. The Wycheproof builds readWYCHEPROOF_SRCS, and the three differential arms andtest/spec_coverage.pyreadDIFF_SRCS, whichbin/diffreads. - The scripts. A script asks make where a variable names what it
links:
bench/aead.shandbench/record.shreadAES_HW_SRCS(print-aes-hw-srcs), andtest/aes-runtime-qemu.shreads the lists its four binaries' rules link (print-aes-runtime-qemu-srcs). The three instruction-count scripts read one list of sources and defines,INSN_SRCSandINSN_DEF(print-insn-lists), where each kept a copy before. - What check builds. The rest of what a script lists is the
script's own choice: the AEAD sources a bench times, or the
modules the Cortex-M3 known answers need.
check-script-buildsrunstest/script-builds.sh, which builds every such program with the host's compiler and runs none:bench/aead.sh --build,bench/record.sh --build,bench/primitives.sh --build,test/qemu-m3.sh --build, andbench/insn_driver.coverINSN_SRCSandINSN_DEF. The step skips when it passed before on the same inputs (INV-37). - Found on the way. The three instruction-count scripts had not
compiled since b309c92, when
record.cbegan to readcfg.hthroughsuite.hand their builds declared no entropy pattern.INSN_DEFdeclaresCH_RAND_EXTERN, which the driver never draws from.bench/record.shhad not compiled since 9cf2028, which gaveaead.c'smacthe multiply's answer whilebench/record_aead.c, which includesaead.c, still called it with seven arguments. - Left as they are.
proof/run.shnames each harness's sources itself, because a harness chooses which callees are stubs, and it rejects a result whose log names a callee with no body.bench/audit-mips.shcompiles one file at a time and links nothing, andtest/quic-builds.shandtest/chacha-builds.shcompile single files on purpose.
Cost: one more step in check, which builds 23 programs in about 45 s of CPU and 9 s of wall time on an M1 Pro whenever a
.cor.hfile changes; INV-40 states the rule, andinv40-aead-calls-hkdfshows the step catches a call a script's list misses. Gain: a source list outside check can no longer miss a source its code calls without check failing first. The differential arms are the exception, becausebin/diff, whose list they share, builds inmake check-slow. - The lanes. A test that check builds keeps the sources it links
in a variable its rule reads, such as
-
A host object holds every fast path beside the portable code and picks among them at init from
ch_cfg.cpu, the caller's description of its CPU, and a device object stays portable (#181, #183, #186). Four build variables choose the speed paths today:AES,CHACHA,WIDEMULandX25519. Entries 81 and 87 moved two of those choices to each session, one field at a time, and colibri buildsAES=runtimewithCHACHA=vector. Camilo ruled on 2026-09-30:- Direction. A host object, for arm64 or x86-64, compiles each fast
path for its own instructions, as
AES=runtimecompilesaes_hw.c, beside the portable code, and picks at init. A device object stays portable.AES=externandWIDEMUL=nativestay as device options. The product variables stay:TRANSPORT,ROLE,TRUST,KEX,RANDandSUITE.AES,CHACHA,WIDEMULandX25519leave the host build, and AVX2 and VAES never become variables. - Discovery. The caller probes the CPU and passes the result in. chapulin still probes nothing.
- Shape.
ch_cfg.cpuis auint32_tof bits. Every init refuses a value withoutCH_CPU_PROBED.CH_CPU_CONSTANT_TIME_AESsays the CPU has the AES and carry-less multiply instructions and states that they run in constant time,CH_CPU_AVX2says it has AVX2, andCH_CPU_VAESthat it has VAES and VPCLMULQDQ on 256-bit registers.CH_CPU_CONSTANT_TIME_MULTIPLYis the caller's statement about the multiply, and PSTATE.DIT and DOITM are the caller's to set. A caller written before a release that adds a bit leaves that bit clear and runs the slower path. A bit for another architecture is refused, not ignored. The field replacesch_cfg.aes_instructionsandch_cfg.widemul.
Camilo then ruled on four points the direction left open: the host test, the field in a device object, the AES timing statement and the bit that picks the wide X25519 field. He named the AES bit
CH_CPU_CONSTANT_TIME_AESso that its name states the claim it carries. The sections below state each ruling and its reason. Once the code lands, this entry amends entries 50, 52, 68, 80, 81, 82, 83, 86 and 87 where they state a speed variable.The host test. A target is a host target when its compiler passes the three probes the Makefile runs today for
AES=runtime,CHACHA=vectorandX25519=wide: it defines__aarch64__or__x86_64__, it defines__ARM_NEONor__SSE2__for a little-endian core, and it defines__SIZEOF_INT128__. Every LP64 compiler for the two architectures passes all three. The Makefile andbuild.zigrun the test, and a host build passes one define,-DCH_CPU_RUNTIME, named afterCH_AES_RUNTIMEandCH_WIDEMUL_RUNTIME, which it replaces. The sources choose on that define and never on the architecture macros, andcpu_cfg.hstops a build that defines it for a target that fails the test. Reason: a firmware tree with its own build, every proof harness,bench/sram.shand the test binaries of the portable code compile the sources without the define, so each gets the code it compiles today. cbmc defines the architecture macros of the machine it runs on,__aarch64__and__ARM_NEONon an arm64 Mac, so sources that chose on those macros would hand CBMC intrinsics, which it cannot read.The product picks the object on a host target. A device client,
ROLE=clientwith a raw or caTRUST, builds the portable object on every target. Entry 44 calls that client a pinned firmware image, and the Makefile calls it a device client when it refusesSUITE=aesgcmfor it (entry 45).TRUST=webpki,ROLE=serverandROLE=bothbuild the host object on a host target. So an arm64 device that runs Linux and pins its server gets the small object throughTRUST, and no seventh variable exists. Reasons: the default build, aTRUST=raw-rsaclient, stays the portable object on every development machine, somake lib, the examples,lint-stack's default budget and every firmware caller see what they see today; and a webpki or server program already needs the clock, the hostname and the buffers of a host. Cost: a raw or ca client on a 64-bit host never runs a fast path, and a webpki or server program on an arm64 device always takes the host object. A check that packages the device object of a server, such as today'sAES=externserver leg, must build for a device target, throughbuild.zig's cross linker or a CI cross lane, or set the result of the host test on its own command line, aslint-trust-separationwould.No field in a device object.
ch_cfg.cpuexists only in a host object, asaes_instructionsexists only underAES=runtime. A device object holds one path per primitive, so the field would choose nothing, and it would add 32 bits to every session's copy ofch_cfg. Required there, it would make every firmware caller's init returnCH_EINVALuntil the caller set it, a failure no compiler reports. Accepted as 0 there, its rule would differ between objects anyway. Cost: a program built for both kinds of object writes the field under#ifdef CH_CPU_RUNTIME. The build record's bit tells a consumer which layout it has, andch_build_matchesrefuses a mismatch.CH_CPU_CONSTANT_TIME_AESstates the timing of the AES instructions. In a host object the bit says that the CPU has the AES and carry-less multiply instructions, and that the caller states they run in constant time on it, in the mode the session's thread runs in. Its name states that claim, as the multiply bit's name states its own. A host object needs noCH_NATIVE_AES, and a host build drops the define. The security argument:- What leaks. AES-GCM leaks its key through timing only if the AES
rounds or the carry-less multiply take a time that depends on their
operands. Neither architecture promises fixed timing outside a mode
its vendor names. Arm's A64 reference lists AESE, AESD, AESMC,
AESIMC, PMULL and PMULL2 as data-independent-time instructions while
PSTATE.DIT is 1. Intel's DOIT list names AESENC, AESDEC, AESIMC,
AESKEYGENASSIST, PCLMULQDQ, VAESENC and VPCLMULQDQ, which hold on Ice
Lake, Gracemont and later parts while the operating system has set
DOITM (
cpu_cfg.h). The same two lists name MADD, UMULH, MUL and MULX, which the multiply bit states. - Who can state it. The statement is about one CPU in one mode. A
host object runs on CPUs its builder never sees: under
AES=runtime,CH_NATIVE_AESalready states the timing of instructions on CPUs nobody named. The caller probes the CPU and sets the mode, so the caller is the one party that can make the statement. Entry 87 moved the multiply's statement to the caller for the same reason. - What stays. chapulin infers nothing. A session whose caller
leaves the bit clear runs ChaCha20 for every traffic key and runs the
table on QUIC's public keys alone (INV-26), as entry 81's absent
answer does. The table never takes a traffic key. A device object
keeps its build statements,
CH_AES_EXTERN_CONSTANT_TIMEandCH_NATIVE_WIDEMUL, because it runs on one part its builder knows. - What is lost. Today a builder who will not state the timing
builds an object in which no session runs AES-GCM. Afterwards only
SUITE=chachadoes that. A caller can also copy a probe's answer into the bit without meaning the claim. The bit's name and its comment incpu_cfg.hstate the claim, andCH_CPU_PROBEDmakes each caller write the field on purpose. - Rejected. Keeping the define would leave a host object with a build line about CPUs nobody has seen, beside a multiply statement it takes at run time from the same vendor lists. Making the multiply bit state the AES timing too would make AES-GCM depend on a statement about the multiply, which colibri does not make today.
The multiply bit picks the wide field.
CH_CPU_CONSTANT_TIME_MULTIPLYstates the 64x64->128 multiply as well. Entry 52 keptCH_NATIVE_MUL128apart fromCH_NATIVE_WIDEMULfor two reasons, and neither holds for a caller's bit. Each test binary runs both fields under both values of the bit, so the test flags'CH_NATIVE_WIDEMULno longer picks a field. And the lists a caller can cite, DIT's and DOIT's, name both widths. Under the bit, X25519 runsx25519_wide.c, which took 34 µs a scalar multiplication on an M1 Pro against 428 µs for the 16-limb field on the native multiply (entry 52). Sox25519_native.cgoes, and a host object runs the 16-limb field on the decomposition alone. Cost: a caller who can state the 32-bit multiply and not the 64-bit one cannot say so. No part known here separates the two.The mapping. Each removed value becomes a bit in a host object:
Today In a host object AES=hwwithCH_NATIVE_AES;AES=runtimewithCH_AES_INSTRUCTIONS_PRESENTCH_CPU_CONSTANT_TIME_AESsetAES=soft;AES=runtimewithCH_AES_INSTRUCTIONS_ABSENTCH_CPU_CONSTANT_TIME_AESclearCHACHA=vectorno bit: NEON or SSE2 in every session WIDEMUL=native;WIDEMUL=runtimewithCH_WIDEMUL_CONSTANT_TIMECH_CPU_CONSTANT_TIME_MULTIPLYsetWIDEMUL=decomposed;WIDEMUL=runtimewithCH_WIDEMUL_NOT_STATEDCH_CPU_CONSTANT_TIME_MULTIPLYclearX25519=widewithCH_NATIVE_MUL128CH_CPU_CONSTANT_TIME_MULTIPLYsetCHACHA=vector WIDEMUL=native, the vector Poly1305CH_CPU_CONSTANT_TIME_MULTIPLYsetA host build refuses
AES=externandWIDEMUL=native, which stay for device objects. A host session never runschacha20.c's loop: every arm64 core has NEON and every x86-64 core SSE2 (entry 82), so no bit turns the vector path off, andchacha20_block, which derives the Poly1305 key, stays the portable function. The AVX2 ChaCha20 and the VAES GCM, in development now, are chosen by the predicates added with them. Those predicates readCH_CPU_AVX2, andCH_CPU_VAESbesideCH_CPU_CONSTANT_TIME_AES, whose statement covers the AES instructions at every width. This entry adds no predicate of its own.What init refuses.
ch_connect,ch_record_init,ch_quic_init,ch_srv_accept,ch_srv_record_init,ch_srv_quic_initandch_srv_checkreturnCH_EINVAL, before they send anything, for a value withoutCH_CPU_PROBEDand for a value with a bit this object does not define for its architecture:CH_CPU_AVX2orCH_CPU_VAESon arm64, or a bit a later release adds. A defined bit for instructions the object never runs, such asCH_CPU_CONSTANT_TIME_AESin a TCP object withoutSUITE=aesgcm, still describes the CPU, and init accepts it. In aSUITE=aesgcmobject withCH_CPU_CONSTANT_TIME_AESclear, init refuses acipher_suiteslist that names an AES-GCM suite, as it does for entry 81's absent answer.How the object chooses. Entries 81 and 87 already built the parts. A target pragma compiles each instruction set's functions for those instructions alone, so the rest of the object runs on any CPU of its architecture. Each file built on the multiply compiles twice, the second time as
<file>_native.c. One branch per operation reads the session's bits, and no function pointer is involved. A host object is compiled for its architecture's base instruction set: a builder who passes-marchfor a newer CPU makes the whole object require that CPU, whatever the bits say.What goes and what merges, counted at 178f791:
What Today After Speed variables, and their values 4, and 11 2, and 4, for device objects Compiler probes for them 4 1, the host test lib-checklegs inmake check29, 12 for a speed value 20 lint-zig-buildconfigurations32, 15 for a speed value 22 Wycheproof binaries 7 3 Test binaries the four variables add 37 about 20 lint-trust-separationrows that name a speed value23 about 12 Native-copy ceilings in lint-wide-multiply110 12 Scripts that test the variables' refusals 3 1 Speed steps in each of CI's arm64 and macOS jobs 4 2 Build-record bits 2 1, CH_BUILD_CPU_RUNTIMEbuild.zigoptions16 12 ch_cfgfields2 1 chapulin.hpptypes and setters4 2 Zig API types and fields 6 3 Timing defines 4 2, for device objects - Nine
lib-checklegs go: the five that name a fast path on a raw client, and the four that repeat a webpki, server or QUIC object the host test now builds. Three stay: a server and colibri's QUIC object, both onSUITE=aesgcm, and the device server onAES=extern. The eight legs of webpki and server products build the host object with no change to their command lines. Ten Zig configurations go the same way. - The three Wycheproof binaries are the default one, the
AES=externone and the host one, which runs once per set of bits that changes a path: four times on arm64, and on x86-64 once more for each ofCH_CPU_AVX2andCH_CPU_VAESonce those paths land. - The five equivalence tests stay. Each loop, session and vector test becomes one host binary that runs under every set of bits.
- The eight 32-bit specs of
lint-wide-multiplynever compile a native copy, andx25519_native.cgoes. - Each of CI's arm64 and macOS jobs keeps one host
lib-checkand one disassembly, which then covers every file a target pragma compiles. The mips job's qemu run stays, and runs the host binaries with the bits clear on a CPU without the instructions. CH_BUILD_CPU_RUNTIMEtakes a new bit, so no record from 0.1.0 matches a host object.CH_NATIVE_AESandCH_NATIVE_MUL128go.CH_NATIVE_WIDEMULandCH_AES_EXTERN_CONSTANT_TIMEstay for device objects.- 132 of the 599 mutants in
test/violations/name a speed variable, its define or one of its files. Most hold code that stays. Each commit re-points or retires the ones that name what it removes.
Migration. Six commits, each through
make check. Each code commit marks its break with!in its header and updates the docs that state what it changes.- The interface:
ch_cfg.cpuand its bits incpu_cfg.h, the host test in the Makefile andbuild.zig,CH_CPU_RUNTIME,CH_BUILD_CPU_RUNTIME, the refusals,Config::cpu(), the Zigcpufields and the tests of the refusals. No path reads the bits yet: the four variables still choose, and no output byte changes. - AES. A host object holds the instructions, and in a QUIC object the
table for public keys, as
AES=runtimedoes, andCH_CPU_CONSTANT_TIME_AESpicks.AES=hw,AES=runtime,aes_instructionsandCH_NATIVE_AESgo. - The multiply. A host object holds the native copies, and the
multiply bit picks.
WIDEMUL=runtimeandch_cfg.widemulgo. - X25519. The wide field joins every host object under the multiply
bit, and
x25519_native.c,X25519andCH_NATIVE_MUL128go. - ChaCha20. The vector path joins every host object, the vector
Poly1305 runs under the multiply bit, and
CHACHAgoes. - CLAUDE.md, with the text Camilo approves.
Each code commit measures the objects it changes with
bench/sram.shandbench/device-ram.sh. The AVX2 and VAES predicates landed before commit 1 (entry 90), and commit 5 makes them read their bits, with the rest of the x86-64 vector paths.- colibri links
TRANSPORT=quic-nonblocking ROLE=both TRUST=webpki SUITE=aesgcm AES=runtime CHACHA=vectorwithCH_NATIVE_AESthroughbuild.zig, and setsaes_instructionsfrom stdx's platform probe (c4milo/stdx#15). From commit 1, every init returnsCH_EINVALuntil colibri setscpu. From commit 2,build.zigno longer declaresAES=runtimeorCH_NATIVE_AES, whichzig buildthen refuses. The probe's answer and the timing claim colibri's build stated withCH_NATIVE_AESmove intoCH_CPU_CONSTANT_TIME_AES. From commit 5 the same holds forCHACHA, and the vector path runs anyway. Its sessions keep their paths: AES-GCM on the instructions underCH_CPU_CONSTANT_TIME_AES, and the decomposition unless colibri sets the multiply bit. stdx's probe must learn AVX2 and VAES before colibri can set those bits. - stompy's object is
TRUST=webpki TRANSPORT=tcp-nonblocking ROLE=bothatTX_RECORD=16384(entry 71). It is a host object, so from commit 1 every init refuses its configuration until stompy setscpu. If it buildsSUITE=aesgcm AES=hwwithCH_NATIVE_AES, as the TCP object entry 80 measured through colibri's h11 client did, commit 2 refuses those values, and withoutCH_CPU_CONSTANT_TIME_AESits sessions offer ChaCha20 alone. From commit 5 its ChaCha20 runs the vector path. - Semver. Item 4 of SemVer 2.0.0 lets anything change in a 0.y.z
release. The series ships as 0.2.0 in
build.zig.zon, tagged once the nightly passes on its last commit, as 0.1.0 was. colibri and stompy pin chapulin by hash, so each moves from 0.1.0 to 0.2.0 in one step, when it chooses.
CLAUDE.md. Three bullets change, with the text Camilo approves. The protocol bullet states the webpki client's suite order on
AES=hwwithCH_NATIVE_AES. The dependency bullet names a variable besidechacha20_vector.[ch],poly1305_vector.[ch], the three implementations behindaes_block.h,ghash_hw.[ch],gcm_hw.[ch]andx25519_wide.[ch]. The constant-time bullet states the widening multiply, theX25519andCHACHAvariables, theAESvariable and the suite's timing defines. Its paragraphs on the two purposes of AES and on the two key types stay.Cost:
- Every host object holds every path. Entry 87 measured the multiply
alone: on arm64 at
-O2, the server object grows from 97,264 to 139,882 bytes when it holds both multiplies, and the AES and vector paths add more.bench/device-ram.shmeasures each commit. - A host session always runs the vector ChaCha20, which no proof covers. The proved loop runs in device objects alone.
- Every host caller writes
ch_cfg.cpu, and colibri and stompy change their builds and their configurations. - A host object's timing statements move from build lines, which a reviewer reads once, into each caller's code.
- A 32-bit target loses two choices:
CHACHA=vectoron 32-bit NEON or SSE2, which no CI leg builds, andWIDEMUL=runtime, which entry 87 measured on mips32r2 and no consumer links.
Gain: one object per product and architecture runs on every CPU of that architecture, at the speed its caller describes. Four variables leave the host build,
make checkpackages nine fewer objects andlint-zig-buildbuilds ten fewer configurations, and an instruction set a later release adds becomes a bit, not a variable.Rejected:
- The host target alone, for every product. The default object, a
raw client, would become the host object on every development
machine. The examples and
lint-stack's default budget would then measure the host object, and a firmware's raw client built on a 64-bit machine would have a field it lacks on its device. - A seventh variable that asks for the device object. It would
bring back a speed choice in the build, which the ruling removes, for
a case
TRUSTalready names. - Choosing on the architecture macros in the sources. The proofs would compile intrinsics, and a firmware tree with its own build could not get the portable object for a 64-bit target.
- Direction. A host object, for arm64 or x86-64, compiles each fast
path for its own instructions, as
-
Every x86-64 object carries an AVX2 ChaCha20 kernel beside its SSE2 path, and VAES and VPCLMULQDQ AES-GCM kernels beside its 128-bit loops; the caller's CPU bits are to pick them, and until
ch_cfg.cpuexists no call runs them. On GitHub's x86-64 runners a 16 KiBrec_sealtook about 6 µs on AES-128-GCM where OpenSSL took 4, and about 20 µs on ChaCha20-Poly1305 in aCHACHA=vector WIDEMUL=nativebuild where OpenSSL took 7.5. chapulin ran SSE2 and the 128-bit AES-NI and PCLMULQDQ there, on CPUs that have AVX2, VAES and VPCLMULQDQ. Under entry 89 a host object picks its paths at init fromch_cfg.cpu, which the caller fills from its own probe, and entry 89 names the bits these kernels need:CH_CPU_AVX2,CH_CPU_VAESandCH_CPU_CONSTANT_TIME_AES. chapulin probes nothing.- The ChaCha20 kernel.
chacha20_avx2.cholds each of the 16 state words of eight consecutive blocks in one 256-bit vector, one block per lane, and runschacha20.c's rounds on the 16 vectors: one group a pass, because AVX2's 16 registers hold one group's 16 words. The rotations by 16 and by 8 move whole bytes, so each is one byte shuffle (VPSHUFB) under a constant order, and the rotations by 12 and 7 shift and OR. A 4x4 transpose within each 128-bit half and VPERM2I128, which joins two halves, turn the 16 vectors into 16 rows of 32 bytes in block order. The pass XORs them into the data from the registers in ascending order, so the output may sit on the input or below it, as entry 86's path allows. The last 1 to 31 bytes pass through a buffer of 32 bytes thatct_wipeclears. - The AES-GCM kernels.
gcm_vaes.crunsgcm_hw.c's three loops two blocks to a 256-bit register. VAESENC runs one AES round on each half of a register, so each round key is loaded into both halves. The counters keep their bytes reversed, so one 32-bit add per pair of blocks is inc32, and one byte shuffle puts each block back in order. GHASH multiplies a pair of blocks by a pair of powers with one VPCLMULQDQ per product, 12 per pass of eight blocks whereghash_vector.htakes 24, adds each sum's two halves, and runsghash_vector.h's reduction unchanged. The powers areghash_vector.h's, paired, and read through a volatile lvalue from the state each call wipes, for the reasonghash_power_atgives. - A step of two passes. The AES rounds run a step at a time, two
of
gcm_hw.h's passes in eight registers. VAESENC's next round on a register waits on its last, and a step of one pass, four registers, left the AES units idle between rounds. In one run on an EPYC 7763 runner that built both in each job, a 16 KiB AES-128-GCM seal took 4.66 µs under gcc and 3.36 under clang with steps of one pass, and 4.07 and 3.20 with steps of two; the open moved from 3.43 to 3.38 and from 3.10 to 3.03. GHASH keeps its pass of eight blocks, so the powers and the reduction stayghash_vector.h's, andgcm.c's split of a message into passes stays as it was. A call with an odd number of passes ends on a step of one pass, whose rounds still run on eight registers; the four it does not use hold keystream no output takes, and the call's wipe clears them. - Target attributes, not flags. Every function in
chacha20_avx2.ccarriestarget("avx2"), and every function ingcm_vaes.ctarget("aes,pclmul,avx2,vaes,vpclmulqdq"), applied by one clang attribute push or one gcc target pragma, as entry 81 applies AES=runtime's. So an object needs no instruction flag for them, and the rest of it runs on any x86-64 CPU, with AES-NI and PCLMULQDQ underAES=hw.chacha20_avx2.cjoins every x86-64CHACHA=vectorobject andgcm_vaes.cevery x86-64AES=hwandAES=runtimeobject, and on arm64 both compile to nothing. - One predicate per path.
chacha20.c'suse_avx2decides whetherchacha20_xorruns the kernel in place ofchacha20_vector.c's SSE2 path, andgcm_hw.c'suse_vaeswhether its three entries hand their blocks to the kernel of the same shape, sogcm.cis unchanged. Both answer 0 untilch_cfg.cpuexists. Thenuse_avx2readsCH_CPU_AVX2, anduse_vaesreadsCH_CPU_VAESandCH_CPU_CONSTANT_TIME_AESand answers 1 only where both are set. Under AES=runtime,ch_cfg.aes_instructionssays only that AES-NI and PCLMULQDQ exist, so it cannot pick the kernels. - Timing. The ChaCha20 kernel is constant time by entry 82's
construction: adds, exclusive-ors, shifts and byte shuffles under
constant orders, with no table, no multiply, and no branch or
address on the key, the nonce, the counter or the data. The GCM
kernels run VAESENC and VPCLMULQDQ under a traffic key in a
SUITE=aesgcm build. Entry 89's
CH_CPU_CONSTANT_TIME_AESstates that the AES instructions and the carry-less multiply run in constant time at every width, and the DOIT list it cites names VAESENC and VPCLMULQDQ;use_vaesis to require that bit.CH_NATIVE_AESstates nothing about the 256-bit forms, and no define makes a kernel run (ct.h). - What tests may ask. Test code asks its CPU through
__builtin_cpu_supportsand CPUID (test/x86_kernels_cpu.h); the library asks nothing. The equivalence tests call the kernels directly, and three binaries and a Wycheproof leg send the library's calls to them through a force-included header that renames the 128-bit entry, while the library's own predicates still answer 0. Each skips a CPU without the instructions, and CI'sx86-64-kernelsjob, underCH_REQUIRE_X86_KERNELS=1, fails on one instead (docs/verification.md, "The x86-64 kernels"). - Left out. An AVX2 Poly1305, which follows
#186's rework of
the multiply files. A 512-bit path: on a Xeon Platinum 8370C runner,
which has AVX-512, OpenSSL sealed 16 KiB of AES-128-GCM in 1.42 µs
where these kernels took 3.29 under clang and 3.64 under gcc. A GHASH
pass of sixteen blocks, which would need sixteen powers of H and a
pass size
gcm.cdoes not know.
Cost:
- 339 lines in
chacha20_avx2.[ch]and 535 ingcm_vaes.[ch], with no harness, because CBMC cannot read an intrinsic. - At
-O2on x86-64,chacha20_avx2.oholds 2,821 bytes of text under gcc 13 and 3,015 under clang 18, andgcm_vaes.o4,542 and 7,542, besidegcm_hw.o's 4,996 and 6,723. Every x86-64CHACHA=vectorobject carries the first, and every x86-64AES=hwandAES=runtimeobject the second. - Stack frames: the ChaCha20 kernel's pass takes 904 bytes under gcc
and 616 under clang; the GCM seal 1,024 and 1,016, the open 1,024
and 984, and counter mode 384 and 376, where
gcm_hw.c's seal takes 576 and 600. - Three binaries and one Wycheproof leg in
make checkon an x86-64 host, and CI'sx86-64-kernelsjob.
Gain, 16 KiB records in µs: the medians bench/record.sh writes, from bench.yml's
record-x86_64job at 6763ce8 on an AMD EPYC 7763 runner, whose one-minute load average ran from 0.44 to 0.97. "Before" is the 128-bit path and "after" the kernels, built in the same job. OpenSSL 3.6.4's figures are the gcc job's; the clang job's differ by 0.02 at most.gcc before gcc after clang before clang after OpenSSL AES-128-GCM seal 6.29 3.69 5.85 3.19 4.05 AES-128-GCM open 6.10 3.39 5.34 3.03 4.12 AES-256-GCM seal 7.06 4.14 6.50 3.61 4.36 AES-256-GCM open 6.92 3.83 5.98 3.44 4.43 ChaCha20-Poly1305 seal 20.19 13.64 21.29 13.65 7.47 ChaCha20-Poly1305 open 19.98 13.39 21.04 13.55 7.47 ChaCha20 alone 13.11 6.51 13.94 6.57 Of the kernels' AES-128-GCM seal, the AEAD took 2.98 µs under gcc and 2.63 under clang, and the record layer 0.71 and 0.57. Poly1305 takes 6.6 to 6.8 µs of the ChaCha20-Poly1305 rows either way, which #186 addresses. A run 40 minutes earlier on the same CPU model, at the same kernels, gave gcc's AES-128-GCM seal 4.07 and clang's 3.20, so the gcc figure moved by 10% from one runner to the next.
- The ChaCha20 kernel.
-
ct_wipecalls the libc'smemsetthrough a volatile function pointer, and the proofs read a byte-loop stub with the same contract.ct_wipestored one byte per iteration through a volatile pointer, and a compiler may not merge volatile stores into wider ones. Since entry 85 a forged AES-GCM open wipes the plaintext it wrote, so refusing a forged 16 KiB record took 8.3 µs on the M1 Pro where a genuine open took 2.6, and every seal and open also wipes its expanded key and the state of its passes once a call.ct_wipenow sits inct_wipe.c, which holdsstatic void *(*const volatile ct_memset)(void *, int, size_t) = memset;
and calls
memsetthrough it when n is not 0.ct.ckeepsct_memeq, andct.hdeclares both.- Why the compiler keeps the call. A volatile object is read each
time the abstract machine reads it (C11 5.1.2.3p6), so the compiler
loads the pointer at every call and cannot tell which function the
load returns. It cannot delete a call to a function it does not
know, and it cannot treat the call as a
memsetwhose stores no later read needs. That holds where it inlinesct_wipeinto a caller whose buffer ends right after the call, as a consumer's link-time optimization does.memsetis the libc's, compiled apart from every caller, so it writes each byte. - What a compiler may still do. It may compare the loaded pointer
with
memsetand drop the call on the branch where the two are equal. gcc and clang make that comparison for an indirect call only from a profile: gcc's value-profile transformations and LLVM's indirect-call promotion read-fprofile-usedata. A profile-guided build of chapulin needs its output read again. - n of 0. C11 7.24.1p2 asks
memsetfor a valid pointer even when n is 0, soct_wipemakes no call then, and a caller with nothing to wipe may pass a null pointer. The branch reads the length, which is public. - The output read. A scratch unit compiled this
ct_wipebeside two functions that each pass a stack buffer of 64 bytes or 16 KiB to a function in another unit, wipe it and return, at-O2and-Os, under Apple clang 21 for arm64 and x86-64 macOS; clang 23.1.2 for arm64 and x86-64 Linux,thumbv7m-none-eabi -mcpu=cortex-m3,mips-linux-musl -march=mips32r2andriscv32-unknown-elf -march=rv32imac; gcc 13.3 for x86-64 and arm64 Linux; the Arm GNU gcc 15.3.1 the m3 lane pins; the mips lane's gcc 12.4; and the riscv32 lane's Bootlin gcc 14.3 at rv32imac and rv32ic. Every clang at both levels, and every gcc at-O2, inlinesct_wipe, and the caller loadsct_memsetand calls through it. gcc at-Oscallsct_wipe.part.0, the out-of-line part ofct_wipeafter the test of n, which does the same. Withct_wipecallingmemsetby name instead, every clang deletes the call at both levels and every gcc at-O2; gcc at-Oskeeps it only because it does not inlinect_wipethere. - The proofs. 84 launch lines in
proof/run.shlinkct.c, and 20 of them bound the loopct_wipe.0. Each now linksproof/ct_wipe_stub.cbeside it, a contract stub that is the loopct_wipewas, so no other harness's formula changed:ct_memeqinct.candct_wipein the stub are the same tokens as the two bodiesct.cheld at a104a3e.proof/ct_harness.cproves the stub writes zero to p[0..n), as it proved of the loop before, and now also that it writes no byte past them.proof/ct_wipe_harness.cprovesct_wipe.cmemory-safe and UB-free, zero over p[0..n) and no other byte written, over a heap buffer of every size CBMC's pointer encoding holds, in 0.5 s and 22 MB. Ofmemsetit proves CBMC's model only. Run one at a time under a 30-minute limit, beside other jobs on the M1 Pro,ctproved 63 properties in 12 s,ct_wipe66 in under a second,aead254 in 6 s,gcm_safety457 in 254 s,handshake_post765 in 119 s andrecord412 in 603 s. At a104a3e the last four prove the same counts, andctproves 59, without the check of the bytes past n. - Tests.
bin/unitwipes n bytes for each n from 0 to 32 between guard bytes, and a null pointer with n of 0.test/poly1305_equiv_vector.ccompilesct_wipe.cinto the unit that holds the vector Poly1305, renamed as that unit renames its other calls, so the compiler can inlinect_wipewhere the call's powers of r end.ct-wipe-plain-memsetcallsmemsetby name, andbin/poly1305_equiv_testthen finds the powers on the stack, under Apple clang 21 and under gcc 13.3 on arm64 and x86-64 Linux. Every other binary compilesct_wipe.cas a unit of its own, where no compiler can delete either body's call, so no other test notices the mutant. - Zig. Zig 0.16's compiler_rt exports a weak
memsetthat stores one byte per iteration: built for aarch64 Linux atReleaseFast, it is a loop ofstrb. A Zig program on Linux that links no libc gets that one, so itsct_wipestill writes every byte through the same volatile pointer, at about the old loop's speed. colibri and stompy get the speed below when they link a libc or export a fastermemsetof their own, which takes the place of the weak one.
Rejected:
memsetand a fence. C11 names no barrier that keeps stores nothing reads afterwards.atomic_signal_fenceorders memory against a signal handler, and whether it keeps amemsetof a buffer whose lifetime then ends is the compiler's choice, not the standard's. Anasmstatement with a memory clobber keeps it, and is GNU C, not C11.- Wider volatile stores. A
uint32_toruint64_tlvalue stored into a byte array breaks the effective-type rule (C11 6.5p7). memset_s,explicit_bzeroormemset_explicit. Annex K is optional, and glibc and musl do not shipmemset_s;explicit_bzerois not in C. C23'smemset_explicitis the standard form of this call, and chapulin builds as C11.- Keeping the loop. It costs what the gain below measures.
- A second body in
ct.cunder__CPROVER__. cbmc defines that macro, soct.ccould have compiled the loop for the proofs and thememsetcall for every build. Camilo ruled on 2026-10-01 that shipped C carries no verification token (docs/proofs.md, "Prior art"), and the stub on the launch line keeps every formula without one.
Cost: one indirect call per wipe, one pointer in read-only data, and a branch on n. On mips32r2
ct.oandct_wipe.otake 120 bytes wherect.otook 100, andct_wipe's frame grew from 0 to 24 (bench/device-ram.sh). Wherememsetis itself a byte loop the call adds a little:bench/insn-mips.shlinks such amemset, and there a 1 KiB AEAD seal went from 81,976 to 82,006 instructions and its stack from 748 to 756 bytes. On the Cortex-M3, whose newlibmemsetstores words, the same seal went from 67,684 to 67,367 instructions.Gain, in paired runs on the M1 Pro under macOS's Apple clang 21, then in the OrbStack VM under clang 18 and gcc 13:
- A forged 16 KiB AES-128-GCM open, timed directly: 8.3, 8.8 and 8.6 µs before, 2.5, 3.0 and 2.8 µs after; at 1 KiB, 877, 909 and 921 ns before, 248, 264 and 275 ns after.
- The record bench's 16 KiB AES-128-GCM seal,
gcm_traffic_seal: 2.59, 2.95 and 2.95 µs before, 2.28, 2.63 and 2.62 µs after; at 1 KiB, 556, 586 and 613 ns before, 248, 283 and 278 ns after. The wipe of the expanded key took 93, 95 and 95 ns, and takes 5.5, 4.7 and 4.7. - ChaCha20-Poly1305's 16 KiB seal did not move past the runs' spread
in the packaged build, and fell by 1% under
CHACHA=vector WIDEMUL=native, whose vector Poly1305 wipes its powers of r on each call; at 1 KiB that build's seal fell by 9%.
The CSVs that docs/performance.md renders were measured again for this entry, and two of their changes predate it.
bench/results-device.csvwas older thanhandshake_auth.c, whose object is 736 bytes at a104a3e where the file held 648. And at a104a3e, before this change as after it, the packaged ChaCha20-Poly1305 record takes 2% to 3% longer than in the runs the tables held at 97826ba, and the VM's clang runs the vector Poly1305 in 4.8 µs where those runs took 3.0. - Why the compiler keeps the call. A volatile object is read each
time the abstract machine reads it (C11 5.1.2.3p6), so the compiler
loads the pointer at every call and cannot tell which function the
load returns. It cannot delete a call to a function it does not
know, and it cannot treat the call as a