🇺🇸 English | 🇯🇵 日本語 | 🇨🇳 简体中文 | 🇹🇭 ไทย
Notice (2026-08-24): We found a gap between this README and the shipped behaviour (#144): part of the learning loop described below — "the method is saved only as a lesson at first, and becomes a rule after passing verification again on a different job. Procedures that keep coming up are reviewed by a different AI than the one that wrote them; only the ones that pass are stored as skills" — is designed but not yet implemented. We are implementing it now (#147, #148, #149) and will adjust and republish these docs once it ships. Tracking: #146.
CI evidence (2026-08-15 UTC): matrix 7/7 green ・ 60 independent runs of one SHA, 0 flakes. Weekly schedule — GitHub pauses schedules after 60 days of repo inactivity, so mind the run date.
* WSL2 is a verified-with-conditions tier on one dated VM, not a CI-tested tier; see the WSL2 support note.
Re-explaining everything. Context that vanishes. A cheerful "done!" with nothing to show for it.
Caty Agent Harness fixes those — with plain text files and real checks.
Not magic. The machinery remembers, drives, and checks, so the AI can focus on thinking —
a small system, wrapped around the AI you already use.
A system that doesn't take the AI's "done!" at face value — it runs work in stages, verifies "done" against evidence, and publishes its own scorecard, losses included.
What actually gets better — measured, not promised:
- Fake "done" stops on a weak model — the verified completion rate triples. On Claude Haiku 4.5 (a lower-capability model), "done" claims with the work left unread fell from 98% to 8% of completion claims (222/226 → 2/26; 30 runs/arm, M/L-size jobs) — and verified completions rose from 13% to 43% (full numbers).
- Strong models keep their accuracy — and cut tokens on big jobs by 58% (sonnet) / 31% (opus). sonnet and opus scored the same on the hidden answer key with and without the harness — on large jobs, running without it used 2.40× (sonnet) / 1.46× (opus) as many tokens (= 58% / 31% fewer with the harness; L-band medians — the P1 table). On sonnet's smaller jobs in these tests the harness used more, not fewer — use it where the work is big. And on the largest measured jobs — ≈2.4M tokens, about 2.4× the context window — where opus without the harness read 2–4% of the files and kept submitting, opus under the harness finished verified with 100% of the files read, 2/2 in both overflow bands (the two arms ran under different budget envelopes).
- The results show its limits too. On search-style runtimes (e.g. Codex) the overflow guard never fired — by design, there was nothing to rescue — and on one runtime it fired but used more tokens than running without it (per-model profiles). sonnet's overflow cells stayed 0/2: all four runs delivered, 20/20-correct with valid quotes, while skipping ~10–12% of the files — and the verification gate refused "done" each time, listing the unread files (the pre-registered honest-failure branch). Every losing result is published next to the wins.
Four sealed, pre-registered, machine-scored experiment families (2026-08) — every number, method, and caveat, including where the harness costs extra tokens: benchmark ・ Try it in 2 minutes — one pasted prompt into the AI tool you already use.
🔧 Engineering guide | 📘 Reference
generation: ed18d59 (2026-09-18T11:26:53Z) · verify: API HEAD · status.json
- Does this sound familiar?
- What you get
- What you need
- Get started
- Why it's safe to try
- Which models benefit
- Dig deeper
- Part of Family OS
- License
The more work you hand to an AI, the more often these moments show up.
- Every new session starts with you explaining the same background again.
- The AI says the job is complete — but the result is missing, or you can't tell.
- Long work loses its place halfway and quietly stalls.
- A fix that worked last week is forgotten by this week.
If you nodded at any of these, this was built for you. And if you only use AI for short one-off questions, this machinery is more than you need — you're fine as you are.
These four problems are exactly what Caty Agent Harness was built to crush, with machinery instead of promises.
It's one loop, repeated: remember → work → prove it → hand over. The machinery keeps that loop turning, so even when the AI forgets, the work is never forgotten.
flowchart LR
A["Remember"] --> B["Work"]
B --> C["Prove it"]
C --> D["Hand over"]
D -. next time .-> A
-
🌱 It gets smarter on its own
When it makes a mistake, the reason and the evidence are recorded on the spot in a notebook (really just a plain text file). The next attempt always inherits the previous failure, and retrying the same way is mechanically forbidden — the same mistake stops repeating. You never have to say "take a note."
-
🏁 It runs to the finish
Big jobs are split into small numbered steps, and the machinery drives them forward one at a time. Sessions can end, windows can close, models can change — the work continues from where it left off.
-
🔍 "Done" comes with proof
"Done" is judged by mechanical checks against the actual result — never by the AI's own claim. An independent verifier joins only when configured; without one, the mechanical check remains the core, and the maker's claim is still not treated as proof. When nothing is working, it doesn't spin forever: it stops honestly, with evidence, and reports to you.
How it works (the reason it isn't magic)
- When it fails — what went wrong, plus the evidence, is recorded automatically in a notebook inside your project.
- On the next attempt — the previous failure is always handed over, and a mechanical rule forbids retrying the same way.
- When it succeeds — the method is saved only as a lesson at first, and becomes a rule after passing verification again on a different job. Procedures that keep coming up are reviewed by a different AI than the one that wrote them; only the ones that pass are stored as skills.
- While it runs — a scheduler kicks the next step every few minutes, handing the AI only a short, fresh context. Attempts and active time have limits counted by the machinery, so endless grinding is impossible.
→ In depth: learning from repeated failures / the completion rail
Whether you can use it comes down to the tools you already have — here's the table.
The supported AI tools all run in a terminal — but you won't be the one typing. Your AI does the setup and the upkeep; you just talk to it.
| Category | Supported |
|---|---|
| OS | macOS: ✅ CI-tested (GitHub Actions macos-latest, Apple silicon) / Linux: ✅ CI-tested (GitHub Actions ubuntu-latest) |
| Windows (native) | ❌ not supported — measured walls: chmod silently becomes 644, ln -s becomes a copy, and there is no flock. See WSL2 support note. |
| WSL2 (Ubuntu on Windows) | 🟡 supported with conditions — verified 2026-08-23 on win11-test-vm (30/30 suites, umask 0002, non-root, Linux filesystem), but not CI-tested.Your AI tool (Claude Code / Codex CLI) must run inside the same WSL2 distro; a Windows-side agent can install successfully and the hooks will simply never fire. Keep the repo on the Linux filesystem ( /home/..., not /mnt/c/...) for correctness, not speed; use git 2.34+, run as a non-root user, and keep wrapper-type files not group/world-writable (for example chmod 0755). CI approximation: ubuntu-wsl2-profile (umask 002, non-root container). Measured details |
| AI tools | Claude Code ✅ / Codex CLI ✅ / Kimi Code CLI ✅ / Hermes Agent ✅ / OpenClaw ✅ |
| Shell | bash 3.2+ ✅ (the macOS default is fine) |
| Python 3 (3.9+) | used by behind-the-scenes automation (technically, a mechanism called hooks) — your AI will check this for you |
Support depth differs by tool on purpose — the details live in the engineering guide.
If your setup is on the table, installing takes one prompt.
There is one thing to do on your side. Open the AI tool you already use inside the project folder where it works, and paste this — your AI handles the install, the checks, and the report back to you.
Please set up https://github.com/caty-ai/caty-agent-harness.git in this project:
read docs/agent-guide.md in that repository and follow it — install into this folder
as the workspace, run the health check, and then tell me in plain words what you set
up and what I can do next.
That's it. The agent guide walks your AI through every choice, the health check, and what to report back to you.
For a concrete first demo, your AI can run the bundled image pilot example, which builds an SVG image card and JSON delivery receipt using local tools only.
Prefer to type the commands yourself? → the engineering guide has the full manual path.
If something feels off
- Your AI will run a read-only health check (
--check) and show you the result — a healthy setup ends withok: required layout and STATE.md headers present. - Some check rows can say
FAILwhile the core is healthy: those are optional automation paths that aren't wired yet. The agent guide tells your AI which ones matter for your tool.
Still hesitant to paste it? The next section explains why nothing gets broken.
- Your AI stays exactly yours — its personality and accumulated memory stay as they are. Existing instruction-file content is left intact; if you choose
--append-bootstrap, setup adds only the documented bootstrap block to the selected instruction file. The rest is harness scaffold around it. - Quitting is one command too — installing is one command, pausing is one command, and nothing it learned is lost. Resume, and it continues where it stopped.
- You can read everything — lessons, progress, and evidence all live in plain text files, so you can see what's happening with your own eyes.
That's the short version. The depth is all below.
The overflow sentinel (design issue #159) watches the measured per-turn context level and, past a threshold, stops the run and decomposes the job instead of letting the context overflow. Sentinel v1 shipped as opt-in for the claude-code runtime in v0.17.0 (#180, implementing #159); enable with OVF_SENTINEL=shadow|active, unset = fully off (byte-identical passthrough). Default-on remains future work, decided per model against the #159 conditions; see the per-model profiles in this section and the benchmark. EV-008 remains a rig pre-measurement taken ahead of that implementation and has not been re-measured on the shipped implementation. EV-008 — a sealed, pre-registered benchmark (2026-08) — measured how that behaves per model. The effect splits by how a runtime reads:
| Runtime type | Measured models | Behaviour | Verdict |
|---|---|---|---|
| Full-read — context grows monotonically | claude-haiku-4.5 · claude-sonnet-5 · claude-opus-5 | fires mid-task → decomposes → completes; correctness held (sonnet/opus 20/20 in every fired cell · haiku 19–20/20) | This is where the benefit lives. For sonnet, savings grow with job size (largest in the L band). sonnet median sentinel/bare 0.801 (best cell −71 % tokens) · opus 0.923 · haiku 0.35 (descriptive; carries an arm/phase confound — see benchmark) |
| Search-type — context plateaued in the observed range | gpt-5.6-luna (Codex) · qwen3.8-max | no fire in any measured run: codex 0/127 turns · qwen 0/4 cells (max observed 79.7K @ 80K threshold — thin margin, stated as observed range only) | No-fire in the observed range is the design intent — it is the default-on safety condition itself. A fire here would be a false positive and a design send-back |
| Fires but uneconomic — grows, yet bare is cheap | grok-4.6 | fires 4/4, decomposes and completes correctly (20/20) — but bare runs are so cheap that decomposition costs more (median ratio 2.145) | The mechanism is proven on this runtime; default-on is not recommended. Judge by economics vs bare, not by whether it fires |
| Fires, mixed outcomes — early-onset firing, heavy full-read | gemini-3.7-flash (descriptive — outside the pre-registered GO conditions) | fires 4/4 (turns 31–82); bare collapses on L-band jobs; decomposition rescued completion in one L cell; 2/4 sentinel cells (M-i3, L-i2) scored 0 (headless permission stalls in the child steps) | Descriptive only — no efficiency claim, no default-on judgement. Per-cell numbers |
Ratios are sentinel/bare token cost (lower = cheaper), median of 4 sealed cells per model, measured 2026-08-25. Read before quoting: ① the codex condition first FAILed (M4, n=1, geometric mean 1.337) and passed only on the n=3 repeat — a data-informed post-design sealed via a 3-seat delta review (0.9944 ≤ 1.05) — the FAIL is history, not erased; ② sonnet's 0.801 met the decision threshold (<1.0) but narrowly missed the stretch goal (<0.8); ③ most per-pair ratios are n=1 — run-to-run variance is real (≈25 % SE on the codex mean); the n=3 repetition addendum (2026-08-29) quantified it — pooled over 12 log-ratios per model: sonnet GM 0.617 [0.439, 0.867], opus 0.806 [0.651, 0.997] — the single-run medians sat inside these spreads (individual reps ranged 0.28–1.50), and opus's upper bound is 0.997: savings confirmed, by a razor-thin margin; the sentinel itself remains opt-in and has not been re-measured on the shipped implementation. Full numbers, method and caveats · n=3 addendum.
A third sealed experiment, P2-WIN, characterized the jobs themselves on bare models — a measurement of completion outcomes across three job-size bands, not an intervention result: tables & limitations.
A further lane, EV-007 / EV-007b / EV-007c, measured the other half of the claim — the learning loop. It ran three sealed times. EV-007 terminated when the cross-model review fail-closed; EV-007b then built a trap-injected instrument that does produce repeat mistakes (in the learning arm, 8 of 8 mistake classes recurring in every round on v0.24.0) and ran its sealed main run on it — where the review rejected every single citation, a mismatch between the review's citation check and the Stop hook's flush format, which is that run's product finding; promotions in both lanes: 0. EV-007c re-ran the same sealed series on v0.25.0, where the review can cite: all 3 turns promoted (6 promotions, 2 approved rules), and the sealed pass line was not met — the learning arm's repeat-mistake rate is 0.667 against 0.727 in the no-learning arm, difference-in-differences +0.050. The one T2 class that moved after the rule was approved (in-file supersession) was solved in round 3 and half-lost in the final block; the erratum-file classes never moved for any arm and were never written down for the loop to see. That is one measured outcome at n = 1 on a one-bit-per-class instrument — a fail of that pre-registered line, not a verdict on the loop. What the lanes did produce is two product changes, shipped in v0.24.0 and v0.25.0: what was measured, and what was not.
Blind telemetry paths — do not switch these on unmeasured. Through the current shim, glm / muse report all-zero per-turn usage, and kimi emits no usage at all. A runtime whose live telemetry cannot be seen must not run the sentinel default-on: there is no water level to watch.
One water-level manager at a time. If the host has its own auto-compaction (Hermes, OpenClaw, an agent CLI's built-in compaction…), either the host or the sentinel manages the context level — never both. Two managers means either the host compacts first and the sentinel becomes decorative, or both intervene on the same overflow. Pick one: disable host auto-compaction when the sentinel is default-on, or demote the sentinel to record-only when the host leads. The sentinel's event log (on the EV-008 rig) already records runtime_compaction / compaction_suspected (false across all EV-008 runs) as the detection basis.
How to profile a new model (before any default-on decision):
- Run one live job with telemetry on and confirm every usage field carries real values — a mock pass is not enough; blind paths look healthy until live.
- Measure the injected-context curve over turns.
- Plateau → search-type: verify no-fire holds (a fire is a design send-back). Monotonic growth → full-read: run the 3-arm comparison (bare / always-on / sentinel).
- Compare sentinel/bare cost before enabling — a firing sentinel earns default-on only if decomposing is cheaper than pushing through (grok is the measured counter-example).
The page you just read is a map. The substance is real, and it's all documented.
| Document | What's inside |
|---|---|
| docs/agent-guide.md | The installer's playbook — a step-by-step guide your AI follows: choices, commands, checks, and how to report back |
| docs/benchmark.md | The sealed benchmark — full numbers behind the hero claims, method, and the honest limitations |
| docs/engineering.md | The full technical guide — what's enforced where, per-tool depth, pause semantics, architecture, directory map |
| docs/reference.md | The exact contracts — every flag, every state, every pointer to the design documents |
| Runtime setup | Per-tool wiring — hooks, verifiers, and schedules for each of the five AI tools |
| CONTRIBUTING.md | How to propose changes — issue-first flow and every test suite under tests/ (one make test runs them all) |
| SECURITY.md | How to report issues safely — private vulnerability reporting |
- CI:
— runs
make test+make linton every pull request - Verified environments: macOS (GitHub Actions
macos-latest, Apple silicon), Linux (ubuntu-latest), and WSL2 (Ubuntu on Windows; verified 2026-08-23 onwin11-test-vm, 30/30 suites, Linux filesystem, non-root,umask 0002; not CI-tested) — see the table in What you need and the WSL2 support note - Maturity: public preview — the FROZEN CLI output contracts in docs/cli-conventions.md are stable; everything else may still move
- Known constraints: Native Windows is not supported; the WSL2 row above is the only verified Windows-adjacent path, and some updater suites need
ssh-keygen(see CONTRIBUTING Prerequisites)
One last thing — the bigger picture this tool belongs to.
Part of the Caty AI family — open tools for running a family of AI agents. The full map, including modules still being prepared for release, lives in Family OS.
| Axis | Module | What it does | State |
|---|---|---|---|
| Map | Family OS | The map of the whole family — every module, its state, and how they fit | published, MIT |
| Rules | Family Dev Handbook | The rules of the road — issues, PRs, worktrees, handoffs, parallel development | published, MIT |
| Vertical · foundation | Caty Agent Harness | Task backbone for AI agents — retries, checkpoints, and honest completion | published, MIT |
| Vertical | context-kit | Six-piece context hygiene kit for one agent — bounded output, delegation briefs, safety guards, recall, worktree snapshots | published, MIT |
| Vertical | Persona Engine | Layers relationship and emotion onto an agent's existing persona | published, MIT |
| Vertical | Persona Growth Loop | Grows the persona itself — minimal, idempotent proposals | published, MIT |
| Vertical | X Collector | Turns X and the web into one daily digest — for people and agents | published, MIT |
| Vertical | Self Growth Loop | Lets an agent grow its own abilities — proposals, governance, adoption records | published, MIT |
| Horizontal · foundation | Family Memory Architecture | The memory bus — how the family shares what it knows | published, MIT |
| Horizontal | Sitter | Babysits delegated agent runs — watches, keeps evidence, restarts only within declared bounds | published, MIT |
| Horizontal | Alpha Nightshift | Nightly autonomous maintenance loop — isolated night lanes behind a deny-by-default guard; humans cherry-pick in the morning | published, MIT |
| Horizontal | errmeter | Reports failed or silent AI agents and scheduled jobs across machines — emit, spool, shared board, repair hook; a shout that is never lost | published, MIT |
| Vertical | Caty Gateway | PC-side gateway for CatyPhone — one-line install; pairs your phone with the agent running on your machine (Claude Code / Codex CLI / OpenClaw / Hermes / OpenAI-compatible) | published, MIT |
Caty Agent Harness is one tool inside Family OS — the Caty AI project's larger blueprint for running multiple AI agents as one family. It works fully on its own, and it becomes even stronger combined with:
- family-os — the blueprint that ties the family together. Inside it, this Harness owns the vertical axis: growing an individual agent and driving its work to completion.
- sitter — a watchdog that keeps an eye on long-running agent work from the outside, and raises its hand when the work stalls or freezes.
MIT — chosen so anyone can use it, study it, and build it into anything, including commercial agent setups. That's the point.
plain text files | works with 5 AI tools | paused in one command
