Python tools for generating, running, and evaluating LLM harnesses.
More Programs or More Rolls? Separating Coverage from Specialization in LLM Harnesses
Ziyang Xu, Haitian Zhong, Hao Zhou, Hao Qin, Chenhan Jin, Te Qi, Shengze Xu, and Tieyong Zeng. arXiv preprint, 2026.
Abstract | PDF | HTML | BibTeX
When should an inference budget buy more programs, and when would more executions of the same program suffice? This study separates answer coverage, repeatable task advantages, and gains from choosing a harness before execution. On 386 complete MATH-500 evaluation tasks, it compares eight generated harnesses plus a baseline against nine byte-identical baseline slots, with three executions per member and a fixed solver.
- Identical programs provide 2.16 percentage points of repeat-averaged oracle headroom, showing that additional coverage can arise from repeated sampling.
- Generated programs show 100 tasks with persistent losses and one with a persistent win relative to the baseline across all three repeats. The sole win is sensitive to answer extraction.
- The frozen pre-execution selector gains 0.00 percentage points. Both populations reach 98.70% oracle coverage at 27 harness executions.
- Stable complementarity remains unresolved at three repeats. Supporting BIRD traces distinguish failures in mechanism implementation, activation, and output validity.
The oracle comparisons use correctness after execution; they are not the accuracy of a deployable selector. The replay budget matches harness executions, not model calls or tokens.
Figure 1 from the published arXiv v1: same-code controls separate coverage from repeatability and useful selection. Persistent win/loss counts require at least one generated member to win/lose against the baseline in all three repeats; the sole win is extraction-sensitive. Intervals are 95% bootstrap confidence intervals. Original figure PDF | Figure provenance.
Use Python 3.10 or newer. From the repository root:
python -m venv .venvActivate the environment with source .venv/bin/activate on Linux/macOS or
.\.venv\Scripts\Activate.ps1 in Windows PowerShell, then install:
python -m pip install -r requirements.txtRun the offline tests and check the source inventory:
python -m pytest -q
python scripts/check_release.pyThe default tests use synthetic inputs and do not make model API calls. Datasets, API credentials, model responses, and runtime outputs are not included.
For optional runtime dependencies:
python -m pip install -r requirements-runtime.txtThe historical collection runtime requires Windows. Some archived utilities also require separately supplied protocol/configuration files and runtime inputs. The default offline test suite is the supported starting point.
| Directory | Contents |
|---|---|
experiment/gsm8k/ |
Math harnesses, generation, collection, and scoring |
experiment/phase2/ |
SQL harness generation, behavioral checks, and scoring |
experiment/revision/ |
Runtime utilities, statistical analysis, and selection |
experiment/diagnostics/ |
Diagnostic methods and synthetic calibration |
external/TTHE/ |
Required TTHE runtime and SQL harness implementations |
review-stage/ |
Selected configuration and synthetic test specifications |
scripts/ |
Source-inventory checks |
release/ |
File checksums and known invalid generated-source list |
assets/ |
Published Figure 1 and its display preview |
Starting points for the paper's evaluation components:
| Component | Code |
|---|---|
| Same-code provenance and repeated outcomes | wp1r_analysis.py, wp1r_closeout.py |
| Complementarity diagnostics and calibration | ccomp_v3.py, calibration_sim.py |
| Pre-execution selection | wp2r_selector.py |
| Execution-matched oracle replay | wp1r_closeout.py |
| SQL harness generation and behavioral checks | experiment/phase2/ |
These analysis sources require separately supplied protocol/configuration files and runtime inputs.
Four intentionally retained, rejected generated programs are listed in
release/known_invalid_sources.json; they are not runtime entry points.
This repository includes code and the published paper's citation and main figure. Benchmark inputs, collected model responses, score matrices, and execution ledgers are not bundled. The offline tests validate code behavior with synthetic inputs; they do not reproduce the paper's measured results.
If you use this code or evaluation methodology, please cite the paper:
@misc{xu2026moreprogramsmorerolls,
title = {More Programs or More Rolls? Separating Coverage from Specialization in {LLM} Harnesses},
author = {Xu, Ziyang and Zhong, Haitian and Zhou, Hao and Qin, Hao and Jin, Chenhan and Qi, Te and Xu, Shengze and Zeng, Tieyong},
year = {2026},
eprint = {2609.35873},
archivePrefix = {arXiv},
primaryClass = {cs.AI},
url = {https://arxiv.org/abs/2609.35873}
}The same entry is available in CITATION.bib. CITATION.cff supplies GitHub's Cite this repository metadata with the paper as the preferred citation.
English is the default maintenance language. See CONTRIBUTING.md. Original code uses the MIT license; bundled dependencies retain their own notices.
