Skip to content

Repository files navigation

Harness Eval

Python tools for generating, running, and evaluating LLM harnesses.

arXiv

Paper

More Programs or More Rolls? Separating Coverage from Specialization in LLM Harnesses

Ziyang Xu, Haitian Zhong, Hao Zhou, Hao Qin, Chenhan Jin, Te Qi, Shengze Xu, and Tieyong Zeng. arXiv preprint, 2026.

Abstract | PDF | HTML | BibTeX

When should an inference budget buy more programs, and when would more executions of the same program suffice? This study separates answer coverage, repeatable task advantages, and gains from choosing a harness before execution. On 386 complete MATH-500 evaluation tasks, it compares eight generated harnesses plus a baseline against nine byte-identical baseline slots, with three executions per member and a fixed solver.

  • Identical programs provide 2.16 percentage points of repeat-averaged oracle headroom, showing that additional coverage can arise from repeated sampling.
  • Generated programs show 100 tasks with persistent losses and one with a persistent win relative to the baseline across all three repeats. The sole win is sensitive to answer extraction.
  • The frozen pre-execution selector gains 0.00 percentage points. Both populations reach 98.70% oracle coverage at 27 harness executions.
  • Stable complementarity remains unresolved at three repeats. Supporting BIRD traces distinguish failures in mechanism implementation, activation, and output validity.

The oracle comparisons use correctness after execution; they are not the accuracy of a deployable selector. The replay budget matches harness executions, not model calls or tokens.

Main figure

Figure 1: Coverage, repeatability, and selection on 386 complete MATH-500 tasks

Figure 1 from the published arXiv v1: same-code controls separate coverage from repeatability and useful selection. Persistent win/loss counts require at least one generated member to win/lose against the baseline in all three repeats; the sole win is extraction-sensitive. Intervals are 95% bootstrap confidence intervals. Original figure PDF | Figure provenance.

Setup

Use Python 3.10 or newer. From the repository root:

python -m venv .venv

Activate the environment with source .venv/bin/activate on Linux/macOS or .\.venv\Scripts\Activate.ps1 in Windows PowerShell, then install:

python -m pip install -r requirements.txt

Usage

Run the offline tests and check the source inventory:

python -m pytest -q
python scripts/check_release.py

The default tests use synthetic inputs and do not make model API calls. Datasets, API credentials, model responses, and runtime outputs are not included.

For optional runtime dependencies:

python -m pip install -r requirements-runtime.txt

The historical collection runtime requires Windows. Some archived utilities also require separately supplied protocol/configuration files and runtime inputs. The default offline test suite is the supported starting point.

Code layout

Directory Contents
experiment/gsm8k/ Math harnesses, generation, collection, and scoring
experiment/phase2/ SQL harness generation, behavioral checks, and scoring
experiment/revision/ Runtime utilities, statistical analysis, and selection
experiment/diagnostics/ Diagnostic methods and synthetic calibration
external/TTHE/ Required TTHE runtime and SQL harness implementations
review-stage/ Selected configuration and synthetic test specifications
scripts/ Source-inventory checks
release/ File checksums and known invalid generated-source list
assets/ Published Figure 1 and its display preview

Starting points for the paper's evaluation components:

Component Code
Same-code provenance and repeated outcomes wp1r_analysis.py, wp1r_closeout.py
Complementarity diagnostics and calibration ccomp_v3.py, calibration_sim.py
Pre-execution selection wp2r_selector.py
Execution-matched oracle replay wp1r_closeout.py
SQL harness generation and behavioral checks experiment/phase2/

These analysis sources require separately supplied protocol/configuration files and runtime inputs.

Four intentionally retained, rejected generated programs are listed in release/known_invalid_sources.json; they are not runtime entry points.

This repository includes code and the published paper's citation and main figure. Benchmark inputs, collected model responses, score matrices, and execution ledgers are not bundled. The offline tests validate code behavior with synthetic inputs; they do not reproduce the paper's measured results.

Citation

If you use this code or evaluation methodology, please cite the paper:

@misc{xu2026moreprogramsmorerolls,
  title = {More Programs or More Rolls? Separating Coverage from Specialization in {LLM} Harnesses},
  author = {Xu, Ziyang and Zhong, Haitian and Zhou, Hao and Qin, Hao and Jin, Chenhan and Qi, Te and Xu, Shengze and Zeng, Tieyong},
  year = {2026},
  eprint = {2609.35873},
  archivePrefix = {arXiv},
  primaryClass = {cs.AI},
  url = {https://arxiv.org/abs/2609.35873}
}

The same entry is available in CITATION.bib. CITATION.cff supplies GitHub's Cite this repository metadata with the paper as the preferred citation.

Contributing and license

English is the default maintenance language. See CONTRIBUTING.md. Original code uses the MIT license; bundled dependencies retain their own notices.

About

Code for More Programs or More Rolls? Separating coverage from specialization in LLM harnesses.

Resources

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages