Trace question
Can every dynamic argument be derived from observable entry state or prior results?
The evaluation does not average real records, public simulators, gold plans, task-only datasets, and executable traces. Every source retains its substrate, denominator, execution status, and claim boundary.
Research thesis
The experiments ask whether provenance can reconstruct a program, whether independent evidence can certify it, whether the resulting intervention preserves source-grounded outcomes, and whether it remains worthwhile against a well-placed manual program.
Can every dynamic argument be derived from observable entry state or prior results?
Does grouped calibration support the registered violation bound, or must the family retire?
Does the admitted intervention improve requests, tokens, latency, and cost without weakening exact outcomes?
Three primary workflow families use real public records, live provider calls, distinct tool vocabularies, and exact source-grounded grading.
Issue-type routing, PR-outcome audit, and backlog-attention routing cover 90 held-out public records. Compiled programs reach 90/90 exact outcomes versus 89/90 baseline; hand-written programs also reach 90/90.
| Family | Exact baseline → compiled | Requests | Interfaces | Tokens | Latency | Cost |
|---|---|---|---|---|---|---|
| Issue type | 30/30 → 30/30 | −50.0% | 0.0% | −39.5% | −51.7% | −32.0% |
| PR outcome | 30/30 → 30/30 | −75.0% | −66.7% | −80.8% | −73.0% | −75.3% |
| Backlog attention | 29/30 → 30/30 | −74.8% | −66.3% | −81.4% | −68.9% | −75.1% |
| Weighted total | 89/90 → 90/90 | −66.6% | −44.2% | −63.1% | −64.2% | −58.7% |
The single backlog baseline miss is not a statistically meaningful quality improvement. Fair manual programs tie learned programs structurally on both new families. The evidence supports workflow-family transfer on one snapshot—not cross-repository, time-forward, or runtime superiority over correct manual code.
All 132 discovery traces choose the same three-read order, but the comment-limit argument varies and is not trace-groundable. GAC refuses the complete region and emits the longest valid prefix: record → labels. The ordinary agent performs the comments read and final rendering.
| Condition | Requests | Tools | Tokens | Wall | Cost | Exact |
|---|---|---|---|---|---|---|
| Unchanged agent | 4.0 | 3.0 | 4,259.4 | 6.16 s | $0.000872 | 30/30 |
| GAC two-read prefix | 2.0 | 3.0 | 2,576.5 | 2.97 s | $0.000593 | 30/30 |
| Hand-written macro | 2.0 | 1.0 | 1,780.8 | 3.25 s | $0.000545 | 30/30 |
The compiler reduces requests 50.0%, total tokens 39.5%, observed wall latency 51.7%, and estimated cost 32.0% versus unchanged. The macro matches request reduction and is better on tools, tokens, and dollars. The paired latency interval between macro and compiler crosses zero; no runtime-superiority claim is licensed.

An earlier aggressive three-read artifact reduces provider requests by 75%, but records one compiler-only factual miss: 17/18 exact versus 18/18 for unchanged and macro. Clean tool replay therefore does not certify the model continuation. A provider-free continuation checker later detects and repairs that retained miss, but it is not a prospective live safety result.
GCS packages the admitted three-read program behind a continuation-pinned projection. Against a provider-visible macro on 12 fresh issues, both pass 12/12; GCS uses one rather than two provider requests, 38.9% fewer tokens, 40.0% lower observed wall latency, and 32.3% lower estimated cost.
A subsequent six-record study gives an independently authored manual program the same pre-model position. GCS and manual both pass 6/6 and tie at one request, one interface, and identical input tokens. Official GEPA 0.1.4 retains its seed after 14 task evaluations; its 59-request optimization overhead is reported separately. These are bounded, exploratory results—not general rankings.
Only NESTFUL and API-Bank retain complete observed values suitable for post-trace compilation. Eight other paths are preserved below for reproducibility, but they do not demonstrate optimizer value and are excluded from the main comparison.
Open the benchmark explorer to search and filter all thirteen rows — every substrate, execution status, effect breakdown, and claim boundary, generated from the evidence files.
| Benchmark | Path | Result | Claim boundary |
|---|---|---|---|
| NESTFUL | Compiler | 1,415 traces; 24 / 12 / 0 held-out; all retire | Provider-free structural compiler evidence |
| API-Bank | Compiler + API replay | 212 traces; 0 / 2 / 0; all retire | Second compiler refusal substrate |
| BFCL v4 | Official checker | 200/200 gold plans valid | No model-quality claim |
| ToolSandbox | Official simulator | 0.982 milestone similarity | Real provider, simulated environment |
| τ²/τ³ | Official simulators | 0/4 reward; 288,757 tokens | Real provider, simulated domains |
| BrowseComp | Live-web subset | 1/3; 28 searches | Hosted-search bypass, not compiler evidence |
| ToolBench / AgentBench | Adapters | Fixtures and tasks normalized | Full data/services gated |
| GAIA / SWE-bench | Preflight / task adapter | Authorization and host gates retained | No score imputed |

Efficiency comparisons are paired by issue. The artifacts report 10,000-sample paired bootstrap intervals and two-sided Wilcoxon signed-rank tests; binary quality uses exact McNemar tests. These secondary tests are descriptive and unadjusted. Zero observed failures is not equivalence: 0/30 still gives a one-sided 95% upper failure bound of approximately 9.5%.
The gate did not demonstrate a risk–coverage frontier. Its score is effectively all-or-none because development contains no unproductive examples. The current evidence demonstrates sample-size refusal and admission mechanics, not selective discrimination among future inputs.
All quantitative tables and figures are regenerated from sealed JSON/CSV artifacts. Dataset revisions, source trees, checksums, task splits, native trace identifiers, model versions, and price assumptions are retained. The core verification path makes no provider calls.
.venv/bin/python paper/scripts/build_artifacts.py
.venv/bin/python paper/scripts/validate_artifacts.py
.venv/bin/python scripts/verify_release.py
.venv/bin/python -m pytest
Exact commands for the external benchmark paths are in the benchmark ledger. The experiment verification report records the independent consistency checks and residual threats.