Evidence is separated by what it can prove.

The evaluation does not average real records, public simulators, gold plans, task-only datasets, and executable traces. Every source retains its substrate, denominator, execution status, and claim boundary.

Research thesis

Historical traces establish recurrence—not admissibility.

The experiments ask whether provenance can reconstruct a program, whether independent evidence can certify it, whether the resulting intervention preserves source-grounded outcomes, and whether it remains worthwhile against a well-placed manual program.

Trace question

Can every dynamic argument be derived from observable entry state or prior results?

Admission question

Does grouped calibration support the registered violation bound, or must the family retire?

Systems question

Does the admitted intervention improve requests, tokens, latency, and cost without weakening exact outcomes?

Experimental design

Three primary workflow families use real public records, live provider calls, distinct tool vocabularies, and exact source-grounded grading.

  • 132 discovery and 30 held-out records per family, balanced across three exact classes
  • Baseline, compiler, and macro order counterbalanced over all six permutations
  • Quality computed independently of tool order from pinned source fields and excerpts
  • Ten repeated test records used for structural and surface determinism analysis
  • Paid provider outputs retained; later oracle corrections never rewrite online evidence
  • Provider optimization overhead separated from held-out deployment measurements

Primary three-family result

Issue-type routing, PR-outcome audit, and backlog-attention routing cover 90 held-out public records. Compiled programs reach 90/90 exact outcomes versus 89/90 baseline; hand-written programs also reach 90/90.

FamilyExact baseline → compiledRequestsInterfacesTokensLatencyCost
Issue type30/30 → 30/30−50.0%0.0%−39.5%−51.7%−32.0%
PR outcome30/30 → 30/30−75.0%−66.7%−80.8%−73.0%−75.3%
Backlog attention29/30 → 30/30−74.8%−66.3%−81.4%−68.9%−75.1%
Weighted total89/90 → 90/90−66.6%−44.2%−63.1%−64.2%−58.7%

The single backlog baseline miss is not a statistically meaningful quality improvement. Fair manual programs tie learned programs structurally on both new families. The evidence supports workflow-family transfer on one snapshot—not full-workflow cross-repository or time-forward generalization, and not runtime superiority over correct manual code. A separate, deliberately narrower two-read extension tests the cross-repository, time-forward question on its own terms (below).

Issue-type family detail

All 132 discovery traces choose the same three-read order, but the comment-limit argument varies and is not trace-groundable. GAC refuses the complete region and emits the longest valid prefix: record → labels. The ordinary agent performs the comments read and final rendering.

ConditionRequestsToolsTokensWallCostExact
Unchanged agent4.03.04,259.46.16 s$0.00087230/30
GAC two-read prefix2.03.02,576.52.97 s$0.00059330/30
Hand-written macro2.01.01,780.83.25 s$0.00054530/30

The compiler reduces requests 50.0%, total tokens 39.5%, observed wall latency 51.7%, and estimated cost 32.0% versus unchanged. The macro matches request reduction and is better on tools, tokens, and dollars. The paired latency interval between macro and compiler crosses zero; no runtime-superiority claim is licensed.

Natural live comparison of requests, tools, tokens, latency, cost, and exact pass rates.
Expanded real-record comparison. All three conditions preserve 30/30 exact contracts; the macro is the strongest practical runtime comparator.

Cross-repository, time-forward extension

A separate frozen-source study asks the narrower question the primary families cannot: does a guarded artifact survive a different repository and a later time window? The task is simplified to an exact two-read PR-outcome contract, so this extends reach, not scope.

A provider-free preflight seals five repositories under a strict time-forward split. Executed discovery reaches 580/580 exact traces (116/116 per repository). Four repositories admit an artifact and complete 120 held-out paired records; pytorch/pytorch retires at compile time under the frozen-candidate exact gate. Baseline, compiled, and fixed-template conditions all pass 120/120.

CohortDiscoveryHeld-out exactRequestsTokensLatencyCost
Frozen five-repository — compiled vs baseline580/580120/120−44.4%−52.4%−49.4%−48.6%
Frozen five-repository — fixed template vs baseline580/580120/120−66.7%−78.6%−68.1%−73.3%
Balanced rerun — compiled vs baseline360/360180/180−66.7%−78.4%−60.7%−72.7%
Balanced rerun — fixed template vs baseline360/360180/180−66.7%−78.4%−63.0%−72.3%

On the frozen cohort the ungated fixed template is more efficient than the learned artifact, which sharpens rather than weakens the claim: the artifact earns its place through automatic discovery, guarded admission, and retirement—not by dominating hand-placed code on a simplified task. That cohort also keeps a principled negative. Across the four completed repositories the artifact compacts all 40 merged and all 40 closed_unmerged held-out pull requests and falls back on all 40 open ones.

A balanced rerun asks whether that open-only fallback is intrinsic. On the three repositories with enough class support for a balanced strict time-forward split (pandas-dev/pandas, psf/requests, pytorch/pytorch), a round-robin preflight seals 120 discovery and 60 held-out pull requests per repository, evenly split across open, merged, and closed_unmerged. All three admit the same two-read artifact, it compacts all 180 held-out records, and the verifier's pr.state hull contains both open and closed everywhere. The earlier fallback pattern and the pytorch/pytorch retirement are properties of the original frozen cohort design, not limits of the guarded runtime.

What this does and does not add. The extension strengthens cross-repository, time-forward evidence for a two-read task. It does not widen the task: both cohorts remain narrower than the three-tool workflow-family studies, so full-workflow generalization is still unestablished.

Prospective gate-frontier study: the pre-declared null

Every registered alpha=.05 gate this paper reports is a step function, so a pre-registered protocol committed a design — before any evidence existed for it — to test that at scale: the same five repositories above, 116 discovery and 60 held-out cases each, three arms compared under identical records, model, cache policy, and ordering. The third arm, a support-only comparator, is operationalized as alpha=1 in an otherwise byte-for-byte identical compile pass — same mining, synthesis, challenge, and frozen-candidate selection — so it isolates the risk budget specifically rather than testing a different mechanism.

Four repositories admit a candidate and complete 240 of the pre-registered 300 held-out pairs (580/580 exact discovery traces); pytorch/pytorch again retires at compile time, reproducing — on an independent, five-times-larger cohort — the same repository's retirement above rather than revealing a new failure mode. On those 240 pairs the learned gate and the support-only gate are statistically indistinguishable on every efficiency metric, because every admitting repository deploys the identical coverage-1.0 threshold.

RepositoryHeld-outExactNonzero coverage in sweepAdmitted
huggingface/datasets6060/600.1087 (upper .375, rejected), 1.01.0
pandas-dev/pandas6060/601.01.0
psf/requests6059/601.01.0
streamlit/streamlit6060/601.01.0
pytorch/pytorch—retirenone (upper 1.0 at every threshold)—

This is the pre-declared null, not a frontier. Reading every point the frozen threshold grid produces, not just the selected one: three of four repositories show a pure two-value step; the fourth's one intermediate value has an exact upper bound (.375) far above the registered budget and is never admissible. No held-out wrong dispatch occurred on either gate. At four times the prior cross-repository scale, against a comparator built to isolate exactly the mechanism in question, the exact-alpha=.05 gate remains a support threshold — this confirms, rather than resolves, the step-gate finding below.

Compilation depth is a safety variable

An earlier aggressive three-read artifact reduces provider requests by 75%, but records one compiler-only factual miss: 17/18 exact versus 18/18 for unchanged and macro. Clean tool replay therefore does not certify the model continuation. A provider-free continuation checker later detects and repairs that retained miss, but it is not a prospective live safety result.

Guarded composite and fair placement

GCS packages the admitted three-read program behind a continuation-pinned projection. Against a provider-visible macro on 12 fresh issues, both pass 12/12; GCS uses one rather than two provider requests, 38.9% fewer tokens, 40.0% lower observed wall latency, and 32.3% lower estimated cost.

A subsequent six-record study gives an independently authored manual program the same pre-model position. GCS and manual both pass 6/6 and tie at one request, one interface, and identical input tokens. Official GEPA 0.1.4 retains its seed after 14 task evaluations; its 59-request optimization overhead is reported separately. These are bounded, exploratory results—not general rankings.

Supplementary benchmark interoperability audit

Only NESTFUL and API-Bank retain complete observed values suitable for post-trace compilation. Eight other paths are preserved below for reproducibility, but they do not demonstrate optimizer value and are excluded from the main comparison.

The stronger refusal example: all 132 GitHub discovery traces read record, labels, and comments, so a full three-read macro looks obvious. GAC rejects that candidate because comments.limit has no consistent trace-grounded expression, emits only the record → labels prefix, and leaves comments plus final rendering to the agent. The admitted prefix passes 92/92 calibration groups at α=.05. Separately, an earlier three-read artifact achieved 45/45 clean tool replays but only 17/18 downstream answers: issue #6602 lost a Markdown URL, and the provider-free continuation guard detected and checked-rendered the miss.

Open the benchmark explorer to search and filter all thirteen rows — every substrate, execution status, effect breakdown, and claim boundary, generated from the evidence files.

BenchmarkPathResultClaim boundary
NESTFULCompiler1,415 traces; 24 / 12 / 0 held-out; all retireProvider-free structural compiler evidence
API-BankCompiler + API replay212 traces; 0 / 2 / 0; all retireSecond compiler refusal substrate
BFCL v4Official checker200/200 gold plans validNo model-quality claim
ToolSandboxOfficial simulator0.982 milestone similarityReal provider, simulated environment
τ²/τ³Official simulators0/4 reward; 288,757 tokensReal provider, simulated domains
BrowseCompLive-web subset1/3; 28 searchesHosted-search bypass, not compiler evidence
ToolBench / AgentBenchAdaptersFixtures and tasks normalizedFull data/services gated
GAIA / SWE-benchPreflight / task adapterAuthorization and host gates retainedNo score imputed
Candidate family supports relative to the exact gate requirement of 92 independent groups.
Certification is the scarce resource. Maximum NESTFUL family support is 26, well below the configured 92-group requirement.

Statistical interpretation

Efficiency comparisons are paired by issue. The artifacts report 10,000-sample paired bootstrap intervals and two-sided Wilcoxon signed-rank tests; binary quality uses exact McNemar tests. These secondary tests are descriptive and unadjusted. Zero observed failures is not equivalence: 0/30 still gives a one-sided 95% upper failure bound of approximately 9.5%.

The gate did not demonstrate a risk–coverage frontier. Its score is effectively all-or-none because development contains no unproductive examples. The current evidence demonstrates sample-size refusal and admission mechanics, not selective discrimination among future inputs. This is no longer only an under-tested limitation: the gate-frontier study above ran a purpose-built, statistically matched comparator at four times the prior cross-repository scale specifically to check it, and confirms the same finding rather than resolving it.

Reproducibility

All quantitative tables and figures are regenerated from sealed JSON/CSV artifacts. Dataset revisions, source trees, checksums, task splits, native trace identifiers, model versions, and price assumptions are retained. The core verification path makes no provider calls.

.venv/bin/python paper/scripts/build_artifacts.py
.venv/bin/python paper/scripts/validate_artifacts.py
.venv/bin/python scripts/verify_release.py
.venv/bin/python -m pytest

Exact commands for the external benchmark paths are in the benchmark audit. The experiment verification report records the independent consistency checks and residual threats.