Ground every value
Every synthesized argument must resolve from entry state or a prior observed result through a bounded transform.
Evidence-licensed workflow specialization
Guarded Agentic Compaction turns recurrent, read-only tool prefixes into inspectable programs—but only when value provenance, effects, compatibility, replay, and finite-sample evidence agree.
The problem
Mature agents often pay an LLM to choose the same evidence path again. A naive optimizer bundles the calls. GAC asks the harder question: what observable evidence permits removing those model boundaries, and when must the optimizer abstain?
Every synthesized argument must resolve from entry state or a prior observed result through a bounded transform.
Unknown effects, writes, approvals, handoffs, errors, incompatible principals, and unsafe runtime positions stop compilation.
When support cannot justify the configured risk bound, the artifact retires and the unchanged agent remains the answer.
The aha moment
A real GitHub trace looks like an obvious full macro:
issue_get_record → issue_get_labels → issue_get_comments(limit=3)
Across 132 discovery traces, this three-read shape recurs with 116 supporting windows. But the comments limit varies and has no consistent trace-grounded expression. GAC rejects the full candidate with ungroundable_slot, emits only record → labels, and leaves comments plus final rendering to the agent. The admitted prefix passes 92/92 calibration groups at α=.05.
Replay can still miss the answer. In the separate 18-record study, tool replay was 45/45 clean, yet issue #6602 lost the URL from a Markdown link and the downstream result was 17/18. GAC’s continuation guard detects the mismatch and checked-renders 18/18 in provider-free replay.
Repetition says what may be compilable. Provenance, continuation scope, and evidence determine whether replacement is justified.
Compiler path
The method is deliberately conservative. Passing one stage never weakens the next; a recurrent sequence can still fail provenance, effects, synthesis, validation, or admission.
Normalize traces, manifests, entry state, outcomes, and isolation keys.
Cut regions at effects, approvals, handoffs, errors, and position boundaries.
Resolve typed provenance and emit only programs in the closed DSL.
Replay grouped traces, challenge contracts, and retain counterexamples.
Freeze the score, apply exact bounds, sign the artifact, or retire it.
Measured evidence
Each row licenses a different claim. Real-record paired experiments test end-to-end intervention; public benchmarks test trace completeness, synthesis, replay, official harness compatibility, or an explicit gate.
| Evidence | Scope | Observed result | Interpretation |
|---|---|---|---|
| Three GitHub workflows | 90 paired held-out public records | Measured Compiled 90/90 vs baseline 89/90; requests −66.6%; cost −58.7% weighted | Efficiency transfers across issue type, PR outcome, and backlog attention on one snapshot. |
| Hand-written programs | Same 90 records | Measured 90/90; fair pre-model programs tie the two new learned programs | Manual code remains the runtime baseline; automatic evidence and lifecycle are the value claim. |
| Cross-repository, time-forward extension | Five frozen repositories, exact two-read task | Measured 580/580 discovery; four repositories complete 120/120 held-out, the fifth retires at compile time | The guarded lifecycle survives a new repository and a later window on a narrower task; a fixed template is still more efficient there. |
| Prospective gate-frontier study | Same five repositories, learned gate vs. support-only comparator | Measured 240/300 pooled held-out pairs on four repositories; learned and support-only gates statistically indistinguishable | The pre-declared null: at 4x the prior scale, against a comparator built to isolate the risk budget, the exact gate remains a support threshold, not a demonstrated frontier. |
| NESTFUL | 1,415 executable traces | Retired 24 pass / 12 abstain / 0 wrong; max support 26 < 92 | Recurrence and sound held-out replay are still insufficient for admission. |
| API-Bank | 212 complete traces | Retired 2 synthesized; both held-out windows abstain | A second compiler substrate reproduces the evidence shortage. |
| Eight supplementary benchmarks | Gold plans, simulations, live web, and real issues | Scoped Adapters, official paths, or prerequisite gates | Interoperability evidence is retained but excluded from the optimizer comparison. |
Why it is different
Agent compilation, meta-tools, prompt evolution, and plan caching already exist. GAC contributes a safety argument for specializing from historical executions.
The paper, browser edition, benchmark explorer, technical deck, sealed result artifacts, and executable claim checks are available from one publication shelf.