Ground every value
Every synthesized argument must resolve from entry state or a prior observed result through a bounded transform.
Evidence-licensed workflow specialization
Guarded Agentic Compaction turns recurrent, read-only tool prefixes into inspectable programs—but only when value provenance, effects, compatibility, replay, and finite-sample evidence agree.
The problem
Mature agents often pay an LLM to choose the same evidence path again. A naive optimizer bundles the calls. GAC asks the harder question: what observable evidence permits removing those model boundaries, and when must the optimizer abstain?
Every synthesized argument must resolve from entry state or a prior observed result through a bounded transform.
Unknown effects, writes, approvals, handoffs, errors, incompatible principals, and unsafe runtime positions stop compilation.
When support cannot justify the configured risk bound, the artifact retires and the unchanged agent remains the answer.
Compiler path
The method is deliberately conservative. Passing one stage never weakens the next; a recurrent sequence can still fail provenance, effects, synthesis, validation, or admission.
Normalize traces, manifests, entry state, outcomes, and isolation keys.
Cut regions at effects, approvals, handoffs, errors, and position boundaries.
Resolve typed provenance and emit only programs in the closed DSL.
Replay grouped traces, challenge contracts, and retain counterexamples.
Freeze the score, apply exact bounds, sign the artifact, or retire it.
Measured evidence
Each row licenses a different claim. Real-record paired experiments test end-to-end intervention; public benchmarks test trace completeness, synthesis, replay, official harness compatibility, or an explicit gate.
| Evidence | Scope | Observed result | Interpretation |
|---|---|---|---|
| Three GitHub workflows | 90 paired held-out public records | Measured Compiled 90/90 vs baseline 89/90; requests −66.6%; cost −58.7% weighted | Efficiency transfers across issue type, PR outcome, and backlog attention on one snapshot. |
| Hand-written programs | Same 90 records | Measured 90/90; fair pre-model programs tie the two new learned programs | Manual code remains the runtime baseline; automatic evidence and lifecycle are the value claim. |
| NESTFUL | 1,415 executable traces | Retired 24 pass / 12 abstain / 0 wrong; max support 26 < 92 | Recurrence and sound held-out replay are still insufficient for admission. |
| API-Bank | 212 complete traces | Retired 2 synthesized; both held-out windows abstain | A second compiler substrate reproduces the evidence shortage. |
| Eight supplementary benchmarks | Gold plans, simulations, live web, and real issues | Scoped Adapters, official paths, or prerequisite gates | Interoperability evidence is retained but excluded from the optimizer comparison. |
Why it is different
Agent compilation, meta-tools, prompt evolution, and plan caching already exist. GAC contributes a safety argument for specializing from historical executions.
The paper, editable slide decks, sealed result artifacts, and executable claim checks are available with the source.