CALIBER
Quickstart
Strategy

CALIBER roadmap

Monthly release train and one- or two-day implementation backlog for the current CALIBER codebase.

ArchitectDecision-makerDeveloperStrategyGA
PrerequisitesCurrent-main capability matrix and first-mile evidence · A named single-tenant deployment owner and supported MLflow provider profile
Reviewed 2026-08-13 · current main branch and GitHub Project #2

This is the execution roadmap for the current CALIBER repository and GitHub Project #2. The v1.0.0 product requirements document defines the product boundary, requirements, and requirement-to-ticket trace; this document owns the delivery sequence, dates, capacity, and release gates. It uses a monthly release train and assumes 1.5 developers: one primary engineering stream plus a half-capacity stream for review, testing, documentation, operations, and release work.

This document is a plan, not a capability claim. The current checkout is an alpha/current-main build. A route, UI page, package, or passing unit test does not make a capability supported production behavior. The release gates below require reproducible evidence.

Operating rules

  • Every developer task is one or two days maximum. A task that grows beyond two days becomes a new child issue with its own title, acceptance criteria, and milestone.
  • A workstream is a coordination issue, not an implementation task. Workstreams stay open while their child tickets are executed and reviewed.
  • Each month has one primary stream and one half-capacity stream. Reserve the remaining capacity for review latency, regressions, pilot findings, release evidence, and operational interruptions.
  • Monthly increments are releasable checkpoints. Formal product tags are created only at M4 (caliber-v0.1.0) and M6 (caliber-v1.0.0) after their gates pass.
  • The first-mile pilot tests current main; it is not a production-readiness or stability claim.
  • Every ticket must state exact scope, deliverables, acceptance criteria, verification commands, dependencies, and out-of-scope work.

Current execution order

flowchart LR
  M1["M1 Sep 2026<br/>Pilot and scope freeze"] --> M2["M2 Oct 2026<br/>Workflow and governance"]
  M2 --> M3["M3 Nov 2026<br/>Contracts, recovery, diagnostics"]
  M3 --> M4["M4 Dec 2026<br/>v0.1.0 controlled-pilot gate"]
  M4 --> M5["M5 Jan 2027<br/>Single-tenant security and runtime safety"]
  M5 --> M6["M6 Feb 2027<br/>Recovery, acceptance, v1.0.0 decision"]
  M6 --> PV["Post-v1 Mar 2027<br/>Breadth and scale decisions"]

The controlling dependency is linear: pilot evidence selects the foundation scope; the foundation makes the supported path repeatable; then the supported single-tenant topology is hardened, operated, recovered, and accepted. A controlled-pilot tag is not a v1 release; it is evidence for the remaining v1 work.

docs/two-week-alpha-plan.md proposed a lighter, earlier checkpoint alongside this milestone; its Days 1–9 executed against current main, with Day 10 (cutting the actual release tag) left to the repo owner (see docs/reports/two-week-alpha-day10-decision.md). Per that plan's own "Relationship to the existing roadmap" section, this is evidence feeding M1's discovery exit criterion below — it does not replace or shorten M1–M6.

v1.0.0 MVP boundary and critical path

v1.0.0 is an enterprise-usable single-tenant product, not a feature inventory or a technical demo. It supports one named, governed prompt-refinement journey: a flagged MLflow trace is verified, diagnosed, evaluated, explicitly applied by an authorized operator, released through a durable intent, and either reconciled or rolled back with auditable evidence. The journey is available through the documented UI and a deliberately small, contract-tested API/SDK/CLI subset.

The supported deployment is one tenant, one active CALIBER process, a durable PostgreSQL metadata store, a named MLflow tracking/profile dependency, object storage where the selected path requires it, external TLS/ingress and network policy, least-privilege service credentials, authenticated role-based users, backup/restore evidence, logs/traces/metrics/readiness, and a human-operated incident and upgrade procedure. The final workload envelope is measured and published; it is not inferred from a test count.

This deliberately excludes multi-tenancy, SSO/SCIM, HA or active-active workers, multi-region disaster recovery, managed hosting, general-purpose untrusted tool/MCP execution, every asset family's lifecycle, and unmeasured horizontal scale. Those are post-v1 decisions. A second process must be rejected or API-only until a later ownership/broker design is implemented.

PhaseCritical pathIssuesExit evidence
M1 · ScopeFreeze the supported tenant, journey, dependency/provider profile, and acceptance evidence#100–#101; #153; #154Approved contract, capability truth, fixture, and release-blocker rules
M2 · FoundationMake the canonical prompt-refinement/release path durable and demonstrable#102–#107; #152Reproducible workflow, explicit release state, and external-effect reconciliation
M3 · ContractsMake the path diagnosable and supportable through agreed interfaces#108–#116; #151API/SDK/CLI contract, failure taxonomy, correlation, readiness, and worker decision
M4 · Controlled pilotProduce a clean, packaged, internal pilot release and close foundation defects#117–#121; #129Controlled-pilot evidence and a decision to enter v1 hardening; not a v1 claim
M5 · v1 hardeningEnforce the single-tenant security/runtime/deployment boundary#122–#125; #130–#135; #157–#159Unsafe configuration refused, one active owner, bounded workload, and recoverable topology
M6 · v1 acceptanceRehearse recovery, run the acceptance matrix, and make the v1.0.0 decision#136–#142Restore/rollback and operator evidence, no unaccepted P0/P1, human go/no-go

Monthly release train

MonthGitHub milestoneRelease/checkpointScopeExit evidence
M1 · Sep 2026M1 2026-09 — First-mile pilot and scope freezeCurrent-main pilot checkpointEpic #78, discovery #100–#101, #153, #154Environment record, feature results, reproducible bugs, capability matrix, and selected M2 scope
M2 · Oct 2026M2 2026-10 — Primary workflow and governanceFoundation incrementWorkstreams #102–#107; #152One seeded workflow is reproducible through UI and API/SDK; release and environment semantics are explicit
M3 · Nov 2026M3 2026-11 — API contracts, recovery, and diagnosticsAutomation/operations incrementWorkstreams #108–#116; #151Supported API/SDK/CLI matrix, bounded failure semantics, request correlation, and readiness diagnosis
M4 · Dec 2026M4 2026-12 — v0.1.0 release candidateControlled-pilot decisionWorkstream #117–#121; #129Clean build/package, pilot evidence, no unaccepted pilot blocker, and human decision to begin v1 hardening
M5 · Jan 2027M5 2027-01 — v1 security and process safetyv1 single-tenant hardeningWorkstreams #67–#70; #122–#135; #157–#159Unsafe defaults rejected, one active loop owner, bounded concurrency, explicit backup inventory, and supported deployment profile
M6 · Feb 2027M6 2027-02 — v1 performance, recovery, and releaseFormal caliber-v1.0.0 decisionWorkstreams #70–#71; #136–#142Measured workload limits, restore/rollback drill, acceptance packet, production-like bug bash, and human go/no-go
Post-v1 · Mar 2027Post-v1.0.0 — 2027-03 reviewDecision checkpoint, not a committed releaseEpic #58; workstreams #73–#75; tasks #143–#149Demand/evidence-based decisions for identity, lifecycle/plugins, managed hosting, and scale

Epic and workstream structure

M1 pilot — Epic #78

M1 is discovery only: use one environment to validate the prompt-refinement journey, fixture assumptions, evidence format, and capability boundary. Record pass/fail/blocked/not-supported results, exact commands, and reproducible defects, then use #100–#101 to establish capability truth and triage rules. M1 does not commit the team to implementing the broader pilot journey inventory or to a production-readiness claim; if the discovery evidence is insufficient, the implementation start moves rather than silently expanding scope.

v0.1.0 foundation and controlled pilot — Epic #56

The foundation is coordinated through #59–#65, #76, and #66. The implementation children are #100–#121, #129, and the attached #151–#154 decisions. The foundation is complete only when one supported workflow is reproducible, the automation subset is honest, failure and diagnostics behavior is actionable, and the clean deployment/release path has evidence.

v1.0.0 single-tenant safety and operations — Epic #57

This is part of the v1 critical path, not optional production expansion. It is coordinated through #67–#71, #77, and Issue #72, with implementation children #122–#125 and #130–#142 plus the attached deployment/recovery work. Any v1 decision requires measured security, one-process ownership, workload, recovery, and operator evidence; feature inventory or test count alone is insufficient. Broker-backed multi-worker delivery is explicitly deferred unless #151 changes the supported topology decision.

Post-v1 decisions — Epic #58

Post-v1 work is deliberately small and decision-led. #143–#149 may produce decision records, prototypes, or narrowly scoped internal seams. Treat the 10-person-day total as a discovery ceiling, not a commitment to complete all seven tickets in March; start only the items justified by the M6 evidence. None of these items can block M6 or be described as current supported product behavior.

Capacity model

The 1.5-developer assumption is a planning limit, not permission to run every ticket in parallel:

Capacity slicePurposeMonthly rule
Primary streamOne critical workflow or safety outcomeOne active workstream; pull its child tickets in dependency order
Half streamTests, docs, pilot triage, diagnostics, or small hardeningOne active supporting workstream; never hide a second major behind it
ReserveReview, regressions, release evidence, incidentsProtect this capacity every month; it is part of the plan

At the start of each month, choose the next child ticket only after checking the parent decision, code seam, fixture, test command, and dependency. At the mid-month review, cut stretch work before cutting release evidence or safety tests.

The implementation estimates currently assigned to executable tickets are:

MilestoneImplementation ticketsEstimated effortCapacity judgement
M1#100–#101; #153; #1544 person-days plus the selected pilot journeysFeasible only if #79–#99 are reduced to the canonical path and exploratory work does not block scope freeze
M2#102–#107; #15210 person-daysComfortable; reserve time for pilot-driven scope changes
M3#108–#116; #15112 person-daysFeasible with a protected diagnostics/review stream
M4#117–#121; #1299 person-daysFeasible only if the bug bash and decision remain release work, not feature expansion
M5#122–#125; #130–#135; #157–#159Re-estimate after #151This is the tightest phase. It cannot fit as written until broker work and duplicate deployment tickets are removed or deferred.
M6#136–#142Re-estimate after M5Feasible only after the topology, RPO/RTO, and workload envelope are concrete.
Post-v1#143–#14910 person-daysDiscovery ceiling; select a subset after #142 rather than committing all work

Codebase grounding

The ticket decomposition follows the current package and test boundaries:

CapabilityImplementation seamsExisting verification anchors
Server and routescaliber/src/caliber/server.py, caliber/src/caliber/routes/caliber/tests/test_routes_*.py, contract smoke tests
Workflow runtimecaliber/src/caliber/workflows/, caliber/src/caliber/orchestrator/workflow compiler/runtime/e2e tests
Release and governancecaliber/src/caliber/release_operations.py, caliber/src/caliber/routes/releases.py, caliber/src/caliber/workflows/deploy_gate.pyrelease, promoter, approval, and deploy-gate tests
Reliability and eventscaliber/src/caliber/events/, worker/runtime modules, effect ledgerevent sequencing, worker edge-case, retry, and effect-ledger tests
Observabilitycaliber/src/caliber/observability/, health, metrics, trace routesreadiness, logging, trace, metrics, and observability route tests
SDK and CLIsdk/caliber-sdk/, sdk/caliber-cli/SDK transport/resource/waiter and CLI tests
UI and pilot journeyscaliber/caliber-ui/src/pages/, Playwright configs, fixture packpage tests, cookbook tests, focused Playwright journeys
Deployment and docsdeploy/, scripts/ci-local.sh, docs sync/build scriptsdeployment, packaging, docs executable-spec, and CI gates

Release gates

M1 exit

  • MVP discovery evidence is captured and approved.
  • The #153/#154 MVP contract is agreed and testable.
  • #100 capability truth is established.
  • #101 triage rules are documented and accepted.

M4 caliber-v0.1.0 controlled-pilot gate

  • #102–#119 are complete or explicitly descoped with evidence.
  • #129 and #120 have no untriaged P0/P1 release blocker.
  • The clean-checkout build, deployment, docs, and rollback path is reproducible.
  • #121 records a human go/no-go decision before any formal tag is published.

M6 caliber-v1.0.0 gate

  • #122–#140 and the selected single-tenant deployment tickets provide dated security, process, performance, recovery, and acceptance evidence.
  • #141 has no unaccepted P0/P1 blocker.
  • #142 records a human go/no-go decision and publishes residual risks.
  • The release notes state the measured supported envelope; they do not imply enterprise identity, managed hosting, multi-region, or unmeasured scale.

Bug and feedback workflow

Testers use the labels area-testing and the applicable release label:

release-v0.1.0 + area-testing   M1–M4 current release work
release-v1.0.0 + area-testing   M5–M6 production-gate testing
help wanted                     Optional volunteer/contributor attention

A bug report must include release commit, environment, fixture, exact steps, expected result, actual result, logs/request ID, severity, and whether the behavior is supported by the capability matrix. A test ticket records the finding; it does not absorb the implementation fix.

Success measures

  • M1 adoption: all named pilot journeys have evidence and no untriaged critical finding.
  • M4 controlled pilot: one supported workflow and automation path are repeatable from a clean checkout with truthful pilot notes.
  • M6 enterprise single-tenant MVP: security, process ownership, measured workload, recovery, acceptance, deployment, and operator evidence support a human decision.
  • Post-v1 discipline: deferred work is selected from production evidence and demand, not from the existence of an attractive architectural seam.

Sources and current limitations

The product completeness report, the pilot fixture pack, the package manifests, CI workflows, and the linked GitHub issues are the evidence sources for this roadmap. Known limitations remain explicit: current-main is alpha; the v1 target is a bounded single-tenant deployment, not a claim of multi-tenant, HA, managed-hosting, multi-region, or broad asset-lifecycle support.

CALIBER : Contextual Adaptive Lifecycle for Intelligent Build, Evaluation, and Refinement — this page is generated from the authoritative Markdown sources in docs/ and the repository-level ARCHITECTURE.md.