feat(baseline): correctness truth chain — catalog CAS log, transform provenance, selection proposal, walker evaluation binding (AGBASE-FR-014..018, T082-T088, T051/T068/T069)
- catalog_revision_log.py: append-only JSONL revision log with If-Match CAS, duplicate
baseline/coordinate rejection, thread-race single-head proof, retire/invalidate
history preservation, publish_failed reconcile on the same revision (T082-T084 core).
- scenario_transform_provenance.py: server-owned 044 ScenarioArtifact chain for
scenario_transform candidates — run/owner/activeness/digest/passed-step/value equality,
forged refs typed-rejected (T086, AGBASE-FR-017).
- baseline_selection.py + schemas: facts-only BaselineSelectionProposal (cross_check ->
checklist -> big-number/closed-period -> filter_sentinel -> lineage_repr), hard 50-entry
budget with recorded exclusions, default+business filter contexts, CAS review with
accept/drop/adjust (T087, AGBASE-FR-016); observatory_entries tier in the catalog
revision schema + resolver smuggle-guard (T088, AGBASE-FR-015).
- transform_dsl.py: bounded structured op-tree DSL (sum_column/difference/ratio/scale,
depth<=4, 10k rows, canonical decimals, byte-stable) wired as derive_value into
bounded_transform (T068, AGSCN-FR-018).
- evaluation_aggregation.py: verdict policy single_shot/best_of_n with strict-majority
quorum, tie/all-error/quorum-miss/below-threshold typed non-pass (T069, AGSCN-FR-022).
- evaluation_binding.py + walker: dependency-ordered binding of the persisted covering
AgentEvaluation into the comparison decision (T051/T046); deferral until the covering
evaluation persists with a deadlock guard (_step_feeds_evaluation); semantic_comparison
scope in the policy mapper — required-mode evidence producers pass rows 1-8, the
semantic overlay binds to comparison steps only (production-chain ordering);
capacity_block.py extracted (INV_7 walker 398 LOC).