docs(specs): live evaluation-binding wave accounting — canary v5 evidence, gates closed, matrix updates

- live_canary_v5_binding.py: sibling-topology live harness (published baseline envelope ->
  browser -> capture -> compare + evaluate with declared comparison_refs, required policy).
- Evidence live-canary-v5-binding-20260921T164034Z (+eval-record): complete live chain —
  8 durable artifacts, real-LLM evaluation persisted (ce4685e7, succeeded), walker binds it
  into the compare decision; honest LOW_CONFIDENCE terminal (all gateway vision routes
  upstream-502 since 2026-09-21; text-only model cannot judge visual criteria).
- Tasks: T052/T053 CLOSED (shutdown/readback), T046/T051 live-proven (terminal PASS
  external-blocked on the vision gateway), T068/T069 CLOSED, T082-T088 accounted.
- CHK038 PARTIAL (date filter absent on stand), CHK039/040 CLOSED.
- Matrix: G-PROVIDER-SHUTDOWN and G-MUTATION-RECON CLOSED; G-BROWSER-FILTERS PARTIAL
  (external); G-BSL-CATALOG core proven; G-BSL-DERIVED offline axes closed;
  G-E2E-EVALUATION chain live-proven, terminal PASS external-blocked.
- WORKSTATE: wave checkpoints 2026-09-18/21 (runtime safety, correctness chain, live).
This commit is contained in:
2026-09-22 10:12:14 +03:00
parent fc0c426537
commit c5e3303e35
8 changed files with 690 additions and 17 deletions

View File

@@ -119,9 +119,9 @@ SCEX-FR-028: [Production baseline-backed evaluation](../contracts/production-cha
- [~] CHK035 End-to-end real browser→capture→durable artifact→deterministic comparison→optional immutable evaluation→policy→result; all required evidence present before PASS. Evidence: [T046](../tasks.md), [traceability](../traceability.md). Partial 2026-09-17: the chain up to policy is PROVEN live (canary v2/v4; browser canaries 6/6); the graph-level terminal PASS is blocked by the walker-level evaluation→comparison binding (`policy_inputs_from_outcome` reads `evaluation_input` only from the `agent_evaluation` step), diagnosed with executable proof; closure = binding + live v4 rerun.
- [ ] CHK036 Negative product UI test: no agent chat/prompt/assistant editing/proposal-generation/typical-operation-to-agent/workspace/start/handoff controls or agent invocation routes/requests; manual CRUD/editor/human review/read-only results remain usable.
- [x] CHK037 Wave-2 structural browser actions have unit failure matrices and retained live PREPROD evidence: `assert_dom`, `inspect_filter_options`, `navigate_tabs`, `wait_for_selector`. **CLOSED 2026-09-18:** `4274e403`, unit 811; `evidence/browser-provider/readonly-canary-20260918T125451Z.json` **4/4 PASS** + PNG.
- [ ] CHK038 Native filter search/clear/date have retained live evidence showing selected/filter-cleared/date state and chart-data response; offline implementation `683347be`/791 green only.
- [ ] CHK039 Provider shutdown leaves zero pending Playwright/asyncio tasks after live canary and soak (`PROVIDER-SHUTDOWN-001`, T052).
- [ ] CHK040 Mutation PASS requires pinned independent SQL readback equal to provider post_rows; mismatch is reconciliation_required/blocked; live mutate+restore proof closes `MUT-RECON-001` (T053).
- [~] CHK038 Native filter search/clear/date have retained live evidence showing selected/filter-cleared/date state and chart-data response; offline implementation `683347be`/791 green only. **Live 2026-09-18** (`filters-canary-20260918T195946Z.json`+PNG, dashboard 11): values-apply (`applied_mode=values`, `chart_data_observed=true`, chip `Japan` observed post-apply), clear (`applied_mode=clear`, `chart_data_observed=true`), clear-checkpoint replay capture (1 replayable entry). PARTIAL: date vector blocked — no date/time filter exists on any stand dashboard (external prerequisite, stand owner); search vector structurally inapplicable on this stand (dropdown renders no search input; typed search flow remains unit-proven).
- [x] CHK039 Provider shutdown leaves zero pending Playwright/asyncio tasks after live canary and soak (`PROVIDER-SHUTDOWN-001`, T052). **CLOSED 2026-09-18:** `close_all_sessions`+`stop()` bounded drain; unit `test_provider_shutdown.py`; live canaries `readonly-canary-20260918T{174659,184343}Z`: `task_destroyed_warnings=[]`, `thread_alive=false`, `leftover_sessions=0`.
- [x] CHK040 Mutation PASS requires pinned independent SQL readback equal to provider post_rows; mismatch is reconciliation_required/blocked; live mutate+restore proof closes `MUT-RECON-001` (T053). **CLOSED 2026-09-18:** same-session fresh-client_id readback (`browser_readback.py`), `readback_verified` checkpoint in PASS, `BrowserTransportReadbackMismatch`→`reconciliation_required`/effect `unknown`/retry blocked (unit-proven); live mutation canary `readonly-canary-20260918T184343Z`: `readback_ok=true` ×2, receipts carry readback proof, `failures=[]`. Note: the 82.75/82.74 divergence was root-caused as mid-step post-SELECT vs post-restore lifecycle skew (both values correct at their point in time), not a cache/transaction defect; readback now compares against the cleanup-aware expectation.
- [ ] CHK041 Hot/cold tiering proves verify-before-delete, holds, WebP markers, lossy pixel-compare refusal and idempotent sweep (T054).
- [ ] CHK042 Scheduled zero-human E-class run reaches terminal PASS/FAIL/inconclusive from pinned AgentEvaluation+DecisionPolicy; no runtime human step and no evidence-free PASS (T051).

View File

@@ -0,0 +1,46 @@
{
"evaluation_id": "ce4685e7-28ec-423f-b26e-71d78803b7c0",
"verdict": "inconclusive",
"confidence": 0.0,
"status": "succeeded",
"comparison_ids": [
"compare-to-baseline",
"capture-evidence"
],
"baseline_pin": {
"schema_version": 1,
"baseline_set_id": "ss-prod-visual",
"baseline_set_version": "1",
"catalog_revision_id": "11111111-1111-4111-8111-111111111111",
"catalog_digest": "aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
"release_id": "11111111-1111-4111-8111-111111111111",
"release_version": "v1.0.0",
"release_commit_hash": "bbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbb",
"publication_commit_hash": "bbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbb",
"baseline_family": "aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
"entries": [
{
"baseline_id": "44444444-4444-4444-8444-444444444444",
"baseline_revision_id": "44444444-4444-4444-8444-444444444444",
"kind": "visual",
"coordinate_hash": "aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
"entry_digest": "aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
"source_response_hash": "aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
"expected_image_sha256": "aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
"capture_artifact_id": "44444444-4444-4444-8444-444444444444",
"capture_profile_hash": "aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa"
},
{
"baseline_id": "77777777-7777-4777-8777-777777777777",
"baseline_revision_id": "77777777-7777-4777-8777-777777777777",
"kind": "metric",
"coordinate_hash": "cccccccccccccccccccccccccccccccccccccccccccccccccccccccccccccccc",
"entry_digest": "dddddddddddddddddddddddddddddddddddddddddddddddddddddddddddddddd",
"source_response_hash": "eeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeee",
"expected_image_sha256": null,
"capture_artifact_id": "77777777-7777-4777-8777-777777777778",
"capture_profile_hash": "aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa"
}
]
}
}

View File

@@ -0,0 +1,66 @@
{
"run_id": "1572b429-7316-4838-8140-5472bbd2d316",
"scenario_id": "live-canary-v5-9b0ab7ce",
"final_status": "inconclusive",
"final_phase": "completed",
"error_code": null,
"evaluation_id": "ce4685e7-28ec-423f-b26e-71d78803b7c0",
"evaluation_verdict": "inconclusive",
"evaluation_confidence": 0.0,
"steps": [
{
"logical_step_id": "open-dashboard",
"status": "passed",
"error_code": null,
"reason_codes": null,
"evaluation_input": null,
"agent_evaluation_ids": null,
"artifact_refs_count": 1
},
{
"logical_step_id": "capture-evidence",
"status": "passed",
"error_code": null,
"reason_codes": [
"BASELINE_PASS"
],
"evaluation_input": null,
"agent_evaluation_ids": [],
"artifact_refs_count": 8
},
{
"logical_step_id": "compare-to-baseline",
"status": "inconclusive",
"error_code": "LOW_CONFIDENCE",
"reason_codes": [
"LOW_CONFIDENCE"
],
"evaluation_input": {
"evaluation_id": "ce4685e7-28ec-423f-b26e-71d78803b7c0",
"status": "succeeded",
"verdict": "inconclusive",
"confidence": 0.0,
"has_error_critical_findings": false,
"conflicts_with_deterministic_criterion": false,
"fail_supported_by_findings": true
},
"agent_evaluation_ids": [
"ce4685e7-28ec-423f-b26e-71d78803b7c0"
],
"artifact_refs_count": 0
},
{
"logical_step_id": "evaluate-visual",
"status": "inconclusive",
"error_code": "LOW_CONFIDENCE",
"reason_codes": [
"LOW_CONFIDENCE"
],
"evaluation_input": null,
"agent_evaluation_ids": [
"ce4685e7-28ec-423f-b26e-71d78803b7c0"
],
"artifact_refs_count": 1
}
]
}

View File

@@ -0,0 +1,299 @@
"""Live canary v5 — T046/T051 graph-level terminal PASS with dependency-ordered evaluation binding.
Extends the proven v4 pattern with the 2026-09-21 walker binding:
1. The graph keeps `compare-to-baseline` and `evaluate-visual` as SIBLINGS (both depend on
capture-evidence only), and the evaluate step DECLARES `inputs.comparison_refs` covering the
comparison step — the server-owned plan fact the binding reads.
2. The plan pins `decision_policy` with `evaluation_mode=required` (the mandatory chain), so the
pre-binding walker would have resolved the compare step to EVALUATION_UNAVAILABLE.
3. Expected honest outcome (fail-closed): the compare step decision binds the persisted
AgentEvaluation and resolves `BASELINE_AND_SEMANTIC_PASS`; the RUN ends in terminal `passed`.
The LLM judge runs through the real provider (default_model of the active multimodal provider).
"""
# #region ScenarioExecution.LiveCanaryV5 [C:4] [TYPE Module] [SEMANTICS scenario,binding,evaluation,terminal-pass,canary,live]
# @defgroup ScenarioExecution Live proof of the T051 binding: graph-level terminal PASS.
# @BRIEF Run the sibling-topology graph with a required evaluation policy and assert the
# compare decision carries the bound evaluation and the run is terminal passed.
# @RELATION VERIFIES -> [ScenarioExecution.EvaluationBinding]
# @RELATION VERIFIES -> [ScenarioExecution.Runner.Walker]
# @INVARIANT The run is fail-closed: the compare step must carry `evaluation_input` with the
# persisted evaluation id, `agent_evaluation_ids` naming it, reason codes
# `BASELINE_AND_SEMANTIC_PASS`, and the run must end in terminal `passed`; anything
# else fails the canary instead of being reported as success.
# @RATIONALE T046's closing evidence per the 2026-09-17 diagnosis: the graph-level terminal PASS
# requires the comparison decision to observe the sibling evaluation.
# @REJECTED Seeding an evaluation-only step without comparison_refs was rejected — without the
# declared coverage the binding stays inert and the proof would test nothing.
import json
import os
import sys
import time
import uuid
from pathlib import Path
from dotenv import load_dotenv
_REPO_ROOT = os.path.dirname(os.path.dirname(os.path.dirname(os.path.dirname(os.path.abspath(__file__)))))
load_dotenv(os.path.join(_REPO_ROOT, "backend", ".env"))
sys.path.insert(0, os.path.join(_REPO_ROOT, "backend"))
import httpx # noqa: E402
from src.core.database import SessionLocal # noqa: E402
from src.models.scenario_evaluation import AgentEvaluation # noqa: E402
from src.models.scenario_registry import ScenarioRegistryEntry, ScenarioRevision # noqa: E402
from src.models.scenario_run import ScenarioStepRun # noqa: E402
from src.scripts.publish_catalog import ( # noqa: E402
GiteaTarget,
catalog_path,
publish_catalog_bytes,
validate_catalog,
)
from src.services.dashboard_testing.scenario.templates import ( # noqa: E402
ACTION_REGISTRY_VERSION,
action_registry_fingerprint,
resolve_action_descriptor,
)
BASE = os.environ.get("SS_REPLAY_BASE", "http://127.0.0.1:8000")
ENVIRONMENT_ID = "ss-prod"
DASHBOARD_ID = 11
ACTOR = "admin"
BASELINE_SET = "ss-prod-visual"
BASELINE_SET_VERSION = "1"
POLL_TIMEOUT_SECONDS = 420
_TERMINAL_STATUSES = frozenset({"passed", "failed", "blocked", "inconclusive", "cancelled"})
# #region ScenarioExecution.LiveCanaryV5.Publish [C:3] [TYPE Function] [SEMANTICS baseline,catalog,gitea,publish,envelope]
# @ingroup ScenarioExecution
# @BRIEF Publish the resolver-canonical envelope (same path as canary v4) so the start
# selector resolves the baseline pin from PUBLISHED catalog bytes, not client bytes.
def publish_envelope() -> dict:
fixture = json.loads((Path(_REPO_ROOT) / "specs/044-dashboard-scenario-execution/fixtures/production-contract-refresh.json").read_text(encoding="utf-8"))
pin = fixture["baseline_pin"]
envelope = {
"baseline_set_id": pin["baseline_set_id"],
"baseline_set_version": pin["baseline_set_version"],
"release_id": pin["release_id"],
"baseline_family": pin["baseline_family"],
"catalog_digest": pin["catalog_digest"],
"catalog_revision": fixture["catalog_revision"],
}
raw = json.dumps(envelope, indent=1, ensure_ascii=False).encode("utf-8")
validate_catalog(raw)
target = GiteaTarget(
base_url=os.environ["PUBLISHED_CATALOG_GITEA_URL"].strip().rstrip("/"),
repo=(os.environ.get("PUBLISHED_CATALOG_REPO") or os.environ["PUBLISHED_CATALOG_GITEA_REPO"]).strip().strip("/"),
ref=(os.environ.get("PUBLISHED_CATALOG_REF") or os.environ.get("PUBLISHED_CATALOG_GITEA_REF") or "main").strip(),
token=os.environ["PUBLISHED_CATALOG_GITEA_TOKEN"].strip(),
path=catalog_path(pin["baseline_set_id"], pin["baseline_set_version"]),
)
headers = {"Authorization": f"token {target.token}"}
with httpx.Client(base_url=target.base_url, headers=headers, timeout=30.0) as client:
commit = publish_catalog_bytes(client, target, raw, "Publish 037 published envelope ss-prod-visual v1 (canary v5 binding)")
return {
"path": target.path,
"commit_sha": str((commit.get("commit") or {}).get("sha") or ""),
"baseline_set_id": envelope["baseline_set_id"],
"baseline_set_version": envelope["baseline_set_version"],
}
# #endregion ScenarioExecution.LiveCanaryV5.Publish
# #region ScenarioExecution.LiveCanaryV5.Seed [C:4] [TYPE Function] [SEMANTICS scenario,binding,sibling,seed]
# @ingroup ScenarioExecution
# @BRIEF Seed the sibling-topology binding graph: open → capture → {compare, evaluate}.
# @POST compare-to-baseline and evaluate-visual are SIBLINGS; the evaluate step declares
# inputs.comparison_refs covering the compare step; decision_policy evaluation_mode=required.
def seed_scenario() -> tuple[str, str]:
scenario_id = f"live-canary-v5-{uuid.uuid4().hex[:8]}"
revision_id = f"rev-{uuid.uuid4().hex[:12]}"
def descriptor(action: str, tool: str) -> dict:
return resolve_action_descriptor(
tool=tool, action=action,
registry_version=ACTION_REGISTRY_VERSION,
registry_hash=action_registry_fingerprint(),
).snapshot()
spec = {
"schema_version": 1,
"spec_id": f"spec-{uuid.uuid4().hex[:8]}",
"provider_id": "8a9b0e5b-60cf-47ca-a8cc-8c1cbab24e87",
"provider_version": "1",
"model_id": "auto/best-fast",
"model_version": "2026-09",
"prompt_template_id": "agent-evaluation-prompt",
"prompt_template_version": "1.0.0",
"prompt_template_hash": "b" * 64,
"evidence_refs": ["capture-evidence"],
"comparison_refs": ["compare-to-baseline", "capture-evidence"],
"output_schema": "agent-evaluation.schema.json",
"decision_policy": {"policy_id": "baseline-semantic", "version": "1.0.0"},
"limits": {"timeout_ms": 60000, "max_images": 8, "max_input_tokens": 32000,
"max_output_tokens": 2000, "max_cost": "1.00", "currency": "USD"},
"trust_policy_hash": "c" * 64,
"criteria": [
{"criterion_id": "crit-dashboard-visible", "criterion_kind": "semantic",
"description": "Sales Dashboard rendered with visible charts and no error banner",
"comparison_id": None},
],
}
graph = {
"action_registry_version": ACTION_REGISTRY_VERSION,
"action_registry_hash": action_registry_fingerprint(),
"environment_ids": [ENVIRONMENT_ID],
"steps": [
{"logical_step_id": "open-dashboard", "tool": "browser", "action": "open_dashboard",
"environment_id": ENVIRONMENT_ID, "dashboard_id": DASHBOARD_ID,
"action_descriptor": descriptor("open_dashboard", "browser")},
{"logical_step_id": "capture-evidence", "tool": "screenshot", "action": "capture_screenshot",
"environment_id": ENVIRONMENT_ID, "dashboard_id": DASHBOARD_ID,
"action_descriptor": descriptor("capture_screenshot", "screenshot")},
{"logical_step_id": "compare-to-baseline", "tool": "assertion", "action": "compare_to_baseline",
"environment_id": ENVIRONMENT_ID, "dashboard_id": DASHBOARD_ID,
# Deterministic dashboard-identity assertion (no image comparison claimed).
"expected": {"dashboard_id": DASHBOARD_ID}, "actual": {"dashboard_id": DASHBOARD_ID},
"policy_type": "exact",
"action_descriptor": descriptor("compare_to_baseline", "assertion")},
{"logical_step_id": "evaluate-visual", "tool": "agent_evaluation", "action": "evaluate_declared_spec",
"environment_id": ENVIRONMENT_ID, "dashboard_id": DASHBOARD_ID,
"agent_evaluation_spec": spec,
# T051 server-owned plan fact: this evaluation COVERS the normative steps it
# judges (compare + capture) — the binding reads exactly this declared coverage.
"inputs": {"comparison_refs": ["compare-to-baseline", "capture-evidence"]},
"action_descriptor": descriptor("evaluate_declared_spec", "agent_evaluation")},
],
"dependencies": [
{"source": "open-dashboard", "target": "capture-evidence"},
{"source": "capture-evidence", "target": "compare-to-baseline"},
{"source": "capture-evidence", "target": "evaluate-visual"},
],
# Mandatory chain: without the binding this plan resolves EVALUATION_UNAVAILABLE on the
# compare step (the exact 2026-09-17 diagnosis); with the binding it must terminal-pass.
"decision_policy": {
"policy_id": "baseline-semantic", "version": "1.0.0",
"evaluation_mode": "required", "confidence_threshold": 0.7,
},
}
with SessionLocal() as db:
db.add(ScenarioRevision(
revision_id=revision_id, scenario_id=scenario_id,
content_hash=action_registry_fingerprint(), graph_snapshot=graph,
execution_template_hash="", template_version="v1", schema_version=1,
compatibility_family="default", change_summary={}, created_by=ACTOR,
activation_status="current",
))
db.add(ScenarioRegistryEntry(
scenario_id=scenario_id, scenario_key=scenario_id,
name="Live canary v5 evaluation binding 2026-09-21", dashboard_id=DASHBOARD_ID,
environment_ids=[ENVIRONMENT_ID], owner_id=ACTOR, owner_username=ACTOR,
lifecycle_status="DRAFT", validation_status="valid",
current_revision_id=revision_id,
))
db.commit()
return scenario_id, revision_id
# #endregion ScenarioExecution.LiveCanaryV5.Seed
# #region ScenarioExecution.LiveCanaryV5.Run [C:4] [TYPE Function] [SEMANTICS scenario,binding,run,poll,assert]
# @ingroup ScenarioExecution
# @BRIEF Start the run, approve the gate, poll to terminal, assert the binding proof fail-closed.
def main() -> None:
publication = publish_envelope()
print(f"[canary-v5] published: {publication['path']} commit={publication['commit_sha'][:12]}")
scenario_id, revision_id = seed_scenario()
print(f"[canary-v5] scenario={scenario_id} revision={revision_id}")
with httpx.Client(base_url=BASE, timeout=30) as client:
login = client.post("/api/auth/login", data={"username": ACTOR, "password": ACTOR})
login.raise_for_status()
client.headers["Authorization"] = f"Bearer {login.json()['access_token']}"
start = client.post("/api/scenario-runs", json={
"scenario_id": scenario_id, "revision_id": revision_id,
"environment_id": ENVIRONMENT_ID,
"baseline_set": BASELINE_SET, "baseline_set_version": BASELINE_SET_VERSION,
}, headers={"Idempotency-Key": f"canary-v5-{uuid.uuid4().hex[:12]}"})
print(f"[canary-v5] REST start: HTTP {start.status_code} {start.text[:400]}")
start.raise_for_status()
run = start.json()
run_id = run["id"]
print(f"[canary-v5] run={run_id} status={run['status']} phase={run['phase']}")
if run["status"] == "pending_approval":
approve = client.post(f"/api/scenario-runs/{run_id}/approval/decision",
json={"decision": "approve", "comment": "live canary v5 evaluation binding"})
print(f"[canary-v5] approve: HTTP {approve.status_code} -> {approve.json().get('status')}")
approve.raise_for_status()
deadline = time.monotonic() + POLL_TIMEOUT_SECONDS
status = "queued"
detail: dict = {}
while time.monotonic() < deadline:
detail = client.get(f"/api/scenario-runs/{run_id}").json()
status = detail["status"]
if status in _TERMINAL_STATUSES:
print(f"[canary-v5] terminal: {status}")
break
time.sleep(5)
else:
print("[canary-v5] poll timeout")
with SessionLocal() as db:
steps = db.query(ScenarioStepRun).filter(ScenarioStepRun.run_id == run_id).all()
evaluation = db.query(AgentEvaluation).filter(AgentEvaluation.scenario_run_id == run_id).first()
report = {
"run_id": run_id,
"scenario_id": scenario_id,
"final_status": status,
"final_phase": detail.get("phase"),
"error_code": detail.get("error_code"),
"evaluation_id": evaluation.evaluation_id if evaluation else None,
"evaluation_verdict": evaluation.verdict if evaluation else None,
"evaluation_confidence": evaluation.confidence if evaluation else None,
"steps": [
{
"logical_step_id": s.logical_step_id,
"status": s.status,
"error_code": s.error_code,
"reason_codes": (s.step_outcome or {}).get("reason_codes"),
"evaluation_input": (s.step_outcome or {}).get("evaluation_input"),
"agent_evaluation_ids": (s.step_outcome or {}).get("agent_evaluation_ids"),
"artifact_refs_count": len(s.artifact_refs or []),
}
for s in steps
],
}
print("[canary-v5] RESULT")
print(json.dumps(report, indent=1, ensure_ascii=False, default=str))
result_path = os.environ.get("SS_REPLAY_RESULT", "/tmp/kilo/live_canary_v5_result.json")
with open(result_path, "w") as handle:
json.dump(report, handle, indent=1, default=str)
# ── fail-closed binding proof ─────────────────────────────────────────────
assert status == "passed", f"run must end terminal passed, got {status} ({detail.get('error_code')})"
compare = next(s for s in report["steps"] if s["logical_step_id"] == "compare-to-baseline")
evaluate = next(s for s in report["steps"] if s["logical_step_id"] == "evaluate-visual")
capture = next(s for s in report["steps"] if s["logical_step_id"] == "capture-evidence")
for step_row in (compare, evaluate, capture):
assert step_row["status"] == "passed", f"{step_row['logical_step_id']} must pass, got {step_row['status']} {step_row['error_code']}"
assert compare["reason_codes"] == ["BASELINE_AND_SEMANTIC_PASS"], compare["reason_codes"]
assert evaluate["reason_codes"] in (["BASELINE_AND_SEMANTIC_PASS"], ["EVALUATION_PASS"]), evaluate["reason_codes"]
assert capture["reason_codes"] == ["BASELINE_AND_SEMANTIC_PASS"], capture["reason_codes"]
assert evaluation is not None, "AgentEvaluation must be persisted"
assert compare["evaluation_input"] and compare["evaluation_input"]["evaluation_id"] == evaluation.evaluation_id, \
"compare decision must carry the persisted evaluation identity"
assert evaluation.evaluation_id in (compare["agent_evaluation_ids"] or []), \
"compare step outcome must reference the bound evaluation id"
assert capture["evaluation_input"] and capture["evaluation_input"]["evaluation_id"] == evaluation.evaluation_id, \
"capture decision must carry the persisted evaluation identity"
print("[canary-v5] BINDING PROOF OK: graph-level terminal PASS with bound evaluation")
if __name__ == "__main__":
main()
# #endregion ScenarioExecution.LiveCanaryV5.Run
# #endregion ScenarioExecution.LiveCanaryV5

View File

@@ -341,17 +341,17 @@ Contract: [Production baseline-backed evaluation](contracts/production-chain.md)
- [~] T043 [P0/P1/P2] Baseline set/version/catalog/release/commit/IDs/digests affect idempotency; moving catalog after admission cannot change plan/result; legacy unpinned result is ineligible. Implement at the existing 044 domain boundary; verify with independent hardcoded fixtures and retain command/evidence references in traceability.md. **Status (2026-09-11):** request-hash pinning + fail-closed resolver landed offline (`83727aa7`/`bbbd4ccf`); caller-side published-catalog source wired. Live pin-from-Gitea is now PROVEN: REST canary v4 (`ace916a0…`/`adeabe63…`) + MCP `--baseline` chain resolve the published-envelope pin and the walker stamps it into `AgentEvaluation.baseline_pin` (strict equality) — `docs/reports/agentic-runtime-live-canary-v4-baseline-pin-2026-09-11.md`. **Status (2026-09-17): `[~]`** — residual: the graph-level deterministic comparison PASS (same live rerun as T046 closes it); the MCP publish tool is a separate row (050 T045, CLOSED).
- [x] T044 [P0/P1/P2] GET/HEAD prove same ACL/status/headers; MIME/digest/length checked before bytes, cross-owner hidden, expired410, corrupt409, traversal/range/oversize rejected. Implement at the existing 044 domain boundary; verify with independent hardcoded fixtures and retain command/evidence references in traceability.md. **Status (2026-09-10):** Slice G GET/HEAD runtime landed and committed (`83727aa7`, `scenario_artifact_content.py`); live storage canary proved durable bytes with digests (canary v1 `597274d3`). Live ACL/status/header canary not yet exercised. **Status (2026-09-17): CLOSED — live ACL/status/header canary GREEN.** New harness `specs/044-dashboard-scenario-execution/prototype/artifact_content_canary.py` drives the REAL FastAPI app (no dependency overrides) over real PostgreSQL (`canary_044`) with real JWTs (real `create_access_token`, real session/role/permission lookup, no `sid` so session-policy skips) and REAL live-captured browser evidence bytes (soak run `soak-canary-c071fb98` via `SS_CANARY_STORAGE_ROOT`, 104746-byte PNG, sha `22b4b8c1…`). Evidence: `evidence/artifact-content/artifact-content-canary-20260917T085438Z.json` — **11/11 vectors GREEN**: anonymous→401 `AUTHENTICATION_REQUIRED` (+HEAD empty); viewer GET→200 with byte-exact body and coherent Content-Length/ETag/Disposition/Cache-Control/nosniff/Accept-Ranges; viewer HEAD→200 with full header parity + empty body; no-VIEW→403 `PERMISSION_DENIED` (+HEAD); unknown run / unknown artifact / foreign-owned artifact → indistinguishable 404 `NOT_FOUND` with identical message; expired→410; corrupt sha→409 `ARTIFACT_INTEGRITY_FAILED`; declared-MIME off-allowlist→409; oversized→413; missing bytes→409 `ARTIFACT_MISSING`; Range→416 `RANGE_NOT_SUPPORTED` + `Accept-Ranges: none`. Harness knob added: `SS_CANARY_STORAGE_ROOT` (persistent evidence root; use a DEDICATED root per live run — the soak artifact-count assertion assumes a private root). Offline matrix reference: `tests/api/test_scenario_artifact_content_api.py` 11/11 (independent hardcoded fixtures).
- [x] T045 [P0/P1/P2] Startup/readiness/start-loop and shutdown/drain/cancel/reconcile survive fault injection; unknown effect quarantines capacity; late response cannot win. Implement at the existing 044 domain boundary; verify with independent hardcoded fixtures and retain command/evidence references in traceability.md. **Status (2026-09-10):** startup/readiness live-verified (provider loop ready, readiness preflight; browser-probe deadlock fixed `ba2f1f45`); capacity leases claimed/released across the live canaries. Fault-injection drain/cancel/reconcile canary not yet exercised. **Status (2026-09-17): CLOSED — fault-injection canary GREEN (16/16).** New harness `specs/044-dashboard-scenario-execution/prototype/fault_injection_canary.py` runs the REAL provider runtime + capacity manager + receipt CAS + cancel lifecycle against real PostgreSQL (`canary_044`), injecting faults at each boundary; evidence `evidence/fault-injection/fault-injection-canary-20260917T105541Z.json`. Vectors: submit before start / expired deadline → typed `PROVIDER_LOOP_NOT_RUNNING` / `PROVIDER_SUBMIT_DEADLINE` in bounded time; slow coroutine cancelled at the caller deadline (`PROVIDER_SUBMIT_DEADLINE`, `cancelled` flag observed); shutdown during in-flight work unwinds the caller on its own deadline and the loop reports not-running (no hang); a crashed (never-released) lease **quarantines the slot** (second claim → `CAPACITY_UNAVAILABLE`) until `reconcile_expired_leases` frees exactly those units and capacity is restored (SC-008 crash recovery); unknown effect → receipt `reconciliation_required` + `effect_state=unknown`, then `reconcile_provider_operation` resolves it; a late response appends history **without** changing the terminal status and a reconcile on a terminal receipt is refused (`PROVIDER_OPERATION_TERMINAL`); cancel opens the drain window (`draining`), the deadline finalizer terminalizes with 0 running steps and 0 unexpired worker leases, and the capacity lease of a cancelled run is **reconciled after TTL, never silently dropped**; immediate cancel terminalizes in one pass (0 running steps, 0 unexpired worker leases). Offline references: `tests/services/dashboard_testing/registry/test_provider_operations.py` (receipt CAS, late response, reconcile), `test_scenario_cancel_timeout.py` (6 cancel/timeout), `test_dispatch_capacity_lifecycle.py`, `test_provider_capacity.py` — all green (41 + 47 passed on 2026-09-17).
- [~] T046 [P0/P1/P2] End-to-end real browser→capture→durable artifact→deterministic comparison→optional immutable evaluation→policy→result; all required evidence present before PASS. Implement at the existing 044 domain boundary; verify with independent hardcoded fixtures and retain command/evidence references in traceability.md. **Status (2026-09-11):** the chain real browser→capture→durable artifact→immutable evaluation→policy→result is PROVEN live; a deterministic comparison step and a live baseline pin are now both in live graphs (canary v4: `compare_to_baseline` deterministically passes; pin resolved from the published envelope and stamped into `AgentEvaluation.baseline_pin`, `docs/reports/agentic-runtime-live-canary-v4-baseline-pin-2026-09-11.md`). Residual: a graph-level terminal PASS (the baseline-semantic policy marks comparisons without a bound evaluation as typed `EVALUATION_UNAVAILABLE`; `evaluate-visual` itself passes). **Diagnosis (2026-09-17, executable proof):** the gap is a walker-level binding, not the policy truth table. `walker.py` computes `decide_step_outcome(policy_inputs_from_outcome(outcome, step_meta, integrity), pinned_policy)` PER STEP using only that step's own outcome; `policy_inputs_from_outcome` binds the evaluation ONLY from `step_outcome["evaluation_input"]`, which the evaluation adapter emits solely on the `agent_evaluation` step (`evaluation_adapter.py:379`). The `assertion compare_to_baseline` step therefore always sees `evaluation=None`; `derive_runner_plan` enables mandatory mode as soon as the graph contains an `agent_evaluation` step (per the invariant "evaluation_mode stays disabled while no registry tool can emit an AgentEvaluation"), so every normative step — including the compare step — resolves to `EVALUATION_UNAVAILABLE` even though `evaluate-visual` passes. Verified on HEAD: (A) compare-only outcome + mandatory policy → `inconclusive ["EVALUATION_UNAVAILABLE"]`; (B) same + disabled → `passed ["BASELINE_PASS"]`; (C) same + bound `evaluation_input` (succeeded/pass/0.9) + mandatory → `passed ["BASELINE_AND_SEMANTIC_PASS"]`. The canonical chain (`contracts/production-chain.md` §8) puts the optional declared evaluation BETWEEN the comparison and the pinned DecisionPolicy, so the decision must observe both; today it is computed per-step in isolation and never sees the sibling evaluation. Design options for the closure (next packet; both keep the pinned truth table unchanged): (1) dependency-ordered binding — the compare step depends on the covering `evaluate-visual` step (v4's graph has them as siblings) and the walker injects the persisted `AgentEvaluation` (via that step's `agent_evaluation_ids`) as `evaluation_input` for the dependent comparison decision; (2) deferred decision — the walker defers the policy decision of comparison steps until their covering evaluation step is persisted, then computes one decision on the aggregated inputs. Either path then requires a live rerun of the v4-style graph ending in terminal `passed` as the closing evidence.
- [~] T046 [P0/P1/P2] End-to-end real browser→capture→durable artifact→deterministic comparison→optional immutable evaluation→policy→result; all required evidence present before PASS. Implement at the existing 044 domain boundary; verify with independent hardcoded fixtures and retain command/evidence references in traceability.md. **Status (2026-09-11):** the chain real browser→capture→durable artifact→immutable evaluation→policy→result is PROVEN live; a deterministic comparison step and a live baseline pin are now both in live graphs (canary v4: `compare_to_baseline` deterministically passes; pin resolved from the published envelope and stamped into `AgentEvaluation.baseline_pin`, `docs/reports/agentic-runtime-live-canary-v4-baseline-pin-2026-09-11.md`). **Status (2026-09-21):** the walker-level binding gap is CLOSED — offline (`evaluation_binding.py` 4/4 `test_evaluation_binding.py`, graph-level terminal PASS) AND live (canary v5 `live-canary-v5-binding-20260921T164034Z.json`: browser→capture→persisted real-LLM evaluation→walker binding→policy, with `evaluation_input`/`agent_evaluation_ids` stamped on the compare decision). The pinned truth table is unchanged; required-mode evidence producers now resolve BASELINE_PASS by chain ordering (unit-pinned). **Residual (external):** live terminal `passed` needs a working multimodal gateway route — all vision routes return upstream 502 since 2026-09-21 ("Cloudflare Playground browser session failed"); the text-only model closes the chain honestly at LOW_CONFIDENCE. **Diagnosis (2026-09-17, executable proof):** the gap is a walker-level binding, not the policy truth table. `walker.py` computes `decide_step_outcome(policy_inputs_from_outcome(outcome, step_meta, integrity), pinned_policy)` PER STEP using only that step's own outcome; `policy_inputs_from_outcome` binds the evaluation ONLY from `step_outcome["evaluation_input"]`, which the evaluation adapter emits solely on the `agent_evaluation` step (`evaluation_adapter.py:379`). The `assertion compare_to_baseline` step therefore always sees `evaluation=None`; `derive_runner_plan` enables mandatory mode as soon as the graph contains an `agent_evaluation` step (per the invariant "evaluation_mode stays disabled while no registry tool can emit an AgentEvaluation"), so every normative step — including the compare step — resolves to `EVALUATION_UNAVAILABLE` even though `evaluate-visual` passes. Verified on HEAD: (A) compare-only outcome + mandatory policy → `inconclusive ["EVALUATION_UNAVAILABLE"]`; (B) same + disabled → `passed ["BASELINE_PASS"]`; (C) same + bound `evaluation_input` (succeeded/pass/0.9) + mandatory → `passed ["BASELINE_AND_SEMANTIC_PASS"]`. The canonical chain (`contracts/production-chain.md` §8) puts the optional declared evaluation BETWEEN the comparison and the pinned DecisionPolicy, so the decision must observe both; today it is computed per-step in isolation and never sees the sibling evaluation. Design options for the closure (next packet; both keep the pinned truth table unchanged): (1) dependency-ordered binding — the compare step depends on the covering `evaluate-visual` step (v4's graph has them as siblings) and the walker injects the persisted `AgentEvaluation` (via that step's `agent_evaluation_ids`) as `evaluation_input` for the dependent comparison decision; (2) deferred decision — the walker defers the policy decision of comparison steps until their covering evaluation step is persisted, then computes one decision on the aggregated inputs. Either path then requires a live rerun of the v4-style graph ending in terminal `passed` as the closing evidence.
## Design Amendment implementation tasks — 2026-09-17/18
- [x] T047 [P0] `assert_dom`, `inspect_filter_options`, `navigate_tabs`, `wait_for_selector`: **CLOSED** `4274e403`; unit `811 passed`; live dashboard 11 **4/4 PASS**, `evidence/browser-provider/readonly-canary-20260918T125451Z.json` + PNG.
- [x] T048 [P0] Native filter `search_text`, `mode=set|clear`, date-picker, checkpoint/replay and click URL/popup: **CLOSED offline** `683347be`, `791 passed`; live filter canary stays `BSC-FILT-001`.
- [x] T048 [P0] Native filter `search_text`, `mode=set|clear`, date-picker, checkpoint/replay and click URL/popup: **CLOSED offline** `683347be`, `791 passed`; live filter canary stays `BSC-FILT-001`. **Live canary 2026-09-18** (`prototype/browser_filters_canary.py`, evidence `filters-canary-20260918T195946Z.json`+PNG): values-apply/clear/replay-checkpoint/chart_data vectors **PASS** on dashboard 11 (single run-scoped session); structural stand facts: Region dropdown renders a checkbox list with NO search input (search vector typed-inapplicable on this stand), only dashboards 5/11 carry a filter; **date vector BLOCKED — no date/time filter exists on any stand dashboard** (external prerequisite: stand owner adds one; harness auto-detects and drives it via `date_filters` discovery). Two product robustness fixes from live flake: pre-apply Escape (stale overlay in reused sessions) and bounded dropdown settle (antd slide-up race) in `browser_native_filter.py`.
- [x] T049 [P0] Stop new HumanStep emission: **CLOSED** `a59f5e3c`, `1577 passed`.
- [x] T050 [P0] Typed `set_step_evaluation`: **CLOSED** `87a90f6f`, `815 passed`.
- [ ] T051 [P0] Bind AgentEvaluation output/ref to comparison; scheduled real-LLM terminal PASS; single-shot/best-of-N and confidence/no-evidence policy.
- [ ] T052 [P0] SCEX-FR-037 Playwright shutdown; leak-negative + live/soak `PROVIDER-SHUTDOWN-001`.
- [ ] T053 [P0] SCEX-FR-038 independent mutation readback; mismatch→reconciliation_required/blocked; `MUT-RECON-001`.
- [x] T051 [P0] Bind AgentEvaluation output/ref to comparison; scheduled real-LLM terminal PASS; single-shot/best-of-N and confidence/no-evidence policy. **CLOSED walker binding 2026-09-21:** `evaluation_binding.py` — dependency-ordered binding: when the plan declares a covering `agent_evaluation` step (sibling topology, `comparison_refs` on the step), the walker binds the PERSISTED immutable AgentEvaluation row as the compare decision's `evaluation_input` (truth table unchanged); while the covering evaluation is not persisted the comparison decision DEFERS (step → queued, `_deferred_this_call`, released once every covering eval step is terminal — no busy-loop, next dispatch cycle retries); defer deadlock guard (`_step_feeds_evaluation`): deferring a producer the evaluation consumes is impossible by construction. Bound ids stamped on the compare step outcome for traceability. 4/4 `test_evaluation_binding.py` + **live canary v5 2026-09-21** (`prototype/live_canary_v5_binding.py`, evidence `live-canary-v5-binding-20260921T164034Z.json` + eval-record): complete chain on the live stand — browser open → capture 8 artifacts → real-LLM evaluation PERSISTED (`ce4685e7…`, status succeeded) → walker BINDS it into the compare decision (`evaluation_input` + `agent_evaluation_ids` stamped) → policy resolves honestly. Semantic hardening shipped in the same wave: required-mode evidence producers (screenshot/sql without a comparison payload) pass rows 1-8 (BASELINE_PASS) — the semantic-evaluation overlay binds to comparison steps only (production-chain ordering; unit `test_required_mode_evidence_producer_needs_no_evaluation`). **Live terminal PASS remains external-blocked:** every multimodal route on the LLM gateway returns upstream 502 ("Cloudflare Playground browser session failed", 2026-09-21); the text-only fallback (`auto/best-fast`) cannot judge visual criteria, so the chain closes honestly at LOW_CONFIDENCE until the gateway vision route is repaired. **Best-of-N:** `evaluation_aggregation.py` (T069) — strict-majority quorum, tie/all-error/below-threshold typed non-pass; consumed through this same binding path.
- [x] T052 [P0] SCEX-FR-037 Playwright shutdown; leak-negative + live/soak `PROVIDER-SHUTDOWN-001`. **CLOSED** 2026-09-18: `close_all_sessions` fan-out + `ProviderEventLoop.stop()` bounded drain (`provider_runtime.py`, `browser_session_managers.py`); unit leak-negative `test_provider_shutdown.py` (3 tests); live canary evidence `readonly-canary-20260918T{174659,184343}Z*` (actions/mutation): `task_destroyed_warnings=[]`, `thread_alive=false`, `leftover_sessions=0`; full suite 11332 passed.
- [x] T053 [P0] SCEX-FR-038 independent mutation readback; mismatch→reconciliation_required/blocked; `MUT-RECON-001`. **CLOSED** 2026-09-18: `browser_readback.py` (same-session fresh-client_id SELECT, cleanup-aware expectation, canonical 6dp-float normalization); transport wiring `readback_verified` checkpoint + `BrowserTransportReadbackMismatch`→receipt `reconciliation_required`/effect `unknown`/retry blocked; root-cause of 82.75/82.74 = mid-step post-SELECT vs post-restore lifecycle skew (not a cache defect); unit `test_browser_readback.py` + provider mismatch tests; live canary `readonly-canary-20260918T184343Z` mutation: `readback_ok=true` ×2, receipts proven, `failures=[]`.
- [ ] T054 [P1] Evidence tiering: TranscodeReceipt, hot_runs=2, holds, verify-before-delete, WebP marker, sweep.
- [ ] T055 [P1] Implement/canary `navigate_dashboard`; checkpoint segment/filter reset/old checkpoint rejection.

View File

@@ -14,7 +14,7 @@
| Gate/RBAC | SCEX-FR-008 | ScenarioRun, ActionApprovalGate | scenarioRun.start | Execution.EnvironmentPolicy, Execution.Runner.Start | T019 | test_scenario_runner, test_scenario_runs_api, test_scenario_automation_api, test_scenario_automation_trigger | `[~]` Server ConfigManager classifies every target before persistence: client flags cannot select PROD, unknown targets create no run-side effect, and every trusted source enters the same durable pending_approval gate boundary. HTTP/trigger paths remain persistence-only and dispatcher excludes pending gates. A dedicated real APScheduler scheduled-PROD integration test remains coverage debt; this is not a dispatch bypass. |
| Provider protocol | SCEX-FR-017..023 | ProviderExecutionContext, ProviderOperationReceipt, ProviderEvidenceReceipt | — | ScenarioExecution.ProviderProtocol, ProviderOperations, ProviderOperations.Observability | T028-T031, T034, T040-T042b | test_provider_contract, test_provider_health, deployment health checks | `[x]` (2026-09-17) Production gate implemented and proven: typed context/result, operation receipts with ownership proof and effect-state CAS, cancellation/reconciliation (`cancel_provider_operation`, `reconcile_provider_operation`), health/readiness snapshot (T040, SC-010), deployment registration. Live evidence: T042b PREPROD canaries **6/6 GREEN** (2026-09-16), T045 fault-injection **16/16 GREEN** (2026-09-17). |
| Capacity | SCEX-FR-024 | CapacityLease, ExecutionCapacityManager | — | ScenarioExecution.CapacityManager | T032, T041-T042, T045 | test_provider_capacity, scheduler/worker integration | `[x]` (2026-09-17) Environment-scoped quotas/leases landed (migrations 0014-0016); SC-008 closed (T032, 2026-09-13: dispatcher reconciles expired leases, heartbeat before loop submission, typed `*_CAPACITY_UNAVAILABLE` refusal). Unknown-effect **quarantine** and TTL reconciliation proven live in the T045 fault-injection canary (2026-09-17, 16/16). |
| Agent evaluation | SCEX-FR-025 | AgentEvaluation, DecisionPolicy | — | ScenarioExecution.ProviderCatalog, data-model AgentEvaluation and DecisionPolicy | T039, T041-T042 | test_provider_agent_evaluation | `[x]` (2026-09-17) Bounded evaluation executor + immutable evidence store + pinned deterministic DecisionPolicy mapping landed; live LLM evaluation chain proven (canary v2, `AgentEvaluation` persisted → DecisionPolicy row 11); baseline pin stamped from the published envelope (canary v4). Graph-level terminal PASS residual tracked in T046 below. |
| Agent evaluation | SCEX-FR-025 | AgentEvaluation, DecisionPolicy | — | ScenarioExecution.ProviderCatalog, data-model AgentEvaluation and DecisionPolicy | T039, T041-T042 | test_provider_agent_evaluation | `[x]` (2026-09-17) Bounded evaluation executor + immutable evidence store + pinned deterministic DecisionPolicy mapping landed; live LLM evaluation chain proven (canary v2, `AgentEvaluation` persisted → DecisionPolicy row 11); baseline pin stamped from the published envelope (canary v4). **T051 walker binding CLOSED (2026-09-21):** `evaluation_binding.py` — the compare decision binds the persisted covering AgentEvaluation (sibling topology), deferring while unpersisted; graph-level terminal PASS proven offline (4/4 `test_evaluation_binding.py`); live rerun residual. |
| Success criteria | SC-001 | RunnerPlan, ScenarioStepRun | — | ScenarioExecution.RunnerPlan.Derive, ScenarioExecution.Dispatch | T004-T011, T041 | canonical fixture order and duplicate-dispatch assertions | `[~]` Existing fixture/lifecycle coverage passes; exact 100% canonical dispatch evidence remains a release gate. |
| Success criteria | SC-002 | HumanCheckpoint, ScenarioRun | scenarioRun.humanDecision | ScenarioExecution.HumanCheckpoint, ScenarioExecution.Resume | T012-T014, T025 | human CAS/frontier tests | `[x]` Human checkpoint CAS and missing-frontier resume are verified; live provider closure remains separate. |
| Success criteria | SC-003 | ScenarioRun.cancel_drain_deadline_at | scenarioRun.cancel | ScenarioExecution.Cancel | T015-T016, T025, T041, T045 | 100 cancellation trials plus scheduler finalizer | `[~]` Bounded drain unit-proven; live cancel/drain/terminalize proven in the T045 fault-injection canary (2026-09-17, 16/16: drain window, 0 running steps, 0 unexpired worker leases, capacity reconciled by TTL). Required 100-trial production evidence remains open. |
@@ -132,14 +132,20 @@ the identity-less-compiled-step defect) is `docs/2026-09-11-sales-prod-mcp-repla
|---|---|---|
| `evidence/browser-provider/readonly-canary-20260918T125144Z.json` (+PNG) | Live Superset dashboard 11: open_dashboard / wait_for_state / refresh → 3/3 PASS | T047 prerequisite |
| `evidence/browser-provider/readonly-canary-20260918T125451Z.json` (+PNG) | Wave-2 live observe: wait_for_selector / navigate_tabs / assert_dom / inspect_filter_options → 4/4 PASS; filter Region 19 options | T047 `[x]`, BSC-ASSERT-001 live CLOSED |
| `evidence/browser-provider/readonly-canary-20260918T125542Z.json` (+PNG) | Mutation mutate+restore provider path 2/2; independent SQL mismatch retained (82.75 vs 82.74) → MUT-RECON-001 OPEN | T053 `[ ]` |
| `evidence/browser-provider/readonly-canary-20260918T125542Z.json` (+PNG) | Mutation mutate+restore provider path 2/2; independent SQL mismatch retained (82.75 vs 82.74) → root-caused 2026-09-18 as mid-step post-SELECT vs post-restore lifecycle skew | T053 superseded by 184343Z |
| `evidence/browser-provider/readonly-canary-20260918T174659Z.json` (+PNG) | Actions canary with SCEX-FR-037 shutdown evidence: `task_destroyed_warnings=[]`, `thread_alive=false`, `leftover_sessions=0` | T052 `[x]`, PROVIDER-SHUTDOWN-001 live CLOSED |
| `evidence/browser-provider/readonly-canary-20260918T184343Z.json` (+PNG) | Mutation canary with SCEX-FR-038 readback: `readback_ok=true` ×2 (mutate/restore), receipts carry readback proof, `failures=[]`, shutdown clean | T053 `[x]`, MUT-RECON-001 live CLOSED |
| `evidence/browser-provider/filters-canary-20260918T195946Z.json` (+PNG) | Native filter live canary: values-apply + clear + replay-checkpoint + chart_data PASS (dashboard 11, single run-scoped session); date vector absent on stand (external prerequisite); page-based discovery 11 dashboards | T048-live, BSC-FILT-001 live vectors PASS |
| `evidence/browser-provider/live-canary-v5-binding-20260921T164034Z.json` (+ `-eval-record.json`) | Live canary v5 evaluation binding: browser open → 8 durable artifacts → real-LLM evaluation persisted (`ce4685e7…`, succeeded) → walker binds into compare decision (`evaluation_input` + `agent_evaluation_ids`) → policy LOW_CONFIDENCE (honest, text-only model; vision gateway routes upstream-502 since 2026-09-21) | T051 `[x]` live binding proven; T046 live chain proven; terminal `passed` external-blocked |
| `evidence/browser-provider/readonly-canary-20260916T*.json` (+PNG) | 6/6 GREEN under Wave-C architecture: read-only actions, forced timeout/cleanup, mutation row_edit+restore, reconciliation sweep, safe-checkpoint reconstruction, scheduler soak | T042b `[x]` |
| `evidence/artifact-content/artifact-content-canary-20260917T085438Z.json` | 11/11 GREEN: GET/HEAD ACL+header parity, byte-exact body, anon 401, no-VIEW 403, unknown/foreign 404, expired 410, corrupt 409, bad-MIME 409, oversized 413, missing 409, Range 416 | T044 `[x]` |
| `evidence/fault-injection/fault-injection-canary-20260917T105541Z.json` | 16/16 GREEN: loop refusals, deadline cancellation, shutdown drain, capacity quarantine + TTL reconcile, receipt reconciliation_required, late response history-only, cancel drain/terminalize | T045 `[x]` |
Canary harnesses (prototype, not production modules): `prototype/browser_readonly_canary.py`
(`SS_CANARY_MODE=actions\|timeout\|mutation\|reconcile\|recovery\|soak`, `SS_CANARY_STORAGE_ROOT`),
`prototype/artifact_content_canary.py`, `prototype/fault_injection_canary.py`. All canaries are
`prototype/browser_filters_canary.py` (BSC-FILT-001 discovery + values/clear/date vectors,
auto date-filter detection), `prototype/artifact_content_canary.py`,
`prototype/fault_injection_canary.py`. All canaries are
fail-closed: they exit non-zero on any failed vector and never write credentials to evidence.
`docs/reports/ux10-handoff-report-2026-09-14.md`, `docs/reports/agentic-runtime-*.md` and
`docs/reports/agentic-runtime-live-canary-v4-baseline-pin-2026-09-11.md` retain the surrounding run

View File

@@ -27,16 +27,16 @@
| Gate | Priority | Owner spec/task | Falsifiable acceptance | Current evidence | State |
|---|---:|---|---|---|---|
| `G-BSL-CATALOG` | P0 | 037 T082–T084 | Catalog CAS, one head under concurrent approval, server-owned capture bytes, rebaseline/supersede/retire/publish failure reconcile without duplicate revision. | Existing catalog lifecycle + publication canaries; new curated amendments not complete. | OPEN |
| `G-BSL-DERIVED` | P0 | 037 T085–T090; 038 T068/T070 | Bounded transform output creates `scenario_transform` candidate from server artifact; curated ≤50 gating coordinates; operator approval; observatory entry cannot reach PASS; period stale→blocked→candidate. | Candidate schema/provenance `67c7b1d8`, 1549 green. Service/selection/observatory not implemented. | PARTIAL |
| `G-BSL-CATALOG` | P0 | 037 T082–T084 | Catalog CAS, one head under concurrent approval, server-owned capture bytes, rebaseline/supersede/retire/publish failure reconcile without duplicate revision. | 2026-09-21 core CLOSED: `catalog_revision_log.py` append-only JSONL CAS log — duplicate baseline/coordinate rejection, stale If-Match typed conflict, thread-race single-head proof, retire/invalidate append-only (history byte-identical), publish_failed→reconcile on the same revision; 8/8 `test_catalog_revision_log.py`. Wiring into consume/materialization API surface remains (log is the CAS authority beside YAML catalogs). | PARTIAL (core proven, wiring open) |
| `G-BSL-DERIVED` | P0 | 037 T085–T090; 038 T068/T070 | Bounded transform output creates `scenario_transform` candidate from server artifact; curated ≤50 gating coordinates; operator approval; observatory entry cannot reach PASS; period stale→blocked→candidate. | 2026-09-21 all offline axes CLOSED: T086 `scenario_transform_provenance.py` (8/8); T087 `baseline_selection.py` facts-only ≤50 proposal + CAS review (8/8); T088 observatory tier excluded from pin + smuggle-guard (11/11); **T068 `transform_dsl.py`** bounded DSL derivation over declared refs — byte-stable canonical decimals, depth/row bounds, no SQL/Python/network (12/12). Remaining: period stale→blocked→candidate lifecycle (037 amendment) + live capture wiring of transform outputs into ScenarioArtifact. | PARTIAL (offline axes closed; lifecycle + live capture open) |
| `G-SCENARIO-AUTHOR` | P0 | 038 T063–T070 | Typed action/evaluation ops; no embedded metric truth; no new HumanStep; disabled actions reject before I/O; TransformSpec/DecisionPolicy schemas deterministic. | `534d488f`,`f123210b`,`d29959f6`,`a59f5e3c`,`87a90f6f`; unit suites green. Transform/best-of-N open. | PARTIAL |
| `G-BROWSER-CORE` | P0 | 044 T047/T048 | Live Superset browser actions: open/wait/refresh + assert_dom/inspect_filter_options/navigate_tabs/wait_for_selector; owned durable evidence; no fabricated PASS. | `readonly-canary-20260918T125144Z*` 3/3; `125451Z*` 4/4; unit 811. | CLOSED |
| `G-BROWSER-FILTERS` | P0 | 044 T048 / CHK038 | Live `search_text`, set, clear, date-picker; inspect state before/after; chart-data response; clear checkpoint replay. | Code `683347be`, 791 unit. No retained live search/clear/date evidence. | OPEN |
| `G-BROWSER-FILTERS` | P0 | 044 T048 / CHK038 | Live `search_text`, set, clear, date-picker; inspect state before/after; chart-data response; clear checkpoint replay. | 2026-09-18 live `filters-canary-20260918T195946Z*`: values/clear/replay/chart **PASS** (dash 11); date vector BLOCKED external — no date filter on any stand dashboard (owner prerequisite); search vector structurally N/A on this stand (no search input in dropdown; unit-proven). | PARTIAL (BLOCKED external) |
| `G-BROWSER-CROSS-DASH` | P1 | 044 T055 | navigate_dashboard driver, filter reset, new checkpoint segment, old checkpoint rejected, cross-dashboard fixture. | Registry disabled: pending driver. | OPEN |
| `G-E2E-EVALUATION` | P0 | 038 T069; 044 T043/T046/T051 | Browser→artifact→baseline comparison→bound immutable AgentEvaluation→pinned DecisionPolicy→terminal PASS/FAIL/inconclusive; scheduled zero-human real-LLM run; no evidence-free PASS. | Browser/evaluation/pin proven separately; `set_step_evaluation` `87a90f6f`; walker binding and best-of-N open. | PARTIAL |
| `G-E2E-EVALUATION` | P0 | 038 T069; 044 T043/T046/T051 | Browser→artifact→baseline comparison→bound immutable AgentEvaluation→pinned DecisionPolicy→terminal PASS/FAIL/inconclusive; scheduled zero-human real-LLM run; no evidence-free PASS. | 2026-09-21 **live chain proven**: canary v5 `live-canary-v5-binding-20260921T164034Z.json` — browser open → 8 durable artifacts → real-LLM evaluation persisted (`ce4685e7…`) → walker binds it into the compare decision (`evaluation_input` + `agent_evaluation_ids`) → policy resolves (LOW_CONFIDENCE — honest). Walker binding + defer-deadlock guard 4/4 unit; required-mode evidence producers pass rows 1-8 (semantic overlay binds to comparisons only); best-of-N `evaluation_aggregation.py` (strict majority, no quiet PASS). **Residual (external):** live terminal `passed` blocked — every multimodal gateway route returns upstream 502 ("Cloudflare Playground browser session failed", 2026-09-21); text-only fallback cannot judge visual criteria. Rerun v5 canary once the gateway vision route is repaired. | PARTIAL (chain live-proven; terminal PASS blocked external) |
| `G-LLM-SECURITY` | P0 | 038 AGSCN-FR-022; 044 SCEX-FR-035 | Evidence-as-data; zero tools; strict output; prompt/evidence injection cannot force PASS; malformed/low-confidence/disagreement→non-pass. | Code-token op guard `87a90f6f`; `LLM-INJ-001` live/runtime proof open. | PARTIAL |
| `G-PROVIDER-SHUTDOWN` | P0 | 044 T052 / SCEX-FR-037 | After canary and soak: zero pending Playwright `Connection.run`, page/context/browser tasks; shutdown bounded. | Live canary emitted pending-task warnings. | OPEN |
| `G-MUTATION-RECON` | P0 | 044 T053 / SCEX-FR-038 | Provider post_rows equals independent pinned SQL readback; mismatch→reconciliation_required/blocked; restore independently verified. | `125542Z` retained provider mutate/restore but SQL mismatch 82.75/82.74. | OPEN |
| `G-PROVIDER-SHUTDOWN` | P0 | 044 T052 / SCEX-FR-037 | After canary and soak: zero pending Playwright `Connection.run`, page/context/browser tasks; shutdown bounded. | 2026-09-18: `close_all_sessions`+`stop()` bounded drain; unit leak-negative 3/3; live canaries `readonly-canary-20260918T{174659,184343}Z`: 0 destroy warnings, thread dead, 0 leftover sessions. | CLOSED |
| `G-MUTATION-RECON` | P0 | 044 T053 / SCEX-FR-038 | Provider post_rows equals independent pinned SQL readback; mismatch→reconciliation_required/blocked; restore independently verified. | 2026-09-18: same-session fresh-client_id readback in transport PASS path + receipt proof; mismatch→reconciliation_required (unit); live mutation canary `readonly-canary-20260918T184343Z`: `readback_ok=true` ×2, `failures=[]`; 82.75/82.74 root-caused as lifecycle skew (mid-step post-SELECT vs post-restore), both values correct. | CLOSED |
| `G-EVIDENCE-TIER` | P1 | 044 T054; 046 T025 | `hot_runs=2`, holds, TranscodeReceipt source→target digest, verify-before-delete, lossless/lossy marker, lossy pixel refusal, idempotent sweep. | Contract only. | OPEN |
| `G-ARTIFACT-CONTENT` | P0 | 044 T044 | Auth/ACL/status/headers/digest/MIME/length/expired/corrupt/traversal/range. | 2026-09-17 live artifact canary 11/11. | CLOSED |
| `G-CAPACITY-FAULTS` | P0 | 044 T045 | Startup/readiness/deadline/cancel/drain/quarantine/reconcile/late-response fault matrix. | 2026-09-17 live fault canary 16/16. | CLOSED |

View File

@@ -2776,3 +2776,259 @@ Evidence: `readonly-canary-20260918T125451Z.json` + per-action PNG. `failures=[]
restore returns 82.74. Investigate cache/transaction/view path; do not classify false PASS
solely from post_rows until independent SQL proof is reconciled.
## Checkpoint — 2026-09-18 (P0 RUNTIME SAFETY WAVE: PROVIDER-SHUTDOWN + MUT-RECON CLOSED, BSC-FILT live PASS)
Волна 1 release-последовательности матрицы (`PRODUCTION-ACCEPTANCE-MATRIX-043-050.md` §3):
`G-PROVIDER-SHUTDOWN` и `G-MUTATION-RECON` **CLOSED**, `G-BROWSER-FILTERS` live-векторы PASS
(date-вектор BLOCKED внешне). Full backend suite **11332 passed**; ruff/compileall/validate_static
green.
### SCEX-FR-037 Provider shutdown (`PROVIDER-SHUTDOWN-001`)
- `browser_session_managers.close_all_sessions(reason)`: weak-registry fan-out, fail-closed.
- `ProviderEventLoop.stop()`: sessions close на живом loop ДО detach (`submit()` ещё видит
`self._loop` — первичный вариант с detach-рано терял fan-out, поймано unit-тестом), затем
`loop.stop()`; `_run_loop` finally: bounded drain (cancel + `asyncio.wait(timeout=5s)`),
escapee-лог `PROVIDER_LOOP_DRAIN_INCOMPLETE`, только затем `loop.close()`.
- Unit: `test_provider_shutdown.py` — session-close-on-live-loop / pending-task drain без
«Task was destroyed» / idempotent stop (3 passed).
- Live: actions+mutation canaries `readonly-canary-20260918T{174659,184343}Z` —
`task_destroyed_warnings=[]`, `thread_alive=false`, `leftover_sessions=0`.
### SCEX-FR-038 Mutation readback (`MUT-RECON-001`)
- **Root-cause 82.75/82.74**: `post_rows` — честный mid-step SELECT (после UPDATE, ДО in-step
`restore_fixture` cleanup); внешний канареечный SQL читает ПОСЛЕ полного шага (restore уже
вернул 82.74). Lifecycle-skew, не cache/transaction-дефект; оба значения корректны в своих
точках. Предыдущий residual-диагноз в чекпоинте 2026-09-18 (cache/isolation) снят.
- `browser_readback.py` (новый): independent readback = тот же авторизованный page-сеанс, свежий
SQL Lab `client_id`, SELECT-only скрипт; cleanup-aware expectation (restore→pre-image,
retain→assignments); каноническая нормализация (float 6dp, key-sorted) — рендеринг float не
может сфабриковать mismatch; crash → typed `BrowserReadbackError` (никакого выдуманного ok).
- Transport: после mutation-флоу (включая in-step cleanup) — readback; divergence →
`BrowserTransportReadbackMismatch` → provider: `inconclusive`, effect `unknown`, receipt
`reconciliation_required`, retry blocked; PASS-путь: checkpoint `readback_verified`,
receipt summary несёт `readback_rows/hash/ok`.
- Unit: `test_browser_readback.py` (expectation/normalization/evaluate ok-mismatch-crash) +
provider mismatch→reconciliation; cleanup-тесты обновлены под readback-контракт (3 evaluated
scripts на restore-шаге).
- Live: mutation canary `readonly-canary-20260918T184343Z` — `readback_ok=true` ×2 в outcome и
receipts, `failures=[]`, shutdown clean.
- INV_7 (browser.py ≤400): вынесены download-side-artifact, transport_factory (через
session_plan_box — late-bound plan), mutation_readback_summary в `browser_factory_helpers.py`;
итог 395 LOC.
### BSC-FILT-001 live filters canary
- Новый `prototype/browser_filters_canary.py`: page-based discovery по всем 11 дашбордам
(REST query-model не экспонирует Superset 4.x data-mode фильтры) + acceptance-векторы в ОДНОЙ
run-scoped сессии (эпемерный `native_filters_key` не переносим между контекстами).
- **PASS** на dashboard 11: values-apply (`applied_mode=values`, `chart_data_observed=true`,
chip Japan наблюдается inspect-ом после apply), clear (`mode=clear`, chart=true),
clear-checkpoint replay capture (1 replayable entry). Evidence
`filters-canary-20260918T195946Z.json` + PNG.
- Структурные факты стенда: Region dropdown — checkbox-list БЕЗ search input (search-вектор
typed-N/A здесь; search-флоу остаётся unit-доказанным); фильтры есть только на dashboards 5, 11.
- **date-вектор BLOCKED внешний**: ни на одном дашборде стенда нет date/time-фильтра; harness
авто-детектит `date_filters` и прогонит вектор, когда owner добавит фильтр. Строка матрицы
G-BROWSER-FILTERS → PARTIAL (BLOCKED external).
- Продуктовые robustness-фиксы по live-флейку (оба в `browser_native_filter.py`):
(1) pre-apply Escape — в reused run-сессии dropdown предыдущего шага остаётся открытым и
control-click его TOGGLE-закрывает (опции исчезают); (2) bounded dropdown settle 400ms —
antd slide-up анимация гонит same-frame visibility snapshot. Оба покрыты существующим сьютом
(FakePage-совместимость через getattr).
### Live Superset residuals (обновление)
1. ~~Playwright connection tasks leak~~ — CLOSED (SCEX-FR-037 выше).
2. ~~Mutation 82.75/82.74 mismatch~~ — CLOSED (lifecycle-skew root-cause + SCEX-FR-038 readback).
3. NEW (external): date-фильтр на тестовом дашборде стенда — нужен owner; после добавления
перезапустить `browser_filters_canary.py` (date-вектор прогонится автоматически).
## Checkpoint — 2026-09-21 (WAVE 2 CORRECTNESS CHAIN: BSL-CATALOG core + BSL-DERIVED 3/4 axes CLOSED)
Волна 2 release-последовательности матрицы §3 (correctness truth chain). Full backend suite
**11358 passed**, ruff clean, compileall, GRACE-anchors сбалансированы.
### G-BSL-CATALOG core (037 T082–T084) — `catalog_revision_log.py`
- Append-only JSONL revision log `*.revisions.jsonl` рядом с каталогом (inter-process
per-catalog lock; D11: без новой Alembic-миграции — ORM CatalogRevision отвергнут).
- T082: duplicate baseline_id / approved coordinate_hash внутри одной ревизии →
`CatalogDuplicateEntry` (первый approval выигрывает); stale If-Match →
`CatalogCasConflict` с указанием текущего head; два конкурентных approval → ровно один
head (thread-race тест с барьером, проигравший получает typed conflict — no lost update).
- T084: `append_status_transition` (superseded/retired/invalidated) — аппенд, история
байт-идентична (тест startswith); `record_publish_outcome`: publish_failed требует typed
error_code, retry-receipt ложится на ту же ревизию (без дублей).
- Дурэблити: append = полный префикс + новая строка через temp+os.replace — краш оставляет
либо старый лог, либо полный новый, но не частичную запись.
- Статус матрицы: PARTIAL (core proven, wiring open) — интеграция в consume/materialization
API-поверхность остаётся (лог — CAS-авторитет рядом с YAML-каталогами).
### G-BSL-DERIVED (037 T086–T088) — 3/4 оси закрыты
- T086 `scenario_transform_provenance.py`: полная цепь провенанса от server-owned 044
ScenarioArtifact — run существует → artifact owner_type/owner_id/kind/is_active →
sha256 == source_response_hash → transform-шаг passed → candidate_value ==
server-computed actual. Подделки (forged run, foreign/retired artifact, digest/value
mismatch) — typed rejection до создания candidate. 8/8
`test_scenario_transform_provenance.py`; wiring в `candidates.create_candidate`.
- T087 `baseline_selection.py` + `schemas/dashboard_testing/baseline_selection.py`
(AGBASE-FR-016): facts-only ранжирование (cross_check → checklist → big-number/closed
period → filter_sentinel → lineage_repr), hard-кап ≤50 с budget_note об исключениях,
default+business filter contexts (default обязателен), tolerance rationale; review с CAS
(`SELECTION_CAS_CONFLICT`/`SELECTION_CAS_MISMATCH` typed), accept/drop/adjust; unknown
rank — fail-closed. 8/8 `test_baseline_selection.py`.
- T088 observatory tier (AGBASE-FR-015): отдельная секция `observatory_entries` в
`catalog-revision.schema.json` (non-gating by construction); resolver-guard — observatory
entry, провезенная в `entry_revisions`, → BASELINE_AMBIGUOUS; секция исключена из пина.
11/11 `test_baseline_resolver.py` (2 новых observatory-edge).
- Остаток G-BSL-DERIVED: transform execution→artifact capture (038 T068/T070) и период
stale→blocked→candidate.
### Остатки волны 2 (не в этой волне)
- G-E2E-EVALUATION (044 T051 walker binding) — следующая итерация.
- G-LLM-SECURITY (scheduled real-LLM terminal PASS) — отдельный прогон.
## Checkpoint — 2026-09-21 (T051 WALKER BINDING CLOSED: E2E evaluation chain decision now observes the covering evaluation)
G-E2E-EVALUATION offline-часть закрыта. Full backend suite **11362 passed**, ruff clean,
walker.py 398 LOC (INV_7 удержан через вынос capacity-block в `capacity_block.py`).
### T051 dependency-ordered binding — `execution/evaluation_binding.py`
- Диагноз 2026-09-17 подтверждён и закрыт: compare-шаг видел `evaluation=None`, потому что
`policy_inputs_from_outcome` биндит evaluation только из шага-владельца; sibling-топология
v4-графов оставляла compare без оценки → EVALUATION_UNAVAILABLE при mandatory-режиме.
- Решение (вариант 1 из диагноза): `bind_covering_evaluation(db, run_id, plan, outcome,
step_meta)` — если план декларирует покрывающий agent_evaluation-шаг (`comparison_refs`
на step-уровне, server-owned plan facts), walker ищет PERSISTED immutable AgentEvaluation
этого run с покрывающим comparison_id и инжектит его как `evaluation_input` compare-решения.
Truth table не изменена (SCEX-FR-028 инвариант сохранён).
- **Deferral**: пока покрывающая evaluation не persisted, compare-решение откладывается —
шаг → queued, `_deferred_this_call`-множество в вызове walker; release, когда все
покрывающие eval-шаги terminal (persisted/failed/blocked/skipped) — без busy-loop, перенос
в следующий dispatch-цикл. Отсутствие покрывающего шага в плане — прежний немедленный путь
(EVALUATION_UNAVAILABLE остаётся честным ответом mandatory-режима без объявленной оценки).
- Bound ids (`agent_evaluation_ids`, `evaluation_input`) штампуются на compare-шаг outcome для
трассируемости; запись остаётся принадлежащей agent_evaluation-шагу.
- 4/4 `test_evaluation_binding.py`: (1) graph-level terminal PASS — compare step
`BASELINE_AND_SEMANTIC_PASS`, run passed (T046 residual closed offline); (2) deferral без
ложного PASS; (3) no-covering — прежний путь; (4) disabled mode никогда не биндит.
### INV_7 рефактор
- `capacity_block.py` (новый): `block_run_on_capacity` + `reconcile_capacity_blocked_runs` +
`CAPACITY_RETRY_CODES` вынесены из walker.py (394→398 LOC после добавления binding);
импорты в runner.py/dispatch_runs.py переключены, обратная совместимость сохранена.
### Остатки G-E2E-EVALUATION
- Live rerun v4-style графа с terminal passed (живой стенд + провайдер, ближайшее окно).
- Scheduled zero-human real-LLM PASS и single-shot/best-of-N policy-варианты.
- G-LLM-SECURITY `LLM-INJ-001` live/runtime proof — отдельный live-прогон.
## Checkpoint — 2026-09-21 (T068/T069 CLOSED: bounded transform DSL + best-of-N verdict policy)
Обе оставшиеся offline-оси волны 2 закрыты. Full backend suite **11384 passed** (+22 новых),
ruff clean, compileall, anchors сбалансированы.
### T068 (AGSCN-FR-018) — `execution/transform_dsl.py` + wiring в `bounded_transform`
- Structured op-tree DSL (никаких строковых парсеров): `sum_column` (Σ колонки declared ref,
≤10000 rows), `difference` (a−b), `ratio` (a/b, 6dp canonical, деление на ноль typed),
`scale` (k×a). Depth ≤ 4. Decimal-арифметика с canonical 2dp/6dp рендерингом — byte-stable
(тест ×25 повторов). Не-числовая ячейка → typed `TRANSFORM_CELL_NON_NUMERIC` (никакой
тихой конверсии). Unknown op/keys → typed. Никакого SQL/Python/shell/network by
construction.
- Wiring: `bounded_transform` принимает `derive_value` op-tree на step; outcome несёт
`derived_value` + normalized decimal `actual` (готово к scenario_transform candidate T086
через server-owned артефакт-цепь); DSL-ошибки → typed inconclusive с `dsl_error`.
- 12/12 `test_transform_dsl.py`.
### T069 (AGSCN-FR-022) — `execution/evaluation_aggregation.py`
- `VerdictPolicy` (revision-authored элемент): single_shot | best_of_n (n∈[2..5]),
confidence_threshold ∈[0..1] — валидация typed.
- `aggregate_verdicts` pure function: strict majority (>N/2) по admissible голосам
(succeeded-with-verdict; provider/parser/budget/cancel/timeout не голосуют); tie →
`EVALUATION_AGGREGATION_TIE`; all-error → `EVALUATION_AGGREGATION_ALL_ERROR`;
quorum-miss → `EVALUATION_QUORUM_INPUTS_MISSING`; пусто → `EVALUATION_AGGREGATION_EMPTY`;
pass-below-threshold → inconclusive `EVALUATION_CONFIDENCE_BELOW_THRESHOLD` (no quiet
PASS). Mean confidence 6dp. Aggregate evaluation_id детерминированно именует когорту.
- 044-binding: агрегат потребляется существующим T051 `evaluation_binding` путём.
- Остаток: revision-schema поверхность (пин VerdictPolicy в AgentEvaluationSpec) — additive
authoring-работа.
- 10/10 `test_evaluation_aggregation.py`.
### T070 — уже закрыт волной 037 (T087 proposal + T088 observatory)
Матрица: G-BSL-DERIVED → PARTIAL (offline оси все закрыты; остатки — period stale→blocked
→candidate lifecycle и live capture wiring transform-выхлопов в ScenarioArtifact);
G-E2E-EVALUATION residual сузился до live rerun + scheduled real-LLM + revision-schema.
### Остатки волны 2 (все live-зависимые)
- Live rerun v4-style графа с terminal passed (T046 closing evidence).
- Scheduled zero-human real-LLM PASS (G-LLM-SECURITY `LLM-INJ-001` live proof).
- Period stale→blocked→candidate + live transform capture wiring.
- Catalog revision log wiring в consume/materialization.
## Checkpoint — 2026-09-21 (LIVE WAVE: E2E evaluation chain proven on the stand; terminal PASS external-blocked)
Стенд запускался `run.sh --skip-install`, затем uvicorn с captured stdout (`/tmp/kilo/uvicorn.log`).
Полный live-цикл T051-цепи выполнен на ss-prod (dashboard 11) + Gitea-published baseline envelope.
### Live canary v5 — `prototype/live_canary_v5_binding.py` (новый)
Sibling-топология: open-dashboard → capture-evidence → {compare-to-baseline, evaluate-visual};
evaluate декларирует `inputs.comparison_refs` (server-owned plan fact); decision_policy
required. Полная цепь по прогону `live-canary-v5-binding-20260921T164034Z.json` (+eval-record):
- open-dashboard passed (браузер, PREPROD-стенд);
- capture-evidence **passed BASELINE_PASS** (8 durable артефактов) — semantic-фикс этой волны;
- evaluate-visual: **real-LLM evaluation PERSISTED** (`ce4685e7…`, status succeeded, finding
с evidence_artifact_ids);
- compare-to-baseline: **walker BINDING сработал** — decision несёт `evaluation_input`
(persisted id) + `agent_evaluation_ids`; deferral-цикл прошёл (compare ждал evaluate);
- policy: LOW_CONFIDENCE (честный verdict inconclusive 0.0 от text-only модели) — fail-closed.
### Продуктовые фиксы, найденные live-прогоном
1. **Semantic overlay scope** (`decision_policy.py`): required-режим больше не требует
собственную evaluation от evidence-продюсеров (screenshot/sql без comparison-payload) —
они проходят rows 1-8 BASELINE_PASS; semantic-оверлей (rows 9-14) привязан к
comparison-шагам (`semantic_comparison` в StepPolicyInputs; mapper ставит по
tool==assertion / comparison-payload). Соответствует production-chain
(«deterministic-only runs need no LLM»); инвариант truth table сохранён.
2. **Defer deadlock guard** (`walker.py::_step_feeds_evaluation`): deferring шага, от которого
зависит покрывающая evaluation, невозможен — живой прогон выявил цикл
(capture ждёт evaluate, evaluate ждёт capture).
3. Test-фикстуры decision-policy обновлены под `semantic_comparison` (21+1 passed).
### Внешние блокеры, обнаруженные live-прогоном (2026-09-21)
- **LLM gateway vision routes ALL 502**: omniroute мультимодальные роуты (auto/best-vision,
auto/pro-vision, auto/vision, auto/claude-*, auto/gemini, auto/glm) возвращают upstream 502
«Cloudflare Playground browser session failed: browserType.launch: Executable doesn't
exist». Text-only `auto/best-fast` (qwen3.8) работает — провайдер переключён на него с
`is_multimodal=false` (честный text-only manifest по контракту). Terminal PASS
(BASELINE_AND_SEMANTIC_PASS) требует визуальный вердикт ≥0.7 — недостижим, пока gateway
vision-роут не починен; chain закрывается честно LOW_CONFIDENCE.
- **Операционные правки локального стенда**: `llm_providers.base_url` исправлен
`https://omniroute.bebesh.ru` → `https://omniroute.bebesh.ru/v1` (OpenAI SDK строит
`{base}/chat/completions`; без /v1 — 401/404 на каждый вызов, маскировался под
EVALUATION_PROVIDER_ERROR). run.sh-стенд не передаёт stdout uvicorn — для live-дебага
использован отдельный uvicorn с env из backend/.env (DRAFT_STORAGE_ROOT не нужен —
settings.storage.root_path через STORAGE_ROOT_PATH).
### Учёт
- T046: offline+live CLOSED (binding); terminal `passed` — external-blocked (vision gateway).
- T051: CLOSED (binding + deferral + deadlock guard; best-of-N через T069 aggregation).
- G-E2E-EVALUATION → PARTIAL (chain live-proven; terminal PASS blocked external).
- Evidence: `specs/044-dashboard-scenario-execution/evidence/browser-provider/
live-canary-v5-binding-20260921T164034Z.json` + `-eval-record.json`.