docs(specs): live canary v2 trace + restore operation_id in browser evidence details

This commit is contained in:
2026-09-10 20:14:49 +03:00
parent 2532702478
commit bf8abff753
4 changed files with 73 additions and 3 deletions

View File

@@ -348,6 +348,7 @@ def build_browser_provider(
"page_url": outcome.page_url,
"action": action,
"effect_state": effect_state,
**({"operation_id": operation_id} if operation_id else {}),
# Evaluation manifest inputs: byte length + MIME let
# ScenarioExecution.EvaluationAdapter.Manifest admit browser evidence.
"artifact_byte_lengths": {

View File

@@ -0,0 +1,69 @@
# Live canary v2 — REST start + server-side binding + live agent_evaluation (Track B) — 2026-09-10
**Scope:** full production chain through the running stand's HTTP surface only:
REST login → `POST /api/scenario-runs` (NO binding in the request) → server-side binding
resolution from deployment settings → PROD gate → approval decision → scheduler dispatch (5s)
→ walker → real Playwright browser → real ScreenshotService (8 captures) → real multimodal
LLM evaluation (`submit_evaluation` → `llm_providers` qwen3.8-flash) → persisted
`AgentEvaluation` row → DecisionPolicy row 11 (LOW_CONFIDENCE) → typed inconclusive.
Script: `specs/044-dashboard-scenario-execution/prototype/live_canary_v2.py` (committed).
Binding: `ss-prod-d11-live-001` (execution principal pinned to the REST actor identity —
the admin user id hash, not the username string).
## Result — live chain COMPLETE (typed inconclusive outcome is the honest verdict)
Run `4eebfab3-6d25-45dd-a3bf-e2fe6f5edf75` (scenario `live-canary-v2-0aeaf975`):
| Step | Tool/Action | Status | Evidence |
|---|---|---|---|
| open-dashboard | browser / open_dashboard (read-only) | passed | `BROWSER_ACTION_EXECUTED`, checkpoint `dashboard_open`, PNG digest, real ss-prod page URL |
| capture-evidence | screenshot / capture_screenshot | inconclusive | `EVALUATION_UNAVAILABLE` (DecisionPolicy row 9 — by design under `evaluation_mode=required`) with 8 durable screenshot refs |
| evaluate-visual | agent_evaluation / evaluate_declared_spec | inconclusive | `LOW_CONFIDENCE` (row 11) with persisted `AgentEvaluation` `d28af0d8-…` (status=succeeded, verdict=inconclusive, raw response bound by digest) |
`AgentEvaluation` row persisted: provider `a13fcc27-…` (openai/qwen3.8-flash via
`https://www.sophnet.com/api/open-apis/v1`), raw response bound by SHA-256, identity fields
server-stamped (spec-pinned provider/model/template/policy hashes; provider text never authorizes).
## Production defects found and fixed during canary v2
1. **Screenshot/browser outcomes carried no per-ref byte lengths** → the evaluation adapter's
input manifest dropped every capture item (`byte_length` required) → the evaluation always
failed `EVALUATION_EVIDENCE_NOT_FOUND` (masked as `EVALUATION_PROVIDER_ERROR`). Fix: both
providers now emit `artifact_byte_lengths` + `artifact_content_types` per ref (commit `3458343d`).
2. **Provider responses cannot satisfy the `AgentEvaluation` schema**: server-owned identity
fields (evaluation_id/operation_id/spec/template/policy hashes, provider+model identity) are
unknowable to the model, and finding payloads carry extra keys. `parse_evaluation_response`
silently swapped such responses for the parser_error fallback whose synthetic manifest can
never pass walker ownership validation — so the live LLM path could never persist a real
record. Fix: `ScenarioExecution.EvaluationAdapter.Normalize` stamps the pinned spec identity,
normalizes findings (unknown keys dropped, rationale folded into message, criterion kind
derived from the pinned spec), and pins the failure shape (commit `3458343d`).
3. **The prompt carried no output contract** — the model had to guess the response shape. The
adapter now states the judge role, the verdict/status semantics, and the evidence-id
copy-verbatim rule (same commit).
## Design observation (not a bug — pinned 038 semantics)
With `evaluation_mode=required` (derived from any `agent_evaluation` step), EVERY
policy-normative step (screenshot/sql_evidence/assertion) without its own `evaluation_input`
resolves to row 9 `EVALUATION_UNAVAILABLE` — capture steps stay inconclusive even when the
capture itself passed. This is the pinned DecisionPolicy contract (walker tests assert it);
multi-step chains that need per-step verdicts must attach evaluation records per normative step
or pin `advisory` explicitly.
## Remaining honest gaps (documented, not fabricated)
1. **Images are not attached to the LLM prompt**: `get_json_completion` sends text only; the
model cannot actually see the captured screenshots (both v2 runs produced honest
"cannot verify"/low-confidence verdicts). Multimodal image attachment in `LLMClient` +
evidence loading is the next bounded task before any visual PASS can be trusted.
2. **T6 live publication still blocked on a valid Gitea PAT** (token rejected 401). The
caller-side source (`published_catalog_source.py`, commit `bbbd4ccf`) is landed with a
fail-closed test matrix; the pin-from-Gitea live trace awaits the token.
3. **T029m MCP live-stand replay / MCPX-FR-030 T044–T046 parity** remain OPEN (spec 050).
## Verification
- Full backend suite re-run at session end (see execution record): 11502+ passed / 0 failed.
- `test_adapter_normalizes_live_provider_response` covers the live response shape regression.

View File

@@ -21,10 +21,10 @@
|---|---|---|
| Runner | Deterministic DAG walker, pinned ActionRegistry, human checkpoint, cancel/retry | Local SQLite/unit profile exists; fail-closed adapters exist |
| Provider loop | FastAPI lifespan owns exactly one `ProviderEventLoop` | Runtime wiring landed and committed 2026-09-10 (`83727aa7`); live stand: loop ready, browser-preflight startup deadlock fixed (`ba2f1f45`) |
| AgentEvaluation | Immutable record + DecisionPolicy mapper; production adapter builds `input_manifest` from prior-step artifacts | Store/parser/executor plus production `evaluation_adapter_from` and walker rows 9–14 landed and committed 2026-09-10. LLM provider configured (qwen3.8-flash multimodal, DB); live evaluation step not yet exercised |
| AgentEvaluation | Immutable record + DecisionPolicy mapper; production adapter builds `input_manifest` from prior-step artifacts | Landed, committed and **live-proven 2026-09-10** (canary v2 run `4eebfab3`: real qwen3.8-flash evaluation, `AgentEvaluation` row persisted, DecisionPolicy row 11 `LOW_CONFIDENCE` fired; trace: `docs/reports/agentic-runtime-live-canary-v2-2026-09-10.md`). OPEN: multimodal image attachment — the prompt is text-only, so visual verdicts cannot be trusted yet |
| Artifact bytes | Authenticated GET/HEAD with MIME/digest/length (Slice G) | Runtime GET/HEAD landed and committed 2026-09-10. Live storage canary PASSED 2026-09-10 (run `597274d3`: 9 `scenario_artifacts` with digests, bytes on disk) |
| Visual chain | Compiler emits browser→capture→comparison→evaluation→policy (Slice E) | One-action templates; live Playwright canary PASSED 2026-09-10 (browser `open_dashboard` + screenshot capture vs ss-prod dashboard 11; comparison/evaluation steps not in the canary graph) |
| Live providers | Registered browser/screenshot + multimodal LLM | Binding `ss-prod-d11-live-001` registered; `/api/ready`: browser/screenshot ready, bindings 1; live canary PASSED (trace: `docs/reports/agentic-runtime-live-canary-2026-09-10.md`) |
| Live providers | Registered browser/screenshot + multimodal LLM | Binding `ss-prod-d11-live-001` resolved **server-side at REST start** (no binding in the request body); `/api/ready`: browser/screenshot ready, bindings 1; canary v2 live trace: `docs/reports/agentic-runtime-live-canary-v2-2026-09-10.md` |
The former agent-driven product-UI flow (dashboard → `/agent` workspace → agent-generated scenario) is SUPERSEDED. 044 is a headless execution engine; agent interaction is external MCP (050). Product UI (045) is read-only evidence plus human approvals.

View File

@@ -91,6 +91,6 @@ Product UI never starts this chain. Manual editor (043) and monitor (045) are eq
- Full `BaselineSelectionPin` survives rerun/rebaseline/result/analytics.
- HumanStep schedules rejected equally over REST/MCP before any durable side effect.
- No agent frontend controls, routes, or requests (T030; negative DOM/route/network still OPEN).
- Live browser/capture canary PASSED 2026-09-10 against ss-prod (binding `ss-prod-d11-live-001`, run `597274d3`: browser `open_dashboard` with checkpoint/page-URL/PNG digest + 8 durable screenshot refs, capacity leases; trace: `docs/reports/agentic-runtime-live-canary-2026-09-10.md`). Still OPEN: live `agent_evaluation` step, provider-version/cost receipts, cancellation receipts.
- Live browser/capture canary PASSED 2026-09-10 against ss-prod (binding `ss-prod-d11-live-001`, run `597274d3`: browser `open_dashboard` with checkpoint/page-URL/PNG digest + 8 durable screenshot refs, capacity leases; trace: `docs/reports/agentic-runtime-live-canary-2026-09-10.md`). Live `agent_evaluation` canary v2 PASSED 2026-09-10 through REST-only start with server-side binding resolution (run `4eebfab3`: `AgentEvaluation` persisted, DecisionPolicy row 11; trace: `docs/reports/agentic-runtime-live-canary-v2-2026-09-10.md`). Still OPEN: multimodal image attachment (text-only prompt — visual verdicts not yet trustworthy), provider-version/cost receipts, cancellation receipts.
T029m live-stand replay of the 2026-09-07 sales scenario remains OPEN. MCPX-FR-030 T044–T046 remain OPEN. Optional approved performance baseline is out of scope.