docs(specs): live canary v2 trace + restore operation_id in browser evidence details
This commit is contained in:
@@ -348,6 +348,7 @@ def build_browser_provider(
|
||||
"page_url": outcome.page_url,
|
||||
"action": action,
|
||||
"effect_state": effect_state,
|
||||
**({"operation_id": operation_id} if operation_id else {}),
|
||||
# Evaluation manifest inputs: byte length + MIME let
|
||||
# ScenarioExecution.EvaluationAdapter.Manifest admit browser evidence.
|
||||
"artifact_byte_lengths": {
|
||||
|
||||
69
docs/reports/agentic-runtime-live-canary-v2-2026-09-10.md
Normal file
69
docs/reports/agentic-runtime-live-canary-v2-2026-09-10.md
Normal file
@@ -0,0 +1,69 @@
|
||||
# Live canary v2 — REST start + server-side binding + live agent_evaluation (Track B) — 2026-09-10
|
||||
|
||||
**Scope:** full production chain through the running stand's HTTP surface only:
|
||||
REST login → `POST /api/scenario-runs` (NO binding in the request) → server-side binding
|
||||
resolution from deployment settings → PROD gate → approval decision → scheduler dispatch (5s)
|
||||
→ walker → real Playwright browser → real ScreenshotService (8 captures) → real multimodal
|
||||
LLM evaluation (`submit_evaluation` → `llm_providers` qwen3.8-flash) → persisted
|
||||
`AgentEvaluation` row → DecisionPolicy row 11 (LOW_CONFIDENCE) → typed inconclusive.
|
||||
|
||||
Script: `specs/044-dashboard-scenario-execution/prototype/live_canary_v2.py` (committed).
|
||||
Binding: `ss-prod-d11-live-001` (execution principal pinned to the REST actor identity —
|
||||
the admin user id hash, not the username string).
|
||||
|
||||
## Result — live chain COMPLETE (typed inconclusive outcome is the honest verdict)
|
||||
|
||||
Run `4eebfab3-6d25-45dd-a3bf-e2fe6f5edf75` (scenario `live-canary-v2-0aeaf975`):
|
||||
|
||||
| Step | Tool/Action | Status | Evidence |
|
||||
|---|---|---|---|
|
||||
| open-dashboard | browser / open_dashboard (read-only) | passed | `BROWSER_ACTION_EXECUTED`, checkpoint `dashboard_open`, PNG digest, real ss-prod page URL |
|
||||
| capture-evidence | screenshot / capture_screenshot | inconclusive | `EVALUATION_UNAVAILABLE` (DecisionPolicy row 9 — by design under `evaluation_mode=required`) with 8 durable screenshot refs |
|
||||
| evaluate-visual | agent_evaluation / evaluate_declared_spec | inconclusive | `LOW_CONFIDENCE` (row 11) with persisted `AgentEvaluation` `d28af0d8-…` (status=succeeded, verdict=inconclusive, raw response bound by digest) |
|
||||
|
||||
`AgentEvaluation` row persisted: provider `a13fcc27-…` (openai/qwen3.8-flash via
|
||||
`https://www.sophnet.com/api/open-apis/v1`), raw response bound by SHA-256, identity fields
|
||||
server-stamped (spec-pinned provider/model/template/policy hashes; provider text never authorizes).
|
||||
|
||||
## Production defects found and fixed during canary v2
|
||||
|
||||
1. **Screenshot/browser outcomes carried no per-ref byte lengths** → the evaluation adapter's
|
||||
input manifest dropped every capture item (`byte_length` required) → the evaluation always
|
||||
failed `EVALUATION_EVIDENCE_NOT_FOUND` (masked as `EVALUATION_PROVIDER_ERROR`). Fix: both
|
||||
providers now emit `artifact_byte_lengths` + `artifact_content_types` per ref (commit `3458343d`).
|
||||
2. **Provider responses cannot satisfy the `AgentEvaluation` schema**: server-owned identity
|
||||
fields (evaluation_id/operation_id/spec/template/policy hashes, provider+model identity) are
|
||||
unknowable to the model, and finding payloads carry extra keys. `parse_evaluation_response`
|
||||
silently swapped such responses for the parser_error fallback whose synthetic manifest can
|
||||
never pass walker ownership validation — so the live LLM path could never persist a real
|
||||
record. Fix: `ScenarioExecution.EvaluationAdapter.Normalize` stamps the pinned spec identity,
|
||||
normalizes findings (unknown keys dropped, rationale folded into message, criterion kind
|
||||
derived from the pinned spec), and pins the failure shape (commit `3458343d`).
|
||||
3. **The prompt carried no output contract** — the model had to guess the response shape. The
|
||||
adapter now states the judge role, the verdict/status semantics, and the evidence-id
|
||||
copy-verbatim rule (same commit).
|
||||
|
||||
## Design observation (not a bug — pinned 038 semantics)
|
||||
|
||||
With `evaluation_mode=required` (derived from any `agent_evaluation` step), EVERY
|
||||
policy-normative step (screenshot/sql_evidence/assertion) without its own `evaluation_input`
|
||||
resolves to row 9 `EVALUATION_UNAVAILABLE` — capture steps stay inconclusive even when the
|
||||
capture itself passed. This is the pinned DecisionPolicy contract (walker tests assert it);
|
||||
multi-step chains that need per-step verdicts must attach evaluation records per normative step
|
||||
or pin `advisory` explicitly.
|
||||
|
||||
## Remaining honest gaps (documented, not fabricated)
|
||||
|
||||
1. **Images are not attached to the LLM prompt**: `get_json_completion` sends text only; the
|
||||
model cannot actually see the captured screenshots (both v2 runs produced honest
|
||||
"cannot verify"/low-confidence verdicts). Multimodal image attachment in `LLMClient` +
|
||||
evidence loading is the next bounded task before any visual PASS can be trusted.
|
||||
2. **T6 live publication still blocked on a valid Gitea PAT** (token rejected 401). The
|
||||
caller-side source (`published_catalog_source.py`, commit `bbbd4ccf`) is landed with a
|
||||
fail-closed test matrix; the pin-from-Gitea live trace awaits the token.
|
||||
3. **T029m MCP live-stand replay / MCPX-FR-030 T044–T046 parity** remain OPEN (spec 050).
|
||||
|
||||
## Verification
|
||||
|
||||
- Full backend suite re-run at session end (see execution record): 11502+ passed / 0 failed.
|
||||
- `test_adapter_normalizes_live_provider_response` covers the live response shape regression.
|
||||
@@ -21,10 +21,10 @@
|
||||
|---|---|---|
|
||||
| Runner | Deterministic DAG walker, pinned ActionRegistry, human checkpoint, cancel/retry | Local SQLite/unit profile exists; fail-closed adapters exist |
|
||||
| Provider loop | FastAPI lifespan owns exactly one `ProviderEventLoop` | Runtime wiring landed and committed 2026-09-10 (`83727aa7`); live stand: loop ready, browser-preflight startup deadlock fixed (`ba2f1f45`) |
|
||||
| AgentEvaluation | Immutable record + DecisionPolicy mapper; production adapter builds `input_manifest` from prior-step artifacts | Store/parser/executor plus production `evaluation_adapter_from` and walker rows 9–14 landed and committed 2026-09-10. LLM provider configured (qwen3.8-flash multimodal, DB); live evaluation step not yet exercised |
|
||||
| AgentEvaluation | Immutable record + DecisionPolicy mapper; production adapter builds `input_manifest` from prior-step artifacts | Landed, committed and **live-proven 2026-09-10** (canary v2 run `4eebfab3`: real qwen3.8-flash evaluation, `AgentEvaluation` row persisted, DecisionPolicy row 11 `LOW_CONFIDENCE` fired; trace: `docs/reports/agentic-runtime-live-canary-v2-2026-09-10.md`). OPEN: multimodal image attachment — the prompt is text-only, so visual verdicts cannot be trusted yet |
|
||||
| Artifact bytes | Authenticated GET/HEAD with MIME/digest/length (Slice G) | Runtime GET/HEAD landed and committed 2026-09-10. Live storage canary PASSED 2026-09-10 (run `597274d3`: 9 `scenario_artifacts` with digests, bytes on disk) |
|
||||
| Visual chain | Compiler emits browser→capture→comparison→evaluation→policy (Slice E) | One-action templates; live Playwright canary PASSED 2026-09-10 (browser `open_dashboard` + screenshot capture vs ss-prod dashboard 11; comparison/evaluation steps not in the canary graph) |
|
||||
| Live providers | Registered browser/screenshot + multimodal LLM | Binding `ss-prod-d11-live-001` registered; `/api/ready`: browser/screenshot ready, bindings 1; live canary PASSED (trace: `docs/reports/agentic-runtime-live-canary-2026-09-10.md`) |
|
||||
| Live providers | Registered browser/screenshot + multimodal LLM | Binding `ss-prod-d11-live-001` resolved **server-side at REST start** (no binding in the request body); `/api/ready`: browser/screenshot ready, bindings 1; canary v2 live trace: `docs/reports/agentic-runtime-live-canary-v2-2026-09-10.md` |
|
||||
|
||||
The former agent-driven product-UI flow (dashboard → `/agent` workspace → agent-generated scenario) is SUPERSEDED. 044 is a headless execution engine; agent interaction is external MCP (050). Product UI (045) is read-only evidence plus human approvals.
|
||||
|
||||
|
||||
@@ -91,6 +91,6 @@ Product UI never starts this chain. Manual editor (043) and monitor (045) are eq
|
||||
- Full `BaselineSelectionPin` survives rerun/rebaseline/result/analytics.
|
||||
- HumanStep schedules rejected equally over REST/MCP before any durable side effect.
|
||||
- No agent frontend controls, routes, or requests (T030; negative DOM/route/network still OPEN).
|
||||
- Live browser/capture canary PASSED 2026-09-10 against ss-prod (binding `ss-prod-d11-live-001`, run `597274d3`: browser `open_dashboard` with checkpoint/page-URL/PNG digest + 8 durable screenshot refs, capacity leases; trace: `docs/reports/agentic-runtime-live-canary-2026-09-10.md`). Still OPEN: live `agent_evaluation` step, provider-version/cost receipts, cancellation receipts.
|
||||
- Live browser/capture canary PASSED 2026-09-10 against ss-prod (binding `ss-prod-d11-live-001`, run `597274d3`: browser `open_dashboard` with checkpoint/page-URL/PNG digest + 8 durable screenshot refs, capacity leases; trace: `docs/reports/agentic-runtime-live-canary-2026-09-10.md`). Live `agent_evaluation` canary v2 PASSED 2026-09-10 through REST-only start with server-side binding resolution (run `4eebfab3`: `AgentEvaluation` persisted, DecisionPolicy row 11; trace: `docs/reports/agentic-runtime-live-canary-v2-2026-09-10.md`). Still OPEN: multimodal image attachment (text-only prompt — visual verdicts not yet trustworthy), provider-version/cost receipts, cancellation receipts.
|
||||
|
||||
T029m live-stand replay of the 2026-09-07 sales scenario remains OPEN. MCPX-FR-030 T044–T046 remain OPEN. Optional approved performance baseline is out of scope.
|
||||
|
||||
Reference in New Issue
Block a user