docs(specs): refresh 017/036/037/038/039 production contracts

Add production contract refresh sections and normative contracts: 017/050
capture-reuse trust, 036 evidence promotion, 037 catalog lifecycle/revision,
038 browser actions/decision policy, 044-047 production chain/evidence/atomic
triage, and 050 MCP interface artifacts. Requirements marked OPEN pending
executable evidence.
This commit is contained in:
2026-09-11 17:27:48 +03:00
parent 6da072b9d4
commit 605c5d0552
67 changed files with 32243 additions and 410 deletions

File diff suppressed because it is too large Load Diff

View File

@@ -48,3 +48,14 @@
- v1 items preserved; v2 items added with CHK-V2 prefix
- Items marked incomplete require spec updates before `/speckit.clarify` or `/speckit.plan`
## Production readiness checklist — 2026-09-08
FR-060: [Capture/client reuse and trust](../contracts/scenario-reuse.md). Historical [x] marks do not close this new production gate; removed frontend components are not current evidence. All rows below implemented=false / OPEN.
- [ ] CHK001 Strict parser rejects duplicate keys, malformed/empty envelopes, unknown verdict/NaN and oversized findings; retains raw secret-safe response digest. Evidence: [T085](../tasks.md), [traceability](../traceability.md).
- [ ] CHK002 JPEG bytes declare image/jpeg; 1024 q60→q30→800 stages re-estimate payload; required visual checks never pass by text-only fallback. Evidence: [T086](../tasks.md), [traceability](../traceability.md).
- [ ] CHK003 Secret exclusion and requested-mask failure are fail-closed; local optional PII policy is recorded; public provider submission remains disabled. Evidence: [T087](../tasks.md), [traceability](../traceability.md).
- [ ] CHK004 Negative product UI test: no agent chat/prompt/assistant editing/proposal-generation/typical-operation-to-agent/workspace/start/handoff controls or agent invocation routes/requests; manual CRUD/editor/human review/read-only results remain usable.
Schema/static success alone is not runtime completion. Optional approved performance baseline is outside scope.

View File

@@ -0,0 +1,19 @@
## @{ LlmAnalysis.ScenarioReuse [C:5] [TYPE ADR]
@BRIEF Reusable capture/client trust, parser and payload contracts for ScenarioRun.
@RATIONALE Keep the mature capture/client primitives while separating ValidationRecord advice from 044 immutable evaluation authority.
@REJECTED Hardcoded image/png for JPEG bytes, malformed output normalized to empty success, logging API-key length and silent text-only visual PASS.
Status: normative target; implemented=false for remaining integration/hardening.
Reuse ScreenshotService login/tab/CDP/media and LLMClient provider selection/JSON transport. 017 owns ValidationPolicy/ValidationRun/ValidationRecord; 044 owns ScenarioRun artifacts/evaluation/outcomes. JPEG inputs and WebP archive are distinct derivative objects with independent hashes; do not delete the ScenarioRun evidence bytes after a ValidationTask archive conversion.
ProviderTrustPolicy v1 is server-owned {trust_class: enterprise_local|external_disabled, approved_endpoint_fingerprint, policy_hash, mask_selectors, secret_redaction_required:true}. Default external_disabled: remote provider config/redirect/DNS drift is blocked before I/O; endpoint locality is rechecked, not inferred from a name or caller flag. Enterprise-local PII exposure remains the explicitly accepted 036/050 residual risk, so PII masking is optional only in that class. Requested masks must actually apply before capture or fail MASK_APPLICATION_FAILED. Credentials/cookies/tokens/query secrets never enter stored bytes, payload, logs or telemetry, including length or suffix. External enablement is outside this refresh and requires an architecture decision plus mandatory masking/redaction and provider audit; env escape hatches alone do not authorize it.
Strict parser v1 accepts exactly one bounded JSON object of the declared schema, explicit verdict and finite confidence [0,1], typed findings with allowlisted evidence refs; unknown envelope, list root, missing verdict, duplicate JSON keys, invalid region/confidence, truncation and excessive depth/bytes return LLM_RESPONSE_INVALID. No prose extraction may manufacture a valid result. Empty findings are legitimate only with a complete valid response. Model text and image text are untrusted evidence, never tools, URLs, policy or instructions. Escape findings on rendering.
Payload admission uses actual MIME from verified artifact and provider-advertised image types. Account for base64/request bytes and model token estimator/version; reserve output tokens. If estimate exceeds 80% context or provider byte/image limits: partition ordered evidence into declared chunks; reduce JPEG quality 60→30 at width1024, re-estimate; width800 quality30, re-estimate; then pinned Path B text fallback only if spec explicitly allows nonvisual diagnostic enrichment. Mandatory visual criterion becomes inconclusive/PAYLOAD_TOO_LARGE, never passed from text. Keep original exact/SSIM image bytes immutable; model derivatives carry source digest, conversion profile and separate digest in manifest. Missing chunk/failed chunk cannot be merged as all-pass.
Retries share one operation deadline/budget: max5 attempts including initial; exponential delays 1/2/4/8s capped by remaining budget and Retry-After. 401/403/schema errors are not retried; 429/5xx/transport failure only within receipt/reconcile rules. No retry after unknown effect, cancel or budget exhaustion.
Provenance: provider+version/model+version/prompt+hash/schema+hash, input and raw-response digests, conversion/chunk manifest, trust policy, estimator, attempt, timestamps, token usage and cost/pricing version. Unreported usage/cost is null with reason, never zero; secrets removed before raw-response persistence. 017 debug raw-response retention remains30d; 044 uses 046 raw-VLM7d and screenshot30d, with audit/reference holds. Success cannot publish until required raw-response artifact receipt is durable.
Open acceptance: JPEG/PNG/WebP MIME matrix; malformed/oversized/empty response; threshold equality; 1024→800 re-estimate/fallback; chunk partial failure; DNS/redirect drift; requested mask failure; no key length logging; bounded retry/cancel; encrypted redacted raw-response expiry and immutable provenance.
## @} LlmAnalysis.ScenarioReuse

View File

@@ -166,3 +166,11 @@ See `contracts/` directory.
8. **Parse URL**: `POST /api/validation-tasks/parse-url` — returns `{dashboard_id, title, filters, tabs}` from URL
9. **Generate Documentation**: `POST /api/tasks` with `plugin_id=llm_documentation` (unchanged)
10. **Generate Commit Message**: `POST /api/git/generate-message` (unchanged)
## Production record contract — 2026-09-08 (FR-060)
CaptureSubmission: source artifact ID/SHA/MIME/length, capture profile/hash, secret-mask receipt, optional PII-mask receipt, trust policy/version, provider/model/prompt versions, payload stages and final token estimate. RawResponseReceipt: owner, raw SHA/ref, parser schema/version/status, usage/cost/pricing provenance; no success on malformed response. Legacy ValidationRecord stays advisory; 044 owns ScenarioRun records.
Normative detail: [Capture/client reuse and trust](contracts/scenario-reuse.md). New fields, CAS transitions and cross-record integrity checks are implemented=false until [production tasks](tasks.md) and [traceability](traceability.md) close with executable evidence. Existing shorter field lists are legacy compatibility projections, not permission to omit the production identity fields.
Frontend contains no agent interaction state, prompt, proposal-generation or workspace/start controls. Human manual editor/review state and read-only evaluation/result artifacts are separate from external MCP authoring. Approved performance baseline is outside scope.

View File

@@ -187,3 +187,14 @@ frontend/src/
| Dual execution paths | Covers visual + data-layer failure modes | Single path leaves blind spots: screenshot-only misses KXD errors, text-only misses visual bugs |
| Multi-chunk screenshots | Long dashboards lose readability in single compressed image | Single full-page at higher resolution costs excessive image tokens |
| Reused URL parser from different feature | Avoids duplicating complex parsing logic for native_filters/permalink | Custom parser would duplicate 300+ lines of already-tested code |
## Production delivery plan — 2026-09-08 (FR-060)
Status: specified, implemented=false; historical unit/prototype/transport results are not current production acceptance.
1. Pin [Capture/client reuse and trust](contracts/scenario-reuse.md) and [data model](data-model.md); write negative fixtures before runtime changes.
2. Implement existing domain boundaries for: CaptureSubmission: source artifact ID/SHA/MIME/length, capture profile/hash, secret-mask receipt, optional PII-mask receipt, trust policy/version, provider/model/prompt versions, payload stages and final token estimate. RawResponseReceipt: owner, raw SHA/ref, parser schema/version/status, usage/cost/pricing provenance; no success on malformed response. Legacy ValidationRecord stays advisory; 044 owns ScenarioRun records.
3. Execute [tasks](tasks.md) T085, T086, T087 and retain reproducible evidence in [traceability](traceability.md), then close [checklist](checklists/requirements.md) individually.
4. Run cross-spec canary only after 037 catalog publication, 038 chain and 044 provider/content/policy gates; use 046 versioned cost/load/SLO limits. Disable admission on rollback, retain pins/receipts/holds; no fallback to stale catalog or synthetic PASS.
Frontend agent prompts, chat, assistant editing, proposal-generation, workspace/start/handoff actions are prohibited. Only external MCP clients interact with agents; frontend provides ordinary manual CRUD/editor, human review/approval and read-only monitoring/evidence. Runtime removal tasks are not closed by this document. Optional approved performance baseline is outside this refresh.

View File

@@ -1,5 +1,12 @@
# Quickstart: LLM Validation v2
> **Refresh 2026-09-08 (production contract):** capture/multimodal client reuse is specified by FR-060 and
> `contracts/scenario-reuse.md` — implemented=false / acceptance OPEN. ValidationRecord remains advisory legacy
> output; ScenarioRun evidence/evaluation authority belongs to 044. The v2 guide below documents the manual
> validation-task UI (permitted surface: manual task CRUD, scheduling, read-only reports) and is historical
> evidence, not production readiness: the 2026-09-08 ss-prod audit found zero registered LLM providers.
> Agent interaction stays external MCP only; no agent controls belong in this UI.
## Prerequisites
1. An API Key from OpenAI, OpenRouter, Kilo, or LiteLLM.

View File

@@ -3,7 +3,7 @@
**Feature Branch**: `017-llm-analysis-plugin`
**Created**: 2026-01-28
**Updated**: 2026-06-07 — v2: task-based flow, dual-path execution, fully implemented
**Status**: Implemented ✅ (~30 files, ~7 880 строк backend + ~4 558 строк frontend)
**Status**: Legacy implementation present; 2026-09-08 capture/client hardening and ScenarioRun integration acceptance OPEN
## v2 Design Decisions
@@ -213,3 +213,11 @@ As an Administrator, I want to configure LLM providers for documentation, git co
- What happens when a dashboard URL contains a permalink key that has expired? (System marks the source as stale, prompts user to paste a fresh URL.)
(End of file — 214 lines)
## Production contract refresh — 2026-09-08
**Frontend boundary (user decision 2026-09-08)**: All agent interaction is external MCP only. Product frontend MUST NOT contain agent chat, prompt/request textarea, assistant editing, typical-operation-to-agent selector, proposal-generation, agent workspace/start or handoff controls/routes. Ordinary manual CRUD/editor, human approval/review, monitoring and read-only evidence/evaluation are permitted. AgentEvaluationCard is read-only, with no prompt/retry-agent/provider controls. Existing agent proposal UI is runtime drift; removal/negative DOM-route-network acceptance remains OPEN in this spec-only change.
**FR-060 — Capture/client reuse and trust**: Capture and multimodal client reuse MUST follow the versioned trust/parser/payload contract. ValidationRecord remains advisory legacy output; ScenarioRun requires 044 immutable evidence/evaluation and policy authority. FR-029c masking applies to mandatory secret regions and requested selectors; enterprise-local PII masking follows the accepted 036/050 policy, while external submission remains disabled pending architecture approval.
Normative contract: [Capture/client reuse and trust](contracts/scenario-reuse.md). New requirements are specified, **implemented=false / acceptance OPEN** until executable evidence closes the linked tasks/checklist/traceability rows. Historical local tests and the manual inconclusive ss-prod run do not prove browser/capture/baseline/LLM production readiness. The refresh scope is the audited P0/P1/P2 agentic E2E and baseline gaps; an approved ExecutionPerformanceBaseline is not introduced.

View File

@@ -1,7 +1,8 @@
# Tasks: LLM Analysis & Documentation Plugins v2
**Feature**: `017-llm-analysis-plugin`
**Status**: v2 Implementation — **complete** ✅ (~30 backend files, ~12 frontend files, ~7 880 + ~4 558 строк)
**Status**: Legacy v2 implementation history; production hardening OPEN (~30 backend files, ~12 frontend files, ~7 880 + ~4 558 строк)
**Spec**: [spec.md](spec.md) | **Plan**: [plan.md](plan.md) | **Data Model**: [data-model.md](data-model.md)
## User Stories & Priorities
@@ -99,7 +100,7 @@ Phase 1 (Setup)
- Sends all images in single `content[]` array → multimodal LLM → JSON result
- [x] T019 [P] Implement `LLMClient.analyze_dashboard_text_batch(payloads, prompt) → dict` in `backend/src/plugins/llm_analysis/service.py`:
- Text-only LLM call with per-dashboard sections → `{dashboards: [{dashboard_id, status, summary, issues}]}`
- [x] T020 [P] Implement `LLMClient._estimate_payload_size()` and `LLMClient._reduce_image_quality()` in `backend/src/plugins/llm_analysis/service.py`:
- [ ] T020 [P] Implement `LLMClient._estimate_payload_size()` and `LLMClient._reduce_image_quality()` in `backend/src/plugins/llm_analysis/service.py`:
- FR-056: estimate → progressive reduction (JPEG quality 60→30, width 1024→800) → Path B fallback
- **Belief-runtime**: `belief_scope("payload_reduction")` + `reason()` before each reduction + `reflect()` after outcome
@@ -214,7 +215,7 @@ Phase 1 (Setup)
- [x] T049 [US3] Update `frontend/src/routes/validation-tasks/[policyId]/runs/[runId]/+page.svelte` — add Path A sections: tab selector (from `tab_screenshots`), per-tab image viewer with zoom, issues table with tab/chart location
**Rejected-path regression test**:
- [x] T050 [US3] Add test: multi-chunk screenshots exceed token limit → quality reduction triggered → Path B fallback when still exceeded — in `backend/tests/services/test_payload_reduction.py`
- [ ] T050 [US3] Add test: multi-chunk screenshots exceed token limit → quality reduction triggered → Path B fallback when still exceeded — in `backend/tests/services/test_payload_reduction.py`
**Verification**: Create task with `screenshot_enabled=true`, 5-tab dashboard → report shows per-tab screenshots + issues.
@@ -349,7 +350,7 @@ Phase 1 (Setup)
| **Phase 7+8** | T056–T063 (provider/prompt), T064–T070 (hub) | After Phase 3, both parallel |
| **Phase 9** | T071–T084 | After all phases |
## Total: 99 tasks across 9 phases — 86% complete
## Historical pre-refresh totals — not current production readiness
| Phase | Tasks | Status |
|-------|-------|--------|
@@ -362,3 +363,15 @@ Phase 1 (Setup)
| 7 — Provider/Prompt | 8 | ✅ 8/8 |
| 8 — Hub Cleanup | 10 | ✅ 9/10 (T070 done) |
| 9 — Polish | 14 | ⬜ 14/14 pending (T071, T075-T084 — e2e tests, semantic audit, full test suite)
## Production readiness — 2026-09-08 (FR-060)
Historical [x] rows above retain only their dated local/transport evidence; they do not prove current production readiness. Reopened rows were contradicted by the audited gaps. Removed frontend/agent paths are historical, not implementation prerequisites. New acceptance is **implemented=false / OPEN**.
Contract: [Capture/client reuse and trust](contracts/scenario-reuse.md).
- [ ] T085 [P0/P1/P2] Strict parser rejects duplicate keys, malformed/empty envelopes, unknown verdict/NaN and oversized findings; retains raw secret-safe response digest. Implement at the existing 017 domain boundary; verify with independent hardcoded fixtures and retain command/evidence references in traceability.md.
- [ ] T086 [P0/P1/P2] JPEG bytes declare image/jpeg; 1024 q60→q30→800 stages re-estimate payload; required visual checks never pass by text-only fallback. Implement at the existing 017 domain boundary; verify with independent hardcoded fixtures and retain command/evidence references in traceability.md.
- [ ] T087 [P0/P1/P2] Secret exclusion and requested-mask failure are fail-closed; local optional PII policy is recorded; public provider submission remains disabled. Implement at the existing 017 domain boundary; verify with independent hardcoded fixtures and retain command/evidence references in traceability.md.
Frontend boundary for this package: manual CRUD/editor, human review/approval, monitoring and read-only evidence/evaluation only; all agent interaction is external MCP. No agent chat/prompt/assistant editing/proposal generation/workspace/start/handoff controls. Runtime removal is OPEN, not performed by this spec refresh. Optional approved performance baseline is outside scope.

View File

@@ -0,0 +1,14 @@
# Traceability — 017-llm-analysis-plugin
## Production acceptance traceability — 2026-09-08
Historical rows above identify prior tests/code only; removed agent UI paths are retired. The following audited gates are **implemented=false / OPEN**, independent of local suite totals.
| Requirement | Domain contract / DTO | Task | Falsifiable acceptance | State |
|---|---|---|---|---|
| FR-060 | [Capture/client reuse and trust](contracts/scenario-reuse.md); [data model](data-model.md) | [T085](tasks.md) | Strict parser rejects duplicate keys, malformed/empty envelopes, unknown verdict/NaN and oversized findings; retains raw secret-safe response digest. | OPEN |
| FR-060 | [Capture/client reuse and trust](contracts/scenario-reuse.md); [data model](data-model.md) | [T086](tasks.md) | JPEG bytes declare image/jpeg; 1024 q60→q30→800 stages re-estimate payload; required visual checks never pass by text-only fallback. | OPEN |
| FR-060 | [Capture/client reuse and trust](contracts/scenario-reuse.md); [data model](data-model.md) | [T087](tasks.md) | Secret exclusion and requested-mask failure are fail-closed; local optional PII policy is recorded; public provider submission remains disabled. | OPEN |
| FR-060; external-MCP-only UI | manual editor/review; read-only evidence | [production tasks](tasks.md) | No frontend agent prompt/chat/assistant editing/proposal generation/workspace/start/handoff routes or requests; human approval remains usable. | OPEN |
Sources: [production gap](../../docs/reports/ss-prod-agentic-e2e-production-gap-2026-09-08.md), [coverage gap](../../docs/reports/ss-prod-agentic-e2e-spec-coverage-2026-09-08.md), [baseline gap](../../docs/reports/ss-prod-agentic-e2e-baseline-gap-2026-09-08.md). Spec schema/static checks prove contract structure only; live canary/runtime closure and optional approved performance baseline are not claimed.

View File

@@ -34,3 +34,14 @@
- [x] CHK016 Traceability maps every functional requirement to contracts, tasks, and tests.
- [x] CHK017 Tasks use exact repository paths, dependency order, and test-first sequencing.
- [x] CHK018 Machine-readable contracts and semantic anchors pass validation.
## Production readiness checklist — 2026-09-08
AGSTAB-FR-014: [Evidence owner and promotion](../../017-llm-analysis-plugin/contracts/scenario-reuse.md). Historical [x] marks do not close this new production gate; removed frontend components are not current evidence. All rows below implemented=false / OPEN.
- [ ] CHK019 Expired/cross-owner draft cannot promote; approved baseline acquires a retention hold before draft cleanup. Evidence: [T052](../tasks.md), [traceability](../traceability.md).
- [ ] CHK020 Changed capture/review/catalog/release between request, decision and consume conflicts with zero YAML/Git side effect. Evidence: [T053](../tasks.md), [traceability](../traceability.md).
- [ ] CHK021 External MCP owner can create/read required AgentRun state; product frontend contains no AgentRunModel/chat/workspace/agent invocation. Evidence: [T054](../tasks.md), [traceability](../traceability.md).
- [ ] CHK022 Negative product UI test: no agent chat/prompt/assistant editing/proposal-generation/typical-operation-to-agent/workspace/start/handoff controls or agent invocation routes/requests; manual CRUD/editor/human review/read-only results remain usable.
Schema/static success alone is not runtime completion. Optional approved performance baseline is outside scope.

View File

@@ -1,34 +1,8 @@
#region AgentTestStabilization.RunUx [C:4] [TYPE ADR] [SEMANTICS ux,agent-run,progress,draft,recovery]
@BRIEF Detailed FSM and recovery contract for the run strip and draft preview.
@RELATION BINDS_TO -> [AgentRuns.Model]
#region AgentTestStabilization.RunUx [C:4] [TYPE Tombstone] [SEMANTICS ux,agent-run,progress,draft,recovery]
@DEPRECATED Former agent-page UX retired by 050; no product frontend agent interaction is allowed.
@REPLACED_BY ScenarioEditor.Modules for manual editor/review; ScenarioRunMonitor.Modules for read-only results and human checkpoints.
@RATIONALE Explicit 2026-09-08 user decision prohibits agent chat/prompt/assistant editing/proposal generation/workspace/start/handoff controls, not merely Gradio transport.
## Run Panel FSM
| State | Visible contract | Actions |
|---|---|---|
| starting | Context, connecting indicator, no fake run id | Cancel chat |
| running | Run id, active/completed stages, last update | Continue chat |
| waiting_input | Required parameter summary | Provide parameters |
| waiting_approval | Exact operation, target paths, warnings | Confirm/deny |
| disconnected | Last known stage marked stale | Recover snapshot |
| completed | All completed/skipped stages and drafts | Preview/download |
| failed | Error code, failed stage, recovery choices | Retry safe analysis or restart |
| cancelled | Explicit cancellation and no side effect claim | Start new run |
## Draft States
- pending: metadata accepted, validation running.
- valid: preview/download and save request allowed.
- warning: preview required; save gate must repeat warnings.
- invalid: preview allowed, save request disabled.
- persisted: target path and persisted timestamp visible.
## UX Tests
1. Stream drop at validate → disconnected → snapshot returns waiting_approval with same drafts.
2. Foreign run event → ignored and logged.
3. Invalid artifact → save disabled with associated reason.
4. Denied save → no persisted marker; cancellation message remains visible.
5. Permission denied → no confirm button and focusable alternatives.
AgentRun, event recovery, artifact ownership and durable gate contracts continue server-side for external MCP clients. Human approval/review cards may display the same gate DTOs, but cannot start or direct an agent. This retired screen model must not be reimplemented or referenced as current test evidence. Negative product DOM/route/network acceptance remains OPEN in tasks.md.
#endregion AgentTestStabilization.RunUx

View File

@@ -1,15 +1,8 @@
#region AgentTestStabilization.UxDecisions [C:3] [TYPE ADR] [SEMANTICS ux,decisions,agent-run]
@BRIEF Final UX decisions for adding scenario-run state to the existing agent page.
#region AgentTestStabilization.UxDecisions [C:3] [TYPE Tombstone] [SEMANTICS ux,decisions,agent-run]
@DEPRECATED Former agent-page UX retired by 050; no product frontend agent interaction is allowed.
@REPLACED_BY ScenarioEditor.Modules for manual editor/review; ScenarioRunMonitor.Modules for read-only results and human checkpoints.
@RATIONALE Explicit 2026-09-08 user decision prohibits agent chat/prompt/assistant editing/proposal generation/workspace/start/handoff controls, not merely Gradio transport.
## Decisions
1. Compose AgentRuns.Model into AgentChat.Model; do not create a second chat client.
2. Keep the run panel in the agent workspace above draft/confirmation content, not in a global drawer.
3. Display the full run id with copy affordance; stage state never depends on prose.
4. Recover by snapshot after a stream gap; never silently restart the run.
5. Draft download is always available only through opaque authenticated URLs and is side-effect free.
6. Warnings are repeated in the approval card so the user reviews the exact write risk.
7. Permission denied is a separate dismiss-only state, consistent with 035.
8. Baseline reason is a required labeled field with inline validation before confirm.
AgentRun, event recovery, artifact ownership and durable gate contracts continue server-side for external MCP clients. Human approval/review cards may display the same gate DTOs, but cannot start or direct an agent. This retired screen model must not be reimplemented or referenced as current test evidence. Negative product DOM/route/network acceptance remains OPEN in tasks.md.
#endregion AgentTestStabilization.UxDecisions

View File

@@ -1,40 +1,8 @@
#region AgentTestStabilization.ScreenModels [C:4] [TYPE ADR] [SEMANTICS ux,screen-models,agent-run]
@BRIEF Screen-model inventory and invariants for recoverable scenario runs inside the existing agent page.
@RELATION DEPENDS_ON -> [AgentRuns.Model]
#region AgentTestStabilization.ScreenModels [C:4] [TYPE Tombstone] [SEMANTICS ux,screen-models,agent-run]
@DEPRECATED Former agent-page UX retired by 050; no product frontend agent interaction is allowed.
@REPLACED_BY ScenarioEditor.Modules for manual editor/review; ScenarioRunMonitor.Modules for read-only results and human checkpoints.
@RATIONALE Explicit 2026-09-08 user decision prohibits agent chat/prompt/assistant editing/proposal generation/workspace/start/handoff controls, not merely Gradio transport.
## Model Inventory
| Model | Ownership | State |
|---|---|---|
| AgentChat.Model | Existing chat, connection, messages, generic HITL | Composes AgentRuns.Model; does not duplicate run atoms |
| AgentRuns.Model | Scenario-run projection and recovery | run id/status, stages, sequence, drafts, pending gate, recovery error |
## AgentRuns.Model Actions
- startFromContext(context): begins starting state; run id arrives from agent_run_started.
- applyMetadata(event): validates run id and monotonic sequence before mutation.
- recover(runId): loads AgentRunSnapshot and atomically replaces projection.
- previewArtifact(id): obtains safe preview; does not mutate repository.
- requestAction(action): asks backend for delegated-action policy; it either records an autonomous AgentAction or returns a bound inline gate.
- decideGate(decision, reason): validates reason locally, then submits authoritative decision.
- reset(): clears only run projection when starting a new conversation/context.
## Invariants
1. Run id cannot change after start without reset.
2. Snapshot replacement is accepted only for the same run id.
3. UI stage derives from event/snapshot fields, never chat text.
4. Sequence gaps trigger degraded/recovery, not speculative stage updates.
5. Invalid drafts cannot enter save request.
6. Baseline approval cannot confirm with blank reason.
7. Ordinary chat with UIContext v1 has no AgentRuns.Model projection.
## L1 Tests
- v1 ordinary chat stays absent.
- v2 started event initializes one run.
- duplicate, stale, foreign-run, and sequence-gap events.
- recovered snapshot restores drafts and pending gate.
- deny clears pending gate without consuming artifacts.
AgentRun, event recovery, artifact ownership and durable gate contracts continue server-side for external MCP clients. Human approval/review cards may display the same gate DTOs, but cannot start or direct an agent. This retired screen model must not be reimplemented or referenced as current test evidence. Negative product DOM/route/network acceptance remains OPEN in tasks.md.
#endregion AgentTestStabilization.ScreenModels

View File

@@ -139,4 +139,12 @@ AgentRunSnapshot contains run fields, ordered stages, recent events, all draft r
- Unpersisted draft bytes: configurable, default 7 days after terminal state.
- Approval decisions: retained with the associated run audit and never silently rewritten.
## Production record contract — 2026-09-08 (AGSTAB-FR-014)
PromotionReceipt: draft artifact owner AgentRun, capture SHA/profile, durable review receipt, target baseline revision and retention hold. ApprovalIntent adds expected catalog revision/release/publication intent and CAS; human decision and consume bind the same immutable hash. ScenarioRun evidence never fabricates AgentRun ownership.
Normative detail: [Evidence owner and promotion](../017-llm-analysis-plugin/contracts/scenario-reuse.md). New fields, CAS transitions and cross-record integrity checks are implemented=false until [production tasks](tasks.md) and [traceability](traceability.md) close with executable evidence. Existing shorter field lists are legacy compatibility projections, not permission to omit the production identity fields.
Frontend contains no agent interaction state, prompt, proposal-generation or workspace/start controls. Human manual editor/review state and read-only evaluation/result artifacts are separate from external MCP authoring. Approved performance baseline is outside scope.
#endregion AgentTestStabilization.DataModel

View File

@@ -3,7 +3,7 @@
**Branch**: 036-agent-test-stabilization | **Date**: 2026-07-13 | **Spec**: [spec.md](./spec.md)
**Input**: [research.md](./research.md), [data-model.md](./data-model.md), [contracts/](./contracts/)
## Summary
## Historical architecture context (removed chat is not an active plan)
Extend the existing Gradio/LangGraph agent with a durable backend-owned AgentRun lifecycle for scenario and investigation workstreams. UIContext v2 carries typed intents without changing the positional Gradio contract. Structured progress and AgentAction metadata drive persistent workspaces; non-delegated repository/baseline actions use payload-bound inline gates.
@@ -78,3 +78,14 @@ frontend/src/lib/
## Complexity Tracking
No planned file needs an exception. If AgentChatModel.svelte.ts crosses its decomposition gate, extract AgentRunModel.svelte.ts as a composed submodel instead of growing the existing model.
## Production delivery plan — 2026-09-08 (AGSTAB-FR-014)
Status: specified, implemented=false; historical unit/prototype/transport results are not current production acceptance.
1. Pin [Evidence owner and promotion](../017-llm-analysis-plugin/contracts/scenario-reuse.md) and [data model](data-model.md); write negative fixtures before runtime changes.
2. Implement existing domain boundaries for: PromotionReceipt: draft artifact owner AgentRun, capture SHA/profile, durable review receipt, target baseline revision and retention hold. ApprovalIntent adds expected catalog revision/release/publication intent and CAS; human decision and consume bind the same immutable hash. ScenarioRun evidence never fabricates AgentRun ownership.
3. Execute [tasks](tasks.md) T052, T053, T054 and retain reproducible evidence in [traceability](traceability.md), then close [checklist](checklists/requirements.md) individually.
4. Run cross-spec canary only after 037 catalog publication, 038 chain and 044 provider/content/policy gates; use 046 versioned cost/load/SLO limits. Disable admission on rollback, retain pins/receipts/holds; no fallback to stale catalog or synthetic PASS.
Frontend agent prompts, chat, assistant editing, proposal-generation, workspace/start/handoff actions are prohibited. Only external MCP clients interact with agents; frontend provides ordinary manual CRUD/editor, human review/approval and read-only monitoring/evidence. Runtime removal tasks are not closed by this document. Optional approved performance baseline is outside this refresh.

View File

@@ -1,15 +1,23 @@
# Quickstart: Agent Test Stabilization
**Purpose**: Implement and verify 036 independently before 037–039.
**Purpose**: Verify 036 server-side run/gate/evidence surfaces independently before 037–039.
> **Refresh 2026-09-08 (production contract):** evidence ownership and promotion (AGSTAB-FR-014) are specified
> with the 017/037/044 chain — implemented=false / acceptance OPEN. The product frontend agent workspace is
> RETIRED: all agent interaction is external MCP only (050); human approval/review cards may display the same
> gate DTOs but cannot start or direct an agent. Frontend `AgentRunModel`/`AgentRunPanel` suites below target
> retired workspace components — treat them as drift-removal candidates, not current acceptance. Negative
> product DOM/route/network acceptance remains OPEN. Commands are historical runtime-gate evidence, not
> production readiness proof.
## Prerequisites
- Backend and agent virtual environments installed.
- PostgreSQL and configured auth/app databases available.
- Frontend dependencies installed.
- Frontend dependencies installed (for retired-workspace drift suites only).
- Test user has dashboard:testing READ and EXECUTE; separate approver has APPROVE.
## Recommended Test Order
## Server-side Test Order (run/gate/evidence surface)
~~~bash
cd agent
@@ -17,34 +25,35 @@ python -m pytest tests/test_agent/test_agent_context.py tests/test_agent/test_ag
cd ../backend
python -m pytest tests/services/agent_runs tests/api/test_agent_runs.py -v
~~~
## Retired-workspace drift suites (removal candidates, not acceptance)
~~~bash
cd ../frontend
npm run test -- --run src/lib/models/__tests__/AgentRunModel.test.ts
npm run test -- --run src/lib/components/agent/__tests__/AgentRunPanel.ux.test.ts
npx playwright test e2e/tests/agent-scenario-run.e2e.js
~~~
## Manual Smoke
## Current Verification Focus (refresh scope)
1. Start backend, standalone agent, and frontend with the repository run scripts.
2. Open a dashboard with env_id selected.
3. Navigate to /agent with UIContext v2 and intent=build_dashboard_test_scenario.
4. Confirm that a run id appears before the first tool card.
5. Simulate or run progress through context and inspect, then reload the page.
6. Recover the same run snapshot and verify no draft crosses conversation/run boundaries.
7. Register one warning draft; preview and download it; verify Git worktree is unchanged.
8. Request save, deny, and verify no target exists.
9. Request again, change the target after confirm, and verify consume returns 409.
10. Confirm the exact request and verify a single write plus consumed gate audit.
11. Repeat with a viewer and verify permission_denied has no confirm button.
12. Open ordinary v1 chat and verify existing 035 behavior.
1. DraftArtifact ownership stays AgentRun-scoped; no fabricated ScenarioRun ownership via aliased IDs.
2. Promotion of capture into a 037 baseline acquires a durable baseline retention hold and a verified review
receipt before draft expiry; no implicit push — publication is explicit per the 037 truth table.
3. Gates bind candidate/review/catalog-CAS/release/publication intent; MCP and REST decisions consume the same
store under fresh ACL at decide/consume.
4. Offline contract chain check (schemas, publication truth table, evaluation evidence binding):
`python specs/044-dashboard-scenario-execution/prototype/validate_contract_refresh.py specs/044-dashboard-scenario-execution/fixtures/production-contract-refresh.json`
5. Negative UI acceptance: no agent chat/prompt/assistant editing/proposal-generation/workspace/start/handoff
control, route or network request exists in product frontend (OPEN).
## Exit Gates
- agent_run_started precedes tool_start in captured stream.
- Snapshot recovery survives Gradio restart.
- agent_run_started precedes tool_start in captured stream (server-side run records).
- Snapshot/gate recovery survives service restart.
- Scenario tool list contains no arbitrary SQL operation.
- Payload/path mutation and replay tests pass.
- Existing 033/035 scoped suites pass.
- Payload/path mutation and replay tests pass; consume decisions reject stale CAS with 409.
- Existing 033/035 scoped suites pass as historical regression only.
- Backend lint, frontend lint/build, and semantic anchor audit pass.
- Drift-removal task for retired agent UI is scheduled with negative DOM/route/network acceptance (OPEN).

View File

@@ -14,7 +14,7 @@
@SEMANTICS: spec, requirements, feature, agent, gradio, langgraph, scenario, artifact, hitl, progress, dashboard-testing
**Feature Branch**: `036-agent-test-stabilization`
**Created**: 2026-07-07 | **Status**: Ready for Implementation
**Created**: 2026-07-07 | **Status**: Partially implemented; production acceptance OPEN (refresh 2026-09-08)
**Input**: "Provide a durable agent runtime for dashboard scenario creation and agentic investigations: stable UI context, explicit workstream intent, structured progress/events, recoverable identifiers, tool-action provenance, and policy-bound approvals."
## User Scenarios
@@ -67,7 +67,7 @@
**Acceptance**:
1. **Given** the agent proposes an action requiring approval **When** it is requested **Then** an inline card lists exact targets, risk, precondition evidence, and rollback/reconciliation before consumption.
2. **Given** the agent proposes baseline approval, production mutation, or automation policy change **When** approval is requested **Then** the inline card requires a user-visible reason and shows provenance.
3. **Given** the user denies a write or approval **When** denial is submitted **Then** no artifact is written or approved, and the chat records an explicit cancellation outcome.
3. **Given** the user denies a write or approval **When** denial is submitted **Then** no artifact is written or approved, and the durable gate audit records an explicit cancellation outcome.
---
@@ -75,7 +75,7 @@
- JWT expires mid-run → next protected operation stops with session-expired recovery and does not retry privileged actions.
- User lacks permission to save artifacts or approve baselines → `permission_denied` appears instead of `confirm_required`.
- Multiple tabs start runs for the same dashboard → each run has an independent `agent_run_id`; artifact preview does not cross-contaminate.
- Existing 033/035 agent chat functions remain available → dashboard test intent does not break ordinary chat, tool cards, or confirmation flows.
- Historical 033/035 chat is removed; durable gate/provenance compatibility remains, no frontend agent entry is restored.
## Requirements
@@ -87,13 +87,13 @@
- **AGSTAB-FR-004**: Draft artifacts MUST be previewable before repository writes and must include type, name, intended path, validation status, warnings, and producing run id.
- **AGSTAB-FR-005**: Repository writes, baseline approvals, production mutations, and automation policy changes MUST obey the bound ActionApprovalGate policy; any required decision captures a user-supplied reason in an inline card.
- **AGSTAB-FR-006**: Permission-denied operations MUST emit `permission_denied` metadata and never present a confirm button for unauthorized actions.
- **AGSTAB-FR-007**: The existing Gradio/LangGraph chat, streaming, confirmation, context, and guardrail flows from 033/035 MUST remain compatible.
- **AGSTAB-FR-007**: Legacy Gradio/LangGraph frontend compatibility is retired. Durable AgentRun, events and gates persist for external MCP clients; product UI MUST NOT host agent interaction.
- **AGSTAB-FR-008**: A live smoke test MUST verify context → stream → tool event → draft artifact preview → confirmation/denial recovery.
- **AGSTAB-FR-009**: [SUPERSEDED 2026-08-24 — see @REJECTED in header] Mandatory masking of screenshot evidence before LLM/VLM analysis is excluded: providers are locally deployed inside the enterprise perimeter, so unmasked captures MAY be submitted directly. Masked derivatives (`register_masked_derivative`) remain an OPTIONAL capability for display/export hygiene and MUST NOT be a precondition for VLM submission, evidence registration, or any gate.
- **AGSTAB-FR-010**: The runtime MUST support typed investigation/revalidation/remediation intents linked to the shared InvestigationCase contract. Events enter Investigation Queue and MUST NOT automatically start an agent chat or tool run.
- **AGSTAB-FR-011**: Every agent tool call MUST persist AgentAction provenance and pass deterministic ACL, environment, capacity, ActionRegistry, and approval-policy checks before execution.
- **AGSTAB-FR-012**: Delegated policy MAY authorize the agent to execute read, diagnostic, authorized fixture mutation, draft, and executable-revision writes autonomously. It MUST NOT bypass deterministic validators, immutable revision creation, or a required ActionApprovalGate.
- **AGSTAB-FR-013**: Persistent case workspaces and inline action cards are the primary agent UX; modal/dialog workflows MUST NOT be required to complete work.
- **AGSTAB-FR-013**: Product frontend uses ordinary manual editor/review and human gate cards only. Agent prompts/chat/assistant editing/proposal-generation/workspace/start/handoff controls are prohibited; agent interaction occurs exclusively in external MCP clients.
### Key Entities
@@ -110,7 +110,7 @@
## Success Criteria
- **SC-001**: 100% of dashboard scenario runs in test fixtures expose a valid `agent_run_id` before the first tool action.
- **SC-002**: Context handoff from dashboard page to `/agent` is verified for at least three dashboard ids with no stale-object reuse.
- **SC-002**: External MCP context round-trip is verified for three dashboard IDs; manual UI editor navigation preserves context and never enters an agent route.
- **SC-003**: Structured progress events render all required stages without text parsing in frontend tests.
- **SC-004**: Repository write and baseline approval attempts are blocked until confirmation in 100% of permissioned test cases.
- **SC-005**: Existing 035 agent context/guardrails tests continue to pass unchanged or with documented fixture updates only.
@@ -131,6 +131,14 @@ Gradio chat retirement is formalized in `specs/050-mcp-interface/spec.md`. Carry
- **Re-targeted**: AGSTAB-FR-003 progress-event UI requirements lose their chat consumer; structured events remain the contract for any future server-side observer. AGSTAB-FR-008 smoke test is superseded by 050 Phase 2 e2e.
- **PII posture (2026-08-24)**: AGSTAB-FR-009 superseded — masking is optional; local-only LLM/VLM deployment inside the enterprise perimeter is the accepted confidentiality boundary for evidence. The mask-miss detection edge case is withdrawn accordingly.
**Status (2026-09-02): done** — реализовано в рамках 050: инструменты и гейты (`specs/050-mcp-interface/tasks.md` T012–T028 [x]), handoff-поверхность (050 T030–T033), демонтаж чата и сервиса `agent/` (050 T040–T041, чекпоинты `specs/WORKSTATE-043-047.md`).
**Historical MCP transport/decommission status (2026-09-02): reported done; not production-readiness evidence** — реализовано в рамках 050: инструменты и гейты (`specs/050-mcp-interface/tasks.md` T012–T028 [x]), handoff-поверхность (050 T030–T033), демонтаж чата и сервиса `agent/` (050 T040–T041, чекпоинты `specs/WORKSTATE-043-047.md`).
## Production contract refresh — 2026-09-08
**Frontend boundary (user decision 2026-09-08)**: All agent interaction is external MCP only. Product frontend MUST NOT contain agent chat, prompt/request textarea, assistant editing, typical-operation-to-agent selector, proposal-generation, agent workspace/start or handoff controls/routes. Ordinary manual CRUD/editor, human approval/review, monitoring and read-only evidence/evaluation are permitted. AgentEvaluationCard is read-only, with no prompt/retry-agent/provider controls. Existing agent proposal UI is runtime drift; removal/negative DOM-route-network acceptance remains OPEN in this spec-only change.
**AGSTAB-FR-014 — Evidence owner and promotion**: DraftArtifact ownership MUST remain AgentRun-scoped. Promoting capture into a 037 baseline acquires a durable baseline retention hold and verified review receipt before draft expiry. ScenarioRun uses 044 owner receipts/content API, never fabricated AgentRun ownership. Gates bind candidate/review/catalog-CAS/release/publication intent, and MCP/REST decisions consume the same store under fresh ACL.
Normative contract: [Evidence owner and promotion](../017-llm-analysis-plugin/contracts/scenario-reuse.md). New requirements are specified, **implemented=false / acceptance OPEN** until executable evidence closes the linked tasks/checklist/traceability rows. Historical local tests and the manual inconclusive ss-prod run do not prove browser/capture/baseline/LLM production readiness. The refresh scope is the audited P0/P1/P2 agentic E2E and baseline gaps; an approved ExecutionPerformanceBaseline is not introduced.
#endregion AgentTestStabilization.Spec

View File

@@ -1,7 +1,7 @@
#region AgentTestStabilization.Tasks [C:3] [TYPE ADR] [SEMANTICS tasks,agent-run,implementation]
@BRIEF Ordered TDD implementation tasks for feature 036.
**Status**: 51/51 completed (100%). Feature branch: `036-agent-test-stabilization`.
**Status**: Historical 51-task implementation record; removed chat tasks are retired, not current acceptance. Production promotion/retention gates OPEN. Feature branch: `036-agent-test-stabilization`.
**Tests**: backend 85 ✅ (agent-runs scope) · frontend 3606 ✅ (historical unit evidence) · E2E 7/7 ✅ against a fresh Docker Compose stack.
**Last updated**: 2026-07-29 — live E2E verified.
@@ -89,7 +89,7 @@
- [x] T050 Verify optional masked derivatives are stored separately and preview ACLs are enforced; unmasked local-provider submission is allowed, while credentials/cookies/tokens are never exposed.
- [x] T051 Audit local-provider evidence handling: unmasked capture is permitted for local LLM/VLM analysis; preview ACLs and secret exclusion remain enforced.
## Implementation Closure — Lifecycle, Idempotency, Recovery, Readiness
## Historical Implementation Record — Lifecycle, Idempotency, Recovery, Readiness
### Lifecycle FSM (AgentRun)
- **States**: CREATED → RUNNING → [WAITING_INPUT | WAITING_APPROVAL] → COMPLETED | FAILED | CANCELLED
@@ -104,7 +104,7 @@
- **Payload hash**: SHA-256 of canonical JSON; backend computes if absent.
- Implemented in `Services.AgentRuns.Service.AppendEvent` and `AgentRunRepository.append_event`.
### Recovery (Frontend)
### Historical Recovery (Frontend retired by 050)
- **AgentRunModel.svelte.ts**: composes with AgentChatModel.svelte.ts; restores run id from route/session-safe state on reload.
- **StreamProcessor**: handles started/progress/drafts/terminal events; gap recovery via `GET /api/agent/runs/{run_id}/events` after reconnection.
- **Recovery tests**: L1 model tests (13) for snapshot, metadata, gate decision, gap recovery.
@@ -126,4 +126,16 @@
T001–T009 → US1 → US2 → US3 → US4 → integration.
Within a phase, test tasks precede implementation. 037 must not begin before the US4 consume contract passes. Phase 8 requires completed US3 (DraftArtifact infrastructure).
## Production readiness — 2026-09-08 (AGSTAB-FR-014)
Historical [x] rows above retain only their dated local/transport evidence; they do not prove current production readiness. Reopened rows were contradicted by the audited gaps. Removed frontend/agent paths are historical, not implementation prerequisites. New acceptance is **implemented=false / OPEN**.
Contract: [Evidence owner and promotion](../017-llm-analysis-plugin/contracts/scenario-reuse.md).
- [ ] T052 [P0/P1/P2] Expired/cross-owner draft cannot promote; approved baseline acquires a retention hold before draft cleanup. Implement at the existing 036 domain boundary; verify with independent hardcoded fixtures and retain command/evidence references in traceability.md.
- [ ] T053 [P0/P1/P2] Changed capture/review/catalog/release between request, decision and consume conflicts with zero YAML/Git side effect. Implement at the existing 036 domain boundary; verify with independent hardcoded fixtures and retain command/evidence references in traceability.md.
- [ ] T054 [P0/P1/P2] External MCP owner can create/read required AgentRun state; product frontend contains no AgentRunModel/chat/workspace/agent invocation. Implement at the existing 036 domain boundary; verify with independent hardcoded fixtures and retain command/evidence references in traceability.md.
Frontend boundary for this package: manual CRUD/editor, human review/approval, monitoring and read-only evidence/evaluation only; all agent interaction is external MCP. No agent chat/prompt/assistant editing/proposal generation/workspace/start/handoff controls. Runtime removal is OPEN, not performed by this spec refresh. Optional approved performance baseline is outside scope.
#endregion AgentTestStabilization.Tasks

View File

@@ -31,4 +31,17 @@
The `trigger` field on AgentRun/VerificationRun gains the additive value `dataset_updated` (producer: `Services.Lineage.Fanout`, 041 US5). Enum extension is data-level and backward compatible; no spec rewrite.
## Production acceptance traceability — 2026-09-08
Historical rows above identify prior tests/code only; removed agent UI paths are retired. The following audited gates are **implemented=false / OPEN**, independent of local suite totals.
| Requirement | Domain contract / DTO | Task | Falsifiable acceptance | State |
|---|---|---|---|---|
| AGSTAB-FR-014 | [Evidence owner and promotion](../017-llm-analysis-plugin/contracts/scenario-reuse.md); [data model](data-model.md) | [T052](tasks.md) | Expired/cross-owner draft cannot promote; approved baseline acquires a retention hold before draft cleanup. | OPEN |
| AGSTAB-FR-014 | [Evidence owner and promotion](../017-llm-analysis-plugin/contracts/scenario-reuse.md); [data model](data-model.md) | [T053](tasks.md) | Changed capture/review/catalog/release between request, decision and consume conflicts with zero YAML/Git side effect. | OPEN |
| AGSTAB-FR-014 | [Evidence owner and promotion](../017-llm-analysis-plugin/contracts/scenario-reuse.md); [data model](data-model.md) | [T054](tasks.md) | External MCP owner can create/read required AgentRun state; product frontend contains no AgentRunModel/chat/workspace/agent invocation. | OPEN |
| AGSTAB-FR-014; external-MCP-only UI | manual editor/review; read-only evidence | [production tasks](tasks.md) | No frontend agent prompt/chat/assistant editing/proposal generation/workspace/start/handoff routes or requests; human approval remains usable. | OPEN |
Sources: [production gap](../../docs/reports/ss-prod-agentic-e2e-production-gap-2026-09-08.md), [coverage gap](../../docs/reports/ss-prod-agentic-e2e-spec-coverage-2026-09-08.md), [baseline gap](../../docs/reports/ss-prod-agentic-e2e-baseline-gap-2026-09-08.md). Spec schema/static checks prove contract structure only; live canary/runtime closure and optional approved performance baseline are not claimed.
#endregion AgentTestStabilization.Traceability

View File

@@ -1,85 +1,28 @@
#region AgentTestStabilization.UxReference [C:3] [TYPE ADR] [SEMANTICS ux,reference,agent,dashboard-testing]
@BRIEF UX reference for durable agent workspaces for scenario creation and investigations.
#region AgentTestStabilization.UxReference [C:3] [TYPE Tombstone] [SEMANTICS ux,reference,agent,dashboard-testing]
@DEPRECATED Former agent-workspace UX reference retired by the 2026-09-08 external-MCP-only decision; no product frontend agent interaction is allowed.
@REPLACED_BY ScenarioEditor.Modules (manual editor/review) and ScenarioRunMonitor.Modules (read-only results, human checkpoints); see specs/036-agent-test-stabilization/contracts/ux/screen-models.md.
@RATIONALE Explicit 2026-09-08 user decision prohibits agent chat, prompt/request input, assistant editing, proposal generation, agent workspace/start and handoff controls in product UI, not merely a Gradio transport change.
**Feature Branch**: `036-agent-test-stabilization`
**Created**: 2026-07-07 | **Status**: Ready for Implementation
**Created**: 2026-07-07 | **Superseded**: 2026-09-08 | **Status**: Historical record; not an active UX contract
## 1. User Persona & Context
## 1. Current boundary (active)
* **Who is the user?**: BI analyst working on scenario creation, an opened investigation, revalidation, or remediation.
* **What is their goal?**: Work in a reliable, recoverable agent thread with evidence, tool actions, and durable outputs.
* **Context**: Browser-based Svelte `/agent` workspace launched from a dashboard page with Superset environment selected.
* AgentRun recovery, draft-artifact ownership, tool-run records and durable ActionApprovalGate contracts continue **server-side** and are consumed by external MCP clients (see `specs/050-mcp-interface/spec.md`).
* Product frontend keeps only: manual CRUD/editor (042/043), human approval/review cards bound to the same gate DTOs, monitoring (045/047) and read-only evidence/evaluation rendering (045). No control may start or direct an agent.
* Existing agent workspace/proposal panels in runtime code are drift; removal plus negative DOM/route/network acceptance remains OPEN in tasks.md. This spec-only refresh changes no runtime code.
## 2. Happy Path Narrative
## 2. Retired interaction (historical narrative)
The analyst starts a scenario workspace or opens an Investigation Queue item. The persistent workspace shows the source context, evidence and agent thread. The agent starts recoverable tool runs, records actions, and may save a validated revision under delegated policy. High-risk actions render a bound inline approval card; denial records a cancellation without side effects.
The 2026-07-07 design placed a persistent Svelte `/agent` workspace on the dashboard page: agent context header with run/intent binding, a stage progress strip, draft-artifact preview with request-save, and inline approval cards. Scenario creation, inspection and draft generation were agent-driven inside the product UI. Every element of that flow is superseded: authoring moved to external MCP clients, and the product UI now loads server-stored proposals as read-only diffs for human review with manual editing (043) and human-in-the-loop approval (036 gates).
## 3. Interface Mockups
The retired mockups (agent context header, progress strip, draft preview, inline gate) are preserved in Git history only; they must not be reimplemented or cited as current test evidence.
### Agent Context Header
## 3. Still-valid server-side UX obligations
```text
┌─────────────────────────────────────────────────────────────────────────────┐
│ 🧪 Dashboard Test Scenario Agent ● connected │
├─────────────────────────────────────────────────────────────────────────────┤
│ Context: FI-0080 | dashboard_id=42 | env=ss-dev │
│ Intent: build_dashboard_test_scenario | run_id: ag-run-20260707-001 │
└─────────────────────────────────────────────────────────────────────────────┘
```
### Progress Strip
```text
[Context ✓] → [Inspect …] → [Scenario] → [Parameters] → [Generate] → [Validate] → [Save]
```
### Draft Artifact Preview
```text
┌──────────────────────── Draft artifacts ────────────────────────────────────┐
│ Run: ag-run-20260707-001 │
│ ✓ scenario.yaml valid │
│ ⚠ baseline-candidates.yaml needs approval reason │
│ ✓ generated-preview.json valid │
│ │
│ [Preview] [Download draft] [Request save] │
└─────────────────────────────────────────────────────────────────────────────┘
```
### Inline ActionApprovalGate
```text
┌──────────────────────── Confirmation required ──────────────────────────────┐
│ Publish baseline candidate │
│ Target: tests/generated/dashboards/fi_0080/ │
│ Evidence: comparison SR-1845, lineage snapshot │
│ Risk: baseline approval │
│ Reconciliation: revert candidate if verification fails │
│ │
│ [Approve] [Deny] [Open diff] │
└─────────────────────────────────────────────────────────────────────────────┘
```
## 4. Error Experience
### Scenario A: Stale or Invalid Context
* **System Response**: Context card turns warning-tone and lists invalid fields.
* **Recovery**: User can return to dashboard, retry context handoff, or continue in non-scenario chat mode.
### Scenario B: Stream Drops During Run
* **System Response**: UI shows `run_id`, last known stage, and reconnect actions.
* **Recovery**: User reconnects and resumes viewing draft status; no write happens during disconnected state.
### Scenario C: Unauthorized Approval
* **System Response**: Permission-denied card with required role and no confirm button.
* **Recovery**: User can download draft only if allowed, or request approval from an authorized role.
## 5. Tone & Voice
* **Style**: Technical, explicit, side-effect aware.
* **Terminology**: Use "run", "draft artifact", "confirmation", "approval reason", and "dashboard context" consistently.
* High-risk decisions (baseline candidate publish, repository save, promotion to 037 baseline) render as human approval cards bound to a durable gate: target paths, evidence refs, risk, reconciliation policy, approve/deny/open-diff. Denial records a cancellation with zero side effects.
* Permission-denied presents the required role and no confirm control; viewer roles never see confirm buttons.
* Recovery semantics survive restarts: run/gate snapshots reload deterministically; no write occurs while a decision is unauthenticated or disconnected.
* Terminology stays explicit and side-effect aware: "run", "draft artifact", "gate", "confirmation", "approval reason", "dashboard context".
#endregion AgentTestStabilization.UxReference

View File

@@ -33,3 +33,14 @@
- [x] CHK015 Traceability maps every functional requirement to contracts, tasks, and tests.
- [x] CHK016 Tasks use exact repository paths, dependency order, and test-first sequencing.
- [x] CHK017 Machine-readable contracts, external references, and semantic anchors pass validation.
## Production readiness checklist — 2026-09-08
AGBASE-FR-014: [CatalogRevision and baseline lifecycle](../contracts/catalog-lifecycle.md). Historical [x] marks do not close this new production gate; removed frontend components are not current evidence. All rows below implemented=false / OPEN.
- [ ] CHK018 Duplicate IDs/approved coordinates reject; stale If-Match conflicts; two concurrent approvals produce one catalog head. Evidence: [T082](../tasks.md), [traceability](../traceability.md).
- [ ] CHK019 Capture bytes/profile/review are server-owned; forged content refs/hashes/approved booleans cannot approve; required raw bytes remain held. Evidence: [T083](../tasks.md), [traceability](../traceability.md).
- [ ] CHK020 Rebaseline supersedes via CAS; retire/invalidate preserve historical pins; explicit Git commit/push failure returns publish_failed and retry reconciles without duplicate revision. Evidence: [T084](../tasks.md), [traceability](../traceability.md).
- [ ] CHK021 Negative product UI test: no agent chat/prompt/assistant editing/proposal-generation/typical-operation-to-agent/workspace/start/handoff controls or agent invocation routes/requests; manual CRUD/editor/human review/read-only results remain usable.
Schema/static success alone is not runtime completion. Optional approved performance baseline is outside scope.

View File

@@ -23,6 +23,7 @@
}
},
"$defs": {
"entryRecord": { "oneOf": [{ "$ref": "#/$defs/entry" }, { "$ref": "#/$defs/visualEntry" }] },
"sha256": { "type": "string", "pattern": "^[a-f0-9]{64}$" },
"filters": {
"type": "object",

View File

@@ -0,0 +1,166 @@
{
"$schema": "https://json-schema.org/draft/2020-12/schema",
"title": "BaselineSelectionPin v1",
"description": "Normative target 2026-09-08; implemented=false. Server resolves approved catalog; client digests never establish authority.",
"type": "object",
"additionalProperties": false,
"required": [
"schema_version",
"baseline_set_id",
"baseline_set_version",
"catalog_revision_id",
"catalog_digest",
"release_id",
"release_version",
"release_commit_hash",
"publication_commit_hash",
"baseline_family",
"entries"
],
"properties": {
"schema_version": {
"const": 1
},
"baseline_set_id": {
"type": "string",
"minLength": 1
},
"baseline_set_version": {
"type": "string",
"minLength": 1
},
"catalog_revision_id": {
"type": "string",
"format": "uuid"
},
"catalog_digest": {
"type": "string",
"pattern": "^[a-f0-9]{64}$"
},
"release_id": {
"type": "string",
"format": "uuid"
},
"release_version": {
"type": "string",
"minLength": 1
},
"release_commit_hash": {
"type": "string",
"pattern": "^[a-f0-9]{40}$"
},
"publication_commit_hash": {
"type": "string",
"pattern": "^[a-f0-9]{40}$"
},
"baseline_family": {
"type": "string",
"pattern": "^[a-f0-9]{64}$"
},
"entries": {
"type": "array",
"items": {
"$ref": "#/$defs/entry"
},
"minItems": 1,
"maxItems": 100,
"uniqueItems": true
}
},
"$defs": {
"entry": {
"type": "object",
"additionalProperties": false,
"required": [
"baseline_id",
"baseline_revision_id",
"kind",
"coordinate_hash",
"entry_digest",
"source_response_hash",
"expected_image_sha256",
"capture_artifact_id",
"capture_profile_hash"
],
"properties": {
"baseline_id": {
"type": "string",
"format": "uuid"
},
"baseline_revision_id": {
"type": "string",
"format": "uuid"
},
"kind": {
"enum": [
"metric",
"visual"
]
},
"coordinate_hash": {
"type": "string",
"pattern": "^[a-f0-9]{64}$"
},
"entry_digest": {
"type": "string",
"pattern": "^[a-f0-9]{64}$"
},
"source_response_hash": {
"type": "string",
"pattern": "^[a-f0-9]{64}$"
},
"expected_image_sha256": {
"anyOf": [
{
"type": "string",
"pattern": "^[a-f0-9]{64}$"
},
{
"type": "null"
}
]
},
"capture_artifact_id": {
"type": "string",
"format": "uuid"
},
"capture_profile_hash": {
"type": "string",
"pattern": "^[a-f0-9]{64}$"
}
},
"allOf": [
{
"if": {
"properties": {
"kind": {
"const": "visual"
}
}
},
"then": {
"properties": {
"expected_image_sha256": {
"type": "string",
"pattern": "^[a-f0-9]{64}$"
}
}
},
"else": {
"properties": {
"expected_image_sha256": {
"type": "null"
}
}
}
}
]
}
},
"$id": "https://superset-tools.local/specs/037-superset-baseline-engine/contracts/baseline-pin.schema.json",
"x-invariants": [
"baseline_id, baseline_revision_id and coordinate_hash unique across entries",
"pin is server-resolved from published approved catalog and release authority",
"all entry digests verified before admission"
]
}

View File

@@ -0,0 +1,28 @@
## @{ BaselineEngine.CatalogRevision [C:5] [TYPE ADR]
@BRIEF Immutable catalog generations, authoritative capture, CAS lifecycle and explicitly published Git truth.
@RATIONALE A catalog-level CAS serializes coordinate ownership and gives ScenarioRun a reproducible selection; separate materialization and publication expose the real Git side effect and retry boundary.
@REJECTED First matching approved entry, silent duplicate-ID skip, caller approval and implicit push on gate consume — these can silently change expected truth or report publication before Git succeeds.
Status: normative target 2026-09-08, implemented=false. Schemas: [catalog-revision.schema.json](catalog-revision.schema.json), [baseline-pin.schema.json](baseline-pin.schema.json). Existing [baseline-catalog.schema.json](baseline-catalog.schema.json) describes v1 entry serialization/import; metric entries do not require visual fingerprints or caller approval. Migration wraps verified entries in CatalogRevision; it never fabricates capture/review evidence for legacy rows.
CatalogRevision identity: UUID + catalog_id + parent_revision_id + SHA-256 of canonical sorted entry revisions; CAS is monotonically increasing integer, exposed as strong ETag. Digest excludes its own field, timestamps and publication status. Baseline identity remains stable baseline_id; each immutable value/policy/provenance version has a fresh baseline_revision_id and entry_digest. Historical payload never changes; statuses and successor edges belong to a new CatalogRevision.
Coordinate uniqueness is enforced transactionally: at most one approved revision per (repository, release, dashboard, kind, chart-or-dataset, result_key, normalized filters including explicit time range, tab, ROI, capture_profile_hash). Irrelevant metric/visual components are explicit null; normalize and hash canonical JSON, not string concatenation. Exact duplicate ID+digest is idempotent; duplicate ID with different bytes is BASELINE_ID_CONFLICT; multiple approved entries are BASELINE_AMBIGUOUS, never first-match. baseline_family is SHA-256 of sorted coordinate identities excluding release and expected-value digests; full catalog/release pins still distinguish changes.
Operations use authenticated READ/EXECUTE/APPROVE permissions from dashboard:testing plus repository/dashboard/environment/evidence ACL. Mutations require Idempotency-Key and If-Match=catalog ETag; missing precondition→428, stale→409 CATALOG_REVISION_CONFLICT, same key changed canonical request→409 IDEMPOTENCY_KEY_REUSED. Errors create no entry/pointer/gate/write side effect.
1. Capture metric: server resolves approved release/environment/query envelope, executes Superset and stores full response bytes plus normalized bounded redacted value. No caller expected/raw digest authority.
2. Capture visual: server captures using approved profile and actual principal through reused ScreenshotService, or imports a server-owned ScenarioArtifact after verifying owner/run/attempt/receipt, bytes/MIME/SHA-256, target/filter/tab/ROI/profile and independent baseline-versus-actual execution identity. A durable confirmed ReviewDisposition bound to capture artifact digest is mandatory before candidate creation.
3. Request/decide/consume: candidate content, capture and review IDs/digests, expected catalog CAS, release, period override reason and target are bound in one durable 036 gate. Revalidate live ACL and all hashes at decide/consume. Consume atomically records a new materialized CatalogRevision and publication intent; working-tree YAML is not a published baseline.
4. Publish is an explicit separately invoked operation with a reviewed gate binding catalog digest, repository, branch and expected branch HEAD. State: materialized→publish_pending→committed→published. Commit exactly the pinned catalog manifest; compare-and-set branch HEAD; push to the configured remote only. Return commit and remote verification receipts. Commit/push failure→publish_failed retaining operation/commit ID; retry reconciles the same commit and never creates duplicate commits or force-pushes. A changed HEAD requires a new reviewed diff/gate. A materialization failure restores only its owned staged write, never concurrent user changes.
5. supersede/rebaseline requires fresh authoritative capture + review + approval and CAS; one new catalog revision atomically marks predecessor superseded and installs one approved successor with supersedes_revision_id. Retire removes future selection; invalidate records evidence/reason and prevents new selection. Closed-period override requires explicit analyst reason and preserves the immutability incident; it cannot erase old expected bytes.
6. An already launched run keeps its pinned historical bytes; invalidation blocks new launches and pending approved-but-undispatched runs on pre-I/O revalidation. Retire/invalidate never rewrites historical run truth.
Only published, approved, unambiguous entries can resolve for production launch. Local diagnostic materialization may be inspected but cannot produce production PASS.
Capture provenance adds media_type, byte_length, sha256, capture_artifact_id, query_envelope_hash/version, explicit time_range, viewport/pixel dimensions, method/readiness/profile/provider/browser versions, target/principal/RLS/filter/layout fingerprints, captured_at and durable review receipt. JPEG/PNG/WebP hashes refer to distinct byte objects; transformations store source→derivative edges, never reuse a digest.
Retention: pin/reference hold protects full metric raw responses and baseline image bytes while an approved entry or retained historical run requires them. Draft 7-day expiry does not apply after promotion; promotion atomically acquires the baseline hold before releasing draft ownership. Retirement releases only its own hold; delete after all holds and 046 retention windows expire. Tombstones retain IDs/digests/actor/reason/deleted_at and deletion receipt. Missing approved bytes makes selection BASELINE_EVIDENCE_UNAVAILABLE; no recapture or current catalog substitution.
Acceptance (open): PostgreSQL concurrent CAS and coordinate conflict; changed duplicate ID; forged review and wrong owner; same-run visual self-comparison; closed-period rebaseline; commit-success/push-failure/restart; changed branch; retention while draft expires; REST/MCP capture/request/decide/consume/publish parity; immutable historical replay.
## @} BaselineEngine.CatalogRevision

View File

@@ -0,0 +1,884 @@
{
"$schema": "https://json-schema.org/draft/2020-12/schema",
"title": "CatalogRevision v1",
"description": "Append-only catalog revision envelope; baseline-catalog.schema.json v1 entries remain an import format. Implemented=false.",
"type": "object",
"additionalProperties": false,
"required": [
"schema_version",
"catalog_revision_id",
"catalog_id",
"parent_revision_id",
"catalog_digest",
"cas_version",
"entry_revisions",
"publication",
"created_at",
"actor_id",
"reason"
],
"properties": {
"schema_version": {
"const": 1
},
"catalog_revision_id": {
"type": "string",
"format": "uuid"
},
"catalog_id": {
"type": "string",
"format": "uuid"
},
"parent_revision_id": {
"anyOf": [
{
"type": "string",
"format": "uuid"
},
{
"type": "null"
}
]
},
"catalog_digest": {
"type": "string",
"pattern": "^[a-f0-9]{64}$"
},
"cas_version": {
"type": "integer",
"minimum": 1
},
"entry_revisions": {
"type": "array",
"items": {
"type": "object",
"additionalProperties": false,
"required": [
"baseline_id",
"baseline_revision_id",
"entry_digest",
"coordinate_hash",
"status",
"supersedes_revision_id",
"entry",
"capture_artifact_id",
"capture_profile_hash",
"review_disposition_id",
"review_status"
],
"properties": {
"baseline_id": {
"type": "string",
"format": "uuid"
},
"baseline_revision_id": {
"type": "string",
"format": "uuid"
},
"entry_digest": {
"type": "string",
"pattern": "^[a-f0-9]{64}$"
},
"coordinate_hash": {
"type": "string",
"pattern": "^[a-f0-9]{64}$"
},
"status": {
"enum": [
"approved",
"superseded",
"retired",
"invalidated"
]
},
"supersedes_revision_id": {
"anyOf": [
{
"type": "string",
"format": "uuid"
},
{
"type": "null"
}
]
},
"entry": {
"oneOf": [
{
"$ref": "#/$defs/entry"
},
{
"$ref": "#/$defs/visualEntry"
}
]
},
"capture_artifact_id": {
"type": "string",
"format": "uuid"
},
"capture_profile_hash": {
"type": "string",
"pattern": "^[a-f0-9]{64}$"
},
"review_disposition_id": {
"anyOf": [
{
"type": "string",
"format": "uuid"
},
{
"type": "null"
}
]
},
"review_status": {
"enum": [
"confirmed",
"not_applicable"
]
}
},
"allOf": [
{
"if": {
"properties": {
"entry": {
"required": [
"kind"
],
"properties": {
"kind": {
"const": "visual"
}
}
}
}
},
"then": {
"properties": {
"review_disposition_id": {
"type": "string",
"format": "uuid"
},
"review_status": {
"const": "confirmed"
}
}
},
"else": {
"properties": {
"review_disposition_id": {
"type": "null"
},
"review_status": {
"const": "not_applicable"
}
}
}
}
]
},
"minItems": 0,
"maxItems": 100,
"uniqueItems": true
},
"publication": {
"type": "object",
"additionalProperties": false,
"required": [
"state",
"publication_id",
"expected_branch_head",
"commit_hash",
"error_code",
"published_receipt_id"
],
"properties": {
"state": {
"enum": [
"materialized",
"publish_pending",
"committed",
"published",
"publish_failed"
]
},
"publication_id": {
"type": "string",
"format": "uuid"
},
"expected_branch_head": {
"type": "string",
"pattern": "^[a-f0-9]{40}$"
},
"commit_hash": {
"anyOf": [
{
"type": "string",
"pattern": "^[a-f0-9]{40}$"
},
{
"type": "null"
}
]
},
"error_code": {
"type": [
"string",
"null"
]
},
"published_receipt_id": {
"anyOf": [
{
"type": "string",
"format": "uuid"
},
{
"type": "null"
}
]
}
},
"allOf": [
{
"if": {
"properties": {
"state": {
"enum": [
"materialized",
"publish_pending"
]
}
}
},
"then": {
"properties": {
"commit_hash": {
"type": "null"
},
"published_receipt_id": {
"type": "null"
}
}
},
"description": "Lifecycle materialized/publish_pending precedes any Git commit; a commit or receipt here is an invalid state combination (catalog-lifecycle.md publish chain)."
},
{
"if": {
"properties": {
"state": {
"const": "committed"
}
}
},
"then": {
"properties": {
"commit_hash": {
"type": "string",
"pattern": "^[a-f0-9]{40}$"
},
"published_receipt_id": {
"type": "null"
}
}
},
"description": "Committed means the pinned manifest commit exists but remote publication is not yet receipted; only published carries a receipt."
},
{
"if": {
"properties": {
"state": {
"const": "published"
}
}
},
"then": {
"properties": {
"commit_hash": {
"type": "string",
"pattern": "^[a-f0-9]{40}$"
},
"published_receipt_id": {
"type": "string",
"format": "uuid"
}
}
}
},
{
"if": {
"properties": {
"state": {
"const": "publish_failed"
}
}
},
"then": {
"properties": {
"published_receipt_id": {
"type": "null"
}
}
},
"description": "publish_failed retains operation/commit identity for reconcile but can never carry a publication receipt."
},
{
"if": {
"properties": {
"state": {
"const": "publish_failed"
}
}
},
"then": {
"properties": {
"error_code": {
"enum": [
"GIT_CAS_CONFLICT",
"GIT_COMMIT_FAILED",
"GIT_PUSH_FAILED",
"GIT_RECONCILE_REQUIRED"
]
}
}
},
"else": {
"properties": {
"error_code": {
"type": "null"
}
}
}
}
]
},
"created_at": {
"type": "string",
"format": "date-time"
},
"actor_id": {
"type": "string",
"minLength": 1
},
"reason": {
"type": "string",
"minLength": 1,
"maxLength": 2000
}
},
"$id": "https://superset-tools.local/specs/037-superset-baseline-engine/contracts/catalog-revision.schema.json",
"$defs": {
"sha256": {
"type": "string",
"pattern": "^[a-f0-9]{64}$"
},
"filters": {
"type": "object",
"additionalProperties": false,
"required": [
"schema_version",
"filters",
"filters_hash"
],
"properties": {
"schema_version": {
"const": 1
},
"filters": {
"type": "array",
"items": {
"type": "object"
}
},
"filters_hash": {
"$ref": "#/$defs/sha256"
}
}
},
"value": {
"$ref": "comparison-types.schema.json#/$defs/value"
},
"policy": {
"$ref": "comparison-types.schema.json#/$defs/metricPolicy"
},
"visualPolicy": {
"$ref": "comparison-types.schema.json#/$defs/visualPolicy"
},
"visualEntry": {
"type": "object",
"additionalProperties": false,
"required": [
"schema_version",
"baseline_id",
"dashboard_id",
"kind",
"release_version",
"release_commit_hash",
"normalized_filters",
"tab_identifier",
"expected_image_sha256",
"source_response_hash",
"captured_at",
"policy",
"status",
"fingerprints",
"provenance",
"approval",
"created_at",
"expected_image_content_ref"
],
"properties": {
"schema_version": {
"const": 1
},
"baseline_id": {
"type": "string",
"format": "uuid"
},
"dashboard_id": {
"type": "integer",
"minimum": 1
},
"release_version": {
"type": "string",
"pattern": "^v\\d+\\.\\d+\\.\\d+(-[a-z0-9.]+)?$"
},
"release_commit_hash": {
"type": "string",
"pattern": "^[a-f0-9]{40}$"
},
"kind": {
"const": "visual"
},
"normalized_filters": {
"$ref": "#/$defs/filters"
},
"tab_identifier": {
"type": "string",
"minLength": 1
},
"region_of_interest": {
"type": [
"object",
"null"
],
"properties": {
"selector": {
"type": "string"
},
"bounds": {
"type": "object",
"required": [
"x",
"y",
"w",
"h"
],
"properties": {
"x": {
"type": "integer"
},
"y": {
"type": "integer"
},
"w": {
"type": "integer",
"minimum": 1
},
"h": {
"type": "integer",
"minimum": 1
}
}
}
}
},
"expected_image_sha256": {
"$ref": "#/$defs/sha256"
},
"expected_image_content_ref": {
"type": "string",
"minLength": 1,
"maxLength": 512
},
"source_response_hash": {
"$ref": "#/$defs/sha256"
},
"content_hash": {
"type": [
"string",
"null"
],
"pattern": "^[a-f0-9]{64}$",
"description": "Dashboard-level content_hash at capture time for FR-013 inheritance."
},
"captured_at": {
"type": "string",
"format": "date-time"
},
"policy": {
"$ref": "#/$defs/visualPolicy"
},
"status": {
"enum": [
"approved",
"superseded",
"retired",
"invalidated"
]
},
"fingerprints": {
"type": "object",
"additionalProperties": false,
"required": [
"query",
"dataset",
"filter",
"layout"
],
"properties": {
"query": {
"$ref": "#/$defs/sha256"
},
"dataset": {
"$ref": "#/$defs/sha256"
},
"filter": {
"$ref": "#/$defs/sha256"
},
"layout": {
"$ref": "#/$defs/sha256"
}
}
},
"provenance": {
"type": "object",
"additionalProperties": false,
"required": [
"environment",
"actor",
"source_artifact_id",
"source_response_hash",
"execution_principal_fingerprint",
"rls_context_hash",
"normalization_version"
],
"properties": {
"environment": {
"type": "string",
"minLength": 1
},
"actor": {
"type": "string",
"minLength": 1
},
"source_artifact_id": {
"type": "string",
"format": "uuid"
},
"source_response_hash": {
"type": "string",
"pattern": "^[a-f0-9]{64}$"
},
"execution_principal_fingerprint": {
"type": "string",
"pattern": "^[a-f0-9]{64}$"
},
"rls_context_hash": {
"type": "string",
"pattern": "^[a-f0-9]{64}$"
},
"normalization_version": {
"type": "string",
"minLength": 1
}
}
},
"approval": {
"$ref": "#/$defs/approval"
},
"immutability": {
"type": [
"object",
"null"
],
"additionalProperties": false,
"properties": {
"enabled": {
"type": "boolean"
},
"period": {
"type": "string"
},
"period_closed_at": {
"type": [
"string",
"null"
],
"format": "date-time",
"description": "ISO-8601 timestamp when the period was formally closed. When null, period is open and immutability checks are skipped."
},
"frozen_at": {
"type": "string",
"format": "date-time"
},
"source_response_hash": {
"$ref": "#/$defs/sha256",
"description": "Authoritative SHA-256 of normalized Superset response/artifact bytes at closure time. Computed server-side."
},
"policy": {
"enum": [
"alert",
"block_publish",
"require_investigation"
]
}
}
},
"created_at": {
"type": "string",
"format": "date-time"
},
"updated_at": {
"type": [
"string",
"null"
],
"format": "date-time"
}
}
},
"entry": {
"type": "object",
"additionalProperties": false,
"required": [
"schema_version",
"baseline_id",
"dashboard_id",
"release_version",
"release_commit_hash",
"result_key",
"normalized_filters",
"expected",
"policy",
"status",
"source_response_hash",
"captured_at",
"provenance",
"created_at"
],
"properties": {
"schema_version": {
"const": 1
},
"baseline_id": {
"type": "string",
"format": "uuid"
},
"dashboard_id": {
"type": "integer",
"minimum": 1
},
"release_version": {
"type": "string",
"pattern": "^v\\d+\\.\\d+\\.\\d+(-[a-z0-9.]+)?$"
},
"release_commit_hash": {
"type": "string",
"pattern": "^[a-f0-9]{40}$"
},
"chart_id": {
"type": [
"integer",
"null"
]
},
"dataset_id": {
"type": [
"integer",
"null"
]
},
"result_key": {
"type": "string",
"minLength": 1
},
"label": {
"type": "string"
},
"normalized_filters": {
"$ref": "#/$defs/filters"
},
"expected": {
"$ref": "#/$defs/value"
},
"source_response_hash": {
"$ref": "#/$defs/sha256"
},
"content_hash": {
"type": [
"string",
"null"
],
"pattern": "^[a-f0-9]{64}$",
"description": "Dashboard-level content_hash at capture time for FR-013 inheritance. Null when not captured."
},
"captured_at": {
"type": "string",
"format": "date-time"
},
"policy": {
"$ref": "#/$defs/policy"
},
"status": {
"enum": [
"approved",
"superseded",
"retired",
"invalidated"
]
},
"immutability": {
"type": [
"object",
"null"
],
"additionalProperties": false,
"properties": {
"enabled": {
"type": "boolean"
},
"period": {
"type": "string"
},
"period_closed_at": {
"type": [
"string",
"null"
],
"format": "date-time",
"description": "ISO-8601 timestamp when the period was formally closed. When null, period is open and immutability checks are skipped."
},
"frozen_at": {
"type": "string",
"format": "date-time"
},
"source_response_hash": {
"$ref": "#/$defs/sha256",
"description": "Authoritative SHA-256 of normalized Superset response/artifact bytes at closure time. Computed server-side."
},
"policy": {
"enum": [
"alert",
"block_publish",
"require_investigation"
]
}
}
},
"provenance": {
"type": "object",
"additionalProperties": false,
"required": [
"environment",
"actor",
"source_artifact_id",
"source_response_hash",
"execution_principal_fingerprint",
"rls_context_hash",
"normalization_version"
],
"properties": {
"environment": {
"type": "string",
"minLength": 1
},
"actor": {
"type": "string",
"minLength": 1
},
"source_artifact_id": {
"type": "string",
"format": "uuid"
},
"source_response_hash": {
"type": "string",
"pattern": "^[a-f0-9]{64}$"
},
"execution_principal_fingerprint": {
"type": "string",
"pattern": "^[a-f0-9]{64}$"
},
"rls_context_hash": {
"type": "string",
"pattern": "^[a-f0-9]{64}$"
},
"normalization_version": {
"type": "string",
"minLength": 1
}
}
},
"created_at": {
"type": "string",
"format": "date-time"
},
"updated_at": {
"type": [
"string",
"null"
],
"format": "date-time"
}
},
"oneOf": [
{
"required": [
"chart_id"
],
"properties": {
"chart_id": {
"type": "integer"
}
}
},
{
"required": [
"dataset_id"
],
"properties": {
"dataset_id": {
"type": "integer"
}
}
}
]
},
"approval": {
"type": "object",
"additionalProperties": false,
"required": [
"by",
"at"
],
"properties": {
"by": {
"type": "string",
"minLength": 1
},
"at": {
"type": "string",
"format": "date-time"
}
}
}
},
"x-invariants": [
"entry wrapper baseline_id/status must equal entry.baseline_id/status",
"baseline_revision_id and approved coordinate_hash are unique",
"catalog digest is server-computed; caller digest is never authority",
"publication receipt binds catalog/release/remote commit",
"publication truth table: materialized/publish_pending → no commit and no receipt; committed → commit, no receipt; published → commit plus receipt; publish_failed → error_code required, receipt never, commit identity retained for reconcile"
]
}

View File

@@ -0,0 +1,435 @@
{
"$schema": "https://json-schema.org/draft/2020-12/schema",
"$id": "https://superset-tools.local/specs/037-superset-baseline-engine/contracts/comparison-types.schema.json",
"title": "Closed comparison types v1",
"description": "Implemented=false. Canonical decimals are strings; tolerances are nonnegative by pattern; range minimum<=maximum and table row width equal to column count are cross-field invariants enforced by 044 validate_contract_refresh.",
"$defs": {
"value": {
"oneOf": [
{
"type": "object",
"additionalProperties": false,
"required": [
"kind",
"canonical"
],
"properties": {
"kind": {
"const": "null"
},
"canonical": {
"type": "null"
}
}
},
{
"type": "object",
"additionalProperties": false,
"required": [
"kind",
"canonical"
],
"properties": {
"kind": {
"const": "boolean"
},
"canonical": {
"type": "boolean"
}
}
},
{
"type": "object",
"additionalProperties": false,
"required": [
"kind",
"canonical"
],
"properties": {
"kind": {
"const": "integer"
},
"canonical": {
"type": "integer"
}
}
},
{
"type": "object",
"additionalProperties": false,
"required": [
"kind",
"canonical"
],
"properties": {
"kind": {
"const": "decimal"
},
"canonical": {
"type": "string",
"pattern": "^-?(0|[1-9][0-9]*)(\\.[0-9]+)?$"
}
}
},
{
"type": "object",
"additionalProperties": false,
"required": [
"kind",
"canonical"
],
"properties": {
"kind": {
"const": "percent"
},
"canonical": {
"type": "string",
"pattern": "^-?(0|[1-9][0-9]*)(\\.[0-9]+)?$"
}
}
},
{
"type": "object",
"additionalProperties": false,
"required": [
"kind",
"canonical"
],
"properties": {
"kind": {
"const": "string"
},
"canonical": {
"type": "string",
"maxLength": 10000
}
}
},
{
"type": "object",
"additionalProperties": false,
"required": [
"kind",
"canonical"
],
"properties": {
"kind": {
"const": "date"
},
"canonical": {
"type": "string",
"format": "date"
}
}
},
{
"type": "object",
"additionalProperties": false,
"required": [
"kind",
"canonical"
],
"properties": {
"kind": {
"const": "datetime"
},
"canonical": {
"type": "string",
"format": "date-time"
}
}
},
{
"type": "object",
"additionalProperties": false,
"required": [
"kind",
"canonical"
],
"properties": {
"kind": {
"const": "table"
},
"canonical": {
"type": "object",
"additionalProperties": false,
"required": [
"columns",
"rows"
],
"properties": {
"columns": {
"type": "array",
"minItems": 1,
"maxItems": 100,
"uniqueItems": true,
"items": {
"type": "string",
"minLength": 1,
"maxLength": 256
}
},
"rows": {
"type": "array",
"maxItems": 10000,
"items": {
"type": "array",
"maxItems": 100,
"items": {
"type": [
"string",
"number",
"boolean",
"null"
]
}
}
}
}
}
}
}
]
},
"metricPolicy": {
"oneOf": [
{
"type": "object",
"additionalProperties": false,
"required": [
"type"
],
"properties": {
"type": {
"const": "exact"
}
}
},
{
"type": "object",
"additionalProperties": false,
"required": [
"type",
"tolerance"
],
"properties": {
"type": {
"const": "absolute_tolerance"
},
"tolerance": {
"type": "string",
"pattern": "^(0|[1-9][0-9]*)(\\.[0-9]+)?$",
"description": "Nonnegative canonical decimal; negative tolerance is rejected at schema level."
}
}
},
{
"type": "object",
"additionalProperties": false,
"required": [
"type",
"tolerance",
"zero_expected"
],
"properties": {
"type": {
"const": "relative_tolerance"
},
"tolerance": {
"type": "string",
"pattern": "^(0|[1-9][0-9]*)(\\.[0-9]+)?$",
"description": "Nonnegative canonical decimal; negative tolerance is rejected at schema level."
},
"zero_expected": {
"enum": [
"absolute",
"reject"
]
}
}
},
{
"type": "object",
"additionalProperties": false,
"required": [
"type",
"minimum",
"maximum"
],
"properties": {
"type": {
"const": "range"
},
"minimum": {
"type": "string",
"pattern": "^-?(0|[1-9][0-9]*)(\\.[0-9]+)?$"
},
"maximum": {
"type": "string",
"pattern": "^-?(0|[1-9][0-9]*)(\\.[0-9]+)?$"
}
}
},
{
"type": "object",
"additionalProperties": false,
"required": [
"type",
"key_columns",
"order_sensitive"
],
"properties": {
"type": {
"const": "row_set"
},
"key_columns": {
"type": "array",
"minItems": 1,
"maxItems": 100,
"uniqueItems": true,
"items": {
"type": "string",
"minLength": 1,
"maxLength": 256
}
},
"order_sensitive": {
"type": "boolean"
}
}
}
]
},
"visualPolicy": {
"oneOf": [
{
"type": "object",
"additionalProperties": false,
"required": [
"type"
],
"properties": {
"type": {
"const": "exact"
}
}
},
{
"type": "object",
"additionalProperties": false,
"required": [
"type",
"ssim_min",
"pixel_diff_threshold"
],
"properties": {
"type": {
"const": "perceptual"
},
"ssim_min": {
"type": "number",
"minimum": 0,
"maximum": 1
},
"pixel_diff_threshold": {
"type": "number",
"minimum": 0,
"maximum": 1
}
}
}
]
},
"metricDelta": {
"type": "object",
"additionalProperties": false,
"required": [
"equal",
"absolute",
"relative",
"missing_rows",
"extra_rows"
],
"properties": {
"equal": {
"type": "boolean"
},
"absolute": {
"anyOf": [
{
"type": "string",
"pattern": "^-?(0|[1-9][0-9]*)(\\.[0-9]+)?$"
},
{
"type": "null"
}
]
},
"relative": {
"anyOf": [
{
"type": "string",
"pattern": "^-?(0|[1-9][0-9]*)(\\.[0-9]+)?$"
},
{
"type": "null"
}
]
},
"missing_rows": {
"type": "integer",
"minimum": 0
},
"extra_rows": {
"type": "integer",
"minimum": 0
}
}
},
"visualValue": {
"type": "object",
"additionalProperties": false,
"required": [
"sha256",
"width",
"height"
],
"properties": {
"sha256": {
"type": "string",
"pattern": "^[a-f0-9]{64}$"
},
"width": {
"type": "integer",
"minimum": 1,
"maximum": 16384
},
"height": {
"type": "integer",
"minimum": 1,
"maximum": 16384
}
}
},
"visualDelta": {
"type": "object",
"additionalProperties": false,
"required": [
"ssim",
"pixel_diff_ratio"
],
"properties": {
"ssim": {
"type": "number",
"minimum": -1,
"maximum": 1
},
"pixel_diff_ratio": {
"type": "number",
"minimum": 0,
"maximum": 1
}
}
}
}
}

View File

@@ -197,4 +197,12 @@ entries: []
Entries sort by chart/dataset identity, result_key, filters_hash, baseline_id. Writer uses atomic temp-file replace inside the resolved repository and only through a consumed 036 gate.
## Production record contract — 2026-09-08 (AGBASE-FR-014)
CatalogRevision/BaselineRevision and BaselineSelectionPin use contracts/catalog-revision.schema.json and contracts/baseline-pin.schema.json. Unique approved coordinate includes dataset/metric/filter/time/RLS/security/normalization and visual capture profile. Baseline ID is stable; revision IDs and catalog digest are immutable. PublicationReceipt binds expected branch head, commit, push state, error and reconciliation key.
Normative detail: [CatalogRevision and baseline lifecycle](contracts/catalog-lifecycle.md). New fields, CAS transitions and cross-record integrity checks are implemented=false until [production tasks](tasks.md) and [traceability](traceability.md) close with executable evidence. Existing shorter field lists are legacy compatibility projections, not permission to omit the production identity fields.
Frontend contains no agent interaction state, prompt, proposal-generation or workspace/start controls. Human manual editor/review state and read-only evaluation/result artifacts are separate from external MCP authoring. Approved performance baseline is outside scope.
#endregion SupersetBaselineEngine.DataModel

View File

@@ -84,3 +84,14 @@ No exception planned. Table normalization must enforce row and payload limits ra
- T081 — Add `GET /verification/history` + `GET /verification/{run_id}` в `api/routes/dashboard_testing/verification.py`, совместимые с `VerificationRunDTO[]`/`VerificationRunDTO`.
**Exit rule**: 039 pipeline views (AGUI-FR-014..016) не могут отображать живые VerificationRun до мерджа T080/T081.
## Production delivery plan — 2026-09-08 (AGBASE-FR-014)
Status: specified, implemented=false; historical unit/prototype/transport results are not current production acceptance.
1. Pin [CatalogRevision and baseline lifecycle](contracts/catalog-lifecycle.md) and [data model](data-model.md); write negative fixtures before runtime changes.
2. Implement existing domain boundaries for: CatalogRevision/BaselineRevision and BaselineSelectionPin use contracts/catalog-revision.schema.json and contracts/baseline-pin.schema.json. Unique approved coordinate includes dataset/metric/filter/time/RLS/security/normalization and visual capture profile. Baseline ID is stable; revision IDs and catalog digest are immutable. PublicationReceipt binds expected branch head, commit, push state, error and reconciliation key.
3. Execute [tasks](tasks.md) T082, T083, T084 and retain reproducible evidence in [traceability](traceability.md), then close [checklist](checklists/requirements.md) individually.
4. Run cross-spec canary only after 037 catalog publication, 038 chain and 044 provider/content/policy gates; use 046 versioned cost/load/SLO limits. Disable admission on rollback, retain pins/receipts/holds; no fallback to stale catalog or synthetic PASS.
Frontend agent prompts, chat, assistant editing, proposal-generation, workspace/start/handoff actions are prohibited. Only external MCP clients interact with agents; frontend provides ordinary manual CRUD/editor, human review/approval and read-only monitoring/evidence. Runtime removal tasks are not closed by this document. Optional approved performance baseline is outside this refresh.

View File

@@ -1,5 +1,18 @@
# Quickstart: Superset Baseline Engine
> **Refresh 2026-09-08 (production contract):** CatalogRevision/CAS/rebaseline/publication (AGBASE-FR-014) are
> normative in `contracts/catalog-lifecycle.md`, `catalog-revision.schema.json`, `baseline-pin.schema.json`,
> `comparison-types.schema.json` — implemented=false / acceptance OPEN. Approve-through-036-gate now records a
> materialized CatalogRevision; publication is an explicit separate operation with the closed truth table
> (materialized→publish_pending→committed→published; publish_failed retains commit identity, never a receipt).
> Working-tree YAML is NOT a published baseline. Every ScenarioRun consumer is bound by a server-resolved
> BaselineSelectionPin. Offline executable check (schemas + truth table + pins, current result
> `PASS: 11 schemas`, `PASS: 5 positive fixture checks; 49 hardcoded negative regressions rejected`):
>
> ```bash
> python specs/044-dashboard-scenario-execution/prototype/validate_contract_refresh.py specs/044-dashboard-scenario-execution/fixtures/production-contract-refresh.json
> ```
## Prerequisite
Complete 036 through approval consume tests. Use checked-in fixtures first; Docker is required only for Superset 4.1.2 integration.
@@ -15,6 +28,11 @@ python -m pytest tests/integration/test_dashboard_testing_superset.py --run-inte
## Independent Smoke
> Steps 8–10 below reflect the pre-refresh gate/YAML approval path and are historical. Under the refreshed
> lifecycle, approval produces a materialized CatalogRevision, and the replay/mutation guard is the CAS on
> catalog_revision_id/cas_version plus the explicit publication receipt — validate through the offline
> contract check above until the runtime chain is implemented.
1. Inspect a fixture dashboard and snapshot the deterministic DashboardQueryModel.
2. Normalize the same date/decimal/list filters in two locale representations; hashes must match.
3. Execute a saved scalar chart via POST /api/v1/chart/data and verify source ids/hash.
@@ -38,3 +56,7 @@ python -m pytest tests/integration/test_dashboard_testing_superset.py --run-inte
## Known Gap (2026-08-07 MVP audit)
Шаги 1–10 проверяют дискретные инструменты напрямую (inspect/normalize/execute/compare/candidate/approve) — они работают. Но пайплайн-автоматизация не замкнута: deploy-хук (`deploy_to_preprod`/`release_create`/`etl_completed`) **не создаёт VerificationRun автоматически**, а GET-эндпоинты `/verification/history` и `/verification/{run_id}` отсутствуют (только POST `/verification-runs`). До T080–T081 (tasks.md Phase 10) релизный цикл не получает автоматические verification-прогоны, а 039 pipeline views не могут загрузить историю.
## Refresh Gap (2026-09-08 production audit)
CatalogRevision/CAS/явная Git-publication, BaselineSelectionPin-привязка всех потребителей и визуальный review-контур (disposition/receipt) специфицированы и офлайн-валидируемы (validate_contract_refresh), но runtime отсутствует: ss-prod E2E зафиксировал отсутствие зарегистрированных browser/screenshot-провайдеров и незакрытую цепочку capture→artifact→catalog. Все строки AGBASE-FR-014 остаются implemented=false / acceptance OPEN; успешные исторические тесты дискретных инструментов не закрывают production-гейты.

View File

@@ -12,7 +12,7 @@
@SEMANTICS: spec, requirements, feature, superset, baseline, chart-data, dataset, filters, normalization, dashboard-testing
**Feature Branch**: `037-superset-baseline-engine`
**Created**: 2026-07-07 | **Status**: Ready for Implementation
**Created**: 2026-07-07 | **Status**: Partially implemented; production acceptance OPEN (refresh 2026-09-08)
**Input**: "Create a Superset-native query and baseline engine for dashboard testing. The system must inspect dashboard query models, map dashboard native filters into Superset chart or dataset query context, execute Superset-side chart or dataset queries without direct SQL, normalize returned metric and table values, compare them with approved baseline catalog entries, and create draft baseline candidates requiring human approval."
## User Scenarios
@@ -84,7 +84,7 @@
- **AGBASE-FR-002**: The engine MUST execute Superset-native chart/dataset query contexts without accepting arbitrary SQL text from the agent.
- **AGBASE-FR-003**: Dashboard native filters MUST be normalized into query filters with target columns, operators, values, scopes, and a deterministic filter hash.
- **AGBASE-FR-004**: Superset result normalization MUST support scalar metrics, big-number charts, table rows, dates, decimals, percentages, empty values, and raw/source metadata.
- **AGBASE-FR-005**: Baseline catalog entries MUST be pinned to a specific `DashboardRelease` via `release_version` and `release_commit_hash`. The release identity replaces per-field fingerprint tracking. Baseline without release pinning is invalid.
- **AGBASE-FR-005**: Baseline catalog entries MUST be pinned to a specific `DashboardRelease` via `release_version` and `release_commit_hash`. Release identity does not remove visual query/dataset/filter/layout staleness checks or capture-profile provenance. Baseline without release pinning is invalid.
- **AGBASE-FR-006**: Each baseline entry MUST record a `source_response_hash` (SHA-256 of the Superset API response at the time the expected value was captured) and `captured_at` (ISO-8601 timestamp). These enable immutability violation detection independent of metric value comparison.
- **AGBASE-FR-007**: For closed-period entries (`immutability.enabled=true`), the system MUST detect immutability violations: when `source_response_hash` changes for the same filters, the status MUST be `immutability_violation` (CRITICAL severity), NOT a stale warning. Automated baseline updates for closed-period entries are forbidden.
- **AGBASE-FR-008**: Comparison output MUST include source, actual value, expected value, diff, status, tolerance rule, and warnings. Status enum MUST include `immutability_violation` as a distinct, critical category separate from `stale_baseline` and `stale_visual_baseline`.
@@ -134,6 +134,14 @@
- Baseline tools (`capture_baseline_candidate`, approval lifecycle, `create_verification_run`) are exposed 1:1 through the MCP catalog (050 Phase 1 parity); no behavioral change to engine semantics.
- The vendored `research/mcp-superset` server is explicitly NOT adopted as an alternative surface: it bypasses this feature's server-side hash computation, release pinning and immutability rules.
**Status (2026-09-02): done** — реализовано в рамках 050: инструменты и гейты (`specs/050-mcp-interface/tasks.md` T012–T028 [x]), handoff-поверхность (050 T030–T033), демонтаж чата и сервиса `agent/` (050 T040–T041, чекпоинты `specs/WORKSTATE-043-047.md`).
**Historical MCP transport/decommission status (2026-09-02): reported done; not production-readiness evidence** — реализовано в рамках 050: инструменты и гейты (`specs/050-mcp-interface/tasks.md` T012–T028 [x]), handoff-поверхность (050 T030–T033), демонтаж чата и сервиса `agent/` (050 T040–T041, чекпоинты `specs/WORKSTATE-043-047.md`).
## Production contract refresh — 2026-09-08
**Frontend boundary (user decision 2026-09-08)**: All agent interaction is external MCP only. Product frontend MUST NOT contain agent chat, prompt/request textarea, assistant editing, typical-operation-to-agent selector, proposal-generation, agent workspace/start or handoff controls/routes. Ordinary manual CRUD/editor, human approval/review, monitoring and read-only evidence/evaluation are permitted. AgentEvaluationCard is read-only, with no prompt/retry-agent/provider controls. Existing agent proposal UI is runtime drift; removal/negative DOM-route-network acceptance remains OPEN in this spec-only change.
**AGBASE-FR-014 — CatalogRevision and baseline lifecycle**: CatalogRevision/BaselineRevision, coordinate uniqueness, duplicate-ID conflict, authoritative visual capture/review, CAS supersede/retire/invalidate/rebaseline and explicit Git publication MUST satisfy the lifecycle contract and schemas. Production baselines MUST be published and retain required source/image bytes; materialized working-tree YAML is insufficient. BaselineSelectionPin binds every ScenarioRun consumer.
Normative contract: [CatalogRevision and baseline lifecycle](contracts/catalog-lifecycle.md). New requirements are specified, **implemented=false / acceptance OPEN** until executable evidence closes the linked tasks/checklist/traceability rows. Historical local tests and the manual inconclusive ss-prod run do not prove browser/capture/baseline/LLM production readiness. The refresh scope is the audited P0/P1/P2 agentic E2E and baseline gaps; an approved ExecutionPerformanceBaseline is not introduced.
#endregion SupersetBaselineEngine.Spec

View File

@@ -70,7 +70,7 @@
- [x] T041 [P] Implement backend/src/services/dashboard_testing/visual_baseline.py: VisualComparisonPolicy (exact + perceptual), layout fingerprint computation, stale_visual_baseline detection.
- [x] T042 [P] Wire visual baseline loading into BaselineEngine.Catalog.Load; extend catalog YAML to support visual entries.
- [x] T043 Implement BaselineEngine.Visual.Compare: digest comparison + perceptual SSIM path; return stale_visual_baseline when layout fingerprint mismatches.
- [x] T044 [P] Implement BaselineEngine.Visual.Candidate: create draft visual candidate from reviewed screenshot artifact with mandatory human disposition.
- [ ] T044 [P] Implement BaselineEngine.Visual.Candidate: create draft visual candidate from reviewed screenshot artifact with mandatory human disposition.
- [x] T045 Add visual baseline approval flow reusing 036 gate; verify approval writes visual entry atomically alongside metric entries.
- [x] T046 Write visual baseline golden fixtures under specs/037-superset-baseline-engine/fixtures/visual/.
- [x] T047 Audit: visual baselines never use metric policies; metric baselines never use visual policies; cross-kind comparison returns inconclusive.
@@ -133,4 +133,16 @@ T001–T005 → US1 → US2 → US3; US4 depends on US3 and completed 036. API i
T001–T005 → US1 → US2 → US3; US4 depends on US3 and completed 036. API integration follows all domain contracts. Phase 7 depends on completed 036 Phase 8 (screenshot evidence artifacts). Phase 10 (T080–T081) closes the pipeline-automation and read-API gap found in the 2026-08-07 audit; it must land before 039 pipeline views can render live VerificationRun data.
## Production readiness — 2026-09-08 (AGBASE-FR-014)
Historical [x] rows above retain only their dated local/transport evidence; they do not prove current production readiness. Reopened rows were contradicted by the audited gaps. Removed frontend/agent paths are historical, not implementation prerequisites. New acceptance is **implemented=false / OPEN**.
Contract: [CatalogRevision and baseline lifecycle](contracts/catalog-lifecycle.md).
- [ ] T082 [P0/P1/P2] Duplicate IDs/approved coordinates reject; stale If-Match conflicts; two concurrent approvals produce one catalog head. Implement at the existing 037 domain boundary; verify with independent hardcoded fixtures and retain command/evidence references in traceability.md.
- [ ] T083 [P0/P1/P2] Capture bytes/profile/review are server-owned; forged content refs/hashes/approved booleans cannot approve; required raw bytes remain held. Implement at the existing 037 domain boundary; verify with independent hardcoded fixtures and retain command/evidence references in traceability.md.
- [ ] T084 [P0/P1/P2] Rebaseline supersedes via CAS; retire/invalidate preserve historical pins; explicit Git commit/push failure returns publish_failed and retry reconciles without duplicate revision. Implement at the existing 037 domain boundary; verify with independent hardcoded fixtures and retain command/evidence references in traceability.md.
Frontend boundary for this package: manual CRUD/editor, human review/approval, monitoring and read-only evidence/evaluation only; all agent interaction is external MCP. No agent chat/prompt/assistant editing/proposal generation/workspace/start/handoff controls. Runtime removal is OPEN, not performed by this spec refresh. Optional approved performance baseline is outside scope.
#endregion SupersetBaselineEngine.Tasks

View File

@@ -31,4 +31,17 @@ The `trigger` field on VerificationRun gains the additive value `dataset_updated
| AGBASE-FR-012 context (pipeline triggers) | BaselineEngine.Verification.Orchestrator | T080 | deploy fixture release → VerificationRun with trigger + outcomes |
| AGUI-FR-015/016 read API | Api.DashboardTesting.VerificationRuns | T081 | history ordering/filters, detail payload, OpenAPI drift |
## Production acceptance traceability — 2026-09-08
Historical rows above identify prior tests/code only; removed agent UI paths are retired. The following audited gates are **implemented=false / OPEN**, independent of local suite totals.
| Requirement | Domain contract / DTO | Task | Falsifiable acceptance | State |
|---|---|---|---|---|
| AGBASE-FR-014 | [CatalogRevision and baseline lifecycle](contracts/catalog-lifecycle.md); [data model](data-model.md) | [T082](tasks.md) | Duplicate IDs/approved coordinates reject; stale If-Match conflicts; two concurrent approvals produce one catalog head. | OPEN |
| AGBASE-FR-014 | [CatalogRevision and baseline lifecycle](contracts/catalog-lifecycle.md); [data model](data-model.md) | [T083](tasks.md) | Capture bytes/profile/review are server-owned; forged content refs/hashes/approved booleans cannot approve; required raw bytes remain held. | OPEN |
| AGBASE-FR-014 | [CatalogRevision and baseline lifecycle](contracts/catalog-lifecycle.md); [data model](data-model.md) | [T084](tasks.md) | Rebaseline supersedes via CAS; retire/invalidate preserve historical pins; explicit Git commit/push failure returns publish_failed and retry reconciles without duplicate revision. | OPEN |
| AGBASE-FR-014; external-MCP-only UI | manual editor/review; read-only evidence | [production tasks](tasks.md) | No frontend agent prompt/chat/assistant editing/proposal generation/workspace/start/handoff routes or requests; human approval remains usable. | OPEN |
Sources: [production gap](../../docs/reports/ss-prod-agentic-e2e-production-gap-2026-09-08.md), [coverage gap](../../docs/reports/ss-prod-agentic-e2e-spec-coverage-2026-09-08.md), [baseline gap](../../docs/reports/ss-prod-agentic-e2e-baseline-gap-2026-09-08.md). Spec schema/static checks prove contract structure only; live canary/runtime closure and optional approved performance baseline are not claimed.
#endregion SupersetBaselineEngine.Traceability

View File

@@ -0,0 +1,184 @@
{
"$schema": "https://json-schema.org/draft/2020-12/schema",
"title": "AgentEvaluationSpec v1",
"description": "Normative ActionRegistry graph extension, implemented=false; no dependency on deferred five-program IR.",
"type": "object",
"additionalProperties": false,
"required": [
"schema_version",
"spec_id",
"provider_id",
"provider_version",
"model_id",
"model_version",
"prompt_template_id",
"prompt_template_version",
"prompt_template_hash",
"evidence_refs",
"comparison_refs",
"tool_allowlist",
"output_schema",
"decision_policy",
"limits",
"trust_policy_hash",
"criteria"
],
"properties": {
"schema_version": {
"const": 1
},
"spec_id": {
"type": "string",
"format": "uuid"
},
"provider_id": {
"type": "string",
"minLength": 1
},
"provider_version": {
"type": "string",
"minLength": 1
},
"model_id": {
"type": "string",
"minLength": 1
},
"model_version": {
"type": "string",
"minLength": 1
},
"prompt_template_id": {
"type": "string",
"minLength": 1
},
"prompt_template_version": {
"type": "string",
"minLength": 1
},
"prompt_template_hash": {
"type": "string",
"pattern": "^[a-f0-9]{64}$"
},
"evidence_refs": {
"type": "array",
"items": {
"type": "string",
"minLength": 1
},
"minItems": 1,
"maxItems": 100
},
"comparison_refs": {
"type": "array",
"items": {
"type": "string",
"minLength": 1
},
"minItems": 1,
"maxItems": 100
},
"tool_allowlist": {
"type": "array",
"maxItems": 0
},
"output_schema": {
"const": "agent-evaluation.schema.json"
},
"decision_policy": {
"$ref": "decision-policy.schema.json"
},
"limits": {
"type": "object",
"additionalProperties": false,
"required": [
"timeout_ms",
"max_images",
"max_input_tokens",
"max_output_tokens",
"max_cost",
"currency"
],
"properties": {
"timeout_ms": {
"type": "integer",
"minimum": 1,
"maximum": 60000,
"default": 60000
},
"max_images": {
"type": "integer",
"minimum": 1,
"maximum": 50
},
"max_input_tokens": {
"type": "integer",
"minimum": 1
},
"max_output_tokens": {
"type": "integer",
"minimum": 1,
"maximum": 8192
},
"max_cost": {
"type": "string",
"pattern": "^[0-9]+(\\.[0-9]+)?$"
},
"currency": {
"type": "string",
"minLength": 1
}
}
},
"trust_policy_hash": {
"type": "string",
"pattern": "^[a-f0-9]{64}$"
},
"criteria": {
"type": "array",
"minItems": 1,
"maxItems": 100,
"items": {
"type": "object",
"additionalProperties": false,
"required": [
"criterion_id",
"criterion_kind",
"description",
"comparison_id"
],
"properties": {
"criterion_id": {
"type": "string",
"minLength": 1,
"maxLength": 128
},
"criterion_kind": {
"enum": [
"semantic",
"deterministic_comparison"
]
},
"description": {
"type": "string",
"minLength": 1,
"maxLength": 2000
},
"comparison_id": {
"type": [
"string",
"null"
],
"maxLength": 128
}
}
}
}
},
"$id": "https://superset-tools.local/specs/038-dashboard-scenario-model/contracts/agent-evaluation-spec.schema.json",
"x-invariants": [
"criterion_id values are unique within criteria",
"closed criterion truth table: deterministic_comparison criteria bind a non-null comparison_id that is a member of comparison_refs; semantic criteria bind comparison_id null and never claim deterministic authority",
"every evaluation finding criterion_id must exist in criteria with matching criterion_kind",
"bounded declared access: evaluation comparison_ids are a subset of comparison_refs and input_manifest artifact ids are a subset of evidence_refs"
]
}

View File

@@ -0,0 +1,411 @@
{
"$schema": "https://json-schema.org/draft/2020-12/schema",
"title": "Immutable AgentEvaluation v1",
"description": "Append-only 044 record, schema owned by 038. Implemented=false; response failures cannot be represented as successful empty findings.",
"type": "object",
"additionalProperties": false,
"required": [
"schema_version",
"evaluation_id",
"scenario_run_id",
"logical_step_id",
"attempt",
"operation_id",
"evaluation_spec_hash",
"provider_id",
"provider_version",
"model_id",
"model_version",
"prompt_template_id",
"prompt_template_version",
"prompt_template_hash",
"output_schema_hash",
"input_manifest_hash",
"input_manifest",
"baseline_pin",
"comparison_ids",
"status",
"verdict",
"confidence",
"findings",
"reason_codes",
"raw_response_artifact_ref",
"raw_response_sha256",
"trust_policy_hash",
"usage",
"started_at",
"finished_at"
],
"properties": {
"schema_version": {
"const": 1
},
"evaluation_id": {
"type": "string",
"format": "uuid"
},
"scenario_run_id": {
"type": "string",
"format": "uuid"
},
"logical_step_id": {
"type": "string",
"format": "uuid"
},
"attempt": {
"type": "integer",
"minimum": 1
},
"operation_id": {
"type": "string",
"format": "uuid"
},
"evaluation_spec_hash": {
"type": "string",
"pattern": "^[a-f0-9]{64}$"
},
"provider_id": {
"type": "string",
"minLength": 1
},
"provider_version": {
"type": "string",
"minLength": 1
},
"model_id": {
"type": "string",
"minLength": 1
},
"model_version": {
"type": "string",
"minLength": 1
},
"prompt_template_id": {
"type": "string",
"minLength": 1
},
"prompt_template_version": {
"type": "string",
"minLength": 1
},
"prompt_template_hash": {
"type": "string",
"pattern": "^[a-f0-9]{64}$"
},
"output_schema_hash": {
"type": "string",
"pattern": "^[a-f0-9]{64}$"
},
"input_manifest_hash": {
"type": "string",
"pattern": "^[a-f0-9]{64}$"
},
"input_manifest": {
"type": "array",
"items": {
"type": "object",
"additionalProperties": false,
"required": [
"artifact_id",
"sha256",
"content_type",
"byte_length",
"role"
],
"properties": {
"artifact_id": {
"type": "string",
"format": "uuid"
},
"sha256": {
"type": "string",
"pattern": "^[a-f0-9]{64}$"
},
"content_type": {
"enum": [
"image/jpeg",
"image/png",
"image/webp",
"application/json"
]
},
"byte_length": {
"type": "integer",
"minimum": 1
},
"role": {
"enum": [
"actual",
"baseline",
"comparison",
"context"
]
}
}
},
"minItems": 1,
"maxItems": 100
},
"baseline_pin": {
"$ref": "../../037-superset-baseline-engine/contracts/baseline-pin.schema.json"
},
"comparison_ids": {
"type": "array",
"items": {
"type": "string",
"format": "uuid"
},
"minItems": 1,
"maxItems": 100
},
"status": {
"enum": [
"succeeded",
"provider_error",
"parser_error",
"budget_exceeded",
"cancelled",
"timed_out"
]
},
"verdict": {
"enum": [
"pass",
"fail",
"inconclusive"
]
},
"confidence": {
"type": "number",
"minimum": 0,
"maximum": 1
},
"findings": {
"type": "array",
"items": {
"type": "object",
"additionalProperties": false,
"required": [
"finding_id",
"severity",
"message",
"evidence_artifact_ids",
"region",
"criterion_id",
"criterion_kind"
],
"properties": {
"finding_id": {
"type": "string",
"minLength": 1
},
"severity": {
"enum": [
"info",
"warning",
"error",
"critical"
]
},
"message": {
"type": "string",
"minLength": 1,
"maxLength": 2000
},
"evidence_artifact_ids": {
"type": "array",
"items": {
"type": "string",
"format": "uuid"
},
"minItems": 1,
"maxItems": 100
},
"region": {
"anyOf": [
{
"type": "object",
"additionalProperties": false,
"required": [
"x",
"y",
"width",
"height"
],
"properties": {
"x": {
"type": "number",
"minimum": 0,
"maximum": 1
},
"y": {
"type": "number",
"minimum": 0,
"maximum": 1
},
"width": {
"type": "number",
"exclusiveMinimum": 0,
"maximum": 1
},
"height": {
"type": "number",
"exclusiveMinimum": 0,
"maximum": 1
}
}
},
{
"type": "null"
}
]
},
"criterion_id": {
"type": "string",
"minLength": 1,
"maxLength": 128
},
"criterion_kind": {
"enum": [
"semantic",
"deterministic_comparison"
]
}
}
},
"minItems": 0,
"maxItems": 100
},
"reason_codes": {
"type": "array",
"items": {
"type": "string",
"minLength": 1
},
"minItems": 0,
"maxItems": 100
},
"raw_response_artifact_ref": {
"type": [
"string",
"null"
]
},
"raw_response_sha256": {
"anyOf": [
{
"type": "string",
"pattern": "^[a-f0-9]{64}$"
},
{
"type": "null"
}
]
},
"trust_policy_hash": {
"type": "string",
"pattern": "^[a-f0-9]{64}$"
},
"usage": {
"type": "object",
"additionalProperties": false,
"required": [
"input_tokens",
"output_tokens",
"cost_amount",
"currency",
"pricing_version"
],
"properties": {
"input_tokens": {
"type": [
"integer",
"null"
],
"minimum": 0
},
"output_tokens": {
"type": [
"integer",
"null"
],
"minimum": 0
},
"cost_amount": {
"type": [
"string",
"null"
],
"pattern": "^[0-9]+(\\.[0-9]+)?$"
},
"currency": {
"type": [
"string",
"null"
]
},
"pricing_version": {
"type": [
"string",
"null"
]
}
}
},
"started_at": {
"type": "string",
"format": "date-time"
},
"finished_at": {
"type": "string",
"format": "date-time"
}
},
"allOf": [
{
"if": {
"properties": {
"status": {
"const": "succeeded"
}
}
},
"then": {
"properties": {
"raw_response_artifact_ref": {
"type": "string",
"minLength": 1
},
"raw_response_sha256": {
"type": "string",
"pattern": "^[a-f0-9]{64}$"
}
}
},
"else": {
"properties": {
"verdict": {
"const": "inconclusive"
},
"confidence": {
"const": 0
},
"findings": {
"maxItems": 0
},
"reason_codes": {
"minItems": 1
}
}
}
}
],
"$id": "https://superset-tools.local/specs/038-dashboard-scenario-model/contracts/agent-evaluation.schema.json",
"x-invariants": [
"evaluation is projected only inside its owning ScenarioRun result: scenario_run_id equals outer run_id and (run, logical_step_id, attempt, operation_id) is unique",
"every input_manifest entry resolves to run-owned result evidence at the same step/attempt with equal sha256, content_type and byte_length; a foreign or absent artifact cannot be accepted",
"raw_response_artifact_ref resolves to run-owned result evidence at the same step/attempt and raw_response_sha256 equals that artifact's sha256; succeeded requires an available/active raw response artifact",
"finding evidence_artifact_ids resolve to result evidence at the same step/attempt",
"comparison_ids resolve to result comparisons at the same step/attempt and baseline_pin equals the run's pinned BaselineSelectionPin exactly"
]
}

View File

@@ -0,0 +1,29 @@
## @{ ScenarioGraph.BrowserActionMap [C:5] [TYPE ADR]
@BRIEF Canonical browser action names and typed driver mapping for registry038.3.0.
@RATIONALE Persist canonical names in newly compiled descriptors while preserving explicit import aliases for old registries.
@REJECTED Accepting an unsupported registry action and discovering transport failure after browser I/O.
Status: normative target, implemented=false. Existing runtime038.1.0 accepts only an incomplete subset; transport support is not inferred from names. Every enabled descriptor MUST have matching input/output schemas and a canaried driver/reconciler; unsupported capability blocks compilation/admission. Aliases are compile-time migrations with new hashes, never runtime fallback.
| Canonical action | Legacy import alias | Required bounded input | Output receipt payload |
|---|---|---|---|
| open_dashboard | same | server dashboard_ref | dashboard_ref, state_hash |
| navigate_tab | none | tab_identifier | tab_identifier, state_hash |
| apply_native_filter | apply_filters, text_filter | normalized_filters_ref (text_filter requires string-valued target) | filters_hash, affected_chart_ids |
| inspect_filter_state | none | declared filter_ids | filters_hash, bounded values_ref |
| apply_table_filter | table_filter | chart_id, column_ref, typed predicate_ref | filters_hash, affected_chart_ids |
| pagination | same | chart_id, positive page, page_size<=1000 | page, row_count, table_ref |
| navigate_dashboard | same | server target_dashboard_ref | dashboard_ref, state_hash |
| extract_table | none | chart_id, max_rows<=10000, max_columns<=100 | table_artifact_ref, row_count |
| scroll_to | none | allowlisted selector_ref | state_hash |
| inspect_columns | none | chart_id | columns_artifact_ref |
| click | none | selector_ref, effect_class=read_only | state_hash |
| select_rows | none | chart_id, bounded row_keys | selection_hash |
| edit_row | row_edit | fixture_ref, affected_keys, patch_ref, mutation_contract | operation_id, before/after evidence, cleanup |
| bulk_edit | same | fixture_ref, affected_keys, patch_ref, mutation_contract | operation_id, before/after evidence, cleanup |
| download | download_xlsx | declared export_ref, format=xlsx | artifact_ref, MIME, SHA-256 |
| refresh | same | dashboard_ref | state_hash, safe_checkpoint_ref |
| wait_for_state | same | registered readiness profile, deadline | state_hash, observed_at |
All *_ref values resolve from pinned graph/server snapshots, not URL/path/code. Browser output payload additionally requires run/step/attempt/operation/binding/principal receipt. URLs, selectors, SQL and arbitrary JS are never supplied as free executable input. Canonical descriptor limit remains120s auth,30s action unless lower action timeout,3 pages,10MiB screenshot,25MiB download; navigation only allowlisted origin/dashboard and same effective principal. Filter readiness must prove actual UI state equals normalized filters and chart scopes; no sleep-only success.
Each action is supported and independently canaried or explicitly disabled with reason; product production GO still requires all applicable original capabilities. Read-only initial canary scope is open/tab/native-filter/wait/refresh/capture; it is not a reduced production feature claim. Non-PROD mutations require fixture lease/rollback and unknown-effect reconciliation; PROD mutations stay forbidden regardless of approval.
Open acceptance: enumerate entire registry through canonical model→compiler→descriptor→driver→reconciler; alias migration pins; schema mismatch and unsupported rejection before I/O; filters/tab fidelity; export ownership; safe reconstruction and cleanup.
## @} ScenarioGraph.BrowserActionMap

View File

@@ -1,6 +1,6 @@
{
"$schema": "https://json-schema.org/draft/2020-12/schema",
"$id": "https://superset-tools.local/schemas/dashboard-test-scenario-v1.json",
"$id": "https://superset-tools.local/specs/038-dashboard-scenario-model/contracts/dashboard-test-scenario.schema.json",
"title": "DashboardTestScenario",
"type": "object",
"additionalProperties": false,
@@ -15,7 +15,6 @@
"objective",
"input_fingerprints",
"parameters",
"verification_program",
"phases",
"steps",
"outputs",
@@ -69,6 +68,10 @@
},
"route": {
"type": "string"
},
"query": {
"type": "object",
"description": "Server-issued authoritative DashboardQueryModel, revalidated by context authority. Caller SQL/query_context is forbidden recursively."
}
}
},
@@ -133,7 +136,7 @@
},
"verification_program": {
"$ref": "#/$defs/verificationProgram",
"description": "Canonical immutable Verification Program IR; included in content_hash."
"description": "Deferred five-program IR; a future compatibility-family migration is required before production use. Current execution uses steps."
},
"phases": {
"type": "array",
@@ -262,28 +265,316 @@
"verificationProgram": {
"type": "object",
"additionalProperties": false,
"required": ["navigation_program", "evidence_program", "transformation_program", "assertion_program", "semantic_evaluation_program"],
"required": [
"navigation_program",
"evidence_program",
"transformation_program",
"assertion_program",
"semantic_evaluation_program"
],
"properties": {
"navigation_program": { "type": "array", "items": { "$ref": "#/$defs/programRef" } },
"evidence_program": { "type": "array", "items": { "oneOf": [{ "$ref": "#/$defs/programRef" }, { "$ref": "#/$defs/sqlEvidenceSpec" }] } },
"transformation_program": { "type": "array", "items": { "$ref": "#/$defs/transformSpec" } },
"assertion_program": { "type": "array", "items": { "$ref": "#/$defs/assertionSpec" } },
"semantic_evaluation_program": { "type": "array", "items": { "$ref": "#/$defs/agentEvaluationSpec" } }
"navigation_program": {
"type": "array",
"items": {
"$ref": "#/$defs/programRef"
}
},
"evidence_program": {
"type": "array",
"items": {
"oneOf": [
{
"$ref": "#/$defs/programRef"
},
{
"$ref": "#/$defs/sqlEvidenceSpec"
}
]
}
},
"transformation_program": {
"type": "array",
"items": {
"$ref": "#/$defs/transformSpec"
}
},
"assertion_program": {
"type": "array",
"items": {
"$ref": "#/$defs/assertionSpec"
}
},
"semantic_evaluation_program": {
"type": "array",
"items": {
"$ref": "#/$defs/agentEvaluationSpec"
}
}
}
},
"programRef": {
"type": "object",
"additionalProperties": false,
"required": [
"logical_step_id"
],
"properties": {
"logical_step_id": {
"type": "string",
"format": "uuid"
}
}
},
"programRef": { "type": "object", "additionalProperties": false, "required": ["logical_step_id"], "properties": { "logical_step_id": { "type": "string", "format": "uuid" } } },
"sqlEvidenceSpec": {
"type": "object", "additionalProperties": false,
"required": ["snippet_id", "logical_step_id", "connection_ref", "database_identity", "sql_template", "sql_hash", "parameter_definitions", "expected_output_schema", "relation_refs", "execution_limits", "generation_provenance", "validation_result"],
"type": "object",
"additionalProperties": false,
"required": [
"snippet_id",
"logical_step_id",
"connection_ref",
"database_identity",
"sql_template",
"sql_hash",
"parameter_definitions",
"expected_output_schema",
"relation_refs",
"execution_limits",
"generation_provenance",
"validation_result"
],
"properties": {
"snippet_id": { "type": "string" }, "logical_step_id": { "type": "string", "format": "uuid" }, "connection_ref": { "type": "string" }, "database_identity": { "type": "string" }, "sql_template": { "type": "string", "minLength": 1 }, "sql_hash": { "$ref": "#/$defs/sha256" },
"parameter_definitions": { "type": "array", "items": { "$ref": "#/$defs/parameter" } }, "expected_output_schema": { "type": "object" }, "relation_refs": { "type": "array", "items": { "type": "string" } }, "execution_limits": { "$ref": "#/$defs/executionLimits" }, "generation_provenance": { "type": "object" }, "validation_result": { "type": "object", "required": ["valid"] }
"snippet_id": {
"type": "string"
},
"logical_step_id": {
"type": "string",
"format": "uuid"
},
"connection_ref": {
"type": "string"
},
"database_identity": {
"type": "string"
},
"sql_template": {
"type": "string",
"minLength": 1
},
"sql_hash": {
"$ref": "#/$defs/sha256"
},
"parameter_definitions": {
"type": "array",
"items": {
"$ref": "#/$defs/parameter"
}
},
"expected_output_schema": {
"type": "object"
},
"relation_refs": {
"type": "array",
"items": {
"type": "string"
}
},
"execution_limits": {
"$ref": "#/$defs/executionLimits"
},
"generation_provenance": {
"type": "object"
},
"validation_result": {
"type": "object",
"required": [
"valid"
]
}
}
},
"executionLimits": {
"type": "object",
"additionalProperties": false,
"required": [
"timeout_ms",
"max_rows",
"max_bytes",
"max_complexity"
],
"properties": {
"timeout_ms": {
"type": "integer",
"minimum": 1
},
"max_rows": {
"type": "integer",
"minimum": 1
},
"max_bytes": {
"type": "integer",
"minimum": 1
},
"max_complexity": {
"type": "integer",
"minimum": 1
}
}
},
"transformSpec": {
"type": "object",
"additionalProperties": false,
"required": [
"logical_step_id",
"version",
"operations"
],
"properties": {
"logical_step_id": {
"type": "string",
"format": "uuid"
},
"version": {
"type": "string"
},
"operations": {
"type": "array",
"items": {
"type": "object",
"required": [
"op"
],
"properties": {
"op": {
"enum": [
"select",
"filter",
"rename",
"cast",
"join",
"group_by",
"sum",
"count",
"distinct",
"coalesce",
"normalize_string",
"normalize_date",
"difference",
"ratio",
"tolerance_compare"
]
}
}
}
}
}
},
"assertionSpec": {
"type": "object",
"additionalProperties": false,
"required": [
"logical_step_id",
"kind",
"left_ref",
"right_ref"
],
"properties": {
"logical_step_id": {
"type": "string",
"format": "uuid"
},
"kind": {
"enum": [
"numeric_equality",
"numeric_tolerance",
"row_comparison",
"column_comparison",
"aggregate_comparison",
"set_equality",
"field_mapping",
"null_fill_check",
"cross_dashboard"
]
},
"left_ref": {
"type": "string"
},
"right_ref": {
"type": "string"
},
"tolerance": {
"type": [
"number",
"null"
]
},
"field_mapping": {
"type": "object"
}
}
},
"agentEvaluationSpec": {
"type": "object",
"additionalProperties": false,
"required": [
"spec_id",
"logical_step_id",
"provider_id",
"model_id",
"prompt_template_id",
"prompt_template_version",
"prompt_template_hash",
"input_manifest",
"evidence_refs",
"tool_allowlist",
"output_schema",
"decision_policy_id"
],
"properties": {
"spec_id": {
"type": "string"
},
"logical_step_id": {
"type": "string",
"format": "uuid"
},
"provider_id": {
"type": "string"
},
"model_id": {
"type": "string"
},
"prompt_template_id": {
"type": "string"
},
"prompt_template_version": {
"type": "string"
},
"prompt_template_hash": {
"$ref": "#/$defs/sha256"
},
"input_manifest": {
"type": "object"
},
"evidence_refs": {
"type": "array",
"items": {
"type": "string"
}
},
"tool_allowlist": {
"type": "array",
"items": {
"type": "string"
}
},
"output_schema": {
"type": "object"
},
"decision_policy_id": {
"type": "string"
}
}
},
"executionLimits": { "type": "object", "additionalProperties": false, "required": ["timeout_ms", "max_rows", "max_bytes", "max_complexity"], "properties": { "timeout_ms": { "type": "integer", "minimum": 1 }, "max_rows": { "type": "integer", "minimum": 1 }, "max_bytes": { "type": "integer", "minimum": 1 }, "max_complexity": { "type": "integer", "minimum": 1 } } },
"transformSpec": { "type": "object", "additionalProperties": false, "required": ["logical_step_id", "version", "operations"], "properties": { "logical_step_id": { "type": "string", "format": "uuid" }, "version": { "type": "string" }, "operations": { "type": "array", "items": { "type": "object", "required": ["op"], "properties": { "op": { "enum": ["select", "filter", "rename", "cast", "join", "group_by", "sum", "count", "distinct", "coalesce", "normalize_string", "normalize_date", "difference", "ratio", "tolerance_compare"] } } } } } },
"assertionSpec": { "type": "object", "additionalProperties": false, "required": ["logical_step_id", "kind", "left_ref", "right_ref"], "properties": { "logical_step_id": { "type": "string", "format": "uuid" }, "kind": { "enum": ["numeric_equality", "numeric_tolerance", "row_comparison", "column_comparison", "aggregate_comparison", "set_equality", "field_mapping", "null_fill_check", "cross_dashboard"] }, "left_ref": { "type": "string" }, "right_ref": { "type": "string" }, "tolerance": { "type": ["number", "null"] }, "field_mapping": { "type": "object" } } },
"agentEvaluationSpec": { "type": "object", "additionalProperties": false, "required": ["spec_id", "logical_step_id", "provider_id", "model_id", "prompt_template_id", "prompt_template_version", "prompt_template_hash", "input_manifest", "evidence_refs", "tool_allowlist", "output_schema", "decision_policy_id"], "properties": { "spec_id": { "type": "string" }, "logical_step_id": { "type": "string", "format": "uuid" }, "provider_id": { "type": "string" }, "model_id": { "type": "string" }, "prompt_template_id": { "type": "string" }, "prompt_template_version": { "type": "string" }, "prompt_template_hash": { "$ref": "#/$defs/sha256" }, "input_manifest": { "type": "object" }, "evidence_refs": { "type": "array", "items": { "type": "string" } }, "tool_allowlist": { "type": "array", "items": { "type": "string" } }, "output_schema": { "type": "object" }, "decision_policy_id": { "type": "string" } } },
"ref": {
"type": "object",
"additionalProperties": false,
@@ -481,16 +772,55 @@
]
},
"mutation_contract": {
"type": ["object", "null"],
"type": [
"object",
"null"
],
"additionalProperties": false,
"required": ["safe_test_fixture_id", "mutation_scope", "allowed_environment", "affected_record_keys", "cleanup_policy", "side_effect_key"],
"required": [
"safe_test_fixture_id",
"mutation_scope",
"allowed_environment",
"affected_record_keys",
"cleanup_policy",
"side_effect_key"
],
"properties": {
"safe_test_fixture_id": { "type": "string" },
"mutation_scope": { "type": "string", "enum": ["controlled_test_data", "external_system"] },
"allowed_environment": { "type": "array", "minItems": 1, "items": { "type": "string" } },
"affected_record_keys": { "type": "array", "minItems": 1, "items": { "type": "string" } },
"cleanup_policy": { "type": "string", "enum": ["rollback", "reconcile", "none"] },
"side_effect_key": { "type": "string" }
"safe_test_fixture_id": {
"type": "string"
},
"mutation_scope": {
"type": "string",
"enum": [
"controlled_test_data",
"external_system"
]
},
"allowed_environment": {
"type": "array",
"minItems": 1,
"items": {
"type": "string"
}
},
"affected_record_keys": {
"type": "array",
"minItems": 1,
"items": {
"type": "string"
}
},
"cleanup_policy": {
"type": "string",
"enum": [
"rollback",
"reconcile",
"none"
]
},
"side_effect_key": {
"type": "string"
}
}
},
"capture_spec": {
@@ -512,8 +842,92 @@
"type": "null"
}
]
},
"agent_evaluation_spec": {
"$ref": "agent-evaluation-spec.schema.json"
},
"decision_policy": {
"$ref": "decision-policy.schema.json"
},
"comparison_spec": {
"type": "object",
"additionalProperties": false,
"required": [
"kind",
"actual_ref",
"baseline_ref"
],
"properties": {
"kind": {
"enum": [
"metric",
"visual"
]
},
"actual_ref": {
"$ref": "#/$defs/ref"
},
"baseline_ref": {
"type": "string",
"pattern": "^baseline\\.[A-Za-z0-9_.-]+$"
}
}
}
}
},
"allOf": [
{
"if": {
"properties": {
"tool": {
"const": "agent_evaluation"
}
}
},
"then": {
"required": [
"agent_evaluation_spec"
],
"properties": {
"action": {
"const": "evaluate_declared_spec"
}
}
}
},
{
"if": {
"properties": {
"tool": {
"const": "screenshot"
}
}
},
"then": {
"required": [
"capture_spec"
],
"properties": {
"capture_spec": {
"$ref": "#/$defs/captureSpec"
}
}
}
},
{
"if": {
"properties": {
"action": {
"const": "compare_to_baseline"
}
}
},
"then": {
"required": [
"comparison_spec"
]
}
}
]
},
"artifactPlan": {
"type": "object",
@@ -859,5 +1273,6 @@
}
}
}
}
},
"description": "Current schema_version=1 uses the ActionRegistry DAG. Five-program decomposition is optional reserved future IR, not a v1 prerequisite. Production extension registry 038.2.0 is normative, implemented=false."
}

View File

@@ -0,0 +1,36 @@
## @{ ScenarioGraph.DecisionPolicyV1 [C:5] [TYPE ADR]
@BRIEF Total baseline-semantic/1.0.0 mapping from deterministic comparison and immutable evaluation to StepOutcome.
@RATIONALE First-match precedence prevents a passing model from erasing deterministic failures; advisory evaluation must not change mandatory-check truth.
@REJECTED Dynamic waiting_human from disagreement — an automated revision must never acquire an undeclared HumanCheckpoint. A manual author may declare a separate checkpoint in the immutable graph.
Status: normative target, implemented=false. Schema: [decision-policy.schema.json](decision-policy.schema.json).
Inputs are the winning attempt, verified evidence manifest, 037 ComparisonResult, optional immutable AgentEvaluation and pinned policy. The mapper is pure; it emits no provider calls or gates.
Rows are evaluated in order. "*" means every value, including missing. "required" means evaluation_mode=required; disabled/advisory ignores model authority. These rows partition the entire input domain.
| Order | Condition (earlier rows false) | StepOutcome | Reason |
|---|---|---|---|
| 1 | Run cancelled or attempt lost/deadline elapsed | no winning outcome; historical record only | CANCELLED/LATE_RESPONSE |
| 2 | Policy/spec/version/manifest identity invalid | blocked | POLICY_INVALID |
| 3 | Required bytes absent, forbidden, corrupt, wrong owner/MIME/digest | blocked | EVIDENCE_UNAVAILABLE/EVIDENCE_CORRUPT |
| 4 | Comparison immutability_violation | failed | IMMUTABILITY_VIOLATION |
| 5 | Comparison fail | failed | BASELINE_MISMATCH; model pass cannot override |
| 6 | Missing/stale/ambiguous/invalidated baseline or unavailable comparison | blocked | BASELINE_MISSING/STALE/AMBIGUOUS/INVALIDATED |
| 7 | Comparison source_error/permission_denied/inconclusive | inconclusive | COMPARISON_INCONCLUSIVE |
| 8 | Comparison pass; disabled/advisory | passed | BASELINE_PASS; model issues remain annotations |
| 9 | Comparison pass; required evaluation missing/provider_error/parser_error/budget_exceeded/cancelled/timed_out | inconclusive | EVALUATION_UNAVAILABLE |
| 10 | Comparison pass; required valid evaluation confidence below threshold | inconclusive | LOW_CONFIDENCE |
| 11 | Comparison pass; required high-confidence inconclusive | inconclusive | MODEL_INCONCLUSIVE |
| 12 | Comparison pass; required evaluation contradicts the same deterministic criterion, or high-confidence pass with error/critical findings, or fail without supporting declared semantic findings | inconclusive | EVALUATION_DISAGREEMENT/EVALUATION_CONTRADICTORY |
| 13 | Comparison pass; required high-confidence fail with a supported declared semantic finding (row 12 false) | failed | SEMANTIC_FAILURE |
| 14 | Comparison pass; required high-confidence pass, no error/critical finding | passed | BASELINE_AND_SEMANTIC_PASS |
| 15 | Any remaining invalid/unrecognized input | blocked | POLICY_INPUT_INVALID |
Row 12 explicitly precedes semantic failure in row 13. Criterion identity is a declared typed reference, never inferred from prose. A claim contradicting the same deterministic comparison is disagreement, not a second oracle. Each finding binds to the evaluation spec, criterion and declared evidence.
Confidence equal to threshold is high. Empty findings plus an explicit schema-valid pass can satisfy row 14; an empty object/list, unrecognized envelope, missing verdict or invalid numbers never can. Deterministic-only baseline checks do not require an LLM provider.
An advisory provider error remains visible, while the authoritative outcome remains deterministic; required semantic checks cannot silently fall back to advisory.
HumanCheckpoint v1 mapping remains confirm→passed, false_positive/inconclusive→inconclusive, pass→passed, fail→failed; disposition cannot mutate this evaluation or its historical outcome.
Acceptance (open): hardcoded cases for all rows, threshold equality, pass/fail disagreement, contradictory findings, empty/malformed output, late attempt and every comparison status; repeated canonical inputs yield identical status/reason codes. No claim of runtime implementation.
## @} ScenarioGraph.DecisionPolicyV1

View File

@@ -0,0 +1,64 @@
{
"$schema": "https://json-schema.org/draft/2020-12/schema",
"title": "DecisionPolicy baseline-semantic/1.0.0",
"description": "Total first-match policy in decision-policy.md; implemented=false.",
"type": "object",
"additionalProperties": false,
"required": [
"policy_id",
"version",
"confidence_threshold",
"evaluation_mode",
"deterministic_hard_failure",
"high_confidence_failure",
"low_confidence",
"disagreement",
"missing_evidence",
"provider_error",
"human_checkpoint"
],
"properties": {
"policy_id": {
"const": "baseline-semantic"
},
"version": {
"const": "1.0.0"
},
"confidence_threshold": {
"type": "number",
"minimum": 0,
"maximum": 1,
"default": 0.7
},
"evaluation_mode": {
"enum": [
"disabled",
"advisory",
"required"
],
"default": "disabled"
},
"deterministic_hard_failure": {
"const": "failed"
},
"high_confidence_failure": {
"const": "failed"
},
"low_confidence": {
"const": "inconclusive"
},
"disagreement": {
"const": "inconclusive"
},
"missing_evidence": {
"const": "blocked"
},
"provider_error": {
"const": "inconclusive"
},
"human_checkpoint": {
"const": "declared_only"
}
},
"$id": "https://superset-tools.local/specs/038-dashboard-scenario-model/contracts/decision-policy.schema.json"
}

View File

@@ -21,7 +21,7 @@
## Readiness for Plan
- [x] CHK009 UI entities are defined for entry action, workspace state, preview, parameters, baselines, artifacts, and confirmations.
- [x] CHK009 UI entities are defined for entry action, manual review surface state, preview, parameters, baselines, artifacts, and confirmations.
- [x] CHK010 Success criteria are measurable through model, component, and browser tests.
- [x] CHK011 UX copy explicitly excludes direct SQL as a validation path.
@@ -33,3 +33,14 @@
- [x] CHK015 Traceability maps every functional requirement to contracts, tasks, and tests.
- [x] CHK016 Tasks use exact repository paths, dependency order, and test-first sequencing.
- [x] CHK017 Semantic anchors and cross-package ownership boundaries pass validation.
## Production readiness checklist — 2026-09-08
AGUI-FR-017: [Execution evidence display ownership](../../045-dashboard-run-monitor/contracts/evidence-ui.md). Historical [x] marks do not close this new production gate; removed frontend components are not current evidence. All rows below implemented=false / OPEN.
- [ ] CHK018 Manual create/editor preserves dashboard context and typed parameters; no agent launch/handoff/prompt/proposal-generation control or route. Evidence: [T082](../tasks.md), [traceability](../traceability.md).
- [ ] CHK019 Authoring evidence preview uses authorized owner refs; human review does not edit evaluation, comparison, baseline or StepOutcome. Evidence: [T083](../tasks.md), [traceability](../traceability.md).
- [ ] CHK020 Pipeline verification views consume real VerificationRun GET DTOs and distinguish missing/blocked/inconclusive from pass; removed workspace tests are not production evidence. Evidence: [T084](../tasks.md), [traceability](../traceability.md).
- [ ] CHK021 Negative product UI test: no agent chat/prompt/assistant editing/proposal-generation/typical-operation-to-agent/workspace/start/handoff controls or agent invocation routes/requests; manual CRUD/editor/human review/read-only results remain usable.
Schema/static success alone is not runtime completion. Optional approved performance baseline is outside scope.

View File

@@ -1,9 +1,8 @@
#region DashboardScenarioUi.Modules [C:5] [TYPE ADR] [SEMANTICS contracts,dashboard-testing,scenario,ux]
@BRIEF Model, API client, entry action, workspace, parameter, baseline, artifact, and confirmation UI contracts.
@BRIEF Manual authoring/review contracts plus tombstones for removed agent frontend modules.
@RATIONALE 039 owns presentation and interaction state while 036–038 own run, graph, and artifact semantics; this boundary prevents UI components from reimplementing backend truth.
@REJECTED Component-local fetching and business state — rejected because independent panels could render different revisions or bypass the shared HITL gate.
@RELATION DEPENDS_ON -> [DashboardScenarioUi.DataModel]
@RELATION DEPENDS_ON -> [AgentRuns.Model]
@RELATION DEPENDS_ON -> [ScenarioGraph.Api]
// #region DashboardTesting.ApiClient [C:3] [TYPE Module] [SEMANTICS dashboard-testing,api,types]
@@ -13,63 +12,36 @@
// @INVARIANT 401/403/409/422/5xx retain typed code/status; no silent empty fallback.
// #endregion DashboardTesting.ApiClient
// #region DashboardTesting.WorkspaceModel [C:5] [TYPE Model] [SEMANTICS dashboard-testing,scenario,workspace,model]
// @defgroup DashboardTesting Screen model for scenario review, resolution, baseline impact, and draft preview.
// @STATE unavailable | bootstrapping | analyzing | needs_parameters | scenario_ready | generating | draft_ready | waiting_approval | saved | failed | disconnected
// @ACTION initialize(context, agentRun) — enter scenario mode for valid v2 intent only.
// @ACTION applyScenarioResponse(response) — atomically replace revision/validation.
// @ACTION updateParameter(name, value) — local declared typed edit.
// @ACTION applyParameterDefinitions() — validate dirty ParameterDefinitions against base revision; launch values remain 045-owned.
// @ACTION generateDraftPack() — request 038 safe template pack.
// @ACTION selectArtifact(id) / loadPreview(id) — side-effect-free preview.
// @ACTION requestSave(ids) / requestBaselineApproval(candidate) — delegate gate creation.
// @INVARIANT AgentRunModel is sole authority for run/stages/gate; workspace never parses chat prose.
// @INVARIANT canRequestSave is false for preview_only, invalid artifact, blockers, or unresolved required input.
// @RELATION DEPENDS_ON -> [DashboardTesting.ApiClient]
// @RELATION DEPENDS_ON -> [AgentRuns.Model]
// @RELATION BINDS_TO -> [DashboardTesting.Workspace]
// @RATIONALE One composed screen model makes cross-panel readiness testable without DOM.
// @REJECTED Independent component fetch/state — causes revision and gate races.
// #region DashboardTesting.WorkspaceModel [C:5] [TYPE Tombstone] [SEMANTICS dashboard-testing,scenario,workspace,model]
@DEPRECATED Removed agent frontend contract; not an active UI dependency.
@REPLACED_BY ScenarioEditor.Modules (manual authoring/review) and ScenarioRunMonitor.DataModel (read-only results).
@RATIONALE Explicit 2026-09-08 user constraint forbids agent workspace/progress interaction in product UI.
// #endregion DashboardTesting.WorkspaceModel
<!-- #region DashboardTesting.EntryAction [C:3] [TYPE Component] [SEMANTICS dashboard-testing,entry,dashboard,context] -->
<!-- @ingroup DashboardTesting -->
<!-- @BRIEF Dashboard header business action that constructs exact UIContext v2 scenario route. -->
<!-- @PRE Dashboard id/name and env are available; actor sees action according to permission projection. -->
<!-- @POST Navigation URL contains scenario intent and existing ordinary AI action remains unchanged. -->
<!-- @POST Navigation opens ordinary manual editor with dashboard/environment context; no AI/agent entry remains. -->
<!-- @RELATION DEPENDS_ON -> [Dashboards.DetailModel] -->
<!-- @UX_STATE ready | missing_environment | permission_denied -->
<!-- @UX_RECOVERY Missing env focuses selector; denial explains required permission. -->
<!-- @UX_TEST three dashboard ids -> no stale id/env across generated hrefs. -->
<!-- #endregion DashboardTesting.EntryAction -->
<!-- #region DashboardTesting.Workspace [C:4] [TYPE Component] [SEMANTICS dashboard-testing,workspace,layout,scenario] -->
<!-- @ingroup DashboardTesting -->
<!-- @RELATION BINDS_TO -> [DashboardTesting.WorkspaceModel] -->
<!-- @RELATION DEPENDS_ON -> [DashboardTesting.Progress] -->
<!-- @RELATION DEPENDS_ON -> [DashboardTesting.ScenarioViews] -->
<!-- @RELATION DEPENDS_ON -> [DashboardTesting.Parameters] -->
<!-- @RELATION DEPENDS_ON -> [DashboardTesting.BaselineImpact] -->
<!-- @RELATION DEPENDS_ON -> [DashboardTesting.ArtifactPreview] -->
<!-- @RELATION DEPENDS_ON -> [DashboardTesting.EvidencePanel] -->
<!-- @UX_STATE Matches model FSM with skeleton/error/recovery for each async boundary. -->
<!-- @UX_REACTIVITY WorkspaceModel atoms/derived -> typed props -> DOM; no component business state. -->
<!-- @UX_RECOVERY Retry source step, switch env, resolve blocker, or return dashboard. -->
<!-- @INVARIANT Workspace is ONLY mounted for manual trigger; pipeline-triggered runs MUST NOT instantiate this component tree. -->
<!-- #region DashboardTesting.Workspace [C:4] [TYPE Tombstone] [SEMANTICS dashboard-testing,workspace,layout,scenario] -->
<!-- @DEPRECATED Agent workspace is prohibited in product UI since 2026-09-08. -->
<!-- @REPLACED_BY ScenarioEditor.Modules: ordinary manual authoring and human review. -->
<!-- #endregion DashboardTesting.Workspace -->
<!-- #region DashboardTesting.Progress [C:3] [TYPE Component] [SEMANTICS dashboard-testing,progress,stages,a11y] -->
<!-- @ingroup DashboardTesting -->
<!-- @RELATION BINDS_TO -> [AgentRuns.Model] -->
<!-- @UX_STATE context | inspect | scenario | parameters | generate | validate | save, each pending/active/completed/blocked/failed. -->
<!-- @UX_FEEDBACK New active stage announced once via aria-live polite. -->
<!-- @UX_RECOVERY Failed/gapped state exposes model recovery action. -->
<!-- @INVARIANT Labels/state come from typed stage DTOs only. -->
<!-- #region DashboardTesting.Progress [C:3] [TYPE Tombstone] [SEMANTICS dashboard-testing,progress,stages,a11y] -->
@DEPRECATED Removed agent frontend contract; not an active UI dependency.
@REPLACED_BY ScenarioEditor.Modules (manual authoring/review) and ScenarioRunMonitor.DataModel (read-only results).
@RATIONALE Explicit 2026-09-08 user constraint forbids agent workspace/progress interaction in product UI.
<!-- #endregion DashboardTesting.Progress -->
<!-- #region DashboardTesting.ScenarioViews [C:4] [TYPE Component] [SEMANTICS dashboard-testing,scenario,graph,coverage] -->
<!-- @ingroup DashboardTesting -->
<!-- @RELATION BINDS_TO -> [DashboardTesting.WorkspaceModel] -->
<!-- @UX_STATE empty | ready | warning | blocked. -->
<!-- @UX_FEEDBACK Summary, phase graph, step table, and all-case coverage reflect one revision. -->
<!-- @UX_RECOVERY Finding links focus affected step/parameter; graph has accessible table fallback. -->
@@ -79,7 +51,6 @@
<!-- #region DashboardTesting.Parameters [C:4] [TYPE Component] [SEMANTICS dashboard-testing,parameters,form,resolution] -->
<!-- @ingroup DashboardTesting -->
<!-- @RELATION BINDS_TO -> [DashboardTesting.WorkspaceModel] -->
<!-- @UX_STATE clean | dirty | invalid | applying | conflict. -->
<!-- @UX_FEEDBACK Inline required/type validation and affected-step count. -->
<!-- @UX_RECOVERY 409 offers reload current revision; reset restores server values. -->
@@ -89,7 +60,6 @@
<!-- #region DashboardTesting.BaselineImpact [C:4] [TYPE Component] [SEMANTICS dashboard-testing,baseline,candidate,approval] -->
<!-- @ingroup DashboardTesting -->
<!-- @RELATION BINDS_TO -> [DashboardTesting.WorkspaceModel] -->
<!-- @UX_STATE approved | stale | missing | candidate | inconclusive | source_error. -->
<!-- @UX_FEEDBACK Source, filters, Verification Program diff, SQL Lab validation, policy, provenance and stale dimensions. -->
<!-- @UX_RECOVERY Discover candidate, keep draft, request approval, or human checkpoint. -->
@@ -98,7 +68,6 @@
<!-- #region DashboardTesting.ArtifactPreview [C:4] [TYPE Component] [SEMANTICS dashboard-testing,artifact,draft,preview] -->
<!-- @ingroup DashboardTesting -->
<!-- @RELATION BINDS_TO -> [DashboardTesting.WorkspaceModel] -->
<!-- @UX_STATE empty | loading | valid | warning | invalid | preview_only | save_eligible. -->
<!-- @UX_FEEDBACK File tree shows intended path, digest short id, validation, warnings, unresolved markers. -->
<!-- @UX_RECOVERY Preview/download are safe; invalid/preview_only explains disabled save. -->
@@ -107,8 +76,6 @@
<!-- #region DashboardTesting.EvidencePanel [C:4] [TYPE Component] [SEMANTICS dashboard-testing,evidence,screenshot,vlm,review] -->
<!-- @ingroup DashboardTesting -->
<!-- @RELATION BINDS_TO -> [DashboardTesting.WorkspaceModel] -->
<!-- @RELATION DEPENDS_ON -> [AgentRuns.DraftList] -->
<!-- @BRIEF View captured screenshots, review VLM findings, and record dispositions. -->
<!-- @UX_STATE empty | loading | ready | reviewing_finding. -->
<!-- Layout: Screenshot viewer (main) + finding list (sidebar) + finding detail card (overlay or inline below viewer). -->
@@ -121,12 +88,14 @@
<!-- #region DashboardTesting.ConfirmationBinding [C:4] [TYPE Component] [SEMANTICS dashboard-testing,confirmation,hitl] -->
<!-- @ingroup DashboardTesting -->
<!-- @RELATION DEPENDS_ON -> [AgentChat.ConfirmationCard] -->
<!-- @RELATION BINDS_TO -> [AgentRuns.Model] -->
<!-- @UX_STATE repository_write | baseline_approval | permission_denied. -->
<!-- @UX_FEEDBACK Exact paths/value/filter fingerprint/warnings; baseline reason required. -->
<!-- @UX_RECOVERY Deny preserves drafts and records cancellation; permission denial has no confirm. -->
<!-- @INVARIANT No second confirmation implementation; reuse 036 gate envelope/card. -->
<!-- @INVARIANT Render the shared 036 durable gate as a human approval card; no agent invocation, prompt or conversation state. -->
<!-- #endregion DashboardTesting.ConfirmationBinding -->
## Active frontend invariant
Only manual CRUD/editor, human review/approval and read-only artifact/result views render. No agent prompt/chat/edit/proposal-generation/typical-operation/start/handoff controls. EvidencePanel is authoring-only; 045 owns immutable execution AgentEvaluationCard, with no retry-agent/provider controls. Negative DOM/route/network acceptance is OPEN.
#endregion DashboardScenarioUi.Modules

View File

@@ -4,7 +4,7 @@
| Decision | Selected | Rejected |
|---|---|---|
| Entry | Business action from dashboard | Tool/artifact dropdown: exposes implementation |
| Surface | Existing /agent route in scenario mode | New disconnected route: duplicates run/chat/recovery |
| Surface | Ordinary manual registry/editor/review | Agent chat/prompt/handoff/workspace: prohibited by explicit user decision 2026-09-08 |
| State | Composed screen model | Component-local fetch/state: revision races |
| Graph | Visual DAG plus full step table | Canvas-only graph: inaccessible and hard to test |
| Parameters | Typed declared form and revision resolve | Free-form chat only: ambiguous and non-validated |

View File

@@ -3,10 +3,10 @@
## Main Sequence
1. Dashboard action opens /agent with UIContext v2.
2. Gradio emits agent_run_started; run panel appears.
3. Agent invokes 037 query-model inspection and emits inspect progress.
4. Agent submits bounded intent to 038 compile; workspace receives ScenarioResponse.
1. Dashboard action opens ordinary manual editor with dashboard/environment context.
2. Product loads stored scenario/validation DTOs through REST; no agent event stream.
3. External MCP clients independently author typed proposals; the UI never requests agent generation.
4. Human review loads server-stored proposal diff; no prompt or assistant operation control.
5. User may edit ParameterDefinitions/defaults through the constrained authoring flow; runtime values are bound only by 044/045 launch preflight.
6. User requests draft pack; 038 registers 036 drafts.
7. User previews/downloads through 036 URLs.

View File

@@ -1,13 +1,13 @@
#region DashboardScenarioUi.WorkspaceUx [C:4] [TYPE ADR] [SEMANTICS ux,dashboard-testing,workspace,scenario,agent]
@BRIEF Layout, state, feedback, recovery, and browser test contract for the AGENT-DRIVEN scenario workspace (manual ad-hoc flow only).
@BRIEF Ordinary manual authoring/review layout; historical WorkspaceUx identifier retained without agent interaction.
@RELATION DEPENDS_ON -> [DashboardScenarioUi.DataModel]
@RATIONALE The agent workspace is the persistent interactive environment where the analyst creates, refines and remediates a test scenario through chat plus structured panels. Pipeline-triggered verification runs use separate views (see release-verification-ux.md).
@RATIONALE Explicit 2026-09-08 user constraint reserves agent interaction for external MCP clients; product frontend owns manual editor and human evidence review only.
@REJECTED Using the agent workspace for pipeline-triggered verification — pipeline views are read-only projections of VerificationRun data and must not require an AgentRun or conversation.
## Desktop Layout
- Header: mode (Dashboard Test Scenario Agent), dashboard/env context, run id, connection indicator.
- Progress: seven named stages (context, inspect, scenario, parameters, generate, validate, save).
- Header: manual editor mode, dashboard/environment context, revision and save status.
- Status: stored validation/save/approval state; no agent progress stream.
- Main two-column area: scenario/steps/coverage (wide) and parameters/baselines (narrow).
- Draft area: file tree left, preview right.
- Evidence area: screenshot viewer + finding list + finding detail card for VLM review.
@@ -17,7 +17,7 @@ At widths below large breakpoint, panels stack in workflow order. Step table and
## Feedback and Recovery
- Context/permission failure occurs before agent analysis begins.
- Context/permission failure occurs before loading editable context.
- Superset API failure offers retry, environment switch, or manual-only scenario shell.
- Finding links focus related step/field.
- Stale baseline offers candidate discovery, never silent replace.
@@ -38,10 +38,10 @@ At widths below large breakpoint, panels stack in workflow order. Step table and
7. Three VLM findings render with correct severity styling; confirmed/dismissed/inconclusive dispositions are visually distinct.
8. Selecting a finding highlights the corresponding region on the screenshot.
9. Provenance footer shows model, prompt version, and hash; stale prompt warning is visible.
10. WorkspaceModel is NOT initialized when trigger != manual; test verifies no state leak.
10. No agent chat/prompt/assistant editing/proposal-generation/typical-operation/workspace/start/handoff control, route or request exists; manual editor/review remains functional.
## Invariant
The agent workspace MUST NOT render for pipeline-triggered VerificationRuns (trigger ∈ {deploy_to_preprod, release_create, scheduled, etl_completed}). Pipeline results are displayed through ReleaseVerification components on deployment/release pages (see release-verification-ux.md).
An agent workspace MUST NOT render for any product route or trigger. Pipeline results are displayed through ReleaseVerification components on deployment/release pages (see release-verification-ux.md).
#endregion DashboardScenarioUi.WorkspaceUx

View File

@@ -1,9 +1,9 @@
#region DashboardScenarioUi.UxDecisions [C:3] [TYPE ADR] [SEMANTICS ux,decisions,dashboard-testing]
@BRIEF Final UI decisions for implementation handoff.
1. Keep ordinary AI and add one clearly labelled scenario-generation action.
2. Scenario mode is determined only by valid UIContext v2 intent.
3. Reuse AgentRun progress/recovery and ConfirmationCard; no duplicate lifecycles.
1. Remove all agent/AI interaction entries; scenario creation opens the ordinary manual editor.
2. Manual editor mode is explicit and dashboard/environment context-bound.
3. Reuse durable human approval DTOs; no AgentRun/chat/workspace dependency.
4. Present business objective, phases, steps, coverage, then tool evidence.
5. ParameterDefinition changes create immutable revisions; runtime ParameterBindings do not and are collected only at launch.
6. Show all unsupported/manual/needs-context cases.

View File

@@ -1,29 +1,9 @@
#region DashboardScenarioUi.ScreenModels [C:4] [TYPE ADR] [SEMANTICS ux,screen-models,dashboard-testing,verification]
@BRIEF Model composition for agent workspace and pipeline verification views. Two independent trees.
@BRIEF Model composition for manual authoring/human review and pipeline verification views. Two independent trees; no agent interaction model is instantiated (2026-09-08 boundary).
## Agent Workspace (manual ad-hoc flow)
## Manual authoring and human review
Active ONLY when the user enters `/agent` with `intent=build_dashboard_test_scenario` and `trigger=manual`.
~~~text
AgentChat.Model
└── AgentRuns.Model transport/recovery/gates/drafts
└── DashboardTesting.WorkspaceModel
├── Scenario views 038 graph/validation/coverage
├── Parameters 038 resolution
├── Baseline impact 037 result/candidate
├── Artifact preview 038 manifest + 036 refs
└── Evidence panel 036 screenshots + 038 findings + dispositions
~~~
AgentRuns.Model and WorkspaceModel are distinct ownership layers, not independent stores. The page creates/composes them once for the route visit.
### Agent Workspace — Test Layers
- L1 AgentRuns.Model: sequence, recovery, gate and drafts.
- L1 WorkspaceModel: FSM, typed parameter drafts, readiness, immutable revision apply, preview lifecycle.
- L2 components: DOM, accessibility, layout, feedback/recovery.
- E2E: dashboard entry action → scenario → deny/confirm.
042 registry and 043 editor own typed scenario/validation, parameter definitions, stored artifact selection and approval projections. No AgentChat.Model, AgentRuns.Model, WorkspaceModel, prompt or agent launch is instantiated. External MCP authoring may persist proposals; frontend only loads their stored diff for human review. L1 tests cover editor CAS and projection; L2 DOM/route/network tests reject agent interaction controls while keeping manual CRUD and approval usable.
## Pipeline Verification Views (automated, no agent)
@@ -56,7 +36,7 @@ Pipeline views MUST render without WorkspaceModel or AgentRun initialization. Co
## Decomposition Gate
- WorkspaceModel: at 400 lines or 40 public methods, split baseline/artifact selection into submodels while preserving facade and invariants.
- Manual editor models follow 043 decomposition; retired WorkspaceModel must not be restored.
- ReleaseVerificationState: at 300 lines, split structure dispositions into a submodel.
#endregion DashboardScenarioUi.ScreenModels

View File

@@ -1,16 +1,16 @@
#region DashboardScenarioUi.DataModel [C:4] [TYPE ADR] [SEMANTICS data-model,ux,scenario,workspace,verification]
@BRIEF Agent workspace projection and pipeline verification state models. Two independent state trees — agent-driven and pipeline-driven.
@BRIEF Manual authoring/review and pipeline verification projection models; no agent frontend state.
@RELATION DEPENDS_ON -> [DashboardScenarioModel.DataModel]
@RELATION DEPENDS_ON -> [AgentTestStabilization.DataModel]
## Agent Workspace — ScenarioWorkspaceState
## Manual Authoring — ScenarioWorkspaceState
Owned by `DashboardTesting.WorkspaceModel`, composed under `AgentChat.Model`. ONLY active when the user enters the agent workspace with `intent=build_dashboard_test_scenario`.
Owned by the 042/043 ordinary editor/review model. Historical WorkspaceModel is retired; no AgentChat/AgentRun subscription initializes this view.
| Atom | Type | Owner/source |
|---|---|---|
| mode | scenario or ordinary chat | UIContext v2 intent |
| scenario | DashboardTestScenario/null | 038 response (via agent) |
| mode | read_only or manual_edit | explicit user editor action |
| scenario | DashboardTestScenario/null | 038/042 server DTO |
| changeRequestContext | ChangeRequestContext/null | authoring input; never guessed by compiler |
| verificationProgram | VerificationProgram/null | 038 content-hashed program preview |
| validation | ScenarioValidationResult/null | 038 response |
@@ -22,23 +22,22 @@ Owned by `DashboardTesting.WorkspaceModel`, composed under `AgentChat.Model`. ON
| selectedArtifactId | UUID/null | local preview |
| preview | loading/content/error | 036 opaque preview |
| domainError | typed code/detail/recovery | API result |
| agentRun | AgentRuns.Model reference | 036 authority |
| evidence | ScreenshotEvidence[] | 036 evidence_captured events |
| vlmFindings | VlmFinding[] | 038 step outputs or adapter results |
| selectedEvidenceId | UUID/null | local selection in EvidencePanel |
| selectedFindingId | string/null | local selection; highlights region on screenshot |
| findingDispositions | map finding_id → disposition + comment | local pending changes |
### Agent Workspace — Derived Values
### Manual Authoring — Derived Values
- workspaceState from intent, AgentRun status/stage, scenario, draft pack, and domain error;
- reviewState from stored scenario, draft pack, validation/save status and domain error;
- canApplyParameterDefinitions when changed definition drafts are valid;
- canGenerateDraft when scenario exists and authoring blockers (for example selectors/baselines) are resolved; runtime ParameterBindings are collected only in 045 launch preflight;
- canRequestSave when draftPack.status=save_eligible, artifacts valid, and no blockers;
- baselineApprovalReady when candidate selected and non-blank reason;
- progress stages from AgentRunModel only.
- progress stages from stored review DTOs only.
### Agent Workspace — Invariants
### Manual Authoring — Invariants
1. Domain status never derives from assistant text.
2. Graph/validation are server-authoritative immutable revisions.
@@ -49,7 +48,7 @@ Owned by `DashboardTesting.WorkspaceModel`, composed under `AgentChat.Model`. ON
7. Switching run/context resets all scenario-local selections and drafts.
8. Evidence and findings are bound to the owning AgentRun and never cross run boundaries.
9. Finding disposition is recorded once per finding; replay of same disposition is idempotent.
10. WorkspaceModel is instantiated for analyst-opened agent intents (creation, investigation, revalidation, remediation, load analysis) and MUST NOT activate for pipeline-triggered runs.
10. Product frontend MUST NOT instantiate agent workspace, chat, prompt, assistant editing, proposal-generation or agent launch/handoff controls. External MCP clients alone interact with agents.
---
@@ -82,4 +81,12 @@ Separate state tree. Lives on deployment/release/dashboard pages. Does NOT requi
5. Badge styling (pass/warn/fail/blocked) always includes icon + text label, never color alone.
6. Immutability violation banner occupies full width, uses CRITICAL severity styling, and cannot be dismissed without an investigation ticket.
## Production record contract — 2026-09-08 (AGUI-FR-017)
ScenarioWorkspaceState is a legacy identifier for manual authoring/review only, without AgentRun/chat/prompt state. EvidencePanel/VlmFindingReviewCard display authoring DTOs and separate auditable human ReviewDisposition. 045 exclusively owns execution EvidenceViewer/AgentEvaluationCard and immutable run outcome projection.
Normative detail: [Execution evidence display ownership](../045-dashboard-run-monitor/contracts/evidence-ui.md). New fields, CAS transitions and cross-record integrity checks are implemented=false until [production tasks](tasks.md) and [traceability](traceability.md) close with executable evidence. Existing shorter field lists are legacy compatibility projections, not permission to omit the production identity fields.
Frontend contains no agent interaction state, prompt, proposal-generation or workspace/start controls. Human manual editor/review state and read-only evaluation/result artifacts are separate from external MCP authoring. Approved performance baseline is outside scope.
#endregion DashboardScenarioUi.DataModel

View File

@@ -2,7 +2,7 @@
**Branch**: 039-dashboard-scenario-ui | **Date**: 2026-07-13 | **Spec**: [spec.md](./spec.md)
## Summary
## Historical architecture context (removed chat is not an active plan)
Add a dashboard-level scenario action and a model-first workspace inside /agent. The UI composes 036 run/recovery/gate state, 037 baseline results, and 038 scenario/draft DTOs. Users review a business flow, resolve typed parameters, inspect coverage and draft artifacts, then explicitly save or approve.
@@ -80,3 +80,14 @@ frontend/src/
## Complexity Tracking
ScenarioWorkspace is decomposed before 400 lines or 40 public model methods. The visual graph must not absorb the accessible step table or parameter logic.
## Production delivery plan — 2026-09-08 (AGUI-FR-017)
Status: specified, implemented=false; historical unit/prototype/transport results are not current production acceptance.
1. Pin [Execution evidence display ownership](../045-dashboard-run-monitor/contracts/evidence-ui.md) and [data model](data-model.md); write negative fixtures before runtime changes.
2. Implement existing domain boundaries for: ScenarioWorkspaceState is a legacy identifier for manual authoring/review only, without AgentRun/chat/prompt state. EvidencePanel/VlmFindingReviewCard display authoring DTOs and separate auditable human ReviewDisposition. 045 exclusively owns execution EvidenceViewer/AgentEvaluationCard and immutable run outcome projection.
3. Execute [tasks](tasks.md) T082, T083, T084 and retain reproducible evidence in [traceability](traceability.md), then close [checklist](checklists/requirements.md) individually.
4. Run cross-spec canary only after 037 catalog publication, 038 chain and 044 provider/content/policy gates; use 046 versioned cost/load/SLO limits. Disable admission on rollback, retain pins/receipts/holds; no fallback to stale catalog or synthetic PASS.
Frontend agent prompts, chat, assistant editing, proposal-generation, workspace/start/handoff actions are prohibited. Only external MCP clients interact with agents; frontend provides ordinary manual CRUD/editor, human review/approval and read-only monitoring/evidence. Runtime removal tasks are not closed by this document. Optional approved performance baseline is outside this refresh.

View File

@@ -1,8 +1,18 @@
# Quickstart: Dashboard Scenario UI
> **Refresh 2026-09-08 (production contract):** execution-evidence display ownership (AGUI-FR-017) is normative
> via `contracts/ux/*` and `../045-dashboard-run-monitor/contracts/evidence-ui.md` — implemented=false /
> acceptance OPEN. The former agent-driven flow (dashboard → `/agent` workspace → agent-generated scenario) is
> SUPERSEDED: agent interaction is external MCP only; product UI owns manual CRUD/editor entry, human
> approval/review, read-only pipeline views and baseline preview with explicit publish. Existing agent panels
> in runtime are drift; removal plus negative DOM/route/network acceptance remains OPEN. Historical suites
> below are runtime-gated evidence, not completion.
## Prerequisites
Stable fixtures and passing contracts from 036, 037, and 038. Backend, agent, and frontend running; one execution user and one approver.
Stable fixtures and passing contracts from 036, 037, and 038. Backend and frontend running; one execution user
and one approver. Authoring proposals (optional) arrive only from an external MCP client (050) and are reviewed
as stored diffs.
## Test Order
@@ -15,31 +25,58 @@ npm run lint
npm run build
~~~
## Manual End-to-End
> `DashboardScenarioWorkspaceModel` and agent-stream E2E scenarios target the retired workspace composition;
> keep them green only as drift-removal regression until the manual editor (043) and pipeline-view suites
> replace them.
1. Open dashboard 42 with environment selected.
2. Click “Создать сценарий тестирования”; verify exact dashboard/env/intent in /agent.
3. Confirm run id appears before analysis tool activity.
4. Review 18-step scenario: phases, step table, graph, all 19 coverage rows, warnings/blockers.
5. Verify relevant steps distinguish Superset API evidence from immutable validated SQL Lab evidence; no runtime SQL edit/rewrite control is present.
## Manual End-to-End (current boundary)
1. Open dashboard 42 with environment selected; verify exact dashboard/env context loads without any agent
route, stream or workspace.
2. Enter the scenario testing workspace: pinned revision loads in manual editor mode (steps/coverage wide
column, parameters/baselines narrow column) with stored validation/save/approval status only.
3. If an external MCP client persisted a proposal, load its stored diff and provenance read-only; accept it
into a manual edit session or discard it. No prompt/start/handoff control may exist.
4. Review the scenario: phases, step table, graph, coverage rows, warnings/blockers.
5. Verify relevant steps distinguish Superset API evidence from immutable validated SQL Lab evidence; no
runtime SQL edit/rewrite control is present.
6. Resolve date, counterparty, and baseline choice; only affected steps update.
7. Generate draft; preview each text artifact and download without Git changes.
7. Validate and preview stored draft artifacts; download without Git changes.
8. Verify preview-only/invalid draft blocks save.
9. Request repository save, inspect exact paths/warnings, deny, and verify no write.
9. Request repository save through the bound 036 gate, inspect exact paths/warnings, deny, and verify no write.
10. Request baseline approval; blank reason must fail locally/server-side.
11. Confirm with reason; verify approved provenance and consumed gate.
12. Reload mid-run/draft and recover the same state.
13. Repeat with unauthorized user; no run or confirm control depending on denied operation.
11. Confirm with reason; verify approved provenance, consumed gate, and that the baseline preview shows
candidate/materialized/published states with an explicit separately-invoked publish action.
12. Reload mid-edit and recover the same revision/draft/gate state.
13. Repeat with unauthorized user; no save/approve/publish control depending on denied operation.
14. Pipeline views on dashboard/release/PREPROD pages render VerificationRun badge/history/StructureDiff from
037 data only — without WorkspaceModel/AgentRun initialization.
15. Negative acceptance (OPEN): DOM/route/network scan proves no agent chat, prompt textarea, assistant
editing, proposal-generation selector, agent workspace/start or handoff control exists.
## Exit Gates
- Correct context for three dashboards and no stale reuse.
- 18-step fixture usable at 1366px and keyboard accessible.
- Parameter readiness projection under 200ms in L1 tests.
- 100% save/approve actions use 036 gate.
- 100% save/approve/publish actions use the 036 gate; publish is explicit and receipted.
- No direct-SQL action/copy.
- No agent interaction control/route/request in product UI; manual editor/review remains fully functional.
- Run results/evidence/evaluation rendering belongs to 045 read-only contracts; 039 never mutates
AgentEvaluation or StepOutcome (authoring review only).
- Frontend tests/lint/build and Playwright smoke pass.
## Known Gap (2026-08-07 MVP audit)
Шаги 1–13 проверяют agent-driven сценарий — работают через события 036. Но `frontend/src/lib/api/dashboard-testing.ts` (scenario compile/validate/resolve + verification) **не импортируется** ни одним модулем: без запущенного агента сценарий не строится. Pipeline views (badge/history/StructureDiff) существуют, но не привязаны к страницам и ждут GET-эндпоинтов 037 (T081) и deploy-хуков 037 (T080). До T054–T058 (tasks.md Phase 9) UI полностью зависит от агент-стрима, а VerificationRun-данные не отображаются.
Шаги 1–13 проверяли agent-driven сценарий через события 036. `frontend/src/lib/api/dashboard-testing.ts`
(scenario compile/validate/resolve + verification) **не импортируется** ни одним модулем: без запущенного
агента сценарий не строится. Pipeline views (badge/history/StructureDiff) существуют, но не привязаны к
страницам и ждут GET-эндпоинтов 037 (T081) и deploy-хуков 037 (T080).
## Refresh Gap (2026-09-08 production audit)
Agent-driven поток больше не является целевым: целевая композиция — manual editor (043) + stored MCP proposal
diff review + read-only pipeline views + 045 evidence rendering. Runtime всё ещё содержит agent-панели (drift),
`dashboard-testing.ts` не подключён, GET-эндпоинты/deploy-хуки 037 отсутствуют. Все строки AGUI-FR-017
implemented=false / acceptance OPEN; негативное UI-acceptance (отсутствие agent-контролов) — отдельная
открытая задача удаления drift.

View File

@@ -1,5 +1,5 @@
#region DashboardScenarioUi.Spec [C:3] [TYPE ADR] [SEMANTICS spec,requirements,ux,scenario,dashboard-testing]
@BRIEF Persistent agent workspace for scenario generation, scenario revision saves, evidence-led remediation, and policy-bound approvals.
@BRIEF Manual scenario authoring/review and pipeline verification views; all agent interaction is external MCP only.
@RELATION DEPENDS_ON -> [Doc.Adr.ADR0001]
@RELATION DEPENDS_ON -> [Doc.Adr.ADR0005]
@RELATION DEPENDS_ON -> [Doc.Adr.ADR0006]
@@ -8,26 +8,26 @@
@RELATION DEPENDS_ON -> [DashboardScenarioModel.Spec]
@RATIONALE Users should approve a business-level dashboard test scenario and required parameters; low-level tool chains are visible for trust but selected by the agent and scenario validator.
@REJECTED Dropdowns such as "Playwright UI tests" vs "SQL checks" vs "XLSX checks" — rejected because each dashboard requires a unique cross-tool program. Validated immutable SQL evidence is visible in the program, not chosen as a loose UI mode.
@REJECTED Persistent chat workspace as the mandatory creation UX — superseded 2026-08-24 by external MCP clients plus non-chat surfaces (042/043) per `specs/050-mcp-interface/spec.md`; the dashboard entry action becomes an explicit handoff surface.
@REJECTED Frontend agent workspace, chat, prompt, assistant editing, proposal-generation and agent handoff/launch controls — explicitly prohibited by user decision 2026-09-08. Human review is not agent invocation.
## Navigation (DSA Indexer keywords)
@SEMANTICS: spec, requirements, feature, ux, agent, scenario, dashboard-testing, artifacts, baseline
**Feature Branch**: `039-dashboard-scenario-ui`
**Created**: 2026-07-07 | **Status**: Ready for Implementation
**Input**: "Provide the user-facing persistent agent workspace. From a dashboard page the analyst creates or remediates a scenario; the agent analyzes context, proposes a graph, collects authoring context, previews artifacts/evidence, saves delegated revisions, and uses inline approval only where policy requires it."
**Created**: 2026-07-07 | **Status**: Partially implemented; production acceptance OPEN (refresh 2026-09-08)
**Input**: Ordinary manual scenario creation/editing, stored artifact review and human approval, with read-only verification results. No embedded agent interactions.
## User Scenarios
### Story 1 — Start Scenario From Dashboard Page (P1)
**Why P1**: The feature must be discoverable from the dashboard the user wants to test and must carry the correct context into the agent.
**Why P1**: The feature must be discoverable from the dashboard the user wants to test and must preserve dashboard context in the manual editor.
**Independent Test**: On a dashboard page, click "Создать сценарий тестирования" and verify `/agent` opens with dashboard context and scenario intent.
**Independent Test**: On a dashboard page, click "Создать сценарий тестирования" and verify the ordinary scenario editor opens with dashboard context and no agent invocation.
**Acceptance**:
1. **Given** a user is viewing a dashboard **When** they click "Создать сценарий тестирования" **Then** `/agent` opens in Dashboard Test Scenario Agent mode with dashboard id, name, env, route, and intent visible.
2. **Given** environment context is missing **When** the user starts the flow **Then** the UI prompts for environment selection before scenario analysis.
1. **Given** a user is viewing a dashboard **When** they click "Создать сценарий тестирования" **Then** the manual scenario editor opens with dashboard id, name, env, route, and intent visible.
2. **Given** environment context is missing **When** the user starts the flow **Then** the UI prompts for environment selection before manual scenario authoring.
3. **Given** the user lacks scenario-generation permission **When** they click the action **Then** a permission-denied recovery path appears and no agent run starts.
---
@@ -39,7 +39,7 @@
**Independent Test**: Feed fixture scenario graph data and verify the UI renders phases, steps, tools, outputs, warnings, blockers, and coverage.
**Acceptance**:
1. **Given** the agent analyzes a dashboard **When** scenario graph is ready **Then** the UI shows goal, phases, step table, dependency graph, tool categories, expected results, coverage, warnings, and blockers.
1. **Given** a stored scenario is available **When** scenario graph is ready **Then** the UI shows goal, phases, step table, dependency graph, tool categories, expected results, coverage, warnings, and blockers.
2. **Given** some checklist cases are not automatable **When** scenario preview renders **Then** they appear as human checkpoints, unsupported, or needs-context with rationale.
3. **Given** the scenario includes source-mart evidence **When** preview renders **Then** the UI shows its validated SqlEvidenceSpec (relation/connection identity, hash, typed parameters, limits and output schema) and distinguishes it from Superset chart API evidence.
@@ -85,7 +85,7 @@
---
### Edge Cases
- Agent analysis fails due to Superset API error → UI shows retry, switch environment, or manual-only scenario recovery.
- Context loading fails → UI offers reload, environment selection, or manual recovery; no agent retry control.
- Scenario has unresolved blockers → generation can continue only for safe preview; executable save is blocked until resolved or converted to human checkpoint.
- Baseline is stale → UI warns and offers discovery candidate flow, not silent update.
- XLSX export is unavailable → UI shows coverage gap and alternate Superset API/UI checks.
@@ -93,13 +93,13 @@
## Requirements
### Functional — Agent Workspace (persistent chat and work surfaces)
### Functional — Manual Authoring and Human Review
These requirements cover the persistent agent workspace where the analyst creates, refines, investigates, and remediates scenarios through chat plus structured work surfaces.
These requirements cover ordinary manual authoring and review. External MCP clients may author stored proposals, but product UI never starts or directs an agent.
- **AGUI-FR-001**: Dashboard pages MUST expose a single business-level action labeled "Создать сценарий тестирования"; UI MUST NOT present low-level artifact/tool choices as the primary entry point.
- **AGUI-FR-002**: `/agent` MUST show Dashboard Test Scenario Agent mode when opened with scenario intent from 036.
- **AGUI-FR-003**: The UI MUST render structured progress stages from agent metadata: context, inspect, scenario, parameters, generate, validate, save.
- **AGUI-FR-002**: Scenario creation MUST open a normal manual editor. `/agent`, agent workspaces, prompt textareas, assistant editing, proposal-generation, typical-operation-to-agent and agent launch/handoff controls MUST be absent from product UI.
- **AGUI-FR-003**: UI MUST render stored validation/save state and human review status without binding to AgentRunModel or a conversation stream.
- **AGUI-FR-004**: The scenario preview MUST show business goal, phases, step list, tool category per step, dependency graph, coverage, warnings, blockers, expected results, and the visible Verification Program (navigation, evidence, transforms, assertions and declared agentic evaluations).
- **AGUI-FR-005**: Authoring MUST support typed ParameterDefinition/default/validation/source edits and show affected steps without restarting the flow. Required parameters without defaults remain save-eligible; 045 alone collects runtime ParameterBindings.
- **AGUI-FR-006**: Baseline UI MUST show approved baseline matches, stale warnings, missing baselines, and draft candidate approval paths.
@@ -112,7 +112,7 @@ These requirements cover the persistent agent workspace where the analyst create
- **AGUI-FR-012**: VLM findings MUST be reviewable with typed disposition controls: confirm (accepts finding as valid), dismiss (marks as false positive), or inconclusive (defers to human checkpoint). Disposition changes MUST be auditable and MUST NOT alter the scenario graph.
- **AGUI-FR-013**: The artifact preview file tree MUST include an `evidence/` branch listing screenshot artifacts and their associated VLM finding files.
These requirements are satisfied through the **DashboardTesting.AgentWorkspace** components. The agent is the interaction surface; structured panels (Parameters, ArtifactPreview, EvidencePanel, ActionTimeline) provide data entry and review within the workspace. The agent may perform delegated durable actions, while policy-gated actions remain inline cards in the same thread.
These requirements belong to the 042 registry and 043 manual editor/review surfaces. Legacy DashboardTesting.AgentWorkspace is retired, not an implementation dependency.
### Functional — Pipeline Verification Views (automated, no agent interaction)
@@ -126,16 +126,7 @@ These requirements are satisfied through the **ReleaseVerification** components.
### Architectural boundary
```text
Agent Workspace (AGUI-FR-001..013) Pipeline Views (AGUI-FR-014..016)
────────────────────────────────── ──────────────────────────────
Entry: «Создать сценарий» → /agent Entry: deployment page, release page,
Interaction: диалог с агентом dashboard page
State: WorkspaceModel + AgentRunModel State: VerificationRun[] + StructureDiff
Trigger: manual (аналитик) Trigger: deploy_to_preprod, release_create,
Data: 038 ScenarioResponse, 036 drafts scheduled, etl_completed (автоматически)
Save: delegated policy / inline gate Save: validate/approve/publish (pipeline gates)
```
Manual authoring (AGUI-FR-001..013) uses registry/editor REST DTOs. Pipeline views (AGUI-FR-014..016) use VerificationRun DTOs. External MCP agents are separate clients; neither UI surface hosts an agent interaction.
### LLM Verification Tooling Reuse (039)
@@ -150,13 +141,13 @@ The EvidencePanel and VlmFindingReviewCard components consume **DTOs only** —
`@REJECTED` Frontend-side LLM/VLM calls or screenshot rendering heuristics — the UI renders typed findings and evidence only; provider calls stay on the backend in reused plugin modules.
### Key Entities — Agent Workspace
### Key Entities — Manual Authoring
- **ScenarioEntryAction**: Dashboard-page action that launches `/agent` with dashboard scenario intent. Only valid in manual (ad-hoc) flow.
- **ScenarioWorkspaceState**: Frontend state machine for agent-driven creation and remediation: context, progress, scenario preview, parameter collection, artifact preview, evidence and inline action states.
- **ScenarioEntryAction**: Dashboard-page entry to the manual scenario editor, preserving dashboard/environment context.
- **ScenarioWorkspaceState**: Legacy identifier for ordinary authoring/review state; no AgentRunModel, chat or prompt state survives.
- **ScenarioPreviewCard**: UI representation of `DashboardTestScenario` from 038.
- **ParameterPanel**: Form for scenario business parameters and validation feedback. Text-input based, with typed validation.
- **BaselineImpactPanel**: UI section showing approved, stale, missing, and candidate baseline statuses within the agent workspace.
- **BaselineImpactPanel**: UI section showing approved, stale, missing, and candidate baseline statuses within manual review.
- **ArtifactPreviewPanel**: UI file tree/content preview for draft generated artifacts. Download is side-effect-free.
- **ScenarioActionCard**: Inline 036 policy/gate card for baseline approval and other non-delegated actions.
- **EvidencePanel**: Panel for viewing captured screenshots, reviewing VLM findings, recording dispositions. Composes screenshot viewer, finding list, finding detail, and disposition controls.
@@ -172,7 +163,7 @@ These entities are independent of the agent workspace. They consume `Verificatio
## Success Criteria
- **SC-001**: Users can launch scenario generation from a dashboard page in one click and see correct context in `/agent` in frontend tests.
- **SC-001**: Dashboard creation opens the manual editor with correct context; route/DOM/network tests prove absence of agent routes, prompts, launch/edit/proposal controls and agent invocations.
- **SC-002**: Scenario preview renders at least 15-step fixture graphs with phases, tools, warnings, and blockers without layout collapse at 1366px width.
- **SC-003**: Parameter updates resolve dependent scenario readiness within 200ms in model tests.
- **SC-004**: Artifact preview blocks save when unresolved markers are present in fixture data.
@@ -181,30 +172,18 @@ These entities are independent of the agent workspace. They consume `Verificatio
- **SC-007**: Pipeline verification views render StructureDiff, metric results, and VerificationRun data without initializing an AgentRun or WorkspaceModel.
- **SC-008**: VerificationStatusBadge appears on dashboard, release, and PREPROD pages within 200ms of data load.
## Implementation Status & MVP Debt (audit 2026-08-07)
## Historical verification and current acceptance
**Facts (code check, not tasks.md):**
- ✅ 13 компонентов dashboard-testing (ScenarioWorkspace, EvidencePanel, ParameterPanel, ArtifactPreview, VlmFindingReviewCard и др.) реализованы и отрендерены в `frontend/src/routes/agent/+page.svelte`.
- ✅ Pipeline-компоненты `VerificationStatusBadge`, `StructureDiffPanel`, `VerificationHistoryList` и API-клиент `getVerificationHistory()`/`getVerificationRun()` **существуют** (`frontend/src/lib/components/dashboards/verification/`, `frontend/src/lib/api/dashboard-testing.ts`).
- 🔴 **Прямой API-слой сценариев не задействован**: `dashboard-testing.ts` (scenario compile/validate/resolve) нигде не импортируется; `DashboardScenarioWorkspaceModel` не выполняет `fetch` — данные только через агентские события (036). Без запущенного агента сценарий не строится.
- 🔴 **Pipeline views не привязаны к страницам**: badge/history/diff-компоненты нигде не отрендерены (dashboard hub / release / PREPROD страницы не используют их), а их GET-эндпоинты (`/verification/history`, `/verification/{runId}`) **отсутствуют на backend** (037 Gap B, только POST `/verification-runs` существует). Проверка «Validate» на PREPROD деплое не создаёт VerificationRun.
The 2026-08-07 workspace component/test claims describe removed code, not current production coverage. T054–T057 historical results remain in tasks/history; T055 screenshot/evaluation closure and T058 deployment binding remain OPEN. The old claim that verification GET endpoints are missing is superseded by the recorded T081 implementation; current E2E binding and ACL must still be verified.
**Закрытие**: задачи T054–T056 (REST-binding, автономный evidence/VLM-путь) + T057–T058 (привязка pipeline views к страницам, verify-action) в `tasks.md` Phase 9. T057–T058 зависят от 037 T080–T081 (deploy-hook triggers + GET read-API). Если продукт сознательно остаётся agent-only, зафиксировать это решение в UX-контракте вместо имплементации.
050 removed chat, but remaining frontend agent proposal/handoff interactions are runtime drift under the explicit 2026-09-08 no-agent-UI requirement. Removal and negative UI acceptance are OPEN; this spec-only refresh does not change runtime.
## Runtime Closure Status (2026-08-07, resolved)
## Production contract refresh — 2026-09-08
- ✅ **T054/T056**: REST-путь сценариев готов — `api/dashboard-testing.ts` (compile/validate/resolve) + `WorkspaceModel.compileFromRest/validateFromRest/resolveFromRest`; `DashboardScenarioWorkspaceModel.rest.test.ts` (3 теста), 79 vitest passed, build OK.
- ✅ **T057**: `DashboardDetailModel.loadVerificationRuns()` + `VerificationHistoryList` привязан на `/dashboards/[id]` (потребляет 037 T081 GET `/verification/history`); `DashboardDetailModel.test.ts` = 67 passed.
- 🟡 **T058 — открыт (deferred)**: verify-action на PREPROD требует `repository_id`, которого нет в dashboard metadata, и отдельной deployment-страницы, которой нет во frontend. Требует dashboard→git-repository linkage + deployment surface. Зафиксировано в tasks.md как известный blocker.
**Frontend boundary (user decision 2026-09-08)**: All agent interaction is external MCP only. Product frontend MUST NOT contain agent chat, prompt/request textarea, assistant editing, typical-operation-to-agent selector, proposal-generation, agent workspace/start or handoff controls/routes. Ordinary manual CRUD/editor, human approval/review, monitoring and read-only evidence/evaluation are permitted. AgentEvaluationCard is read-only, with no prompt/retry-agent/provider controls. Existing agent proposal UI is runtime drift; removal/negative DOM-route-network acceptance remains OPEN in this spec-only change.
## Drift Amendment — MCP Interface (2026-08-24)
**AGUI-FR-017 — Execution evidence display ownership**: Post-chat authoring review belongs to 042/043; agent authoring belongs exclusively to external MCP clients. 039 EvidencePanel/VlmFindingReviewCard are authoring-preview contracts only; 045 EvidenceViewer/AgentEvaluationCard MUST render immutable ScenarioRun results from 044 DTOs. ReviewDisposition cannot modify AgentEvaluation or StepOutcome. Removed workspace components are historical, not implemented runtime evidence. Baseline preview shows candidate/materialized/published states and an explicit publish action.
Gradio chat retirement is formalized in `specs/050-mcp-interface/spec.md`.
- **Superseded**: AGUI-FR-001..013 as a mandatory in-product chat workspace; scenario preview, parameters and artifact review remain available in non-chat surfaces (042 registry detail, 043 editor).
- **Entry action**: «Создать сценарий тестирования» opens a HandoffSurface — connection instructions plus a copyable prompt carrying dashboard context (`objectType/objectId/envId/route/intent`) for an external MCP client; it never links to `/agent`.
- **Unchanged**: pipeline verification views (AGUI-FR-014..016), Svelte 5 conventions for remaining surfaces, DTO-only component contracts.
**Status (2026-09-02): done** — реализовано в рамках 050: инструменты и гейты (`specs/050-mcp-interface/tasks.md` T012–T028 [x]), handoff-поверхность (050 T030–T033), демонтаж чата и сервиса `agent/` (050 T040–T041, чекпоинты `specs/WORKSTATE-043-047.md`).
Normative contract: [Execution evidence display ownership](../045-dashboard-run-monitor/contracts/evidence-ui.md). New requirements are specified, **implemented=false / acceptance OPEN** until executable evidence closes the linked tasks/checklist/traceability rows. Historical local tests and the manual inconclusive ss-prod run do not prove browser/capture/baseline/LLM production readiness. The refresh scope is the audited P0/P1/P2 agentic E2E and baseline gaps; an approved ExecutionPerformanceBaseline is not introduced.
#endregion DashboardScenarioUi.Spec

View File

@@ -87,7 +87,7 @@
- [x] T054 [P] Bind `DashboardScenarioWorkspaceModel` to `frontend/src/lib/api/dashboard-testing.ts`: add `compileScenario()` / `validateScenario()` / `resolveScenario()` actions invoked from `ParameterPanel` and preview refresh, with `permission_denied`/`api_error` mapped to existing `setDomainError`.
@POST: preview can be refreshed from REST without re-running the agent; existing event-driven path remains the primary entry.
**DONE**: added compileScenario/validateScenario/resolveScenario to api client + WorkspaceModel.compileFromRest/validateFromRest/resolveFromRest; verified by `DashboardScenarioWorkspaceModel.rest.test.ts` (3 tests).
- [x] T055 [P] Evidence/VLM autonomous path: when evidence artifacts exist, allow `captureScenarioScreenshot` + `analyzeScenarioScreenshot` via the API client from the EvidencePanel (fallback when no live agent stream), showing `STALE_PROMPT` recovery.
- [ ] T055 [P] Evidence/VLM autonomous path: when evidence artifacts exist, allow `captureScenarioScreenshot` + `analyzeScenarioScreenshot` via the API client from the EvidencePanel (fallback when no live agent stream), showing `STALE_PROMPT` recovery.
**DONE (REST surface)**: capture/vlm/disposition endpoints exist on backend (038 T057/T058); frontend REST functions compileScenario/validateScenario/resolveScenario added; autonomous evidence/VLM binding in EvidencePanel deferred — see T058-adjacent note below.
- [x] T056 [P] Frontend integration test: `frontend/src/lib/models/__tests__/DashboardScenarioWorkspaceModel.api.test.ts` mocking `$lib/api/dashboard-testing` — compile → parameter fill → recompile → disposition, asserting no agent stream required for preview refresh.
**DONE**: `DashboardScenarioWorkspaceModel.rest.test.ts` (compile→scenario_ready, validate→blockers, api failure→api_error).
@@ -102,4 +102,16 @@
T001–T005 → US1 → run/progress → scenario views → parameters/baselines → drafts → confirmation → E2E. Tests precede each implementation group. Phase 8 depends on 036 Phase 8 (evidence_captured events), 038 Phase 8 (typed VlmFindings with disposition), and completed Phase 5 (ArtifactPreview) of this spec. Phase 9 (T054–T056) closes the REST-binding gap found in the 2026-08-07 audit — optional if product accepts agent-only operation. Phase 9 pipeline tasks (T057–T058) depend on 037 Phase 10 (T080–T081) providing deploy-hook triggers and GET read-API.
## Production readiness — 2026-09-08 (AGUI-FR-017)
Historical [x] rows above retain only their dated local/transport evidence; they do not prove current production readiness. Reopened rows were contradicted by the audited gaps. Removed frontend/agent paths are historical, not implementation prerequisites. New acceptance is **implemented=false / OPEN**.
Contract: [Execution evidence display ownership](../045-dashboard-run-monitor/contracts/evidence-ui.md).
- [ ] T082 [P0/P1/P2] Manual create/editor preserves dashboard context and typed parameters; no agent launch/handoff/prompt/proposal-generation control or route. Implement at the existing 039 domain boundary; verify with independent hardcoded fixtures and retain command/evidence references in traceability.md.
- [ ] T083 [P0/P1/P2] Authoring evidence preview uses authorized owner refs; human review does not edit evaluation, comparison, baseline or StepOutcome. Implement at the existing 039 domain boundary; verify with independent hardcoded fixtures and retain command/evidence references in traceability.md.
- [ ] T084 [P0/P1/P2] Pipeline verification views consume real VerificationRun GET DTOs and distinguish missing/blocked/inconclusive from pass; removed workspace tests are not production evidence. Implement at the existing 039 domain boundary; verify with independent hardcoded fixtures and retain command/evidence references in traceability.md.
Frontend boundary for this package: manual CRUD/editor, human review/approval, monitoring and read-only evidence/evaluation only; all agent interaction is external MCP. No agent chat/prompt/assistant editing/proposal generation/workspace/start/handoff controls. Runtime removal is OPEN, not performed by this spec refresh. Optional approved performance baseline is outside scope.
#endregion DashboardScenarioUi.Tasks

View File

@@ -34,4 +34,17 @@
| AGUI-FR-014 (PREPROD Verify) | ReleaseVerification components + POST verification-runs | T057, T058 | verify action creates run; CRITICAL structure blocks Validate |
| AGUI-FR-015/016 (history/badge) | VerificationStatusBadge, VerificationHistoryList, StructureDiffPanel | T057 | badge within 200ms; history/diff from live API (depends 037 T081) |
## Production acceptance traceability — 2026-09-08
Historical rows above identify prior tests/code only; removed agent UI paths are retired. The following audited gates are **implemented=false / OPEN**, independent of local suite totals.
| Requirement | Domain contract / DTO | Task | Falsifiable acceptance | State |
|---|---|---|---|---|
| AGUI-FR-017 | [Execution evidence display ownership](../045-dashboard-run-monitor/contracts/evidence-ui.md); [data model](data-model.md) | [T082](tasks.md) | Manual create/editor preserves dashboard context and typed parameters; no agent launch/handoff/prompt/proposal-generation control or route. | OPEN |
| AGUI-FR-017 | [Execution evidence display ownership](../045-dashboard-run-monitor/contracts/evidence-ui.md); [data model](data-model.md) | [T083](tasks.md) | Authoring evidence preview uses authorized owner refs; human review does not edit evaluation, comparison, baseline or StepOutcome. | OPEN |
| AGUI-FR-017 | [Execution evidence display ownership](../045-dashboard-run-monitor/contracts/evidence-ui.md); [data model](data-model.md) | [T084](tasks.md) | Pipeline verification views consume real VerificationRun GET DTOs and distinguish missing/blocked/inconclusive from pass; removed workspace tests are not production evidence. | OPEN |
| AGUI-FR-017; external-MCP-only UI | manual editor/review; read-only evidence | [production tasks](tasks.md) | No frontend agent prompt/chat/assistant editing/proposal generation/workspace/start/handoff routes or requests; human approval remains usable. | OPEN |
Sources: [production gap](../../docs/reports/ss-prod-agentic-e2e-production-gap-2026-09-08.md), [coverage gap](../../docs/reports/ss-prod-agentic-e2e-spec-coverage-2026-09-08.md), [baseline gap](../../docs/reports/ss-prod-agentic-e2e-baseline-gap-2026-09-08.md). Spec schema/static checks prove contract structure only; live canary/runtime closure and optional approved performance baseline are not claimed.
#endregion DashboardScenarioUi.Traceability

View File

@@ -1,21 +1,21 @@
#region DashboardScenarioUi.UxReference [C:4] [TYPE ADR] [SEMANTICS ux,reference,scenario,dashboard-testing]
@BRIEF UX reference for the Dashboard Test Scenario Agent workspace and dashboard entry point.
@BRIEF UX reference for manual scenario authoring/review, baseline preview/publish and read-only pipeline verification views; agent interaction is external MCP only.
@RELATION DEPENDS_ON -> [DashboardScenarioUi.Spec]
@RATIONALE The user experience is centered on approving a business scenario, not selecting implementation technologies.
@REJECTED Primary dropdown options for Playwright, SQL, and XLSX were rejected because they expose implementation details and conflict with unique cross-tool scenario generation.
@RATIONALE The user experience is centered on approving a business scenario and reviewing stored evidence, not selecting implementation technologies and not directing agents.
@REJECTED The 2026-07-07 agent-workspace flow (dashboard entry into `/agent`, agent progress strip, agent-generated proposal stream) is superseded by user decision 2026-09-08: frontend agent chat/prompt/assistant editing/proposal-generation/workspace/start/handoff controls are prohibited; existing runtime agent panels are drift pending removal. Primary dropdown options for Playwright, SQL, and XLSX remain rejected because they expose implementation details.
**Feature Branch**: `039-dashboard-scenario-ui`
**Created**: 2026-07-07 | **Status**: Ready for Implementation
**Created**: 2026-07-07 | **Refreshed**: 2026-09-08 | **Status**: Aligned with production contract refresh; agent-workspace UX retired (historical version preserved in Git)
## 1. User Persona & Context
* **Who is the user?**: QA engineer, dashboard owner, or analyst responsible for repeatable dashboard validation.
* **What is their goal?**: Generate a unique dashboard test scenario, provide business parameters, inspect generated artifacts, and approve durable changes.
* **Context**: Svelte UI in superset-tools; entry from dashboard page into `/agent` with scenario intent.
* **What is their goal?**: Create/edit a dashboard test scenario manually or review a stored external-MCP proposal, provide business parameters, inspect artifacts, approve baselines, and monitor pipeline verification results.
* **Context**: Svelte UI in superset-tools. Entry from the dashboard page into the scenario testing workspace (manual editor mode per `contracts/ux/dashboard-scenario-ux.md`) and read-only pipeline views on dashboard/release/PREPROD pages. Agent authoring happens exclusively in external MCP clients (050); the product UI never starts or directs an agent.
## 2. Happy Path Narrative
The user opens a dashboard and clicks "Создать сценарий тестирования". The agent analyzes the dashboard, presents a proposed scenario flow with browser, Superset API, XLSX, assertion, evidence, and report steps. The user fills required business parameters, previews generated artifacts and baseline impacts, then confirms save or keeps the artifacts as draft.
The user opens a dashboard and enters the scenario testing workspace. The workspace loads the pinned scenario revision in manual editor mode: steps/coverage, parameters and baseline panels, draft artifact tree with preview. If an external MCP client persisted a proposal, the UI loads its stored diff and provenance read-only for human review; the user accepts it into a manual edit session or discards it. After parameters are resolved and artifacts preview valid, the user requests save/approval through the bound 036 gate. Baseline preview shows candidate/materialized/published states with an explicit publish action (publish is separately invoked per 037 catalog lifecycle). Execution results, evidence bytes and evaluation findings are rendered read-only by 045 (EvidenceViewer/AgentEvaluationCard), never edited from 039.
## 3. Interface Mockups
@@ -27,25 +27,26 @@ The user opens a dashboard and clicks "Создать сценарий тест
├─────────────────────────────────────────────────────────────────────────────┤
│ [Открыть в Superset] [Валидация] [Документация] │
│ │
│ [Создать сценарий тестирования] │
│ [Сценарий тестирования] ● verification: pass (read-only badge) │
└─────────────────────────────────────────────────────────────────────────────┘
```
### Agent Workspace
### Manual Editor Workspace (no agent stream)
```text
┌─────────────────────────────────────────────────────────────────────────────┐
│ 🧪 Dashboard Test Scenario Agent ● connected │
│ 🧪 Сценарий тестирования — manual editor revision r7 | saved ✓ │
├─────────────────────────────────────────────────────────────────────────────┤
│ Context: FI-0080 | dashboard_id=42 | env=ss-dev │
│ Progress: [Context ✓] → [Inspect ✓] → [Scenario …] → [Parameters] → [Save] │
│ Status: validation ✓ | save: idle | approvals: 1 pending │
│ Stored MCP proposal: scenario-r123-proposal.diff [Review diff] [Discard] │
└─────────────────────────────────────────────────────────────────────────────┘
```
### Scenario Preview
```text
┌──────────────────────── Предложенный сценарий ──────────────────────────────┐
┌──────────────────────── Сценарий (revision r7) ─────────────────────────────┐
│ Цель: проверить фильтры, метрику и XLSX выгрузку │
│ Steps: 18 | Tools: browser, Superset API, XLSX, assertions, report │
│ │
@@ -56,7 +57,7 @@ The user opens a dashboard and clicks "Создать сценарий тест
│ 4 Скачать XLSX browser xlsx.file │
│ 5 Сравнить с baseline assertion pass/fail/inconclusive │
│ │
│ [Открыть граф] [Заполнить параметры] [Сгенерировать draft] │
│ [Открыть граф] [Заполнить параметры] [Валидировать] │
└─────────────────────────────────────────────────────────────────────────────┘
```
@@ -77,29 +78,44 @@ The user opens a dashboard and clicks "Создать сценарий тест
### Artifact Preview
```text
┌──────────────────────── Generated draft ────────────────────────────────────┐
│ dashboard_42_test_scenario/ │
┌──────────────────── Stored draft artifacts (read-only preview) ─────────────┐
│ Revision: r7 (candidate) │
│ ├── scenario.yaml ✓ valid │
│ ├── runner.plan.json ✓ valid │
│ ├── xlsx_assertions.py ⚠ needs review │
│ ├── report_template.md ✓ valid │
│ └── baseline_candidates.yaml ⚠ approval required │
│ │
│ [Preview file] [Download draft] [Save to repository] │
│ [Preview file] [Download draft] [Request save via gate] │
└─────────────────────────────────────────────────────────────────────────────┘
```
### Baseline Approval
### Baseline Preview & Approval
```text
┌──────────────────────── Baseline approval ──────────────────────────────────┐
│ Metric: Просроченная ДЗ │
│ Filters: Дата=2026-05-29, Контрагент=АСК │
│ Candidate: 1 234 567.89 from Superset API │
│ Candidate: 1 234 567.89 from Superset API state: candidate │
│ Existing baseline: none │
│ Reason required: [Initial approved QA baseline________________________] │
│ │
│ [Confirm approval] [Keep draft] [Deny] │
├─────────────────────────────────────────────────────────────────────────────┤
│ Publication: candidate → materialized → publish_pending → committed → │
│ published [Publish explicitly] (working-tree materialization is NOT │
│ publication; receipts bind catalog digest/branch/commit) │
└─────────────────────────────────────────────────────────────────────────────┘
```
### Pipeline Verification Views (read-only)
```text
┌──────────────────── Dashboard detail / Release / PREPROD pages ─────────────┐
│ ReleaseVerification.StatusBadge pass | warn | fail | blocked | pending │
│ ReleaseVerification.HistoryList last N VerificationRuns (037) │
│ ReleaseVerification.StructureDiffPanel critical diff blocks Validate │
│ metric results (037 ComparisonResult[]) | visual findings (037 VlmFinding[]) │
└─────────────────────────────────────────────────────────────────────────────┘
```
@@ -115,14 +131,19 @@ The user opens a dashboard and clicks "Создать сценарий тест
* **System Response**: Save action is disabled for executable artifacts; blockers are grouped by step.
* **Recovery**: Fill parameter, provide selector hint, convert to human checkpoint, or remove step.
### Scenario C: User Leaves Mid-Generation
### Scenario C: User Leaves Mid-Edit
* **System Response**: On return, `/agent` displays recoverable run id and latest draft state.
* **Recovery**: Resume preview, discard draft, or restart analysis.
* **System Response**: On return, the workspace reloads the same pinned revision, stored draft/proposal state and pending gates; stale context shows a timestamp with a recover action.
* **Recovery**: Resume preview/edit, discard draft, or re-run validation. No agent run is resumed because none exists in product UI.
### Scenario D: Stale or Invalidated Baseline
* **System Response**: Baseline panel shows stale/invalidated state with the pinned revision still displayed.
* **Recovery**: Candidate discovery and explicit rebaseline through the 037 lifecycle; never silent replacement.
## 5. Tone & Voice
* **Style**: Clear, operational, confidence-aware.
* **Terminology**: Use "scenario", "verification program", "draft", "artifact", "baseline", "Superset API", "SQL Lab evidence", "human checkpoint". Never imply that SQL is editable or generated at run time.
* **Terminology**: Use "scenario", "revision", "verification program", "stored draft/proposal", "artifact", "baseline", "candidate/materialized/published", "Superset API", "SQL Lab evidence", "human checkpoint". Never imply that SQL is editable or generated at run time, and never imply that the UI can start, prompt or steer an agent.
#endregion DashboardScenarioUi.UxReference

View File

@@ -0,0 +1,184 @@
openapi: 3.1.0
info:
title: ScenarioArtifact protected content
version: 1.0.0
description: Normative implemented=false. Server verifies ownership and complete MIME/digest/length before headers.
security: [{bearerAuth: []}]
paths:
/api/scenario-runs/{run_id}/artifacts/{artifact_id}/content:
parameters:
- {name: run_id, in: path, required: true, schema: {type: string, format: uuid}}
- {name: artifact_id, in: path, required: true, schema: {type: string, format: uuid}}
- {name: Range, in: header, schema: {type: string}, description: Unsupported; reject 416 after ACL.}
get:
operationId: scenarioArtifact.content
responses:
'200':
description: Complete bytes; image limit 10485760, other allowlisted MIME limit 26214400.
headers: &verifiedHeaders
Content-Type: {schema: {type: string, enum: [image/jpeg, image/png, image/webp, application/json, application/pdf, application/vnd.openxmlformats-officedocument.spreadsheetml.sheet]}}
Content-Length: {schema: {type: integer, minimum: 1, maximum: 26214400}}
Content-Disposition: {schema: {type: string}, description: 'attachment; filename is a sanitized server-generated artifact UUID plus MIME-derived extension, never a storage path.'}
ETag: {schema: {type: string}, description: Quoted lowercase SHA-256.}
Content-Digest: {schema: {type: string}, description: 'sha-256=:base64 digest bytes:'}
Cache-Control: {schema: {const: 'private, no-store'}}
X-Content-Type-Options: {schema: {const: nosniff}}
Accept-Ranges: {schema: {const: none}}
content:
image/jpeg: {schema: {type: string, format: binary}}
image/png: {schema: {type: string, format: binary}}
image/webp: {schema: {type: string, format: binary}}
application/json: {schema: {type: string, format: binary}}
application/pdf: {schema: {type: string, format: binary}}
application/vnd.openxmlformats-officedocument.spreadsheetml.sheet: {schema: {type: string, format: binary}}
'401': {$ref: '#/components/responses/Error401'}
'403': {$ref: '#/components/responses/Error403'}
'404': {$ref: '#/components/responses/Error404'}
'409': {$ref: '#/components/responses/Error409'}
'410': {$ref: '#/components/responses/Error410'}
'413': {$ref: '#/components/responses/Error413'}
'416': {$ref: '#/components/responses/Error416'}
'503': {$ref: '#/components/responses/Error503'}
head:
operationId: scenarioArtifact.head
description: Identical ACL, integrity checks, status and headers as GET; no body on success or error. Conditional headers do not bypass checks.
responses:
'200':
description: Verified metadata with no response body.
headers: *verifiedHeaders
'401':
description: 'Same AUTHENTICATION_REQUIRED status and typed Error-Code header component as GET; no body.'
headers: {Error-Code: {$ref: '#/components/headers/ErrorCodeAuth'}}
'403':
description: 'Same PERMISSION_DENIED status and typed Error-Code header component as GET; no body.'
headers: {Error-Code: {$ref: '#/components/headers/ErrorCodePermission'}}
'404':
description: 'Same NOT_FOUND status and typed Error-Code header component as GET; no body.'
headers: {Error-Code: {$ref: '#/components/headers/ErrorCodeNotFound'}}
'409':
description: 'Same ARTIFACT_INTEGRITY_FAILED/ARTIFACT_MISSING status and typed Error-Code header component as GET; no body.'
headers: {Error-Code: {$ref: '#/components/headers/ErrorCodeIntegrity'}}
'410':
description: 'Same ARTIFACT_EXPIRED status and typed Error-Code header component as GET; no body.'
headers: {Error-Code: {$ref: '#/components/headers/ErrorCodeExpired'}}
'413':
description: 'Same ARTIFACT_TOO_LARGE status and typed Error-Code header component as GET; no body.'
headers: {Error-Code: {$ref: '#/components/headers/ErrorCodeTooLarge'}}
'416':
description: 'Same RANGE_NOT_SUPPORTED status and typed Error-Code header component as GET; no body.'
headers: {Error-Code: {$ref: '#/components/headers/ErrorCodeRange'}}
'503':
description: 'Same ARTIFACT_STORAGE_UNAVAILABLE status and typed Error-Code header component as GET; no body.'
headers: {Error-Code: {$ref: '#/components/headers/ErrorCodeStorage'}}
components:
securitySchemes:
bearerAuth: {type: http, scheme: bearer}
headers:
ErrorCodeAuth:
description: Status-specific typed error discriminator shared by GET and HEAD; untyped Error-Code headers are forbidden.
schema: {type: string, enum: [AUTHENTICATION_REQUIRED]}
ErrorCodePermission:
schema: {type: string, enum: [PERMISSION_DENIED]}
ErrorCodeNotFound:
schema: {type: string, enum: [NOT_FOUND]}
ErrorCodeIntegrity:
schema: {type: string, enum: [ARTIFACT_INTEGRITY_FAILED, ARTIFACT_MISSING]}
ErrorCodeExpired:
schema: {type: string, enum: [ARTIFACT_EXPIRED]}
ErrorCodeTooLarge:
schema: {type: string, enum: [ARTIFACT_TOO_LARGE]}
ErrorCodeRange:
schema: {type: string, enum: [RANGE_NOT_SUPPORTED]}
ErrorCodeStorage:
schema: {type: string, enum: [ARTIFACT_STORAGE_UNAVAILABLE]}
schemas:
Error:
type: object
additionalProperties: false
required: [code, message, retryable, correlation_id]
properties:
code: {type: string, enum: [AUTHENTICATION_REQUIRED, PERMISSION_DENIED, NOT_FOUND, ARTIFACT_INTEGRITY_FAILED, ARTIFACT_MISSING, ARTIFACT_EXPIRED, ARTIFACT_TOO_LARGE, RANGE_NOT_SUPPORTED, ARTIFACT_STORAGE_UNAVAILABLE]}
message: {type: string, maxLength: 1000}
retryable: {type: boolean}
correlation_id: {type: string, maxLength: 128}
responses:
Error401:
description: 'AUTHENTICATION_REQUIRED; no artifact bytes or raw paths. Foreign child is 404 after parent ACL; expired owner-visible tombstone is 410.'
headers:
Error-Code: {$ref: '#/components/headers/ErrorCodeAuth'}
content:
application/json:
schema:
allOf:
- {$ref: '#/components/schemas/Error'}
- {properties: {code: {enum: [AUTHENTICATION_REQUIRED]}, retryable: {const: false}}}
Error403:
description: 'PERMISSION_DENIED; no artifact bytes or raw paths. Foreign child is 404 after parent ACL; expired owner-visible tombstone is 410.'
headers:
Error-Code: {$ref: '#/components/headers/ErrorCodePermission'}
content:
application/json:
schema:
allOf:
- {$ref: '#/components/schemas/Error'}
- {properties: {code: {enum: [PERMISSION_DENIED]}, retryable: {const: false}}}
Error404:
description: 'NOT_FOUND; no artifact bytes or raw paths. Foreign child is 404 after parent ACL; expired owner-visible tombstone is 410.'
headers:
Error-Code: {$ref: '#/components/headers/ErrorCodeNotFound'}
content:
application/json:
schema:
allOf:
- {$ref: '#/components/schemas/Error'}
- {properties: {code: {enum: [NOT_FOUND]}, retryable: {const: false}}}
Error409:
description: 'ARTIFACT_INTEGRITY_FAILED or ARTIFACT_MISSING; no artifact bytes or raw paths. Foreign child is 404 after parent ACL; expired owner-visible tombstone is 410.'
headers:
Error-Code: {$ref: '#/components/headers/ErrorCodeIntegrity'}
content:
application/json:
schema:
allOf:
- {$ref: '#/components/schemas/Error'}
- {properties: {code: {enum: [ARTIFACT_INTEGRITY_FAILED, ARTIFACT_MISSING]}, retryable: {const: false}}}
Error410:
description: 'ARTIFACT_EXPIRED; no artifact bytes or raw paths. Foreign child is 404 after parent ACL; expired owner-visible tombstone is 410.'
headers:
Error-Code: {$ref: '#/components/headers/ErrorCodeExpired'}
content:
application/json:
schema:
allOf:
- {$ref: '#/components/schemas/Error'}
- {properties: {code: {enum: [ARTIFACT_EXPIRED]}, retryable: {const: false}}}
Error413:
description: 'ARTIFACT_TOO_LARGE; no artifact bytes or raw paths. Foreign child is 404 after parent ACL; expired owner-visible tombstone is 410.'
headers:
Error-Code: {$ref: '#/components/headers/ErrorCodeTooLarge'}
content:
application/json:
schema:
allOf:
- {$ref: '#/components/schemas/Error'}
- {properties: {code: {enum: [ARTIFACT_TOO_LARGE]}, retryable: {const: false}}}
Error416:
description: 'RANGE_NOT_SUPPORTED; no artifact bytes or raw paths. Foreign child is 404 after parent ACL; expired owner-visible tombstone is 410.'
headers:
Error-Code: {$ref: '#/components/headers/ErrorCodeRange'}
content:
application/json:
schema:
allOf:
- {$ref: '#/components/schemas/Error'}
- {properties: {code: {enum: [RANGE_NOT_SUPPORTED]}, retryable: {const: false}}}
Error503:
description: 'ARTIFACT_STORAGE_UNAVAILABLE; no artifact bytes or raw paths. Foreign child is 404 after parent ACL; expired owner-visible tombstone is 410.'
headers:
Error-Code: {$ref: '#/components/headers/ErrorCodeStorage'}
content:
application/json:
schema:
allOf:
- {$ref: '#/components/schemas/Error'}
- {properties: {code: {enum: [ARTIFACT_STORAGE_UNAVAILABLE]}, retryable: {const: true}}}

View File

@@ -0,0 +1,30 @@
## @{ ScenarioExecution.ProductionChain [C:5] [TYPE ADR]
@BRIEF Authoritative baseline-backed browser and semantic execution, bytes delivery, lifecycle and result ownership.
@RATIONALE Reuse existing ScreenshotService, LLMClient and 037 comparators through bounded ScenarioRun adapters; their older ValidationRecord/DraftArtifact owners cannot become run truth by aliasing IDs.
@REJECTED Calling DashboardValidationPlugin._execute_path_a as the ScenarioRun orchestrator, relabeling model status as StepOutcome, or resolving current baseline after launch.
Status: normative target, implemented=false; current screenshot adapter exists but ss-prod bindings=0, loop=not_started and LLM providers=0. Passing fake-byte tests are historical seam evidence only.
Canonical visual chain: pinned 042 ScenarioRevision → 044 baseline preflight/RunnerPlan → browser open/tab/filter/wait → ScreenshotService capture → durable ScenarioArtifact/EvidenceReceipt → 037 exact/SSIM comparison using pinned baseline → immutable ComparisonResult → optional declared AgentEvaluation → pinned DecisionPolicy → StepOutcome → result/SSE/045/047. Metric branch uses the same pin and 037 Superset-native normalization/comparison before optional evaluation. Artifact registration is atomic provider output, not an agent-authored digest. Missing semantic provider cannot disable mandatory evaluation; deterministic-only runs need no LLM.
ScenarioBaselineResolver accepts revision baseline refs, requested baseline_set/version and server target/principal; resolves [BaselineSelectionPin](../../037-superset-baseline-engine/contracts/baseline-pin.schema.json) from published catalog bytes before run/gate/lease/I/O. No candidate/raw expected value, stale fingerprints, mixed release, missing bytes or ambiguous coordinate is runnable. Error codes BASELINE_MISSING, BASELINE_STALE, BASELINE_AMBIGUOUS, BASELINE_NOT_PUBLISHED, BASELINE_EVIDENCE_UNAVAILABLE map to 422; changed pin/CAS maps to 409. No silent latest/default selection; absent baseline_set is legal only for a graph with no baseline refs, represented by null pin.
Canonical request identity includes resolved BaselineSelectionPin, submitted selector/version, scenario/revision/content hash, typed params, target/release, principal, trigger, execution toggles, policy/binding/registry hashes. Replays resolve the recorded pin, not a moving catalog: same key+same selector and pinned intent returns original run; changing selector/version or any canonical field→409. A new key resolves current explicitly requested generation. Same baseline_set name whose generation changed cannot return a different run under an old key.
RunnerPlan, ScenarioRun, StepOutcome evidence, full result, comparison, evaluation input manifest and analytics carry the exact pin; baseline_revision heuristic from first mapping is forbidden. baseline_set_version is immutable; mutable aliases must resolve to a version before idempotency lookup.
AnalyticsContextKey v1 = SHA-256 of UTF-8 canonical JSON object with exactly environment_class, compatibility_family, baseline_family, dashboard_release_id, dashboard_fingerprint, dataset_lineage_fingerprint, execution_principal_fingerprint. Values use sorted keys, no whitespace, explicit null; baseline_family from the pin. Model instability adds spec/model/prompt/output-schema hashes in a separate grouping key.
### ProviderRuntime lifecycle owner
FastAPI lifespan is sole owner per worker process. Ordered startup: validate trusted config/limits → start exactly one ProviderEventLoop and await its ready handshake → resolve credentials/bindings and register providers → verify storage, browser binary, principal/RLS and cancellation capability → enable dispatcher/scheduler admission. Empty bindings is valid application liveness, not live-execution readiness. Duplicate start is idempotent; failed loop startup gives ready=false/PROVIDER_LOOP_START_TIMEOUT with zero I/O/admissions and rolls back only resources acquired here.
Shutdown: disable admission → persist cancel/drain intents for active operation IDs → provider cancel/reconcile → flush receipts/outbox and release/quarantine leases → close sessions/contexts/temp files → stop/join loop. Default startup timeout 5s; drain 30s plus one 5s tick. Multiple workers share durable capacity; no cross-loop context.
Cancel screenshot, browser and LLM by operation_id. stopped/completed/unknown receipts are durable; unknown blocks retry/PASS and retains quarantined capacity until cleanup/reconcile. After deadline close run inconclusive (read-only) or blocked (possible mutation) with reconciliation_required; reconciliation work remains queued, never pretend resource release. Late response cannot win after cancellation, lease loss or timeout. New safe retry creates new attempt/evaluation/lease; old snapshots stay immutable.
### ScenarioArtifactContentService
Protected GET/HEAD /api/scenario-runs/{run_id}/artifacts/{artifact_id}/content. Current principal must pass scenario-result:view plus scenario/dashboard/environment/RLS and artifact ACL on every request. Unauthenticated→401; known parent permission denied→403; foreign/nonexistent child→404 indistinguishably. Baseline evidence may be exposed only by a run-owned authorized reference receipt retaining source-owner linkage; arbitrary AgentRun IDs cannot read bytes.
Resolve opaque artifact ID server-side; no caller path/URL/storage key. Verify receipt ownership, positive length, complete stored bytes SHA-256 and magic-byte MIME before headers or first byte. Allow image/jpeg,image/png,image/webp,application/json,application/pdf and application/vnd.openxmlformats-officedocument.spreadsheetml.sheet for declared kinds only; reject MIME mismatch or corruption→409 ARTIFACT_INTEGRITY_FAILED with no body bytes. Version 1 supports full bounded objects only (image max 10MiB, other max 25MiB); Range→416, Accept-Ranges:none; oversized→413.
GET 200 bytes and HEAD 200 empty body use identical authorization/integrity checks and Content-Type, Content-Length, ETag (quoted SHA-256), Content-Digest (sha-256 base64), Cache-Control:private,no-store, X-Content-Type-Options:nosniff. JSON/raw-response downloads use attachment with safe generated filename; never inline HTML/SVG. Verify before streaming via bounded spool; v1 ignores conditional cache requests so HEAD/GET never bypass digest/ACL checks.
Expired/deleted retained metadata→410 ARTIFACT_EXPIRED after ACL; missing storage for nonexpired metadata→409 ARTIFACT_MISSING; provider/storage outage→503; all errors redact paths, secrets, query bodies and provider credentials. Retention shares 046 holds and 037 baseline holds. UI uses authenticated fetch→Blob URL revoked on navigation; URL carries no bearer or credential.
AgentEvaluationStore is append-only, unique (run, logical_step_id, attempt, operation_id); durable raw response (secret-sanitized/encrypted) precedes successful evaluation publication. Full request/evidence/model/prompt/schema/policy hashes plus actual timings/usage are retained. Invalid response persists typed failure, never empty success. Result DTO is 044-owned; 045 renders it, 039 owns authoring review only.
Acceptance (open): real JPEG/PNG/WebP bytes GET/HEAD; wrong owner/RLS; corrupt/expired/missing/range; lifespan failures and shutdown mid-capture/LLM; PostgreSQL lease contention; baseline pin survives revision/catalog edits; required/advisory/disabled policy cases; UI reload and event replay with exact immutable IDs.
## @} ScenarioExecution.ProductionChain

View File

@@ -0,0 +1,811 @@
{
"$schema": "https://json-schema.org/draft/2020-12/schema",
"title": "ScenarioResult evidence projection v1",
"description": "044 authority consumed by045/047/050; implemented=false.",
"type": "object",
"additionalProperties": false,
"required": [
"schema_version",
"run_id",
"scenario_revision_id",
"scenario_content_hash",
"status",
"analytics_context_key",
"baseline_pin",
"evidence",
"comparisons",
"evaluations",
"step_outcomes",
"declared_step_ids"
],
"properties": {
"schema_version": {
"const": 1
},
"run_id": {
"type": "string",
"format": "uuid"
},
"scenario_revision_id": {
"type": "string",
"format": "uuid"
},
"scenario_content_hash": {
"type": "string",
"pattern": "^[a-f0-9]{64}$"
},
"status": {
"enum": [
"passed",
"failed",
"inconclusive",
"blocked",
"cancelled"
]
},
"analytics_context_key": {
"type": "string",
"pattern": "^[a-f0-9]{64}$"
},
"baseline_pin": {
"anyOf": [
{
"$ref": "../../037-superset-baseline-engine/contracts/baseline-pin.schema.json"
},
{
"type": "null"
}
]
},
"evidence": {
"type": "array",
"items": {
"$ref": "#/$defs/artifact"
},
"maxItems": 1000
},
"comparisons": {
"type": "array",
"items": {
"$ref": "#/$defs/comparison"
},
"maxItems": 1000
},
"evaluations": {
"type": "array",
"items": {
"$ref": "../../038-dashboard-scenario-model/contracts/agent-evaluation.schema.json"
},
"maxItems": 1000
},
"step_outcomes": {
"type": "array",
"items": {
"$ref": "#/$defs/outcome"
},
"maxItems": 1000,
"minItems": 1
},
"declared_step_ids": {
"type": "array",
"minItems": 1,
"maxItems": 100,
"uniqueItems": true,
"items": {
"type": "string",
"format": "uuid"
}
}
},
"$defs": {
"artifact": {
"type": "object",
"additionalProperties": false,
"required": [
"artifact_id",
"owner_type",
"owner_id",
"logical_step_id",
"attempt",
"operation_id",
"content_url",
"content_type",
"byte_length",
"sha256",
"capture_metadata",
"retention_class",
"expires_at",
"availability",
"active"
],
"properties": {
"artifact_id": {
"type": "string",
"format": "uuid"
},
"owner_type": {
"const": "scenario_run"
},
"owner_id": {
"type": "string",
"format": "uuid"
},
"logical_step_id": {
"type": "string",
"format": "uuid"
},
"attempt": {
"type": "integer",
"minimum": 1
},
"operation_id": {
"type": "string",
"format": "uuid"
},
"content_url": {
"type": "string",
"pattern": "^/api/scenario-runs/[^/]+/artifacts/[^/]+/content$"
},
"content_type": {
"enum": [
"image/jpeg",
"image/png",
"image/webp",
"application/json",
"application/pdf",
"application/vnd.openxmlformats-officedocument.spreadsheetml.sheet"
]
},
"byte_length": {
"type": "integer",
"minimum": 1,
"maximum": 26214400
},
"sha256": {
"type": "string",
"pattern": "^[a-f0-9]{64}$"
},
"capture_metadata": {
"anyOf": [
{
"type": "object",
"additionalProperties": false,
"required": [
"capture_profile_hash",
"viewport",
"tab_identifier",
"filters_hash",
"captured_at",
"browser_version",
"provider_version",
"query_model_hash",
"execution_principal_fingerprint",
"mask_policy_hash",
"masking_applied"
],
"properties": {
"capture_profile_hash": {
"type": "string",
"pattern": "^[a-f0-9]{64}$"
},
"viewport": {
"type": "object",
"additionalProperties": false,
"required": [
"width",
"height"
],
"properties": {
"width": {
"type": "integer",
"minimum": 1,
"maximum": 16384
},
"height": {
"type": "integer",
"minimum": 1,
"maximum": 16384
}
}
},
"tab_identifier": {
"type": "string",
"minLength": 1,
"maxLength": 256
},
"filters_hash": {
"type": "string",
"pattern": "^[a-f0-9]{64}$"
},
"captured_at": {
"type": "string",
"format": "date-time"
},
"browser_version": {
"type": "string",
"minLength": 1,
"maxLength": 256
},
"provider_version": {
"type": "string",
"minLength": 1,
"maxLength": 256
},
"query_model_hash": {
"type": "string",
"pattern": "^[a-f0-9]{64}$"
},
"execution_principal_fingerprint": {
"type": "string",
"pattern": "^[a-f0-9]{64}$"
},
"mask_policy_hash": {
"type": "string",
"pattern": "^[a-f0-9]{64}$"
},
"masking_applied": {
"type": "boolean"
}
}
},
{
"type": "null"
}
]
},
"retention_class": {
"type": "string",
"minLength": 1
},
"expires_at": {
"anyOf": [
{
"type": "string",
"format": "date-time"
},
{
"type": "null"
}
]
},
"availability": {
"enum": [
"available",
"expired",
"missing",
"corrupt"
]
},
"active": {
"type": "boolean"
}
},
"allOf": [
{
"if": {
"properties": {
"content_type": {
"enum": [
"image/jpeg",
"image/png",
"image/webp"
]
}
}
},
"then": {
"properties": {
"byte_length": {
"maximum": 10485760
},
"capture_metadata": {
"type": "object"
}
}
},
"else": {
"properties": {
"capture_metadata": {
"type": "null"
}
}
}
}
]
},
"comparison": {
"type": "object",
"additionalProperties": false,
"required": [
"comparison_id",
"logical_step_id",
"attempt",
"baseline_id",
"baseline_revision_id",
"catalog_digest",
"comparison_kind",
"status",
"actual",
"expected",
"delta",
"policy",
"actual_artifact_id",
"expected_artifact_id",
"actual_sha256",
"expected_sha256",
"reason_codes",
"created_at"
],
"properties": {
"comparison_id": {
"type": "string",
"format": "uuid"
},
"logical_step_id": {
"type": "string",
"format": "uuid"
},
"attempt": {
"type": "integer",
"minimum": 1
},
"baseline_id": {
"anyOf": [
{
"type": "string",
"format": "uuid"
},
{
"type": "null"
}
]
},
"baseline_revision_id": {
"anyOf": [
{
"type": "string",
"format": "uuid"
},
{
"type": "null"
}
]
},
"catalog_digest": {
"anyOf": [
{
"type": "string",
"pattern": "^[a-f0-9]{64}$"
},
{
"type": "null"
}
]
},
"comparison_kind": {
"enum": [
"metric",
"visual"
]
},
"status": {
"enum": [
"pass",
"fail",
"inconclusive",
"missing_baseline",
"stale_baseline",
"stale_visual_baseline",
"immutability_violation",
"permission_denied",
"source_error",
"ambiguous_baseline",
"invalidated_baseline"
]
},
"actual": {
"anyOf": [
{
"$ref": "../../037-superset-baseline-engine/contracts/comparison-types.schema.json#/$defs/value"
},
{
"$ref": "../../037-superset-baseline-engine/contracts/comparison-types.schema.json#/$defs/visualValue"
},
{
"type": "null"
}
]
},
"expected": {
"anyOf": [
{
"$ref": "../../037-superset-baseline-engine/contracts/comparison-types.schema.json#/$defs/value"
},
{
"$ref": "../../037-superset-baseline-engine/contracts/comparison-types.schema.json#/$defs/visualValue"
},
{
"type": "null"
}
]
},
"delta": {
"anyOf": [
{
"$ref": "../../037-superset-baseline-engine/contracts/comparison-types.schema.json#/$defs/metricDelta"
},
{
"$ref": "../../037-superset-baseline-engine/contracts/comparison-types.schema.json#/$defs/visualDelta"
},
{
"type": "null"
}
]
},
"policy": {
"anyOf": [
{
"$ref": "../../037-superset-baseline-engine/contracts/comparison-types.schema.json#/$defs/metricPolicy"
},
{
"$ref": "../../037-superset-baseline-engine/contracts/comparison-types.schema.json#/$defs/visualPolicy"
},
{
"type": "null"
}
]
},
"actual_artifact_id": {
"anyOf": [
{
"type": "string",
"format": "uuid"
},
{
"type": "null"
}
]
},
"expected_artifact_id": {
"anyOf": [
{
"type": "string",
"format": "uuid"
},
{
"type": "null"
}
]
},
"actual_sha256": {
"anyOf": [
{
"type": "string",
"pattern": "^[a-f0-9]{64}$"
},
{
"type": "null"
}
]
},
"expected_sha256": {
"anyOf": [
{
"type": "string",
"pattern": "^[a-f0-9]{64}$"
},
{
"type": "null"
}
]
},
"reason_codes": {
"type": "array",
"items": {
"type": "string",
"minLength": 1
},
"maxItems": 1000
},
"created_at": {
"type": "string",
"format": "date-time"
}
},
"allOf": [
{
"if": {
"properties": {
"comparison_kind": {
"const": "metric"
}
}
},
"then": {
"properties": {
"actual": {
"anyOf": [
{
"$ref": "../../037-superset-baseline-engine/contracts/comparison-types.schema.json#/$defs/value"
},
{
"type": "null"
}
]
},
"expected": {
"anyOf": [
{
"$ref": "../../037-superset-baseline-engine/contracts/comparison-types.schema.json#/$defs/value"
},
{
"type": "null"
}
]
},
"delta": {
"anyOf": [
{
"$ref": "../../037-superset-baseline-engine/contracts/comparison-types.schema.json#/$defs/metricDelta"
},
{
"type": "null"
}
]
},
"policy": {
"anyOf": [
{
"$ref": "../../037-superset-baseline-engine/contracts/comparison-types.schema.json#/$defs/metricPolicy"
},
{
"type": "null"
}
]
}
}
}
},
{
"if": {
"properties": {
"comparison_kind": {
"const": "visual"
}
}
},
"then": {
"properties": {
"actual": {
"anyOf": [
{
"$ref": "../../037-superset-baseline-engine/contracts/comparison-types.schema.json#/$defs/visualValue"
},
{
"type": "null"
}
]
},
"expected": {
"anyOf": [
{
"$ref": "../../037-superset-baseline-engine/contracts/comparison-types.schema.json#/$defs/visualValue"
},
{
"type": "null"
}
]
},
"delta": {
"anyOf": [
{
"$ref": "../../037-superset-baseline-engine/contracts/comparison-types.schema.json#/$defs/visualDelta"
},
{
"type": "null"
}
]
},
"policy": {
"anyOf": [
{
"$ref": "../../037-superset-baseline-engine/contracts/comparison-types.schema.json#/$defs/visualPolicy"
},
{
"type": "null"
}
]
}
}
}
},
{
"if": {
"properties": {
"status": {
"enum": [
"pass",
"fail",
"immutability_violation"
]
}
}
},
"then": {
"properties": {
"actual": {
"type": "object"
},
"expected": {
"type": "object"
},
"delta": {
"type": "object"
},
"policy": {
"type": "object"
}
}
}
},
{
"if": {
"properties": {
"status": {
"enum": [
"pass",
"fail",
"immutability_violation"
]
}
}
},
"then": {
"properties": {
"baseline_id": {
"type": "string"
},
"baseline_revision_id": {
"type": "string"
},
"catalog_digest": {
"type": "string"
},
"actual_artifact_id": {
"type": "string"
},
"expected_artifact_id": {
"type": "string"
},
"actual_sha256": {
"type": "string"
},
"expected_sha256": {
"type": "string"
}
}
}
}
]
},
"outcome": {
"type": "object",
"additionalProperties": false,
"required": [
"logical_step_id",
"attempt",
"status",
"comparison_ids",
"agent_evaluation_ids",
"deterministic_evidence_refs",
"decision_policy_id",
"decision_policy_version",
"reason_codes",
"decided_at"
],
"properties": {
"logical_step_id": {
"type": "string",
"format": "uuid"
},
"attempt": {
"type": "integer",
"minimum": 1
},
"status": {
"enum": [
"passed",
"failed",
"inconclusive",
"blocked",
"cancelled",
"skipped"
]
},
"comparison_ids": {
"type": "array",
"items": {
"type": "string",
"format": "uuid"
},
"maxItems": 1000
},
"agent_evaluation_ids": {
"type": "array",
"items": {
"type": "string",
"format": "uuid"
},
"maxItems": 1000
},
"deterministic_evidence_refs": {
"type": "array",
"items": {
"type": "string",
"minLength": 1
},
"maxItems": 1000
},
"decision_policy_id": {
"type": [
"string",
"null"
]
},
"decision_policy_version": {
"type": [
"string",
"null"
]
},
"reason_codes": {
"type": "array",
"items": {
"type": "string",
"minLength": 1
},
"maxItems": 1000
},
"decided_at": {
"type": "string",
"format": "date-time"
}
}
}
},
"$id": "https://superset-tools.local/specs/044-dashboard-scenario-execution/contracts/result-evidence.schema.json",
"allOf": [
{
"if": {
"properties": {
"status": {
"const": "passed"
}
}
},
"then": {
"properties": {
"step_outcomes": {
"minItems": 1,
"items": {
"properties": {
"status": {
"const": "passed"
}
}
}
}
}
}
}
],
"x-invariants": [
"artifact.owner_id equals outer run_id; content_url embeds exact run_id and artifact_id",
"outcome logical IDs exactly cover declared_step_ids; unique winning attempt per ID",
"all evaluation run IDs equal outer run_id; outcome comparison/evaluation/artifact references resolve within projection and match step/attempt",
"evaluation comparison_ids resolve within projection comparisons at the same step/attempt; evaluation baseline_pin equals the outer pin; (run, step, attempt, operation) is unique per evaluation",
"evaluation input_manifest and raw_response artifacts resolve within projection evidence and match owner, run, step, attempt, SHA-256, MIME and byte_length; succeeded evaluations require available/active raw response; finding evidence artifacts resolve at the same step/attempt",
"baseline comparison IDs/digests match baseline_pin entries; no first-entry fallback; metric policies carry nonnegative tolerances, range minimum<=maximum, and table rows equal column width",
"passed requires all declared steps passed and required evidence verified; candidate/client hashes cannot establish authority"
]
}

View File

@@ -0,0 +1,618 @@
{
"fixture_version": "2026-09-08",
"baseline_pin": {
"schema_version": 1,
"baseline_set_id": "ss-prod-visual",
"baseline_set_version": "1",
"catalog_revision_id": "11111111-1111-4111-8111-111111111111",
"catalog_digest": "aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
"release_id": "11111111-1111-4111-8111-111111111111",
"release_version": "v1.0.0",
"release_commit_hash": "bbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbb",
"publication_commit_hash": "bbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbb",
"baseline_family": "aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
"entries": [
{
"baseline_id": "44444444-4444-4444-8444-444444444444",
"baseline_revision_id": "44444444-4444-4444-8444-444444444444",
"kind": "visual",
"coordinate_hash": "aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
"entry_digest": "aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
"source_response_hash": "aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
"expected_image_sha256": "aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
"capture_artifact_id": "44444444-4444-4444-8444-444444444444",
"capture_profile_hash": "aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa"
},
{
"baseline_id": "77777777-7777-4777-8777-777777777777",
"baseline_revision_id": "77777777-7777-4777-8777-777777777777",
"kind": "metric",
"coordinate_hash": "cccccccccccccccccccccccccccccccccccccccccccccccccccccccccccccccc",
"entry_digest": "dddddddddddddddddddddddddddddddddddddddddddddddddddddddddddddddd",
"source_response_hash": "eeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeee",
"expected_image_sha256": null,
"capture_artifact_id": "77777777-7777-4777-8777-777777777778",
"capture_profile_hash": "aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa"
}
]
},
"catalog_revision": {
"schema_version": 1,
"catalog_revision_id": "11111111-1111-4111-8111-111111111111",
"catalog_id": "11111111-1111-4111-8111-111111111111",
"parent_revision_id": null,
"catalog_digest": "aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
"cas_version": 1,
"entry_revisions": [
{
"baseline_id": "44444444-4444-4444-8444-444444444444",
"baseline_revision_id": "44444444-4444-4444-8444-444444444444",
"entry_digest": "aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
"coordinate_hash": "aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
"status": "approved",
"supersedes_revision_id": null,
"entry": {
"schema_version": 1,
"baseline_id": "44444444-4444-4444-8444-444444444444",
"dashboard_id": 12,
"kind": "visual",
"release_version": "v1.0.0",
"release_commit_hash": "bbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbb",
"normalized_filters": {
"schema_version": 1,
"filters": [],
"filters_hash": "aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa"
},
"tab_identifier": "overview",
"expected_image_sha256": "aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
"expected_image_content_ref": "baselines/ss-prod-visual/44444444-4444-4444-8444-444444444444.png",
"source_response_hash": "aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
"captured_at": "2026-09-08T12:00:00Z",
"policy": {
"type": "perceptual",
"ssim_min": 0.99,
"pixel_diff_threshold": 0.01
},
"status": "approved",
"fingerprints": {
"query": "aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
"dataset": "aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
"filter": "aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
"layout": "aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa"
},
"provenance": {
"environment": "ss-prod",
"actor": "capture-service",
"source_artifact_id": "44444444-4444-4444-8444-444444444444",
"source_response_hash": "aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
"execution_principal_fingerprint": "aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
"rls_context_hash": "aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
"normalization_version": "1"
},
"approval": {
"by": "reviewer",
"at": "2026-09-08T12:00:00Z"
},
"created_at": "2026-09-08T12:00:00Z"
},
"capture_artifact_id": "44444444-4444-4444-8444-444444444444",
"capture_profile_hash": "aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
"review_disposition_id": "44444444-4444-4444-8444-444444444446",
"review_status": "confirmed"
},
{
"baseline_id": "77777777-7777-4777-8777-777777777777",
"baseline_revision_id": "77777777-7777-4777-8777-777777777777",
"entry_digest": "dddddddddddddddddddddddddddddddddddddddddddddddddddddddddddddddd",
"coordinate_hash": "cccccccccccccccccccccccccccccccccccccccccccccccccccccccccccccccc",
"status": "approved",
"supersedes_revision_id": null,
"entry": {
"schema_version": 1,
"baseline_id": "77777777-7777-4777-8777-777777777777",
"dashboard_id": 12,
"chart_id": 34,
"release_version": "v1.0.0",
"release_commit_hash": "bbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbb",
"result_key": "revenue_by_region",
"normalized_filters": {
"schema_version": 1,
"filters": [],
"filters_hash": "aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa"
},
"expected": {
"kind": "table",
"canonical": {
"columns": [
"region",
"revenue",
"orders"
],
"rows": [
[
"north",
1000,
10
],
[
"south",
2000,
20
]
]
}
},
"source_response_hash": "eeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeee",
"captured_at": "2026-09-08T12:00:00Z",
"policy": {
"type": "row_set",
"key_columns": [
"region"
],
"order_sensitive": true
},
"status": "approved",
"provenance": {
"environment": "ss-prod",
"actor": "capture-service",
"source_artifact_id": "77777777-7777-4777-8777-777777777778",
"source_response_hash": "eeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeee",
"execution_principal_fingerprint": "aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
"rls_context_hash": "aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
"normalization_version": "1"
},
"created_at": "2026-09-08T12:00:00Z"
},
"capture_artifact_id": "77777777-7777-4777-8777-777777777778",
"capture_profile_hash": "aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
"review_disposition_id": null,
"review_status": "not_applicable"
}
],
"publication": {
"state": "published",
"publication_id": "11111111-1111-4111-8111-111111111111",
"expected_branch_head": "bbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbb",
"commit_hash": "bbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbb",
"error_code": null,
"published_receipt_id": "11111111-1111-4111-8111-111111111111"
},
"created_at": "2026-09-08T12:00:00Z",
"actor_id": "reviewer",
"reason": "Approved source fixture"
},
"result": {
"schema_version": 1,
"run_id": "11111111-1111-4111-8111-111111111111",
"scenario_revision_id": "22222222-2222-4222-8222-222222222222",
"scenario_content_hash": "aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
"status": "passed",
"analytics_context_key": "aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
"baseline_pin": {
"schema_version": 1,
"baseline_set_id": "ss-prod-visual",
"baseline_set_version": "1",
"catalog_revision_id": "11111111-1111-4111-8111-111111111111",
"catalog_digest": "aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
"release_id": "11111111-1111-4111-8111-111111111111",
"release_version": "v1.0.0",
"release_commit_hash": "bbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbb",
"publication_commit_hash": "bbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbb",
"baseline_family": "aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
"entries": [
{
"baseline_id": "44444444-4444-4444-8444-444444444444",
"baseline_revision_id": "44444444-4444-4444-8444-444444444444",
"kind": "visual",
"coordinate_hash": "aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
"entry_digest": "aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
"source_response_hash": "aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
"expected_image_sha256": "aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
"capture_artifact_id": "44444444-4444-4444-8444-444444444444",
"capture_profile_hash": "aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa"
},
{
"baseline_id": "77777777-7777-4777-8777-777777777777",
"baseline_revision_id": "77777777-7777-4777-8777-777777777777",
"kind": "metric",
"coordinate_hash": "cccccccccccccccccccccccccccccccccccccccccccccccccccccccccccccccc",
"entry_digest": "dddddddddddddddddddddddddddddddddddddddddddddddddddddddddddddddd",
"source_response_hash": "eeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeee",
"expected_image_sha256": null,
"capture_artifact_id": "77777777-7777-4777-8777-777777777778",
"capture_profile_hash": "aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa"
}
]
},
"declared_step_ids": [
"22222222-2222-4222-8222-222222222222"
],
"evidence": [
{
"artifact_id": "33333333-3333-4333-8333-333333333333",
"owner_type": "scenario_run",
"owner_id": "11111111-1111-4111-8111-111111111111",
"logical_step_id": "22222222-2222-4222-8222-222222222222",
"attempt": 1,
"operation_id": "11111111-1111-4111-8111-111111111111",
"content_url": "/api/scenario-runs/11111111-1111-4111-8111-111111111111/artifacts/33333333-3333-4333-8333-333333333333/content",
"content_type": "image/jpeg",
"byte_length": 1024,
"sha256": "aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
"capture_metadata": {
"capture_profile_hash": "aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
"viewport": {
"width": 1920,
"height": 1200
},
"tab_identifier": "overview",
"filters_hash": "aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
"captured_at": "2026-09-08T12:00:00Z",
"browser_version": "pinned",
"provider_version": "1",
"query_model_hash": "aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
"execution_principal_fingerprint": "aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
"mask_policy_hash": "aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
"masking_applied": true
},
"retention_class": "screenshot",
"expires_at": null,
"availability": "available",
"active": true
},
{
"artifact_id": "33333333-3333-4333-8333-333333333334",
"owner_type": "scenario_run",
"owner_id": "11111111-1111-4111-8111-111111111111",
"logical_step_id": "22222222-2222-4222-8222-222222222222",
"attempt": 1,
"operation_id": "ffffffff-ffff-4fff-8fff-ffffffffffff",
"content_url": "/api/scenario-runs/11111111-1111-4111-8111-111111111111/artifacts/33333333-3333-4333-8333-333333333334/content",
"content_type": "application/json",
"byte_length": 512,
"sha256": "eeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeee",
"capture_metadata": null,
"retention_class": "source_response",
"expires_at": null,
"availability": "available",
"active": true
},
{
"artifact_id": "66666666-6666-4666-8666-666666666666",
"owner_type": "scenario_run",
"owner_id": "11111111-1111-4111-8111-111111111111",
"logical_step_id": "22222222-2222-4222-8222-222222222222",
"attempt": 1,
"operation_id": "55555555-5555-4555-8555-555555555556",
"content_url": "/api/scenario-runs/11111111-1111-4111-8111-111111111111/artifacts/66666666-6666-4666-8666-666666666666/content",
"content_type": "application/json",
"byte_length": 2048,
"sha256": "6666666666666666666666666666666666666666666666666666666666666666",
"capture_metadata": null,
"retention_class": "evaluation_raw",
"expires_at": null,
"availability": "available",
"active": true
}
],
"comparisons": [
{
"comparison_id": "44444444-4444-4444-8444-444444444444",
"logical_step_id": "22222222-2222-4222-8222-222222222222",
"attempt": 1,
"baseline_id": "44444444-4444-4444-8444-444444444444",
"baseline_revision_id": "44444444-4444-4444-8444-444444444444",
"catalog_digest": "aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
"comparison_kind": "visual",
"status": "pass",
"actual": {
"sha256": "aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
"width": 1920,
"height": 1200
},
"expected": {
"sha256": "aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
"width": 1920,
"height": 1200
},
"delta": {
"ssim": 1,
"pixel_diff_ratio": 0
},
"policy": {
"type": "perceptual",
"ssim_min": 0.99,
"pixel_diff_threshold": 0.01
},
"actual_artifact_id": "33333333-3333-4333-8333-333333333333",
"expected_artifact_id": "44444444-4444-4444-8444-444444444444",
"actual_sha256": "aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
"expected_sha256": "aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
"reason_codes": [
"BASELINE_MATCH"
],
"created_at": "2026-09-08T12:00:00Z"
},
{
"comparison_id": "88888888-8888-4888-8888-888888888888",
"logical_step_id": "22222222-2222-4222-8222-222222222222",
"attempt": 1,
"baseline_id": "77777777-7777-4777-8777-777777777777",
"baseline_revision_id": "77777777-7777-4777-8777-777777777777",
"catalog_digest": "aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
"comparison_kind": "metric",
"status": "pass",
"actual": {
"kind": "table",
"canonical": {
"columns": [
"region",
"revenue",
"orders"
],
"rows": [
[
"north",
1000,
10
],
[
"south",
2000,
20
]
]
}
},
"expected": {
"kind": "table",
"canonical": {
"columns": [
"region",
"revenue",
"orders"
],
"rows": [
[
"north",
1000,
10
],
[
"south",
2000,
20
]
]
}
},
"delta": {
"equal": true,
"absolute": null,
"relative": null,
"missing_rows": 0,
"extra_rows": 0
},
"policy": {
"type": "row_set",
"key_columns": [
"region"
],
"order_sensitive": true
},
"actual_artifact_id": "33333333-3333-4333-8333-333333333334",
"expected_artifact_id": "77777777-7777-4777-8777-777777777778",
"actual_sha256": "eeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeee",
"expected_sha256": "eeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeee",
"reason_codes": [
"BASELINE_MATCH"
],
"created_at": "2026-09-08T12:00:00Z"
}
],
"evaluations": [
{
"schema_version": 1,
"evaluation_id": "55555555-5555-4555-8555-555555555555",
"scenario_run_id": "11111111-1111-4111-8111-111111111111",
"logical_step_id": "22222222-2222-4222-8222-222222222222",
"attempt": 1,
"operation_id": "55555555-5555-4555-8555-555555555556",
"evaluation_spec_hash": "5555555555555555555555555555555555555555555555555555555555555555",
"provider_id": "llm-provider-a",
"provider_version": "1",
"model_id": "vision-model",
"model_version": "2026-08",
"prompt_template_id": "agent-evaluation-prompt",
"prompt_template_version": "1.0.0",
"prompt_template_hash": "7777777777777777777777777777777777777777777777777777777777777777",
"output_schema_hash": "8888888888888888888888888888888888888888888888888888888888888888",
"input_manifest_hash": "9999999999999999999999999999999999999999999999999999999999999999",
"input_manifest": [
{
"artifact_id": "33333333-3333-4333-8333-333333333333",
"sha256": "aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
"content_type": "image/jpeg",
"byte_length": 1024,
"role": "actual"
},
{
"artifact_id": "33333333-3333-4333-8333-333333333334",
"sha256": "eeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeee",
"content_type": "application/json",
"byte_length": 512,
"role": "comparison"
}
],
"baseline_pin": {
"schema_version": 1,
"baseline_set_id": "ss-prod-visual",
"baseline_set_version": "1",
"catalog_revision_id": "11111111-1111-4111-8111-111111111111",
"catalog_digest": "aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
"release_id": "11111111-1111-4111-8111-111111111111",
"release_version": "v1.0.0",
"release_commit_hash": "bbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbb",
"publication_commit_hash": "bbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbb",
"baseline_family": "aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
"entries": [
{
"baseline_id": "44444444-4444-4444-8444-444444444444",
"baseline_revision_id": "44444444-4444-4444-8444-444444444444",
"kind": "visual",
"coordinate_hash": "aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
"entry_digest": "aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
"source_response_hash": "aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
"expected_image_sha256": "aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
"capture_artifact_id": "44444444-4444-4444-8444-444444444444",
"capture_profile_hash": "aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa"
},
{
"baseline_id": "77777777-7777-4777-8777-777777777777",
"baseline_revision_id": "77777777-7777-4777-8777-777777777777",
"kind": "metric",
"coordinate_hash": "cccccccccccccccccccccccccccccccccccccccccccccccccccccccccccccccc",
"entry_digest": "dddddddddddddddddddddddddddddddddddddddddddddddddddddddddddddddd",
"source_response_hash": "eeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeee",
"expected_image_sha256": null,
"capture_artifact_id": "77777777-7777-4777-8777-777777777778",
"capture_profile_hash": "aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa"
}
]
},
"comparison_ids": [
"44444444-4444-4444-8444-444444444444",
"88888888-8888-4888-8888-888888888888"
],
"status": "succeeded",
"verdict": "pass",
"confidence": 0.93,
"findings": [
{
"finding_id": "finding-1",
"severity": "info",
"message": "Dashboard layout and rendered chart regions visually match the pinned screenshot baseline.",
"evidence_artifact_ids": [
"33333333-3333-4333-8333-333333333333"
],
"region": {
"x": 0,
"y": 0,
"width": 1,
"height": 1
},
"criterion_id": "crit-visual-layout",
"criterion_kind": "semantic"
},
{
"finding_id": "finding-2",
"severity": "info",
"message": "Metric row set equals the pinned baseline response; no missing or extra rows.",
"evidence_artifact_ids": [
"33333333-3333-4333-8333-333333333334"
],
"region": null,
"criterion_id": "crit-metric-rowset",
"criterion_kind": "deterministic_comparison"
}
],
"reason_codes": [
"EVALUATION_PASS"
],
"raw_response_artifact_ref": "66666666-6666-4666-8666-666666666666",
"raw_response_sha256": "6666666666666666666666666666666666666666666666666666666666666666",
"trust_policy_hash": "bbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbb",
"usage": {
"input_tokens": 1200,
"output_tokens": 300,
"cost_amount": "0.01",
"currency": "USD",
"pricing_version": "1"
},
"started_at": "2026-09-08T12:00:00Z",
"finished_at": "2026-09-08T12:00:05Z"
}
],
"step_outcomes": [
{
"logical_step_id": "22222222-2222-4222-8222-222222222222",
"attempt": 1,
"status": "passed",
"comparison_ids": [
"44444444-4444-4444-8444-444444444444",
"88888888-8888-4888-8888-888888888888"
],
"agent_evaluation_ids": [
"55555555-5555-4555-8555-555555555555"
],
"deterministic_evidence_refs": [
"33333333-3333-4333-8333-333333333333",
"33333333-3333-4333-8333-333333333334"
],
"decision_policy_id": "baseline-semantic",
"decision_policy_version": "1.0.0",
"reason_codes": [
"BASELINE_PASS"
],
"decided_at": "2026-09-08T12:00:00Z"
}
]
},
"evaluation_spec": {
"schema_version": 1,
"spec_id": "55555555-5555-4555-8555-555555555557",
"provider_id": "llm-provider-a",
"provider_version": "1",
"model_id": "vision-model",
"model_version": "2026-08",
"prompt_template_id": "agent-evaluation-prompt",
"prompt_template_version": "1.0.0",
"prompt_template_hash": "7777777777777777777777777777777777777777777777777777777777777777",
"evidence_refs": [
"33333333-3333-4333-8333-333333333333",
"33333333-3333-4333-8333-333333333334"
],
"comparison_refs": [
"44444444-4444-4444-8444-444444444444",
"88888888-8888-4888-8888-888888888888"
],
"tool_allowlist": [],
"output_schema": "agent-evaluation.schema.json",
"decision_policy": {
"policy_id": "baseline-semantic",
"version": "1.0.0",
"confidence_threshold": 0.7,
"evaluation_mode": "required",
"deterministic_hard_failure": "failed",
"high_confidence_failure": "failed",
"low_confidence": "inconclusive",
"disagreement": "inconclusive",
"missing_evidence": "blocked",
"provider_error": "inconclusive",
"human_checkpoint": "declared_only"
},
"limits": {
"timeout_ms": 60000,
"max_images": 4,
"max_input_tokens": 32000,
"max_output_tokens": 2000,
"max_cost": "1.00",
"currency": "USD"
},
"trust_policy_hash": "bbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbb",
"criteria": [
{
"criterion_id": "crit-visual-layout",
"criterion_kind": "semantic",
"description": "Rendered dashboard layout and chart regions visually match the pinned screenshot baseline.",
"comparison_id": null
},
{
"criterion_id": "crit-metric-rowset",
"criterion_kind": "deterministic_comparison",
"description": "Metric row set for revenue_by_region equals the pinned baseline response.",
"comparison_id": "88888888-8888-4888-8888-888888888888"
}
]
}
}

View File

@@ -0,0 +1,147 @@
"""Live exercise: 050 T045 — the curated MCP `publish_baseline_catalog` operation end-to-end.
Chain: OAuth/PKCE MCP session → tools/call `publish_baseline_catalog` (envelope + expected branch
head) → `approval_required` gate → `decide_approval` (RUN_PROD admin) → scheduler dispatch → the 037
publication worker commits with branch-head CAS → durable `published` state with receipts.
Plus the moved-HEAD failure canary: a bogus expected head fails typed with a durable `publish_failed`
row (operation retained for reconcile).
"""
import argparse
import json
import os
import sys
import time
import uuid
from pathlib import Path
sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
from live_mcp_replay import LiveMcpReplay, ReplayError # noqa: E402
_REPO_ROOT = os.path.dirname(os.path.dirname(os.path.dirname(os.path.dirname(os.path.abspath(__file__)))))
_REF = os.environ.get("PUBLISHED_CATALOG_REF", "master")
_REPO = os.environ.get("PUBLISHED_CATALOG_REPO", "busya/ss-tools")
def _envelope() -> dict:
fixture = json.loads((Path(_REPO_ROOT) / "specs/044-dashboard-scenario-execution/fixtures/production-contract-refresh.json").read_text(encoding="utf-8"))
pin = fixture["baseline_pin"]
return {
"baseline_set_id": pin["baseline_set_id"],
"baseline_set_version": pin["baseline_set_version"],
"release_id": pin["release_id"],
"baseline_family": pin["baseline_family"],
"catalog_digest": pin["catalog_digest"],
"catalog_revision": fixture["catalog_revision"],
}
def _current_head(repo: str, ref: str) -> str | None:
import httpx
headers = {"Authorization": f"token {os.environ['PUBLISHED_CATALOG_GITEA_TOKEN']}"}
url = f"{os.environ['PUBLISHED_CATALOG_GITEA_URL']}/api/v1/repos/{repo}/branches/{ref}"
response = httpx.get(url, headers=headers, timeout=15.0)
if response.status_code != 200:
return None
commit = response.json().get("commit") or {}
return commit.get("id") or commit.get("sha")
# #region ScenarioExecution.LiveMcpPublish.Run [C:4] [TYPE Function] [SEMANTICS mcp,publish,gate,cas,receipts]
# @ingroup ScenarioExecution
# @BRIEF Drive the gated publish operation live: approval_required → decide → published (+ moved-HEAD canary).
# @PRE The stand is up; the MCP admin principal holds scenario RUN_PROD; the deployment env is configured.
# @POST Returns evidence with both operation rows; a missing `published`/`publish_failed` state fails.
def run_publish(moved_head: bool = False) -> dict:
replay = LiveMcpReplay()
evidence: dict = {"base": replay.http.base_url}
user_jwt = replay.login()
replay.oauth_pkce_token(user_jwt)
replay.mcp("initialize", {"protocolVersion": "2025-06-18", "capabilities": {},
"clientInfo": {"name": "t045-publish", "version": "1.0"}})
replay.mcp("notifications/initialized", notify=True)
evidence["tools_listed"] = len(replay.mcp("tools/list", {}).get("tools", []))
envelope = _envelope()
expected_head = ("f" * 64) if moved_head else _current_head(_REPO, _REF)
if not expected_head:
raise ReplayError("branch head could not be observed via the Gitea API")
evidence["expected_head"] = expected_head[:12]
arguments = {"envelope": envelope, "expected_branch_head": expected_head,
"idempotency_key": f"t045-publish-{uuid.uuid4().hex[:12]}"}
gated = replay.call("publish_baseline_catalog", arguments)
evidence["gate"] = {"status": gated.get("status"), "gate_id": gated.get("gate_id"),
"risk_level": gated.get("risk_level")}
if gated.get("status") != "approval_required":
raise ReplayError(f"publish was not gated: {json.dumps(gated)[:300]}")
decided = replay.call("decide_approval", {"gate_id": gated["gate_id"], "decision": "approve"})
evidence["decision"] = {"status": decided.get("status"), "gate_id": decided.get("gate_id")}
if decided.get("status") != "approved":
raise ReplayError(f"gate decision did not approve: {json.dumps(decided)[:300]}")
operation = _await_operation(gated["gate_id"], timeout=120)
evidence["operation"] = operation
if moved_head:
if operation["state"] != "publish_failed" or operation["error_code"] != "PUBLISH_HEAD_MOVED":
raise ReplayError(f"moved head did not fail typed: {json.dumps(operation)[:300]}")
elif operation["state"] != "published":
raise ReplayError(f"publication did not reach published: {json.dumps(operation)[:300]}")
return evidence
# #endregion ScenarioExecution.LiveMcpPublish.Run
# #region ScenarioExecution.LiveMcpPublish.Support [C:3] [TYPE Function] [SEMANTICS publish,poll,state]
# @ingroup ScenarioExecution
# @BRIEF Await the scheduler dispatch and project the durable publication operation by its key.
def _await_operation(gate_id: str, timeout: int = 120) -> dict:
from src.core.database import SessionLocal # noqa: E402
from src.models.publication_operation import PublicationOperation # noqa: E402
from src.models.scenario_approval import ActionApprovalGate # noqa: E402
from src.models.auth import McpToolInvocationRecord # noqa: E402
deadline = time.monotonic() + timeout
idempotency_key: str | None = None
while time.monotonic() < deadline:
with SessionLocal() as db:
gate = db.get(ActionApprovalGate, gate_id)
if gate is not None:
record = db.get(McpToolInvocationRecord, gate.owner_id)
if record is not None:
idempotency_key = (record.continuation_payload or {}).get("idempotency_key")
if record.dispatch_status in ("completed", "failed") and idempotency_key:
operation = (
db.query(PublicationOperation)
.filter(PublicationOperation.idempotency_key == idempotency_key)
.one_or_none()
)
if operation is not None:
return {"operation_id": operation.id, "state": operation.state,
"commit_sha": operation.commit_sha, "error_code": operation.error_code,
"published_receipt": operation.published_receipt, "attempts": operation.attempts}
time.sleep(5)
raise ReplayError(f"dispatch did not settle in {timeout}s (gate={gate_id})")
# #endregion ScenarioExecution.LiveMcpPublish.Support
def main() -> None:
parser = argparse.ArgumentParser(description="T045 live MCP publish exercise.")
parser.add_argument("--moved-head", action="store_true",
help="run the moved-HEAD failure canary instead of the publish happy path")
args = parser.parse_args()
try:
evidence = run_publish(moved_head=args.moved_head)
evidence["result"] = "ok"
except ReplayError as exc:
evidence = {"result": "blocked", "error": str(exc)}
print("[t045] RESULT")
print(json.dumps(evidence, indent=1, ensure_ascii=False, default=str))
result_path = os.environ.get("SS_REPLAY_RESULT", "/tmp/kilo/live_mcp_publish_result.json")
with open(result_path, "w") as handle:
json.dump(evidence, handle, indent=1, default=str)
if __name__ == "__main__":
main()

View File

@@ -0,0 +1,362 @@
#!/usr/bin/env python3
# #region SpecRefresh.Validation [C:5] [TYPE Module] [SEMANTICS specs,schema,validation]
# @BRIEF Offline schema/reference and cross-record fixture validation; not runtime readiness proof.
# @INVARIANT Fixtures assert hardcoded false-accept regressions independently of production services.
# @RATIONALE JSON Schema cannot compare nested owner IDs, pinned digests or declared-step closure.
# @REJECTED Treating structural schema success as authorization or evidence-byte verification.
from copy import deepcopy
from decimal import Decimal
import json
from pathlib import Path
from urllib.parse import urldefrag, urljoin
from jsonschema import Draft202012Validator, FormatChecker, ValidationError
from referencing import Registry, Resource
ROOT = Path(__file__).resolve().parents[3]
PREFIXES = ("017-", "036-", "037-", "038-", "039-", "042-", "043-", "044-", "045-", "046-", "047-", "050-")
SCHEMAS = {}
RESOURCES = {}
for package in (ROOT / "specs").iterdir():
if package.name.startswith(PREFIXES):
for path in package.rglob("*.schema.json"):
data = json.loads(path.read_text())
Draft202012Validator.check_schema(data)
SCHEMAS[path] = data
resource = Resource.from_contents(data)
RESOURCES[path.as_uri()] = resource
RESOURCES["https://superset-tools.local/" + path.relative_to(ROOT).as_posix()] = resource
if "$id" in data:
RESOURCES[data["$id"]] = resource
REGISTRY = Registry().with_resources(RESOURCES.items())
# #region SpecRefresh.Validation.Schema [C:3] [TYPE Function]
# @BRIEF Validate draft-2020-12 with local-only resolution and format checks.
def schema_validate(instance, relative):
path = ROOT / relative
validator = Draft202012Validator(
SCHEMAS[path], registry=REGISTRY, format_checker=FormatChecker()
)
validator.validate(instance)
# #endregion SpecRefresh.Validation.Schema
# #region SpecRefresh.Validation.Require [C:2] [TYPE Function]
# @BRIEF Fail a named cross-field invariant without production imports.
def require(condition, message):
if not condition:
raise ValueError(message)
# #endregion SpecRefresh.Validation.Require
# #region SpecRefresh.Validation.Pin [C:3] [TYPE Function]
# @BRIEF Reject duplicate selection coordinates even when records differ.
def validate_pin(pin):
schema_validate(pin, "specs/037-superset-baseline-engine/contracts/baseline-pin.schema.json")
for field in ("baseline_id", "baseline_revision_id", "coordinate_hash"):
values = [entry[field] for entry in pin["entries"]]
require(len(values) == len(set(values)), "duplicate pin " + field)
# #endregion SpecRefresh.Validation.Pin
# #region SpecRefresh.Validation.Catalog [C:3] [TYPE Function]
# @BRIEF Verify wrapper identity/status and unique approved catalog coordinates.
def validate_catalog(catalog):
schema_validate(catalog, "specs/037-superset-baseline-engine/contracts/catalog-revision.schema.json")
entries = catalog["entry_revisions"]
for item in entries:
for field in ("baseline_id", "status"):
require(item[field] == item["entry"][field], "catalog wrapper mismatch: " + field)
for field in ("baseline_revision_id",):
values = [item[field] for item in entries]
require(len(values) == len(set(values)), "duplicate catalog revision")
coordinates = [item["coordinate_hash"] for item in entries if item["status"] == "approved"]
require(len(coordinates) == len(set(coordinates)), "duplicate approved coordinate")
# #endregion SpecRefresh.Validation.Catalog
# #region SpecRefresh.Validation.Result [C:5] [TYPE Function]
# @BRIEF Check owner/URL, winning closure and reference identity in immutable result projection.
# @PRE Structural schema validation precedes cross-field comparisons; no caller authority inferred.
def validate_result(result):
schema_validate(result, "specs/044-dashboard-scenario-execution/contracts/result-evidence.schema.json")
if result["baseline_pin"] is not None:
validate_pin(result["baseline_pin"])
run_id = result["run_id"]
artifacts = {item["artifact_id"]: item for item in result["evidence"]}
comparisons = {item["comparison_id"]: item for item in result["comparisons"]}
evaluations = {item["evaluation_id"]: item for item in result["evaluations"]}
for field, index in (("evidence", artifacts), ("comparisons", comparisons), ("evaluations", evaluations)):
require(len(index) == len(result[field]), "duplicate result identity: " + field)
for artifact_id, artifact in artifacts.items():
require(artifact["owner_id"] == run_id, "foreign artifact owner")
expected_url = f"/api/scenario-runs/{run_id}/artifacts/{artifact_id}/content"
require(artifact["content_url"] == expected_url, "foreign artifact content URL")
outcomes = result["step_outcomes"]
ids = [item["logical_step_id"] for item in outcomes]
require(len(ids) == len(set(ids)), "duplicate winning logical step")
require(set(ids) == set(result["declared_step_ids"]), "incomplete declared-step closure")
for evaluation in evaluations.values():
require(evaluation["scenario_run_id"] == run_id, "foreign evaluation run")
validate_evaluations(result, artifacts, comparisons, evaluations)
for outcome in outcomes:
identity = (outcome["logical_step_id"], outcome["attempt"])
for field, index in (("comparison_ids", comparisons), ("agent_evaluation_ids", evaluations)):
for ref in outcome[field]:
require(ref in index, "unresolved outcome ref")
require((index[ref]["logical_step_id"], index[ref]["attempt"]) == identity, "foreign step attempt")
for ref in outcome["deterministic_evidence_refs"]:
require(ref in artifacts, "unresolved evidence ref")
artifact = artifacts[ref]
require((artifact["logical_step_id"], artifact["attempt"]) == identity, "foreign artifact attempt")
if outcome["status"] == "passed":
require(artifact["availability"] == "available" and artifact["active"], "passed unverified evidence")
validate_comparisons(result, artifacts)
# #endregion SpecRefresh.Validation.Result
# #region SpecRefresh.Validation.Comparisons [C:4] [TYPE Function]
# @BRIEF Resolve exact baseline and actual evidence identities without map-order heuristics.
def validate_comparisons(result, artifacts):
pin = result["baseline_pin"]
entries = {} if pin is None else {entry["baseline_id"]: entry for entry in pin["entries"]}
for comparison in result["comparisons"]:
validate_comparison_values(comparison)
if comparison["status"] not in ("pass", "fail", "immutability_violation"):
continue
require(pin is not None, "comparison without baseline pin")
entry = entries.get(comparison["baseline_id"])
require(entry is not None, "unresolved baseline")
for left, right in (("baseline_revision_id", "baseline_revision_id"), ("comparison_kind", "kind"),
("expected_artifact_id", "capture_artifact_id")):
require(comparison[left] == entry[right], "baseline identity mismatch")
require(comparison["catalog_digest"] == pin["catalog_digest"], "catalog digest mismatch")
expected_sha = entry["expected_image_sha256"] if entry["kind"] == "visual" else entry["source_response_hash"]
require(comparison["expected_sha256"] == expected_sha, "expected digest mismatch")
actual = artifacts.get(comparison["actual_artifact_id"])
require(actual is not None and actual["sha256"] == comparison["actual_sha256"], "actual evidence mismatch")
require((actual["logical_step_id"], actual["attempt"]) ==
(comparison["logical_step_id"], comparison["attempt"]), "comparison attempt mismatch")
# #endregion SpecRefresh.Validation.Comparisons
# #region SpecRefresh.Validation.ComparisonValues [C:3] [TYPE Function]
# @ingroup SpecRefresh
# @BRIEF Enforce nonnegative tolerances, ordered ranges and canonical table row width.
# @REJECTED Negative tolerance values and inverted range bounds silently widening deterministic comparison outcomes.
def validate_comparison_values(comparison):
policy = comparison["policy"]
if isinstance(policy, dict):
if policy.get("type") in ("absolute_tolerance", "relative_tolerance"):
require(Decimal(policy["tolerance"]) >= 0, "negative tolerance")
if policy.get("type") == "range":
require(Decimal(policy["minimum"]) <= Decimal(policy["maximum"]), "inverted tolerance range")
for field in ("actual", "expected"):
value = comparison[field]
if isinstance(value, dict) and value.get("kind") == "table":
columns = value["canonical"]["columns"]
for row in value["canonical"]["rows"]:
require(len(row) == len(columns), "non-canonical table row width")
# #endregion SpecRefresh.Validation.ComparisonValues
# #region SpecRefresh.Validation.EvaluationEvidence [C:4] [TYPE Function]
# @ingroup SpecRefresh
# @BRIEF Bind evaluation manifest, raw response and finding artifacts to run-owned projection evidence.
# @PRE Evidence index comes from the immutable result projection; caller-declared digests never establish authority.
# @REJECTED Accepting evaluation artifacts by bare reference without owner/run/step/attempt/digest/MIME/length match.
def validate_evaluation_evidence(evaluation, artifacts, run_id):
identity = (evaluation["logical_step_id"], evaluation["attempt"])
succeeded = evaluation["status"] == "succeeded"
for item in evaluation["input_manifest"]:
artifact = artifacts.get(item["artifact_id"])
require(artifact is not None, "unresolved manifest artifact")
require(artifact["owner_id"] == run_id, "foreign manifest artifact owner")
require((artifact["logical_step_id"], artifact["attempt"]) == identity, "foreign manifest artifact attempt")
for field in ("sha256", "content_type", "byte_length"):
require(item[field] == artifact[field], "manifest artifact mismatch: " + field)
if succeeded:
require(artifact["availability"] == "available" and artifact["active"], "succeeded evaluation consumed unverified manifest artifact")
raw_ref = evaluation["raw_response_artifact_ref"]
if raw_ref is not None:
artifact = artifacts.get(raw_ref)
require(artifact is not None, "unresolved raw-response artifact")
require(artifact["owner_id"] == run_id, "foreign raw-response artifact owner")
require((artifact["logical_step_id"], artifact["attempt"]) == identity, "foreign raw-response artifact attempt")
require(evaluation["raw_response_sha256"] == artifact["sha256"], "raw-response digest mismatch")
if succeeded:
require(artifact["availability"] == "available" and artifact["active"], "succeeded evaluation without verified raw response")
for finding in evaluation["findings"]:
for ref in finding["evidence_artifact_ids"]:
artifact = artifacts.get(ref)
require(artifact is not None, "unresolved finding evidence artifact")
require((artifact["logical_step_id"], artifact["attempt"]) == identity, "foreign finding evidence attempt")
# #endregion SpecRefresh.Validation.EvaluationEvidence
# #region SpecRefresh.Validation.Evaluations [C:4] [TYPE Function]
# @ingroup SpecRefresh
# @BRIEF Verify evaluation identity uniqueness, exact pinned baseline equality and comparison membership.
def validate_evaluations(result, artifacts, comparisons, evaluations):
require(result["baseline_pin"] is not None or not evaluations, "evaluation without run baseline pin")
identities = set()
for evaluation in evaluations.values():
identity = (result["run_id"], evaluation["logical_step_id"], evaluation["attempt"], evaluation["operation_id"])
require(identity not in identities, "duplicate evaluation identity")
identities.add(identity)
if result["baseline_pin"] is not None:
require(evaluation["baseline_pin"] == result["baseline_pin"], "evaluation pin mismatch")
for ref in evaluation["comparison_ids"]:
require(ref in comparisons, "unresolved evaluation comparison")
require((comparisons[ref]["logical_step_id"], comparisons[ref]["attempt"]) ==
(evaluation["logical_step_id"], evaluation["attempt"]), "foreign evaluation comparison attempt")
validate_evaluation_evidence(evaluation, artifacts, result["run_id"])
# #endregion SpecRefresh.Validation.Evaluations
# #region SpecRefresh.Validation.EvaluationSpec [C:4] [TYPE Function]
# @ingroup SpecRefresh
# @BRIEF Close the criterion truth table: unique IDs and kind-bound comparison membership within declared refs.
# @REJECTED Deterministic authority inferred for semantic criteria; deterministic criteria without declared comparison.
def validate_evaluation_spec(spec):
schema_validate(spec, "specs/038-dashboard-scenario-model/contracts/agent-evaluation-spec.schema.json")
ids = [criterion["criterion_id"] for criterion in spec["criteria"]]
require(len(ids) == len(set(ids)), "duplicate criterion id")
refs = set(spec["comparison_refs"])
for criterion in spec["criteria"]:
bound = criterion["comparison_id"] is not None
require(bound == (criterion["criterion_kind"] == "deterministic_comparison"), "criterion/comparison binding mismatch")
if bound:
require(criterion["comparison_id"] in refs, "criterion comparison not a declared member")
# #endregion SpecRefresh.Validation.EvaluationSpec
# #region SpecRefresh.Validation.ResultCriteria [C:4] [TYPE Function]
# @ingroup SpecRefresh
# @BRIEF Match result evaluations against the declared spec: bounded access, criterion existence and kind parity.
def validate_result_criteria(result, spec):
criteria = {item["criterion_id"]: item for item in spec["criteria"]}
refs = set(spec["comparison_refs"])
evidence_refs = set(spec["evidence_refs"])
for evaluation in result["evaluations"]:
declared = set(evaluation["comparison_ids"])
require(declared <= refs, "evaluation comparisons exceed declared spec scope")
for item in evaluation["input_manifest"]:
require(item["artifact_id"] in evidence_refs, "manifest artifact outside declared spec scope")
for criterion in spec["criteria"]:
if criterion["criterion_kind"] == "deterministic_comparison":
require(criterion["comparison_id"] in declared, "deterministic criterion outside evaluation comparisons")
for finding in evaluation["findings"]:
criterion = criteria.get(finding["criterion_id"])
require(criterion is not None, "finding references undeclared criterion")
require(finding["criterion_kind"] == criterion["criterion_kind"], "finding criterion kind mismatch")
# #endregion SpecRefresh.Validation.ResultCriteria
# #region SpecRefresh.Validation.Refs [C:3] [TYPE Function]
# @BRIEF Resolve every JSON schema reference locally, including JSON pointers.
def check_refs(node, base):
if isinstance(node, dict):
base = urljoin(base, node.get("$id", ""))
if "$ref" in node:
target, fragment = urldefrag(urljoin(base, node["$ref"]))
require(target in RESOURCES, "missing schema resource " + target)
REGISTRY.resolver(base_uri=base).lookup(node["$ref"])
for value in node.values():
check_refs(value, base)
elif isinstance(node, list):
for value in node:
check_refs(value, base)
# #endregion SpecRefresh.Validation.Refs
# #region SpecRefresh.Validation.Fixtures [C:3] [TYPE Function]
# @BRIEF Reproduce independent review false-accept cases with fixed expected rejection.
def run_fixtures():
fixtures = json.loads((ROOT / "specs/044-dashboard-scenario-execution/fixtures/production-contract-refresh.json").read_text())
validate_pin(fixtures["baseline_pin"])
validate_catalog(fixtures["catalog_revision"])
validate_result(fixtures["result"])
spec = fixtures["evaluation_spec"]
validate_evaluation_spec(spec)
validate_result_criteria(fixtures["result"], spec)
cases = []
def reject(name, key, mutate, validate):
instance = deepcopy(fixtures[key])
mutate(instance)
try:
validate(instance)
except (ValueError, ValidationError):
cases.append(name)
return
raise AssertionError("FALSE ACCEPT: " + name)
def pair_result(instance):
validate_result_criteria(instance, spec)
def pair_spec(instance):
validate_result_criteria(fixtures["result"], instance)
foreign = "99999999-9999-4999-8999-999999999999"
# catalog publication truth table and entry review branches (P1-1 / P1-4)
reject("published-null-commit", "catalog_revision", lambda x: x["publication"].update(commit_hash=None), validate_catalog)
reject("published-no-receipt", "catalog_revision", lambda x: x["publication"].update(published_receipt_id=None), validate_catalog)
reject("failed-no-error", "catalog_revision", lambda x: x["publication"].update(state="publish_failed"), validate_catalog)
reject("materialized-with-commit", "catalog_revision", lambda x: x["publication"].update(state="materialized"), validate_catalog)
reject("materialized-with-receipt", "catalog_revision", lambda x: x["publication"].update(state="materialized", commit_hash=None), validate_catalog)
reject("committed-with-receipt", "catalog_revision", lambda x: x["publication"].update(state="committed"), validate_catalog)
reject("catalog-wrapper-status-mismatch", "catalog_revision", lambda x: x["entry_revisions"][0].update(status="retired"), validate_catalog)
reject("duplicate-catalog-revision", "catalog_revision", lambda x: x["entry_revisions"].append(deepcopy(x["entry_revisions"][0])), validate_catalog)
reject("visual-review-not-confirmed", "catalog_revision", lambda x: x["entry_revisions"][0].update(review_status="not_applicable"), validate_catalog)
reject("visual-review-no-disposition", "catalog_revision", lambda x: x["entry_revisions"][0].update(review_disposition_id=None), validate_catalog)
reject("metric-review-bound", "catalog_revision", lambda x: x["entry_revisions"][1].update(review_disposition_id="44444444-4444-4444-8444-444444444446"), validate_catalog)
# baseline pin discrimination
reject("visual-no-digest", "baseline_pin", lambda x: x["entries"][0].update(expected_image_sha256=None), validate_pin)
reject("metric-with-image", "baseline_pin", lambda x: x["entries"][0].update(kind="metric"), validate_pin)
reject("duplicate-coordinate", "baseline_pin", lambda x: x["entries"].append(deepcopy(x["entries"][0])), validate_pin)
# result projection owner/closure/evidence
reject("empty-passed", "result", lambda x: x.update(step_outcomes=[]), validate_result)
reject("foreign-owner", "result", lambda x: x["evidence"][0].update(owner_id=foreign), validate_result)
reject("foreign-url", "result", lambda x: x["evidence"][0].update(content_url="/api/scenario-runs/foreign/artifacts/foreign/content"), validate_result)
reject("arbitrary-comparison", "result", lambda x: x["comparisons"][0].update(actual={"anything": True}), validate_result)
reject("oversize-image", "result", lambda x: x["evidence"][0].update(byte_length=10485761), validate_result)
reject("open-capture-provenance", "result", lambda x: x["evidence"][0]["capture_metadata"].update(secret="forbidden"), validate_result)
reject("incomplete-closure", "result", lambda x: x["declared_step_ids"].append(foreign), validate_result)
reject("wrong-baseline-digest", "result", lambda x: x["comparisons"][0].update(expected_sha256="f" * 64), validate_result)
reject("unresolved-comparison", "result", lambda x: x["step_outcomes"][0].update(comparison_ids=[foreign]), validate_result)
# numeric comparison value invariants (P1-2)
reject("negative-tolerance", "result", lambda x: x["comparisons"][1].update(policy={"type": "absolute_tolerance", "tolerance": "-1"}), validate_result)
reject("inverted-range", "result", lambda x: x["comparisons"][1].update(policy={"type": "range", "minimum": "5", "maximum": "1"}), validate_result)
reject("row-width-mismatch", "result", lambda x: x["comparisons"][1]["actual"]["canonical"]["rows"][0].append("extra"), validate_result)
# evaluation evidence ownership (P0-1)
reject("evaluation-foreign-run", "result", lambda x: x["evaluations"][0].update(scenario_run_id=foreign), validate_result)
reject("duplicate-evaluation", "result", lambda x: x["evaluations"].append(deepcopy(x["evaluations"][0])), validate_result)
reject("evaluation-pin-mismatch", "result", lambda x: x["evaluations"][0]["baseline_pin"].update(catalog_digest="e" * 64), validate_result)
reject("evaluation-unresolved-comparison", "result", lambda x: x["evaluations"][0].update(comparison_ids=[foreign]), validate_result)
reject("evaluation-foreign-manifest-artifact", "result", lambda x: x["evaluations"][0]["input_manifest"][0].update(artifact_id=foreign), validate_result)
reject("evaluation-manifest-digest-mismatch", "result", lambda x: x["evaluations"][0]["input_manifest"][0].update(sha256="c" * 64), validate_result)
reject("evaluation-manifest-mime-mismatch", "result", lambda x: x["evaluations"][0]["input_manifest"][0].update(content_type="image/png"), validate_result)
reject("evaluation-manifest-length-mismatch", "result", lambda x: x["evaluations"][0]["input_manifest"][0].update(byte_length=2048), validate_result)
reject("evaluation-manifest-attempt-mismatch", "result", lambda x: x["evidence"][1].update(attempt=2), validate_result)
reject("evaluation-foreign-raw-response", "result", lambda x: x["evaluations"][0].update(raw_response_artifact_ref=foreign), validate_result)
reject("evaluation-raw-digest-mismatch", "result", lambda x: x["evaluations"][0].update(raw_response_sha256="a" * 64), validate_result)
reject("evaluation-raw-step-mismatch", "result", lambda x: x["evidence"][2].update(logical_step_id=foreign), validate_result)
reject("evaluation-raw-unavailable", "result", lambda x: x["evidence"][2].update(availability="expired"), validate_result)
reject("finding-unresolved-evidence", "result", lambda x: x["evaluations"][0]["findings"][0].update(evidence_artifact_ids=[foreign]), validate_result)
# evaluation criteria truth table (P0-2)
reject("duplicate-criterion-id", "evaluation_spec", lambda x: x["criteria"].append(deepcopy(x["criteria"][0])), validate_evaluation_spec)
reject("criterion-comparison-not-member", "evaluation_spec", lambda x: x["criteria"][1].update(comparison_id="77777777-7777-4777-8777-777777777777"), validate_evaluation_spec)
reject("deterministic-criterion-unbound", "evaluation_spec", lambda x: x["criteria"][1].update(comparison_id=None), validate_evaluation_spec)
reject("semantic-criterion-bound", "evaluation_spec", lambda x: x["criteria"][0].update(comparison_id="44444444-4444-4444-8444-444444444444"), validate_evaluation_spec)
reject("finding-unknown-criterion", "result", lambda x: x["evaluations"][0]["findings"][0].update(criterion_id="ghost"), pair_result)
reject("finding-kind-mismatch", "result", lambda x: x["evaluations"][0]["findings"][0].update(criterion_kind="deterministic_comparison"), pair_result)
reject("manifest-outside-spec", "result", lambda x: x["evaluations"][0]["input_manifest"][0].update(artifact_id="66666666-6666-4666-8666-666666666666"), pair_result)
reject("evaluation-comparison-outside-spec", "result", lambda x: x["evaluations"][0].update(comparison_ids=[foreign]), pair_result)
reject("deterministic-criterion-outside-evaluation", "evaluation_spec", lambda x: (x["comparison_refs"].append("77777777-7777-4777-8777-777777777777"), x["criteria"][1].update(comparison_id="77777777-7777-4777-8777-777777777777")), pair_spec)
print(f"PASS: 5 positive fixture checks; {len(cases)} hardcoded negative regressions rejected")
# #endregion SpecRefresh.Validation.Fixtures
if __name__ == "__main__":
for path, schema in SCHEMAS.items():
check_refs(schema, path.as_uri())
print(f"PASS: {len(SCHEMAS)} schemas and all JSON references")
run_fixtures()
# #endregion SpecRefresh.Validation

View File

@@ -0,0 +1,16 @@
## @{ ScenarioRunMonitor.EvidenceUi [C:4] [TYPE ADR]
@BRIEF Typed ScenarioRun evidence, baseline comparison and immutable evaluation display.
@RATIONALE 045 owns execution evidence display; 039 authoring disposition is a different owner and must not rewrite run truth.
@REJECTED AgentRun DraftList or ValidationRecord as ScenarioResult DTO; model verdict as the primary result badge.
Status: normative target, implemented=false.
Canonical DTO: 044 ScenarioExecutionResult/ScenarioRunEvent from [openapi.yaml](../../044-dashboard-scenario-execution/contracts/openapi.yaml). RunConfiguration sends baseline_set and baseline_set_version as top-level fields, not params.launch_config. Mandatory comparison/evaluation steps cannot be toggled off. Present resolved catalog/release/commit/IDs/digests before PROD approval and preserve the same pin on reload.
EvidenceViewer (canonical component name; EvidencePanel is its compatibility label) receives artifact_id, authenticated content_url, MIME/bytes/digest, capture metadata, retention/expires_at and active/historical marker. It fetches via protected GET/HEAD, creates a Blob URL and revokes on context change/unmount; no token in URL, direct provider or storage access.
UX states: idle→loading→ready|forbidden|expired|missing|corrupt|error. Each has text+icon, keyboard focus, accessible image description and capture metadata; expired shows tombstone/provenance and no misleading retry, corrupt disables download/preview and links investigation. Network error permits bounded retry; 401 requests login. Historical attempts are explicitly marked and never mixed into active evidence.
AgentEvaluationCard is read-only: evaluation/provider/model/prompt/schema hashes, input manifest, baseline pin, deterministic ComparisonResult, semantic verdict/confidence/findings and policy-derived StepOutcome are separate fields. Show unavailable/malformed/budget/low-confidence/disagreement reasons. Render text escaped; no confirm/dismiss here. HumanCheckpointPanel remains the sole runtime disposition control and uses canonical v1 labels/outcomes.
Result must display baseline_set/version, catalog_revision/digest, release and publication commits, entry IDs/revisions/source/image digests; expected/actual/delta/comparison policy; capture and model transformations are distinguished. Legacy runs missing the pin show provenance_unavailable, never infer from current catalog.
Comparison aligns logical_step_id and step_content_hash, exposes baseline_revision_changed and expected-image/source-value changes plus latency_delta_ms=(finished-started)b−a. Null duration/value yields unavailable, not0. Compare numeric canonical decimals only for compatible kind/unit/filter/target/principal; return non_comparable reasons for changed target, principal, baseline family, schema or missing evidence. A baseline revision change may show side-by-side values, but cannot imply flakiness or product regression.
SSE replay uses monotonic sequence and attempt IDs; duplicate events are idempotent, gaps trigger authorized detail/result reload. Completed snapshot cannot be overwritten by late old-attempt events.
Open acceptance: all UX states, keyboard and narrow viewport, JPEG authenticated display/download, reload/reconnect, old attempt, low confidence/disagreement, baseline change/noncomparability, no fake AgentRun dependency.
## @} ScenarioRunMonitor.EvidenceUi

View File

@@ -0,0 +1,18 @@
## @{ ScenarioAutomation.ProductionOperations [C:5] [TYPE ADR]
@BRIEF Identical REST/MCP admission, durable trigger identity, retention holds and measured rollout.
@RATIONALE Disabled invalid rules are still durable automation configuration and must pass the same eligibility checks as enabled rules.
@REJECTED REST disabled-config bypass, anonymous operational reads, independent provider quotas and deletion of baseline-held artifacts.
Status: normative target, implemented=false; DEF-02/SEC-01 remain open.
All schedule/trigger create/update/enable/disable and direct triggers invoke one service policy; user/MCP/service principals are resolved live. A human-containing, candidate/unactivated, archived, invalid-context or unsupported-provider revision is ineligible even enabled=false; rejection precedes any schedule/gate/run/notification/queue write. Delete of an owned existing invalid rule remains allowed for cleanup. REST and MCP return identical semantic code, field errors and side-effect counts.
Six reads (schedules, trigger rules, policies, notifications, metrics, retention) require authenticated VIEW/READ and per-object scenario/dashboard/environment ACL; anonymous→401 even empty, lacking permission→403, foreign object→404. Pagination never leaks totals of inaccessible objects.
Rules pin baseline_set_id/version and target/revision policy. A due event atomically resolves activated revision+BaselineSelectionPin, verifies eligibility and stores immutable canonical request hash plus source_event_id; duplicate events/restarts reuse the same run/gate. Changed baseline set or canonical request under same idempotency key→409. Candidate save never changes a schedule target. PROD eligible run is pending_approval until exact request/binding/policy approval. Transition record, queue and notification outbox are one transaction; dispatcher queued→running CAS is sole I/O authority.
Use existing APScheduler semantics; no second scheduler. Persistent retries and restart must preserve source identity. Capacity is shared across ScenarioRun/AgentRun/VerificationRun/LoadRun and bounded globally per environment.
Retention defaults: screenshots/heavy artifacts30d, raw VLM7d, step metrics90d, run metadata180d, audit365d; approved-baseline references and active operations add holds. Deletion is mark→eligible-after-all-holds→delete bytes→verify absent→tombstone/audit; retry same deletion ID idempotently. Failure remains deletion_pending; never report deleted while bytes survive. Preserve analytics minimum window independent of metadata pruning.
Operational acceptance profile ops-v1 (proposed thresholds, implemented=false): dispatch overhead p95<100ms for100 steps, API metadata p95<200ms, policy p95<100ms, durable trigger admission<1s excluding waiting approval/capacity, cancel drain30s+5s. At default concurrency (DEV/PREPROD2 browser contexts, PROD1), no lease oversubscription; starvation test every eligible workload makes progress within2 completed older leases when equal priority. Queue p95<30s under offered load below measured service capacity; overload returns capacity_blocked and no I/O.
Measure 5/15/50 tabs, warm/cold browser, 401/403/429/5xx, disk failure, crash/restart, duplicate trigger, timeout/cancel/unknown reconciliation. Record queue/login/capture/compare/LLM latency, bytes, model/prompt, input/output tokens, actual/estimated cost with pricing version (null if unreported), retries and verdict reasons without secrets. Cost cap is pinned nonnegative decimal + currency per evaluation/run; reserve worst-case admitted cost and reconcile actual. Missing estimator/pricing blocks cost-limited dispatch.
Canary progression: deterministic-only fixture→real PREPROD baseline comparison→declared semantic local provider→readonly approved PROD canary→bounded load. Require100% ownership/pin/policy/security cases, zero leaks/unknown winning outcomes and all SLO thresholds; failed gate disables new affected-provider admissions, drains/reconciles active operations, retains evidence and rolls back binding/version. No automatic baseline updates. Approved ExecutionPerformanceBaseline/QualityBaseline is explicitly out of scope; SLO acceptance profile is not approved expected product data.
## @} ScenarioAutomation.ProductionOperations

View File

@@ -0,0 +1,15 @@
## @{ ScenarioAnalytics.AtomicTriage [C:5] [TYPE ADR]
@BRIEF Atomic queue/case/episode transitions and exact baseline-aware analytics.
@RATIONALE A disposed case left queued (DEF-04) allows duplicate work and false recurrence; one transaction owns all projections.
@REJECTED scenario_id:environment_id grouping, triage rewriting historical result, and queue updates after committing case disposition.
Status: normative target, implemented=false.
Open item using expected queue version: lock current episode/item, authorize snapshot evidence, create or return one case, record immutable source/evidence snapshot and set queue state=case_opened atomically. Unique active case per queue item. Legacy queued maps to new on migration only; API returns canonical queue state enum.
Disposition CAS validates expected case+queue+episode versions and live ACL; accepted requires nonblank analyst accepted-risk/wont-fix/duplicate reason, resolved requires verified evidence and completed reconciliation. In the same PostgreSQL transaction append decision audit, update case+TriageRecord, mark queue resolved and close FailureEpisode. Queue reads after commit cannot return the old row as actionable. Same request key/digest replays; changed digest/stale versions→409 with zero partial projection.
Concurrent new signal wins either before closure (included in the version check, stale disposition rejected) or after closure (fresh episode/item); never silently discard it. Active-episode duplicate signals increase count once per source ID; opening/disposition never auto-start AgentRun or MCP activity.
Consume exact stored 044 AnalyticsContextKey v1 and BaselineSelectionPin. Key absent/malformed or unverifiable→analytics_ineligible with count/reason; no substitute environment-only key. Baseline_family/compatibility/logical step/target/principal context must match. Retain exact baseline set/revision/entry/source/image/catalog/release/commit pins in contributing-run provenance; display revision changes, do not silently merge incomparable baselines.
Deterministic flaky denominator includes only comparable final pass/fail deterministic comparisons: last30 eligible, post-failure pass, >=2 transitions, failure ratio strictly between0.05 and0.50. Cancelled/infra/blocked/inconclusive excluded with counts. Model variance is separate, keyed by evaluation_spec_hash/provider/model/prompt/schema hashes plus context; report disagreement/low-confidence/instability without altering deterministic denominator. Insufficient sample→unknown, not healthy. Health uses30d contextual history; retain supporting metrics90d and audit365d per046.
Run-to-run duration deltas and rolling quality/SLO telemetry are projections, not an approved PerformanceBaseline. No baseline candidate/approval authority is granted to analytics.
Open acceptance: accepted item absent from actionable queue; stale/concurrent disposition; signal concurrent with closure; recurrence new episode; RLS-forbidden evidence; exact context vectors differing in one field; baseline change; model-only variance excluded; missing key and insufficient sample.
## @} ScenarioAnalytics.AtomicTriage

View File

@@ -0,0 +1,12 @@
# Requirements checklist — 050-mcp-interface
## Production readiness checklist — 2026-09-08
MCPX-FR-030: [Contract-complete public parity](../contracts/modules.md). Historical [x] marks do not close this new production gate; removed frontend components are not current evidence. All rows below implemented=false / OPEN.
- [ ] CHK001 REST/MCP lifecycle/read/auth errors and disabled automation validation are identical; service principal cannot decide human gate. Evidence: [T044](../tasks.md), [traceability](../traceability.md).
- [ ] CHK002 consume/publish failures return typed errors/pending state with no legacy fallback; every prerequisite is externally MCP-reachable. Evidence: [T045](../tasks.md), [traceability](../traceability.md).
- [ ] CHK003 Fresh external-client chain preserves authoritative context and complete baseline pin without raw ORM/REST repair; no frontend agent controls/routes/requests. Evidence: [T046](../tasks.md), [traceability](../traceability.md).
- [ ] CHK004 Negative product UI test: no agent chat/prompt/assistant editing/proposal-generation/typical-operation-to-agent/workspace/start/handoff controls or agent invocation routes/requests; manual CRUD/editor/human review/read-only results remain usable.
Schema/static success alone is not runtime completion. Optional approved performance baseline is outside scope.

View File

@@ -0,0 +1,60 @@
## @{ McpInterface.Modules [C:5] [TYPE ADR]
@BRIEF Versioned public parity for the audited production chain; implemented=false, acceptance OPEN.
@PRE Every call authenticates a principal, checks live object/environment ACL and validates bounded typed input before I/O.
@POST REST and MCP call the same service transaction and return the same domain result/error; transport wrapping never converts failure to success.
@INVARIANT Client hashes, paths and raw image bytes never establish baseline authority.
@RATIONALE Existing consume fallback and transport validation divergence make a successful tool envelope insufficient evidence of durable completion.
@REJECTED Inferring approval/publication from working-tree YAML, duplicating services per transport, or exposing an agent prompt/launch UI.
### Catalog and error contract
Catalog version `050.2.0` adds the schemas in [tool-contracts.schema.json](tool-contracts.schema.json).
Existing names are retained where present; new names are normative additions, not a runtime catalog claim.
Every tool has a strict input schema, output schema, permission, principal class and idempotency classification.
Read tools require authenticated object ACL; list filtering never substitutes invocation authorization.
Errors: `invalid_arguments` (422), `unauthenticated` (401), `permission_denied` (403),
`not_found` (404), `conflict` (409), `gone` (410), `limit_exceeded` (413/429),
`precondition_required` (428), `provider_unavailable` (503).
MCP returns `isError=true` and the stable code; REST returns the corresponding status and identical details.
Approval required is a typed non-success pending result, never `completed`. No transport fallback invokes legacy approve after consume failure.
### Service ownership and curated operations
| Operation | Domain boundary | Required authority / completion |
|---|---|---|
| capture_baseline, capture_visual_baseline | 037 authoritative source/capture | baseline write; registered server capture receipt, no arbitrary client content_ref |
| review_visual_baseline | 037 durable ReviewDisposition | human USER reviewer; immutable capture SHA/profile bound receipt |
| request_baseline_approval | 036 gate + 037 candidate | owner ACL; binds candidate/review/catalog CAS/release/publication intent |
| decide_approval | 036 durable gate | human USER; fresh ACL, reason, CAS, no side effect on denial |
| consume_baseline_approval | 037 catalog CAS | exact approved intent; materialized receipt, not published success |
| publish_baseline_catalog | 037 publication worker | explicit authorized Git intent; expected branch head; commit/push/reconcile receipt |
| rebaseline, retire_baseline, invalidate_baseline | 037 lifecycle | expected catalog revision, reason, gate where policy requires; historical pins unchanged |
| start_scenario_run | 044 admission | explicit revision/env/baseline selector; server resolution before idempotency/gates |
| get_scenario_run, get_scenario_run_result | 044 read projection | run ACL; exact stored pin/result, bounded pages |
| get_scenario_artifact | 044 evidence metadata | run/artifact owner ACL; typed protected content ref, never raw storage path |
| create/update schedule, webhook, CI dispatch | 046 automation | same REST validation even disabled; no HumanStep target |
| open/dispose investigation | 047 case transaction | case/queue/episode CAS atomic; external agent is pull-only |
All result evidence follows [044 result DTO](../../044-dashboard-scenario-execution/contracts/result-evidence.schema.json).
Artifact content is fetched with the same authorized [GET/HEAD contract](../../044-dashboard-scenario-execution/contracts/artifact-content.openapi.yaml).
MCP metadata refs are not bearer capabilities: expired authorization cannot fetch bytes.
Immutable baseline pin follows [037 schema](../../037-superset-baseline-engine/contracts/baseline-pin.schema.json).
Publication retries reuse publication_id and reconcile remote commit before retrying push; do not create duplicate revisions.
### Browser, LLM and frontend boundary
MCP invokes existing 038/044 browser/capture/provider services; canonical browser mapping is
[038 browser-actions](../../038-dashboard-scenario-model/contracts/browser-actions.md).
Only an external MCP client conducts agent authoring. Product frontend has manual CRUD/editor,
human approval/review, monitoring and read-only evidence/evaluation. No agent chat, prompt,
assistant editing, typical-operation-to-agent, proposal generation, agent workspace/start or handoff controls.
AgentEvaluationCard is read-only; no retry-agent/provider/prompt buttons.
An external agent may submit a new validated revision; it cannot mutate historical evaluation or baseline from a finding.
### Acceptance
Fixtures must prove anonymous and cross-owner rejection, user/service decision distinction,
stale CAS and changed-intent replay rejection, publication failure/retry without false success,
run pin preservation across rebaseline, REST/MCP disabled automation parity, and absence of agent frontend controls.
Provider canary readiness and optional approved performance baseline are not implied by this catalog.
## @} McpInterface.Modules

View File

@@ -0,0 +1,385 @@
{
"$schema": "https://json-schema.org/draft/2020-12/schema",
"title": "Audited MCP operation inputs 050.2.0",
"description": "Normative additions, implemented=false. Server derives principal, digests and pin. Does not replace unaffected legacy catalog tool schemas.",
"oneOf": [
{
"type": "object",
"additionalProperties": false,
"required": [
"schema_version",
"idempotency_key",
"tool",
"environment_id",
"release_id",
"source_response_handle",
"coordinate_handle",
"capture_profile_id"
],
"properties": {
"schema_version": {
"const": "050.2.0"
},
"idempotency_key": {
"type": "string",
"minLength": 1,
"maxLength": 256
},
"tool": {
"enum": [
"capture_baseline",
"capture_visual_baseline"
]
},
"environment_id": {
"type": "string",
"minLength": 1,
"maxLength": 256
},
"release_id": {
"type": "string",
"minLength": 1,
"maxLength": 256
},
"source_response_handle": {
"type": "string",
"minLength": 1,
"maxLength": 256
},
"coordinate_handle": {
"type": "string",
"minLength": 1,
"maxLength": 256
},
"capture_profile_id": {
"type": "string",
"minLength": 1,
"maxLength": 256
}
}
},
{
"type": "object",
"additionalProperties": false,
"required": [
"schema_version",
"idempotency_key",
"tool",
"candidate_id",
"capture_artifact_id",
"capture_sha256",
"disposition",
"reason"
],
"properties": {
"schema_version": {
"const": "050.2.0"
},
"idempotency_key": {
"type": "string",
"minLength": 1,
"maxLength": 256
},
"tool": {
"const": "review_visual_baseline"
},
"candidate_id": {
"type": "string",
"minLength": 1,
"maxLength": 256
},
"capture_artifact_id": {
"type": "string",
"minLength": 1,
"maxLength": 256
},
"capture_sha256": {
"type": "string",
"pattern": "^[a-f0-9]{64}$"
},
"disposition": {
"enum": [
"approve",
"reject",
"inconclusive"
]
},
"reason": {
"type": "string",
"minLength": 1,
"maxLength": 2000
}
}
},
{
"type": "object",
"additionalProperties": false,
"required": [
"schema_version",
"idempotency_key",
"tool",
"candidate_id",
"baseline_id",
"expected_catalog_revision_id",
"release_id",
"gate_id",
"reason"
],
"properties": {
"schema_version": {
"const": "050.2.0"
},
"idempotency_key": {
"type": "string",
"minLength": 1,
"maxLength": 256
},
"tool": {
"enum": [
"request_baseline_approval",
"consume_baseline_approval",
"publish_baseline_catalog",
"rebaseline",
"retire_baseline",
"invalidate_baseline"
]
},
"candidate_id": {
"type": [
"string",
"null"
],
"maxLength": 256
},
"baseline_id": {
"type": [
"string",
"null"
],
"maxLength": 256
},
"expected_catalog_revision_id": {
"type": "string",
"minLength": 1,
"maxLength": 256
},
"release_id": {
"type": "string",
"minLength": 1,
"maxLength": 256
},
"gate_id": {
"type": [
"string",
"null"
],
"maxLength": 256
},
"reason": {
"type": "string",
"minLength": 1,
"maxLength": 2000
}
}
},
{
"type": "object",
"additionalProperties": false,
"required": [
"schema_version",
"idempotency_key",
"tool",
"scenario_id",
"revision_id",
"environment_id",
"baseline_set_id",
"baseline_set_version",
"parameters"
],
"properties": {
"schema_version": {
"const": "050.2.0"
},
"idempotency_key": {
"type": "string",
"minLength": 1,
"maxLength": 256
},
"tool": {
"const": "start_scenario_run"
},
"scenario_id": {
"type": "string",
"minLength": 1,
"maxLength": 256
},
"revision_id": {
"type": "string",
"minLength": 1,
"maxLength": 256
},
"environment_id": {
"type": "string",
"minLength": 1,
"maxLength": 256
},
"baseline_set_id": {
"type": "string",
"minLength": 1,
"maxLength": 256
},
"baseline_set_version": {
"type": "string",
"minLength": 1,
"maxLength": 128
},
"parameters": {
"type": "object",
"maxProperties": 100,
"additionalProperties": {
"type": [
"string",
"number",
"boolean",
"null"
]
}
}
}
},
{
"type": "object",
"additionalProperties": false,
"required": [
"schema_version",
"tool",
"run_id"
],
"properties": {
"schema_version": {
"const": "050.2.0"
},
"tool": {
"enum": [
"get_scenario_run",
"get_scenario_run_result"
]
},
"run_id": {
"type": "string",
"minLength": 1,
"maxLength": 256
}
}
},
{
"type": "object",
"additionalProperties": false,
"required": [
"schema_version",
"tool",
"run_id",
"artifact_id"
],
"properties": {
"schema_version": {
"const": "050.2.0"
},
"tool": {
"const": "get_scenario_artifact"
},
"run_id": {
"type": "string",
"minLength": 1,
"maxLength": 256
},
"artifact_id": {
"type": "string",
"minLength": 1,
"maxLength": 256
}
}
}
],
"$defs": {
"error": {
"type": "object",
"additionalProperties": false,
"required": [
"code",
"message",
"retryable",
"correlation_id"
],
"properties": {
"code": {
"enum": [
"invalid_arguments",
"unauthenticated",
"permission_denied",
"not_found",
"conflict",
"gone",
"limit_exceeded",
"precondition_required",
"provider_unavailable"
]
},
"message": {
"type": "string",
"maxLength": 2000
},
"retryable": {
"type": "boolean"
},
"correlation_id": {
"type": "string",
"minLength": 1,
"maxLength": 256
}
}
},
"result": {
"type": "object",
"additionalProperties": false,
"required": [
"status",
"operation_id",
"resource_id",
"gate_id",
"baseline_pin"
],
"properties": {
"status": {
"enum": [
"completed",
"approval_required",
"pending",
"failed"
]
},
"operation_id": {
"type": "string",
"minLength": 1,
"maxLength": 256
},
"resource_id": {
"type": [
"string",
"null"
]
},
"gate_id": {
"type": [
"string",
"null"
]
},
"baseline_pin": {
"$ref": "../../037-superset-baseline-engine/contracts/baseline-pin.schema.json"
}
}
}
},
"$id": "https://superset-tools.local/specs/050-mcp-interface/contracts/tool-contracts.schema.json"
}

View File

@@ -0,0 +1,51 @@
## @{ McpInterface.DataModel [C:4] [TYPE ADR]
@BRIEF MCP projection and durable operation identity; implemented=false for refresh additions.
@INVARIANT MCP owns no second baseline, run, gate or analytics truth store.
## ToolDefinition
`{name, catalog_version, input_schema, output_schema, permission, principal_classes,
read_only, idempotency_required, deprecated_since?}`. Strict schemas are versioned;
tools/list is permission-filtered. Invocation rechecks live ACL regardless of list history.
[Tool schemas](contracts/tool-contracts.schema.json) define the audited additions.
## ToolInvocation / AgentAction
`{invocation_id, principal_id, delegator_id?, tool_name, catalog_version,
canonical_arguments_digest, idempotency_key?, domain_operation_id, gate_id?,
status, stable_error_code?, created_at, completed_at?}`.
Unique mutation identity is principal + tool + idempotency key; changed canonical
intent conflicts. Retry returns the original domain receipt, never repeats a commit.
Secrets, raw cookies and storage paths are excluded from arguments/audit.
Bounded error detail includes correlation ID without secret payload.
## BaselineSelectionPin
Exact [037 BaselineSelectionPin](../037-superset-baseline-engine/contracts/baseline-pin.schema.json)
is server-resolved from approved published catalog entries. Caller selectors are not pins.
Admission includes resolved set/version/catalog/release/commit/entry IDs and digests in
request identity; persist unchanged in RunnerPlan/run/result and analytics.
Authoring sessions store baseline constraints, not launch bindings.
## Approval and publication receipt
036 ActionApprovalGate binds actor/delegator, object, canonical intent digest,
candidate/review receipt, expected catalog revision, release and publication intent.
User decision is immutable and CAS-protected. Consume reloads the exact approval.
037 PublicationReceipt owns materialized/publish_pending/committed/published/publish_failed;
MCP reports those states verbatim. A pending push is not completed publication.
## Evidence and context projections
044 result/evidence DTOs are authoritative; MCP reads typed bounded metadata and protected
content refs. ScenarioRun and AgentRun ownership cannot be interchanged.
047 uses the stored full AnalyticsContextKey, not environment-only fallback.
External clients may read authoring workspace state; frontend has no workspace/chat/prompt
model or agent request action. Human ReviewDisposition is separate from immutable evaluation.
## Migration
Legacy success envelopes without a durable receipt are non-authoritative and must be
reported as unverified, not backfilled as published. Existing run pins missing required
provenance remain legacy/ineligible for cross-run comparison until rerun. No fabricated hashes.
## @} McpInterface.DataModel

View File

@@ -0,0 +1,38 @@
# Implementation plan — MCP production parity
Status: specified, implemented=false for 2026-09-08 refresh. Historical transport tests
in tests/test-report-2026-09-01.md do not close production E2E.
## Dependency order
1. Pin strict catalog/input/output schemas and negative REST/MCP parity fixtures.
2. Implement 037 catalog CAS, authoritative capture/review and explicit Git publication.
3. Implement 038 canonical chain and 044 provider lifecycle/content/result contracts.
4. Route curated MCP tools through those same services; remove consume fallback.
5. Preserve server baseline pin in run idempotency and 046 automation; implement 047 atomic triage.
6. Remove frontend agent interaction controls/routes. Retain manual editor, human approval
and read-only result cards. Connection/security administration does not launch agents.
7. Run least-privilege external-client fixture from fresh context without REST/ORM repair,
then staged production canary under 046 load/cost/SLO gates.
## Scope and owners
MCP adapter owns validation/auth/projection, not a new browser, LLM, baseline or run store.
[Modules](contracts/modules.md), [data model](data-model.md), [schemas](contracts/tool-contracts.schema.json)
are normative. Runtime targets are the existing backend MCP tool modules and domain services;
this refresh edits no runtime code.
## Release gates
Anonymous/cross-owner reads rejected; service principals cannot approve; stale CAS and
changed-intent replay create no partial state; failed publication never returns success;
full baseline pin survives rerun/rebaseline/result/analytics; disabled HumanStep schedules
reject equally over REST/MCP; no agent frontend controls or requests in DOM/route/network tests.
Live browser/capture/evaluation canary must retain provider versions, evidence digests,
cost and cancellation receipts. Optional approved performance baseline is outside scope.
## Rollback
Disable audited capabilities and new dispatch before rollback; preserve immutable receipts,
pins and retention holds. Drain/cancel/reconcile active operations under 044. Do not silently
fall back to unpinned baselines, inline client payloads, legacy agent UI or fabricated results.

View File

@@ -0,0 +1,34 @@
# Traceability — 050-mcp-interface
## Production acceptance traceability — 2026-09-08
Historical rows above identify prior tests/code only; removed agent UI paths are retired. The following audited gates are **implemented=false / OPEN**, independent of local suite totals.
| Requirement | Domain contract / DTO | Task | Falsifiable acceptance | State |
|---|---|---|---|---|
| MCPX-FR-030 | [Contract-complete public parity](contracts/modules.md); [data model](data-model.md) | [T044](tasks.md) | REST/MCP lifecycle/read/auth errors and disabled automation validation are identical; service principal cannot decide human gate. | OPEN |
| MCPX-FR-030 | [Contract-complete public parity](contracts/modules.md); [data model](data-model.md) | [T045](tasks.md) | consume/publish failures return typed errors/pending state with no legacy fallback; every prerequisite is externally MCP-reachable. | OPEN (consume path + CLI publisher live-proven; MCP `publish_baseline_catalog` + 037 publication-worker contract missing) |
| MCPX-FR-030 | [Contract-complete public parity](contracts/modules.md); [data model](data-model.md) | [T046](tasks.md) | Fresh external-client chain preserves authoritative context and complete baseline pin without raw ORM/REST repair; no frontend agent controls/routes/requests. | CLOSED (MCP `--baseline` chain carries the plan pin without raw ORM; REST canary proves the record stamping) |
| MCPX-FR-030; external-MCP-only UI | manual editor/review; read-only evidence | [production tasks](tasks.md) | No frontend agent prompt/chat/assistant editing/proposal generation/workspace/start/handoff routes or requests; human approval remains usable. | OPEN |
Sources: [production gap](../../docs/reports/ss-prod-agentic-e2e-production-gap-2026-09-08.md), [coverage gap](../../docs/reports/ss-prod-agentic-e2e-spec-coverage-2026-09-08.md), [baseline gap](../../docs/reports/ss-prod-agentic-e2e-baseline-gap-2026-09-08.md). Spec schema/static checks prove contract structure only; live canary/runtime closure and optional approved performance baseline are not claimed.
## T029m live external replay — evidence 2026-09-11
Gate row `E2E-EXT-002` **CLOSED**. Full external MCP chain on the live stand, zero non-MCP seeding:
- Client: `specs/044-dashboard-scenario-execution/prototype/live_mcp_replay.py` (committed; transport-only adaptation of `tests/test_mcp_client_flow_http.py` OAuth/PKCE + `tests/test_mcp_initial_scenario_e2e.py` tool chain).
- Run: `backend/.venv/bin/python specs/044-dashboard-scenario-execution/prototype/live_mcp_replay.py` against `http://127.0.0.1:8000` → run `110a6517-e433-4860-9d77-231521acba19`, scenario `e2687f03-f6d3-45f7-b114-d27b023ae085`, revision `fd126da0-eb22-4da6-be73-387589762689`, agent_run `a6168c7d-aa93-439e-8c2e-161f585b211f`.
- Outcome: `register_draft_pack` `context_authority=verified`; PROD start → one durable gate → approved; live `capture_screenshot` passed (8 durable refs); `apply_native_filter` typed `BROWSER_ACTION_NOT_SUPPORTED`; terminal `inconclusive` (honest fail-closed).
- Full trace + defects fixed: `docs/2026-09-11-sales-prod-mcp-replay.md`. The baseline-pinned variant is now proven: `live_mcp_replay.py --baseline` resolves a full `runner_plan.baseline_pin` from the published catalog through the MCP start (no raw ORM); T045 remains OPEN only for the MCP publish tool.
## T045/T046 published-catalog baseline pin — evidence 2026-09-11
T046 baseline-pin **CLOSED**; T045 **OPEN** (MCP publish tool missing):
- Publisher: `backend/src/scripts/publish_catalog.py` (create=POST / update=PUT-with-sha; validates the resolver-canonical envelope; typed auth/conflict/unavailable; offline `tests/scripts/test_publish_catalog.py`).
- Published: `catalogs/ss-prod-visual/v1.json` @ `busya/ss-tools` — commits `e41d6ad2` (create), `3bb2c63e` (update-sha proof), `4d5b7e4c`/`3f782574` (resolver-canonical envelope).
- Publish-failure canary: bogus token → `PUBLISH_AUTH_FAILED` (HTTP 401), exit 1, no partial state.
- REST canary v4 (`ace916a0…`, re-run `adeabe63…`): selector → pin from published bytes → strict `runner_plan.baseline_pin == AgentEvaluation.baseline_pin` (walker stamping).
- MCP `--baseline` chain (`live_mcp_replay.py --baseline`, no raw ORM): compile baseline capability → start with published selector → full `runner_plan.baseline_pin` (set/version/digest).
- Trace: `docs/reports/agentic-runtime-live-canary-v4-baseline-pin-2026-09-11.md`. T045 residual: MCPX-FR-030 curated MCP `publish_baseline_catalog` on the 037 publication-worker contract.