Files
ss-tools/specs-036-050-20260907-111314.md
busya 605c5d0552 docs(specs): refresh 017/036/037/038/039 production contracts
Add production contract refresh sections and normative contracts: 017/050
capture-reuse trust, 036 evidence promotion, 037 catalog lifecycle/revision,
038 browser actions/decision policy, 044-047 production chain/evidence/atomic
triage, and 050 MCP interface artifacts. Requirements marked OPEN pending
executable evidence.
2026-09-11 17:27:48 +03:00

1.8 MiB
Raw Permalink Blame History

Another LLM created this feature package. Your task: conduct an independent orthogonal spec review. Evaluate readiness, find contradictions, gaps, implementation risks, and prepare a structured report with corrections. Focus on spec review, not rewriting the implementation.

================================================================================ Artifact order: spec → ux_reference → checklists → UX (alternatives → decisions → screen-models → api-ux → per-screen → design-tokens) → plan → research → data-model → modules → openapi → contracts → quickstart → traceability → tasks → prototype. Non-mergeable files (.json/.py/binary) are skipped. Features: 036-agent-test-stabilization, 037-superset-baseline-engine, 038-dashboard-scenario-model, 039-dashboard-scenario-ui, 040-dashboard-load-testing, 041-dataset-lineage-blast-radius, 042-dashboard-scenario-registry, 043-dashboard-scenario-editor, 044-dashboard-scenario-execution, 045-dashboard-run-monitor, 046-dashboard-scenario-automation, 047-dashboard-scenario-analytics, 048-rls-management-workspace, 049-idm-account-integration, 050-mcp-interface Generated: 2026-09-07T11:13:14.611664

================================================================================ FEATURE: 036-agent-test-stabilization Files: 22


SPEC — Feature Specification

Source: spec.md

#region AgentTestStabilization.Spec [C:3] [TYPE ADR] [SEMANTICS spec,requirements,agent,test-scenarios,stabilization] @BRIEF Stabilize the durable agent runtime for scenario creation and evidence-led investigation, revalidation, and remediation workstreams. @RELATION DEPENDS_ON -> [Doc.Adr.ADR0001] @RELATION DEPENDS_ON -> [Doc.Adr.ADR0003] @RELATION DEPENDS_ON -> [Doc.Adr.ADR0005] @RELATION DEPENDS_ON -> [Doc.Adr.ADR0006] @RELATION DEPENDS_ON -> [Doc.Adr.ADR0008] @RATIONALE Dashboard test scenario generation produces durable artifacts and baseline approval flows, so the agent must be a recoverable execution surface rather than a best-effort chat stream. @REJECTED Starting dashboard scenario generation directly on top of unstabilized chat behavior — rejected because failures would be indistinguishable across Gradio transport, context passing, artifact preview, and HITL gates. @REJECTED Long-term Gradio chat as the primary agent interaction surface — rejected 2026-08-24: durable-runtime contracts (AgentRun, AgentAction provenance, ActionApprovalGate) persist unchanged, but the conversational transport retires in favor of external MCP clients per specs/050-mcp-interface/spec.md. @REJECTED AGSTAB-FR-009 mandatory pre-VLM screenshot masking — rejected 2026-08-24: all LLM/VLM providers run locally inside the enterprise trust perimeter; PII exposure to local models is an accepted residual risk. Masking becomes an optional capture option for display hygiene, never a gate for analysis.

Navigation (DSA Indexer keywords)

@SEMANTICS: spec, requirements, feature, agent, gradio, langgraph, scenario, artifact, hitl, progress, dashboard-testing

Feature Branch: 036-agent-test-stabilization
Created: 2026-07-07 | Status: Ready for Implementation Input: "Provide a durable agent runtime for dashboard scenario creation and agentic investigations: stable UI context, explicit workstream intent, structured progress/events, recoverable identifiers, tool-action provenance, and policy-bound approvals."

User Scenarios

Story 1 — Dashboard Context Run Start (P1)

Why P1: Dashboard scenario generation must start from a precise dashboard context, environment, and intent without relying on user prose.

Independent Test: Open /agent from a dashboard route with query context and verify the agent run exposes the exact dashboard object, env, route, and scenario intent in debug/status metadata.

Acceptance:

  1. Given a user clicks "Создать сценарий тестирования" on a dashboard When /agent opens Then UIContext contains objectType=dashboard, objectId, objectName, envId, route, contextVersion, and intent=build_dashboard_test_scenario.
  2. Given the user opens /agent from two different dashboards in sequence When each run starts Then the second run uses the new dashboard context and never reuses stale objectId or envId.
  3. Given malformed or oversized context is supplied When the agent run starts Then invalid context is rejected or discarded with visible recovery and audit metadata, not injected into the agent prompt.

Story 2 — Recoverable Long Agent Runs (P1)

Why P1: Scenario generation may inspect Superset, produce draft artifacts, and wait for user approval; it cannot be a fragile one-shot response.

Independent Test: Start a simulated long scenario run, disconnect/reload /agent, and verify the user can see run status and draft artifact references by run_id.

Acceptance:

  1. Given a scenario run starts When the backend accepts the request Then the UI receives a stable agent_run_id correlated with conversation id and dashboard context.
  2. Given the Gradio stream drops during a run When the user reconnects Then the latest run status, progress stage, and recoverable draft references are visible.
  3. Given no first observable event arrives within the configured timeout When timeout elapses Then the UI shows actionable recovery instead of endless thinking.

Story 3 — Structured Progress and Draft Artifacts (P2)

Why P2: The UI must render scenario progress and artifact preview from structured events, not parse natural-language chat text.

Independent Test: Run a fixture-based scenario generation stub and verify progress stages and draft artifacts render from metadata events only.

Acceptance:

  1. Given a run progresses through context, inspect, scenario, parameters, generate, validate, and save stages When events stream Then the UI updates each stage from structured metadata.
  2. Given draft artifacts are produced When the agent reports them Then the user can preview artifact names, types, warnings, and validation status before any repository write.
  3. Given an artifact contains warnings When preview renders Then warnings are visible and block silent save.

Story 4 — HITL Write and Approval Gates (P2)

Why P2: Risky actions need payload-bound authorization, while delegated low-risk actions remain autonomous and audited.

Independent Test: Trigger a repository-write or baseline-approval action and verify confirm_required appears with target path, operation type, risks, and denial path.

Acceptance:

  1. Given the agent proposes an action requiring approval When it is requested Then an inline card lists exact targets, risk, precondition evidence, and rollback/reconciliation before consumption.
  2. Given the agent proposes baseline approval, production mutation, or automation policy change When approval is requested Then the inline card requires a user-visible reason and shows provenance.
  3. Given the user denies a write or approval When denial is submitted Then no artifact is written or approved, and the chat records an explicit cancellation outcome.

Edge Cases

  • JWT expires mid-run → next protected operation stops with session-expired recovery and does not retry privileged actions.
  • User lacks permission to save artifacts or approve baselines → permission_denied appears instead of confirm_required.
  • Multiple tabs start runs for the same dashboard → each run has an independent agent_run_id; artifact preview does not cross-contaminate.
  • Existing 033/035 agent chat functions remain available → dashboard test intent does not break ordinary chat, tool cards, or confirmation flows.

Requirements

Functional

  • AGSTAB-FR-001: /agent MUST accept and validate dashboard UIContext with explicit intent=build_dashboard_test_scenario while preserving backward compatibility with existing 035 uicontext payload position.
  • AGSTAB-FR-002: Every long-running agent operation for dashboard testing MUST expose a stable agent_run_id correlated with conversation id, user id, dashboard context, and current progress stage.
  • AGSTAB-FR-003: The agent runtime MUST emit structured progress metadata for scenario-oriented stages; UI MUST NOT infer these stages by parsing prose.
  • AGSTAB-FR-004: Draft artifacts MUST be previewable before repository writes and must include type, name, intended path, validation status, warnings, and producing run id.
  • AGSTAB-FR-005: Repository writes, baseline approvals, production mutations, and automation policy changes MUST obey the bound ActionApprovalGate policy; any required decision captures a user-supplied reason in an inline card.
  • AGSTAB-FR-006: Permission-denied operations MUST emit permission_denied metadata and never present a confirm button for unauthorized actions.
  • AGSTAB-FR-007: The existing Gradio/LangGraph chat, streaming, confirmation, context, and guardrail flows from 033/035 MUST remain compatible.
  • AGSTAB-FR-008: A live smoke test MUST verify context → stream → tool event → draft artifact preview → confirmation/denial recovery.
  • AGSTAB-FR-009: [SUPERSEDED 2026-08-24 — see @REJECTED in header] Mandatory masking of screenshot evidence before LLM/VLM analysis is excluded: providers are locally deployed inside the enterprise perimeter, so unmasked captures MAY be submitted directly. Masked derivatives (register_masked_derivative) remain an OPTIONAL capability for display/export hygiene and MUST NOT be a precondition for VLM submission, evidence registration, or any gate.
  • AGSTAB-FR-010: The runtime MUST support typed investigation/revalidation/remediation intents linked to the shared InvestigationCase contract. Events enter Investigation Queue and MUST NOT automatically start an agent chat or tool run.
  • AGSTAB-FR-011: Every agent tool call MUST persist AgentAction provenance and pass deterministic ACL, environment, capacity, ActionRegistry, and approval-policy checks before execution.
  • AGSTAB-FR-012: Delegated policy MAY authorize the agent to execute read, diagnostic, authorized fixture mutation, draft, and executable-revision writes autonomously. It MUST NOT bypass deterministic validators, immutable revision creation, or a required ActionApprovalGate.
  • AGSTAB-FR-013: Persistent case workspaces and inline action cards are the primary agent UX; modal/dialog workflows MUST NOT be required to complete work.

Key Entities

  • DashboardScenarioUIContext: Versioned context payload identifying the dashboard, environment, source route, and scenario-generation intent.
  • AgentRun: Recoverable execution unit for long-running agent work; owns progress stage, events, draft artifact references, and terminal status. trigger field identifies the invocation context: manual, deploy_to_preprod, scheduled, etl_completed, release_create.
  • VerificationRun: Record linking an AgentRun to a DashboardRelease (037). Captures which verification categories (metric, visual, structure, content_integrity) were executed and their outcomes. Created at PREPROD deploy, release creation, scheduled checks, and ETL-triggered runs.
  • ScenarioProgressEvent: Structured metadata event representing a stage or step in scenario generation.
  • DraftArtifactRef: Previewable reference to generated content that is not yet persisted as an approved repository artifact.
  • ScreenshotEvidence: A DraftArtifact of kind screenshot_evidence carrying capture metadata (viewport, filters_hash, tab identifier, readiness policy, masking config), a content-addressable digest, and an opaque preview URL. Never contains raw storage paths.
  • CaptureMeta: Structured metadata attached to screenshot artifacts: viewport dimensions, device scale factor, capture method, tab/filter context hash, readiness strategy, browser version, and optional applied masking selectors (empty when masking is unused).
  • ActionApprovalGate: Payload-bound authorization envelope for actions that policy does not delegate; rendered inline in its owning work surface.
  • InvestigationCase / AgentAction: Shared 036–047 case and delegated-tool-action contracts defined in contracts/investigation-cases.md.

Success Criteria

  • SC-001: 100% of dashboard scenario runs in test fixtures expose a valid agent_run_id before the first tool action.
  • SC-002: Context handoff from dashboard page to /agent is verified for at least three dashboard ids with no stale-object reuse.
  • SC-003: Structured progress events render all required stages without text parsing in frontend tests.
  • SC-004: Repository write and baseline approval attempts are blocked until confirmation in 100% of permissioned test cases.
  • SC-005: Existing 035 agent context/guardrails tests continue to pass unchanged or with documented fixture updates only.

Implementation Status & Confirmed Reuse (audit 2026-08-07)

Facts (code check, not tasks.md):

  • ✅ Работает: AgentRun/AgentRunEvent/DraftArtifact/ApprovalGate persistence, agent/src/ss_tools/agent/_run_tracker.py создаёт run'ы через httpx к backend API, HITL gates (confirm/deny/expiry/replay-guard) реализованы, Services.AgentRuns.Evidence (register_screenshot_draft/register_masked_derivative) подключено.
  • ✅ Подтверждённое переиспользование: evidence-адаптер зовёт Plugin.Service.ScreenshotService, Plugin.Service.LLMClient и Plugin.Service.RedactionService (relation'ы зафиксированы в contracts/evidence.md). Второй Playwright/LLM-клиент запрещён.
  • 🟡 Зависимость: 038 Phase 10 (T057–T059) и 040 Phase 9 (T075–T079) потребляют 036 evidence bridge и ApprovalGate; до их мерджа screenshot/VLM-ветки 038 остаются runtime-заглушками, хотя 036-слой готов.

Drift Amendment — MCP Interface (2026-08-24)

Gradio chat retirement is formalized in specs/050-mcp-interface/spec.md. Carry-over mapping:

  • Preserved: AgentRun, AgentRunEvent, DraftArtifact, ApprovalGate, AgentAction provenance, deterministic pre-execution checks (AGSTAB-FR-011/012) — MCP tool invocations produce the same server-side records.
  • Retired with the chat: Gradio transport, streaming progress rendering, LangGraph interrupt_before confirmation loop and _confirmation.py; HITL continues through durable gates rendered by web surfaces or MCP decision tools.
  • Re-targeted: AGSTAB-FR-003 progress-event UI requirements lose their chat consumer; structured events remain the contract for any future server-side observer. AGSTAB-FR-008 smoke test is superseded by 050 Phase 2 e2e.
  • PII posture (2026-08-24): AGSTAB-FR-009 superseded — masking is optional; local-only LLM/VLM deployment inside the enterprise perimeter is the accepted confidentiality boundary for evidence. The mask-miss detection edge case is withdrawn accordingly.

Status (2026-09-02): done — реализовано в рамках 050: инструменты и гейты (specs/050-mcp-interface/tasks.md T012–T028 [x]), handoff-поверхность (050 T030–T033), демонтаж чата и сервиса agent/ (050 T040–T041, чекпоинты specs/WORKSTATE-043-047.md).

#endregion AgentTestStabilization.Spec


UX REFERENCE — Interaction Narrative

Source: ux_reference.md

#region AgentTestStabilization.UxReference [C:3] [TYPE ADR] [SEMANTICS ux,reference,agent,dashboard-testing] @BRIEF UX reference for durable agent workspaces for scenario creation and investigations.

Feature Branch: 036-agent-test-stabilization Created: 2026-07-07 | Status: Ready for Implementation

1. User Persona & Context

  • Who is the user?: BI analyst working on scenario creation, an opened investigation, revalidation, or remediation.
  • What is their goal?: Work in a reliable, recoverable agent thread with evidence, tool actions, and durable outputs.
  • Context: Browser-based Svelte /agent workspace launched from a dashboard page with Superset environment selected.

2. Happy Path Narrative

The analyst starts a scenario workspace or opens an Investigation Queue item. The persistent workspace shows the source context, evidence and agent thread. The agent starts recoverable tool runs, records actions, and may save a validated revision under delegated policy. High-risk actions render a bound inline approval card; denial records a cancellation without side effects.

3. Interface Mockups

Agent Context Header

┌─────────────────────────────────────────────────────────────────────────────┐
│ 🧪 Dashboard Test Scenario Agent                         ● connected        │
├─────────────────────────────────────────────────────────────────────────────┤
│ Context: FI-0080 | dashboard_id=42 | env=ss-dev                             │
│ Intent: build_dashboard_test_scenario | run_id: ag-run-20260707-001         │
└─────────────────────────────────────────────────────────────────────────────┘

Progress Strip

[Context ✓] → [Inspect …] → [Scenario] → [Parameters] → [Generate] → [Validate] → [Save]

Draft Artifact Preview

┌──────────────────────── Draft artifacts ────────────────────────────────────┐
│ Run: ag-run-20260707-001                                                     │
│  ✓ scenario.yaml                         valid                               │
│  ⚠ baseline-candidates.yaml              needs approval reason               │
│  ✓ generated-preview.json                valid                               │
│                                                                             │
│ [Preview] [Download draft] [Request save]                                    │
└─────────────────────────────────────────────────────────────────────────────┘

Inline ActionApprovalGate

┌──────────────────────── Confirmation required ──────────────────────────────┐
│ Publish baseline candidate                                                    │
│ Target: tests/generated/dashboards/fi_0080/                                  │
│ Evidence: comparison SR-1845, lineage snapshot                               │
│ Risk: baseline approval                                                       │
│ Reconciliation: revert candidate if verification fails                       │
│                                                                             │
│ [Approve] [Deny] [Open diff]                                                 │
└─────────────────────────────────────────────────────────────────────────────┘

4. Error Experience

Scenario A: Stale or Invalid Context

  • System Response: Context card turns warning-tone and lists invalid fields.
  • Recovery: User can return to dashboard, retry context handoff, or continue in non-scenario chat mode.

Scenario B: Stream Drops During Run

  • System Response: UI shows run_id, last known stage, and reconnect actions.
  • Recovery: User reconnects and resumes viewing draft status; no write happens during disconnected state.

Scenario C: Unauthorized Approval

  • System Response: Permission-denied card with required role and no confirm button.
  • Recovery: User can download draft only if allowed, or request approval from an authorized role.

5. Tone & Voice

  • Style: Technical, explicit, side-effect aware.
  • Terminology: Use "run", "draft artifact", "confirmation", "approval reason", and "dashboard context" consistently.

#endregion AgentTestStabilization.UxReference


CHECKLISTS — Requirements Quality — requirements.md

Source: checklists/requirements.md

Requirements Checklist: 036 Agent Test Stabilization

Purpose: Validate specification quality and readiness for planning. Created: 2026-07-07 Feature: specs/036-agent-test-stabilization/spec.md

Spec Completeness

  • CHK001 User stories are independently testable.
  • CHK002 Functional requirements avoid implementation code while defining observable behavior.
  • CHK003 No unresolved [NEEDS CLARIFICATION] markers remain.
  • CHK004 Edge cases include stale context, dropped stream, permission denial, and backwards compatibility.
  • CHK004a PII posture is explicit: all configured MCP and LLM/VLM clients are local; masking is optional display/export hygiene and is not a VLM or evidence gate. Credentials, cookies, tokens, and unrelated secrets remain prohibited in provider payloads and telemetry.

Constitution Coverage

  • CHK005 Semantic contract and ADR relations are present in spec header.
  • CHK006 Decision memory records why stabilization precedes scenario generation.
  • CHK007 External orchestrator boundary is preserved.
  • CHK008 RBAC/HITL requirements are explicit for writes and approvals.
  • CHK009 Svelte UX remains model-driven and structured-event based.

Readiness for Plan

  • CHK010 Required outputs are limited to agent runtime stabilization artifacts.
  • CHK011 Feature explicitly excludes Superset query baseline engine and scenario generation logic.
  • CHK012 Success criteria are measurable by backend/frontend tests and one live smoke flow.

Implementation Package

  • CHK013 Research and implementation plan resolve all design decisions.
  • CHK014 Data model, module, event, OpenAPI, and UX contracts are present.
  • CHK015 Quickstart defines a safe vertical-slice verification path.
  • CHK016 Traceability maps every functional requirement to contracts, tasks, and tests.
  • CHK017 Tasks use exact repository paths, dependency order, and test-first sequencing.
  • CHK018 Machine-readable contracts and semantic anchors pass validation.

UX ALTERNATIVES — Design Space Explored

Source: contracts/ux/alternatives.md

#region AgentTestStabilization.Alternatives [C:3] [TYPE ADR] [SEMANTICS ux,alternatives,agent-run] @BRIEF Rejected UX/runtime compositions for recoverable dashboard-testing runs.

Considered Alternatives

  1. Chat transcript as run state — rejected because localized prose cannot provide durable sequence, recovery, or audit semantics.
  2. Frontend-only run tracker — rejected because reload, disconnect, or Gradio restart would lose authoritative state.
  3. A second confirmation modal for scenario writes — rejected because the existing confirmation card already owns HITL presentation and must bind to the backend gate.
  4. Persist drafts directly in the target repository — rejected because preview and download must remain side-effect free until approval is consumed.
  5. Hide ordinary dashboard AI behind the scenario action — rejected because general chat and scenario generation are distinct business intents and both remain available.

Selected Composition

Backend AgentRun is authoritative; the standalone agent persists structured events before streaming; AgentRunModel projects snapshots; existing ConfirmationCard renders bound gates.

#endregion AgentTestStabilization.Alternatives


UX DECISIONS — Final Choices

Source: contracts/ux/decisions.md

#region AgentTestStabilization.UxDecisions [C:3] [TYPE ADR] [SEMANTICS ux,decisions,agent-run] @BRIEF Final UX decisions for adding scenario-run state to the existing agent page.

Decisions

  1. Compose AgentRuns.Model into AgentChat.Model; do not create a second chat client.
  2. Keep the run panel in the agent workspace above draft/confirmation content, not in a global drawer.
  3. Display the full run id with copy affordance; stage state never depends on prose.
  4. Recover by snapshot after a stream gap; never silently restart the run.
  5. Draft download is always available only through opaque authenticated URLs and is side-effect free.
  6. Warnings are repeated in the approval card so the user reviews the exact write risk.
  7. Permission denied is a separate dismiss-only state, consistent with 035.
  8. Baseline reason is a required labeled field with inline validation before confirm.

#endregion AgentTestStabilization.UxDecisions


UX SCREEN MODELS — Model Inventory

Source: contracts/ux/screen-models.md

#region AgentTestStabilization.ScreenModels [C:4] [TYPE ADR] [SEMANTICS ux,screen-models,agent-run] @BRIEF Screen-model inventory and invariants for recoverable scenario runs inside the existing agent page. @RELATION DEPENDS_ON -> [AgentRuns.Model]

Model Inventory

Model Ownership State
AgentChat.Model Existing chat, connection, messages, generic HITL Composes AgentRuns.Model; does not duplicate run atoms
AgentRuns.Model Scenario-run projection and recovery run id/status, stages, sequence, drafts, pending gate, recovery error

AgentRuns.Model Actions

  • startFromContext(context): begins starting state; run id arrives from agent_run_started.
  • applyMetadata(event): validates run id and monotonic sequence before mutation.
  • recover(runId): loads AgentRunSnapshot and atomically replaces projection.
  • previewArtifact(id): obtains safe preview; does not mutate repository.
  • requestAction(action): asks backend for delegated-action policy; it either records an autonomous AgentAction or returns a bound inline gate.
  • decideGate(decision, reason): validates reason locally, then submits authoritative decision.
  • reset(): clears only run projection when starting a new conversation/context.

Invariants

  1. Run id cannot change after start without reset.
  2. Snapshot replacement is accepted only for the same run id.
  3. UI stage derives from event/snapshot fields, never chat text.
  4. Sequence gaps trigger degraded/recovery, not speculative stage updates.
  5. Invalid drafts cannot enter save request.
  6. Baseline approval cannot confirm with blank reason.
  7. Ordinary chat with UIContext v1 has no AgentRuns.Model projection.

L1 Tests

  • v1 ordinary chat stays absent.
  • v2 started event initializes one run.
  • duplicate, stale, foreign-run, and sequence-gap events.
  • recovered snapshot restores drafts and pending gate.
  • deny clears pending gate without consuming artifacts.

#endregion AgentTestStabilization.ScreenModels


UX API CONTRACT — Endpoints & Shapes

Source: contracts/ux/api-ux.md

#region AgentTestStabilization.ApiUx [C:3] [TYPE ADR] [SEMANTICS ux,api,agent-run,recovery] @BRIEF UX mapping for Gradio metadata and Agent Run REST recovery. @RELATION DEPENDS_ON -> [AgentTestStabilization.Events] @RELATION DEPENDS_ON -> [AgentRuns.Api]

Interaction Map

Trigger API/event Immediate UI Failure recovery
Scenario first send agent_run_started Show run id and context stage If absent by 60s, actionable timeout
Progress scenario_progress Update named stage Sequence gap triggers snapshot GET
Draft registered draft_artifacts Replace draft inventory by id Keep prior inventory and show recover
Reload/drop GET /api/agent/runs/{id} Restore authoritative projection 403 ownership; 404 expired/not found
Save request extended confirm_required Show exact paths/warnings Deny is terminal for operation, not run
Unauthorized request permission_denied Dismiss-only access card No confirm control

Status Handling

  • 401: session expired; stop privileged action and request login.
  • 403: show required permission; never retry automatically.
  • 409: projection/gate conflict; recover snapshot before another action.
  • 413/422: show field/artifact validation details.
  • 5xx: preserve run id and last snapshot; offer retry GET, not duplicate POST.

Accessibility

Progress uses an ordered list with aria-current=step. New stage announcements are polite. Confirmation focus moves into the card and returns to the invoking control after deny/consume.

#endregion AgentTestStabilization.ApiUx


UX DESIGN — Per-Screen Contracts — design-tokens.md

Source: contracts/ux/design-tokens.md

#region AgentTestStabilization.DesignTokens [C:2] [TYPE ADR] [SEMANTICS ux,design-tokens,agent-run] @BRIEF Semantic token and shared-component mapping for run progress, drafts, and recovery.

Purpose Required tokens/components
Panel/card bg-surface-card, border-border, text-text
Active/info stage bg-primary-light, text-primary, Icon
Completed bg-success-light, text-success
Warning/waiting bg-warning-light, text-warning
Failed/invalid bg-destructive-light, border-destructive-ring, text-destructive
Stale/disconnected bg-surface-muted, text-text-muted
Actions Button from $lib/ui; primary/secondary/destructive variants
Loading Skeleton or Spinner from $lib/ui

Raw color families and page-level raw buttons are forbidden. Focus rings must use shared component behavior and remain visible in warning/destructive surfaces.

#endregion AgentTestStabilization.DesignTokens


UX DESIGN — Per-Screen Contracts — agent-run-ux.md

Source: contracts/ux/agent-run-ux.md

#region AgentTestStabilization.RunUx [C:4] [TYPE ADR] [SEMANTICS ux,agent-run,progress,draft,recovery] @BRIEF Detailed FSM and recovery contract for the run strip and draft preview. @RELATION BINDS_TO -> [AgentRuns.Model]

Run Panel FSM

State Visible contract Actions
starting Context, connecting indicator, no fake run id Cancel chat
running Run id, active/completed stages, last update Continue chat
waiting_input Required parameter summary Provide parameters
waiting_approval Exact operation, target paths, warnings Confirm/deny
disconnected Last known stage marked stale Recover snapshot
completed All completed/skipped stages and drafts Preview/download
failed Error code, failed stage, recovery choices Retry safe analysis or restart
cancelled Explicit cancellation and no side effect claim Start new run

Draft States

  • pending: metadata accepted, validation running.
  • valid: preview/download and save request allowed.
  • warning: preview required; save gate must repeat warnings.
  • invalid: preview allowed, save request disabled.
  • persisted: target path and persisted timestamp visible.

UX Tests

  1. Stream drop at validate → disconnected → snapshot returns waiting_approval with same drafts.
  2. Foreign run event → ignored and logged.
  3. Invalid artifact → save disabled with associated reason.
  4. Denied save → no persisted marker; cancellation message remains visible.
  5. Permission denied → no confirm button and focusable alternatives.

#endregion AgentTestStabilization.RunUx


PLAN — Implementation Plan

Source: plan.md

Implementation Plan: Agent Test Stabilization

Branch: 036-agent-test-stabilization | Date: 2026-07-13 | Spec: spec.md Input: research.md, data-model.md, contracts/

Summary

Extend the existing Gradio/LangGraph agent with a durable backend-owned AgentRun lifecycle for scenario and investigation workstreams. UIContext v2 carries typed intents without changing the positional Gradio contract. Structured progress and AgentAction metadata drive persistent workspaces; non-delegated repository/baseline actions use payload-bound inline gates.

Technical Context

Language/Version: Python 3.11+ agent package; Python 3.13+ backend; TypeScript/Svelte 5 frontend
Dependencies: Gradio, LangGraph/LangChain, httpx; FastAPI 0.126, Pydantic 2, SQLAlchemy; Svelte 5.56, Vitest 4
Storage: PostgreSQL for run/event/gate records; configured backend draft-storage root for draft bytes
Testing: pytest in agent/ and backend/; Vitest L1/L2; Playwright smoke
Performance Goals: run id before first tool; snapshot GET p95 under 200ms; progress render under 100ms
Constraints: no Gradio positional break; no repository mutation before approval; dual identity auth; ordinary chat compatible
Scale: 10 concurrent chats, 100 events and 50 draft references per run, 30-day terminal-run retention

Constitution Check

Principle Gate
Semantic contracts C3+ modules/functions receive hierarchical anchors and test edges
Decision memory Durable owner, UIContext v2, event protocol, draft storage, and bound gate decisions are recorded
External orchestrator Agent and backend remain outside Superset
Module discipline New run service is decomposed into repository, event, artifact, and approval modules
RBAC READ/EXECUTE/WRITE/APPROVE checks occur in backend; agent filtering is defense in depth
Svelte 5 AgentChat model extension and new panels are runes/model-first
TDD Contract tests precede C3+ implementation; replay and SQL exclusions have tests
Attention Contract IDs share agent-run or dashboard-testing semantics and stay bounded

Gate result: PASS. No constitution exception is required.

Project Structure

agent/src/ss_tools/agent/
├── _context.py
├── _run_tracker.py                # NEW backend run API client
├── _tool_filter.py
└── app.py

backend/src/
├── api/routes/agent_runs.py       # NEW REST surface
├── models/agent_run.py
├── schemas/agent_run.py
└── services/agent_runs/
    ├── service.py
    ├── repository.py
    ├── artifacts.py
    └── approvals.py

frontend/src/lib/
├── models/AgentChatTypes.ts
├── models/AgentChatModel.svelte.ts
├── models/AgentChat.StreamProcessor.svelte.ts
└── components/agent/
    ├── AgentRunPanel.svelte
    └── DraftArtifactList.svelte

Delivery Phases

  1. Contract and migration foundation: DTOs, DB migration, repository, permissions.
  2. Run API: ownership-scoped create/snapshot/event/draft/gate endpoints.
  3. Agent integration: UIContext v2, tracker, scenario allow-list, events.
  4. Frontend model/UI: typed FSM, reconnection snapshot, draft and approval rendering.
  5. Verification: unit/contract/integration tests, ordinary-chat regression, live smoke.

Cross-Spec Boundary

  • 036 does not inspect dashboards or compare metrics.
  • 037 registers draft baseline candidates and requests gates through 036.
  • 038 registers draft scenario artifacts through 036.
  • 039 consumes 036 event/snapshot DTOs and must not create a second run state machine.

Complexity Tracking

No planned file needs an exception. If AgentChatModel.svelte.ts crosses its decomposition gate, extract AgentRunModel.svelte.ts as a composed submodel instead of growing the existing model.


RESEARCH — Technical Decisions

Source: research.md

#region AgentTestStabilization.Research [C:4] [TYPE ADR] [SEMANTICS research,agent,dashboard-testing,run,hitl] @BRIEF Phase 0 decisions for recoverable dashboard-scenario agent runs, structured events, draft artifacts, and bound approvals. @RELATION DEPENDS_ON -> [AgentTestStabilization.Spec] @RELATION DEPENDS_ON -> [AgentChat.GradioApp] @DEPRECATED Historical edge: target removed in 731aaaa8 (050 T041 chat decommission); the /agent route now renders [AgentChat.HandoffSurface]. @RELATION DEPENDS_ON -> [AgentChat.Model] @RATIONALE The existing 033/035 chat path already streams and confirms tools, but it does not provide a durable business-run aggregate or recoverable draft inventory. @REJECTED Treating conversation history or LangGraph checkpoints as the AgentRun system of record — rejected because neither exposes an owned, queryable lifecycle for progress, drafts, and approval decisions.

1. Existing System Fact Check

Concern Existing implementation Planning consequence
Gradio submit AgentChatModel.svelte.ts submits eight values; serialized uicontext is last Do not add or reorder positional inputs
Context validation agent/src/ss_tools/agent/_context.py accepts context version 1 and dashboard/dataset/migration types Add backward-compatible v2 intent validation
Stream processing AgentChat.StreamProcessor.svelte.ts dispatches metadata by type Add typed metadata variants; never parse assistant prose
HITL _confirmation.py emits confirm_required and LangGraph resumes confirm/deny Extend the envelope; preserve ordinary tool confirmation
Recovery Conversation messages persist; active scenario progress/drafts do not Add durable backend AgentRun aggregate
Dashboard entry DashboardHeader.svelte already builds an ordinary /agent dashboard link 039 adds a separate business action with explicit intent
SQL tool Agent registry contains superset_execute_sql Scenario intent pipeline must exclude it

2. Durable Run Ownership

  • Decision: FastAPI/backend owns AgentRun, AgentRunEvent, DraftArtifact, and ApprovalGate; the standalone agent writes through authenticated internal REST calls.
  • Rationale: Backend already owns users, conversations, RBAC, repository resolution, and audit persistence. Gradio workers may restart and cannot be the durable authority.
  • Alternative rejected: In-memory run registry in the agent — loses state on restart and cannot safely enforce ownership.
  • Alternative rejected: Reuse generic TaskManager as the sole store — AgentRun includes waiting-for-user and draft/approval semantics, while a task is execution monitoring. A later implementation may link task_id, but must not collapse the aggregates.

3. UIContext Compatibility

  • Decision: v1 remains valid for ordinary chat. v2 adds intent, with the only initial value build_dashboard_test_scenario.
  • Invariant: v2 scenario intent requires objectType=dashboard, numeric objectId, non-empty envId, and a dashboard route.
  • Wire rule: uicontext_str stays the final Gradio positional argument. No agent_run_id positional input is added; the backend emits agent_run_started.
  • Security: Context remains informational metadata and is validated before prompt injection. Unknown fields are rejected.

4. Structured Event Protocol

  • Decision: Add agent_run_started, scenario_progress, draft_artifacts, and agent_run_terminal metadata types.
  • Stage order: context → inspect → scenario → parameters → generate → validate → save.
  • Monotonicity: A stage may repeat with a higher sequence, but cannot regress except when the event explicitly declares retry_of_sequence.
  • Recovery: Frontend obtains the authoritative snapshot with GET /api/agent/runs/{run_id} after reload or stream loss.
  • Timeout: Existing 60-second first-activity guard remains; agent_run_started counts as first activity and must be emitted before any tool action.

5. Draft Artifact Safety

  • Decision: Draft content is stored outside the target repository under backend-managed draft storage; frontend receives opaque artifact_id, preview metadata, and safe download URLs.
  • Invariant: intended_path is relative, normalized, contains no parent traversal, and is not written until a bound approval executes.
  • Alternative rejected: Writing generated files to the repository and calling them drafts — that already mutates Git state and defeats preview-before-write.

6. Bound HITL Approval

  • Decision: An ApprovalGate binds operation, actor, run, target paths, and canonical payload SHA-256. Confirm/deny records an immutable decision. Execution recalculates the hash.
  • Baseline approval: Requires a non-blank reason and permission dashboard:testing APPROVE.
  • Repository save: Requires permission dashboard:testing WRITE.
  • Unauthorized action: Emits permission_denied directly; no gate and no confirm control.
  • Resurrection ban: A confirmed gate cannot be replayed, retargeted, or used after expiry.

7. Source Touch Map

Area Planned files
Agent context/events agent/src/ss_tools/agent/_context.py, app.py, new _run_tracker.py
Agent tool filtering agent/src/ss_tools/agent/_tool_filter.py, tools.py
Durable run API new backend/src/api/routes/agent_runs.py, backend/src/services/agent_runs/
Persistence new models/migration under backend/src/models/ and backend/alembic/versions/
Frontend model AgentChatTypes.ts, AgentChatModel.svelte.ts, AgentChat.StreamProcessor.svelte.ts
Frontend UI new frontend/src/lib/components/agent/AgentRunPanel.svelte and draft list
Tests agent pytest, backend API/service pytest, frontend L1/L2 vitest, one Playwright smoke

8. Resolved Scope

036 supplies the execution and approval substrate only. Superset query truth is 037, scenario graph semantics are 038, and the full business workspace is 039.

#endregion AgentTestStabilization.Research


DATA MODEL — Entities & Relations

Source: data-model.md

#region AgentTestStabilization.DataModel [C:4] [TYPE ADR] [SEMANTICS data-model,agent-run,event,draft,approval] @BRIEF Canonical entities, validation rules, lifecycle transitions, and ownership rules for feature 036. @RELATION DEPENDS_ON -> [AgentTestStabilization.Spec] @RELATION DEPENDS_ON -> [AgentTestStabilization.Research]

DashboardScenarioUIContext

Field Type Rule
objectType literal dashboard Required for scenario intent
objectId numeric string Required, 1–20 digits
objectName string/null Max 256 chars
envId string Required, max 128 chars
route string Must begin /dashboards/, max 512 chars
contextVersion 1 or 2 v1 ordinary chat; v2 scenario extension
intent enum/null build_dashboard_test_scenario, investigate_failure, investigate_staleness, revalidate_scenario, analyze_load, remediate_scenario, manage_automation

Compatibility: v1 payloads without intent remain valid. A typed agentic intent with v1 is invalid; an unknown intent is invalid.

AgentRun

Field Type Notes
id UUID Stable public agent_run_id
conversation_id UUID/string Existing conversation correlation
user_id string Owner; immutable
intent enum Typed workstream; starts with dashboard scenario build and also supports investigation/revalidation/remediation intents
trigger enum manual, investigation_case, deploy_to_preprod, release_create, release_approve, release_publish, post_publish, scheduled, etl_completed
dashboard_id, environment_id string Immutable context correlation
context_snapshot JSON Validated UIContext v2
status enum See FSM
current_stage enum Highest accepted scenario stage
last_sequence int Monotonic per run
error_code, error_detail nullable string Sanitized terminal/degraded information
created_at, updated_at, finished_at timestamp UTC

intent is a typed workstream: build_dashboard_test_scenario | investigate_failure | investigate_staleness | revalidate_scenario | analyze_load | remediate_scenario | manage_automation. An AgentRun may be linked to an InvestigationCase (047) and is one recoverable tool-execution session inside that case; it is never the case's lifecycle owner. The shared queue/case/action contract is contracts/investigation-cases.md.

DelegatedAuthorityPolicy is the server-owned versioned authority source for AgentAction evaluation; every action records its policy snapshot/decision. InvestigationSignal is the outbox envelope joining deterministic producers to 047 Queue. Neither an agent nor a client may decide authorization, emit a synthetic policy decision, or create a Queue item by prose.

AgentRun FSM

CREATED → RUNNING ↔ WAITING_INPUT
                  ↔ WAITING_APPROVAL
                  → COMPLETED
                  → FAILED
                  → CANCELLED

Terminal states are immutable. Reconnection reads state; it never restarts execution implicitly.

AgentRunEvent

Field Rule
run_id Existing owned run
sequence Positive, unique per run, strictly increasing
event_type run_started, progress, drafts_updated, approval_requested, approval_resolved, terminal
stage context, inspect, scenario, parameters, generate, validate, or save
status pending, active, completed, blocked, failed, or skipped
payload Typed, redacted JSON; max 64 KiB
occurred_at UTC timestamp

Duplicate (run_id, sequence) is idempotent only if canonical payload hashes match; otherwise it is a conflict.

DraftArtifact

Field Rule
id UUID
run_id Parent AgentRun
kind scenario, runner_plan, report_template, baseline_candidate, evidence_manifest, screenshot_evidence, other
name Display name, max 255
intended_path Relative POSIX path; no absolute path, parent traversal, NUL, or symlink traversal
content_ref Opaque backend storage key; never exposed as filesystem path
sha256 Lowercase 64-character digest
validation_status valid, warning, invalid, pending
warnings Bounded list of structured code/message objects
persisted_at Null until approved write succeeds
capture_meta Null unless kind=screenshot_evidence; see ScreenshotEvidence below

Draft downloads are read-only and do not change repository state.

ScreenshotEvidence

When kind is screenshot_evidence, the artifact carries a capture_meta block:

Field Type Rule
viewport {width, height} Required; fixed capture dimensions
device_scale_factor number Default 1
tab_identifier string/null Active tab or viewport label
filter_context_hash sha256 Match with 037 NormalizedFilterContext.filters_hash
readiness_strategy string canvas_stabilized, network_idle, or fixed_wait
readiness_timeout_ms integer Max wait before capture
capture_method string cdp, full_page, or region
mask_selectors string[] CSS selectors applied before capture; empty if none
browser_version string/null Captured when available
captured_at ISO-8601 Server timestamp of capture

Masking is optional display/export hygiene because all configured MCP and LLM/VLM clients are local deployments inside the enterprise trust perimeter. When requested, a masked derivative is stored as a separate DraftArtifact of kind other with name containing the original artifact id and _masked suffix. PII may remain in local provider payloads; credentials, cookies, tokens, and unrelated secrets are never transmitted or persisted in telemetry.

ActionApprovalGate

Field Rule
id UUID
owner_type agent_run, scenario_run, verification_run, or load_run
owner_id UUID of the owning run/intent
operation repository_write, baseline_approval, scenario_execution, load_execution, scenario_revision_write, automation_policy_write, controlled_test_data_mutation, or prod_mutation
request_hash SHA-256 of canonical operation, targets, artifact hashes, and baseline payload
target_paths Normalized relative paths
risk_level guarded or dangerous
required_permission dashboard:testing WRITE or APPROVE
status pending, confirmed, denied, consumed, expired
reason_required true for baseline approval
reason Required non-blank when applicable, max 2000
actor_id, decided_at, expires_at Immutable decision audit

Approval FSM

PENDING → CONFIRMED → CONSUMED
        → DENIED
        → EXPIRED

Execution requires CONFIRMED status, non-expired gate, unchanged request_hash, and a fresh object-level authorization check. Consumption is atomic with the durable side effect. A gate is rendered as an inline card in its case/chat or work surface, never as a blocking modal. ApprovalGate is a legacy name only; all new contracts use ActionApprovalGate. /agent/runs/{runId}/approval-gates remains an adapter route that fixes owner_type=agent_run.

Snapshot DTO

AgentRunSnapshot contains run fields, ordered stages, recent events, all draft references, and the current pending gate. It never contains raw draft bytes, JWTs, internal storage paths, or another user's run.

Retention

  • Active and waiting runs are retained.
  • Terminal run metadata and audit: minimum 30 days.
  • Unpersisted draft bytes: configurable, default 7 days after terminal state.
  • Approval decisions: retained with the associated run audit and never silently rewritten.

#endregion AgentTestStabilization.DataModel


CONTRACTS — Module & Function Contracts

Source: contracts/modules.md

#region AgentTestStabilization.Modules [C:5] [TYPE ADR] [SEMANTICS contracts,agent-run,dashboard-testing,hitl] @BRIEF Implementation contracts for durable agent runs, scenario events, drafts, approvals, and frontend recovery. @RELATION DEPENDS_ON -> [AgentTestStabilization.DataModel] @RELATION DEPENDS_ON -> [AgentTestStabilization.Events] @RATIONALE Durable state and side-effect authorization are split from Gradio streaming so worker restarts cannot erase audit truth. @REJECTED Persisting authoritative run state only in Gradio or frontend memory — rejected because both are restartable clients.

Backend

#region AgentRuns.Api [C:4] [TYPE Module] [SEMANTICS agent-run,api,ownership,rbac]

@defgroup AgentRuns Ownership-scoped REST surface for scenario run recovery and internal event ingestion.

@LAYER API

@RELATION DEPENDS_ON -> [Services.AgentRuns.Service]

@RELATION DEPENDS_ON -> [Schemas.AgentRun]

@INVARIANT Browser reads require run ownership or admin permission; internal writes require service identity plus propagated user identity.

@REJECTED Trusting user_id from request JSON — actor identity must come from validated auth.

#endregion AgentRuns.Api

#region AgentRuns.Service.Create [C:4] [TYPE Function] [SEMANTICS agent-run,create,context]

@ingroup AgentRuns

@BRIEF Create one durable run for a validated dashboard-scenario context.

@PRE UIContext v2 is valid; actor has dashboard:testing EXECUTE.

@POST Persists CREATED then RUNNING run and returns id before any scenario tool action.

@SIDE_EFFECT Writes AgentRun and initial event in one transaction.

@DATA_CONTRACT CreateAgentRunRequest -> AgentRunSnapshot

@TEST_EDGE duplicate_idempotency_key -> same actor/context returns same run.

@TEST_EDGE context_actor_mismatch -> 403 and no row.

@TEST_EDGE invalid_v1_scenario_context -> 422.

#endregion AgentRuns.Service.Create

#region AgentRuns.Service.AppendEvent [C:5] [TYPE Function] [SEMANTICS agent-run,event,sequence,state]

@ingroup AgentRuns

@BRIEF Append a typed event and advance run lifecycle under monotonic sequence rules.

@PRE Run is non-terminal; event validates against event_type payload schema.

@POST Sequence increases exactly once; status/stage projection matches the appended event.

@SIDE_EFFECT Inserts AgentRunEvent and updates AgentRun atomically.

@INVARIANT A duplicate sequence is idempotent only when its canonical hash matches.

@DATA_CONTRACT AppendEventRequest -> AgentRunEventResponse

@TEST_EDGE stage_regression -> 409 without mutation.

@TEST_EDGE same_sequence_same_hash -> existing event returned.

@TEST_EDGE same_sequence_different_hash -> 409 conflict.

@RATIONALE Transactional projection prevents the snapshot and event log from disagreeing after a crash.

@REJECTED Client-side sequence authority — concurrent writers could fork history.

#endregion AgentRuns.Service.AppendEvent

#region AgentRuns.Artifacts.Register [C:5] [TYPE Function] [SEMANTICS agent-run,draft,artifact,storage]

@ingroup AgentRuns

@BRIEF Store draft bytes outside the repository and register safe preview metadata.

@PRE Run is owned and active; path and size pass policy; provided digest matches bytes.

@POST Draft is retrievable by opaque id; target repository remains unchanged.

@SIDE_EFFECT Writes draft storage and DraftArtifact record.

@DATA_CONTRACT RegisterDraftRequest + bytes -> DraftArtifactRef

@RATIONALE Opaque draft references separate previewable generated content from repository state and prevent path disclosure before HITL approval.

@INVARIANT content_ref and filesystem path are never returned to clients.

@TEST_EDGE traversal_path -> 422 before write.

@TEST_EDGE digest_mismatch -> draft bytes removed and 422.

@TEST_EDGE invalid_artifact_warning -> registered as invalid but never saveable.

@REJECTED Temporary write inside the Git worktree — it dirties the repository before approval.

#endregion AgentRuns.Artifacts.Register

#region AgentRuns.Approvals.Request [C:5] [TYPE Function] [SEMANTICS agent-run,hitl,approval,hash]

@ingroup AgentRuns

@BRIEF Create a one-shot approval gate bound to exact operation inputs.

@PRE Actor has permission to request the operation; all referenced drafts belong to the run.

@POST PENDING gate stores canonical request_hash and expiry; run enters WAITING_APPROVAL.

@SIDE_EFFECT Writes ApprovalGate and approval_requested event.

@DATA_CONTRACT ApprovalRequest -> ApprovalGateView

@INVARIANT Unauthorized calls return permission_denied semantics and create no gate.

@RATIONALE Gate creation binds exact draft ownership and operation inputs before any decision can be recorded; this keeps the later consume step auditable.

@REJECTED Creating a reusable unbound confirmation token — rejected because it could be replayed against another run, artifact, or baseline candidate.

@TEST_EDGE foreign_artifact -> 403.

@TEST_EDGE missing_baseline_reason_requirement -> gate created with reason_required=true.

#endregion AgentRuns.Approvals.Request

#region AgentRuns.Approvals.Decide [C:5] [TYPE Function] [SEMANTICS agent-run,hitl,decision,audit]

@ingroup AgentRuns

@BRIEF Record immutable confirmation or denial for a pending gate.

@PRE Gate is pending, unexpired, owned by actor; reason present when required.

@POST Gate becomes CONFIRMED or DENIED exactly once and decision event is appended.

@SIDE_EFFECT Writes decision audit; denial returns run to RUNNING or CANCELLED according to operation policy.

@DATA_CONTRACT ApprovalDecisionRequest -> ApprovalGateView

@RATIONALE Decisions are immutable audit facts; recording exactly one decision prevents ambiguous approval state under retries or multiple tabs.

@REJECTED Last-write-wins gate decisions — rejected because a late denial/confirmation could silently override the actor's earlier decision.

@TEST_EDGE blank_required_reason -> 422.

@TEST_EDGE repeated_decision -> 409.

@TEST_EDGE expired_gate -> EXPIRED and no confirmation.

#endregion AgentRuns.Approvals.Decide

#region AgentRuns.Approvals.Consume [C:5] [TYPE Function] [SEMANTICS agent-run,hitl,consume,side-effect]

@ingroup AgentRuns

@BRIEF Execute the exact approved write and atomically consume its gate.

@PRE Gate confirmed; fresh RBAC passes; actor and recalculated request_hash match.

@POST Side effect succeeds once and gate is CONSUMED, or neither is committed.

@SIDE_EFFECT Writes repository artifacts or approved baseline plus audit event.

@INVARIANT No retargeting, replay, partial path set, or post-confirm payload mutation.

@TEST_INVARIANT Approval_Request_Binding -> VERIFIED_BY: payload_change, target_change, replay.

@TEST_EDGE payload_change -> 409 and no write.

@TEST_EDGE permission_revoked_after_confirm -> 403 and gate not consumed.

@TEST_EDGE replay_consumed_gate -> 409 and no second write.

@RATIONALE Rechecking hash and RBAC at consume time closes the confirmation-to-execution gap.

@REJECTED Confirming only a human-readable prompt — the executable arguments could change after review.

#endregion AgentRuns.Approvals.Consume

Standalone Agent

#region AgentRuns.Gradio.Emit [C:4] [TYPE Function] [SEMANTICS agent-run,gradio,stream,event]

@ingroup AgentRuns

@BRIEF Persist then emit typed run metadata through the existing Gradio stream.

@PRE AgentRun exists; event accepted by backend.

@POST Emitted metadata carries authoritative run_id and sequence.

@SIDE_EFFECT Backend event write followed by Gradio yield.

@INVARIANT Persistence precedes emission; a dropped stream remains recoverable.

@TEST_EDGE backend_event_failure -> emit recoverable error, never unpersisted progress.

#endregion AgentRuns.Gradio.Emit

#endregion AgentTestStabilization.Modules


CONTRACTS — Remaining — agent-pipeline.md

Source: contracts/agent-pipeline.md

#region AgentTestStabilization.AgentPipeline [C:4] [TYPE ADR] [SEMANTICS contracts,agent-run,pipeline,tools] @BRIEF Agent-side tracker, context validation, scenario tool filtering, and verification trigger contracts. @RELATION DEPENDS_ON -> [AgentTestStabilization.DataModel] @RELATION DEPENDS_ON -> [AgentRuns.Api]

#region AgentRuns.Tracker [C:4] [TYPE Module] [SEMANTICS agent-run,agent,http,events]

@defgroup AgentRuns Agent-side typed client for backend run lifecycle APIs.

@LAYER Service

@RELATION DEPENDS_ON -> [AgentRuns.Api]

@INVARIANT Every call carries service JWT and current user JWT; logs redact both.

@RATIONALE Small client isolates retry/idempotency behavior from the large Gradio handler.

#endregion AgentRuns.Tracker

#region AgentRuns.Context.ValidateV2 [C:3] [TYPE Function] [SEMANTICS agent-run,context,intent,validation]

@ingroup AgentRuns

@BRIEF Validate UIContext v1/v2 and scenario-intent cross-field rules.

@POST Returns normalized context; v1 ordinary chat remains compatible.

@TEST_EDGE v1_without_intent -> accepted; v2_scenario_dashboard -> accepted; scenario_dataset -> rejected; unknown_intent -> rejected.

#endregion AgentRuns.Context.ValidateV2

#region AgentRuns.ToolPipeline.Scenario [C:4] [TYPE Function] [SEMANTICS agent-run,tools,scenario,no-sql]

@ingroup AgentRuns

@BRIEF Restrict scenario-intent tools to dashboard inspection, scenario, artifact, and approval operations.

@PRE RBAC filtering has already run.

@POST Tool list excludes arbitrary SQL operations.

@SIDE_EFFECT Emits pipeline_result audit metadata.

@INVARIANT No raw credentialed/arbitrary SQL tool reaches build_dashboard_test_scenario; bounded authoring SqlEvidenceSpec reaches 038 compiler validation only.

@TEST_INVARIANT No_Direct_SQL -> VERIFIED_BY: scenario_tool_list, replayed_sql_call.

@TEST_EDGE hallucinated_sql_tool_call -> invocation guard rejects.

@REJECTED Relying only on system prompt to avoid SQL — tool availability is enforceable.

#endregion AgentRuns.ToolPipeline.Scenario

#region AgentRuns.Verification.CreateRun [C:4] [TYPE Function] [SEMANTICS agent-run,verification,release,dashboard-testing]

@ingroup AgentRuns

@BRIEF Create an AgentRun triggered by release pipeline or scheduler, linked to VerificationRun.

@PRE Trigger valid; actor has dashboard testing execute; release triggers have DashboardRelease.

@POST AgentRun and linked VerificationRun are created atomically.

@SIDE_EFFECT Writes AgentRun and VerificationRun.

@DATA_CONTRACT CreateAgentRunRequest -> AgentRunSnapshot + VerificationRun

@RELATION DEPENDS_ON -> [AgentRuns.Service.Create]

@TEST_EDGE scheduled without release -> run without release_id.

@RATIONALE Trigger enum makes verification scope explicit and recoverable.

@REJECTED Creating verification outside AgentRun lifecycle — recovery and gate semantics would diverge.

#endregion AgentRuns.Verification.CreateRun

#endregion AgentTestStabilization.AgentPipeline


CONTRACTS — Remaining — agent-runs.openapi.yaml

Source: contracts/agent-runs.openapi.yaml

openapi: 3.1.0 info: title: Agent Run Recovery API version: 0.1.0 description: Durable run, event, draft, and approval substrate for dashboard scenario agent flows. servers:

  • url: /api paths: /agent/runs: post: operationId: createAgentRun security: [{ bearerAuth: [] }, { serviceAndUserAuth: [] }] requestBody: required: true content: application/json: schema: { $ref: '#/components/schemas/CreateAgentRunRequest' } responses: '201': description: Created content: { application/json: { schema: { $ref: '#/components/schemas/AgentRunSnapshot' } } } '403': { $ref: '#/components/responses/Forbidden' } '422': { $ref: '#/components/responses/Invalid' } /investigation-cases/{caseId}/agent-runs: post: operationId: createInvestigationAgentRun summary: Start an explicitly opened case's recoverable agent workstream security: [{ bearerAuth: [] }] parameters: - name: caseId in: path required: true schema: { type: string, format: uuid } requestBody: required: true content: application/json: schema: { $ref: '#/components/schemas/CreateAgentRunRequest' } responses: '201': { description: Created for an existing InvestigationCase } '403': { $ref: '#/components/responses/Forbidden' } '409': { description: Case is terminal or request idempotency conflicts } /investigation-cases/{caseId}/actions: post: operationId: executeInvestigationAction summary: Evaluate delegated authority and execute or gate one canonical AgentAction security: [{ bearerAuth: [] }, { serviceAndUserAuth: [] }] parameters: - name: caseId in: path required: true schema: { type: string, format: uuid } requestBody: required: true content: application/json: schema: { $ref: '#/components/schemas/AgentActionRequest' } responses: '202': description: Delegated action accepted or ActionApprovalGate required content: { application/json: { schema: { $ref: '#/components/schemas/AgentActionDecision' } } } '403': { $ref: '#/components/responses/Forbidden' } '409': { description: Case terminal, policy changed, or side-effect key replay conflict } '422': { $ref: '#/components/responses/Invalid' } /agent/runs/{runId}: get: operationId: getAgentRun security: [{ bearerAuth: [] }] parameters: - $ref: '#/components/parameters/RunId' responses: '200': description: Ownership-scoped authoritative snapshot content: { application/json: { schema: { $ref: '#/components/schemas/AgentRunSnapshot' } } } '403': { $ref: '#/components/responses/Forbidden' } '404': { description: Not found } /agent/runs/{runId}/events: post: operationId: appendAgentRunEvent security: [{ serviceAndUserAuth: [] }] parameters: - $ref: '#/components/parameters/RunId' requestBody: required: true content: application/json: schema: { $ref: '#/components/schemas/AppendEventRequest' } responses: '201': description: Appended content: { application/json: { schema: { $ref: '#/components/schemas/AgentRunEvent' } } } '409': { description: Sequence, stage, or terminal-state conflict } '422': { $ref: '#/components/responses/Invalid' } /agent/runs/{runId}/artifacts: post: operationId: registerDraftArtifact security: [{ serviceAndUserAuth: [] }] parameters: - $ref: '#/components/parameters/RunId' requestBody: required: true content: multipart/form-data: schema: type: object required: [metadata, file] properties: metadata: $ref: '#/components/schemas/RegisterDraftMetadata' file: type: string format: binary responses: '201': description: Draft stored outside repository content: { application/json: { schema: { $ref: '#/components/schemas/DraftArtifactRef' } } } '413': { description: Draft too large } '422': { $ref: '#/components/responses/Invalid' } /agent/runs/{runId}/approval-gates: post: operationId: requestApprovalGate security: [{ serviceAndUserAuth: [] }] parameters: - $ref: '#/components/parameters/RunId' requestBody: required: true content: application/json: schema: { $ref: '#/components/schemas/ApprovalRequest' } responses: '201': description: Bound pending gate content: { application/json: { schema: { $ref: '#/components/schemas/ApprovalGateView' } } } '403': { $ref: '#/components/responses/Forbidden' } /agent/runs/{runId}/approval-gates/{gateId}/decision: post: operationId: decideApprovalGate security: [{ bearerAuth: [] }] parameters: - $ref: '#/components/parameters/RunId' - name: gateId in: path required: true schema: { type: string, format: uuid } requestBody: required: true content: application/json: schema: type: object required: [decision] properties: decision: { type: string, enum: [confirm, deny] } reason: { type: [string, 'null'], maxLength: 2000 } responses: '200': description: Immutable decision recorded content: { application/json: { schema: { $ref: '#/components/schemas/ApprovalGateView' } } } '409': { description: Gate already decided or expired } '422': { $ref: '#/components/responses/Invalid' } /agent/runs/{runId}/approval-gates/{gateId}/consume: post: operationId: consumeApprovalGate description: Internal execution endpoint; recalculates request hash and RBAC before one side effect. security: [{ serviceAndUserAuth: [] }] parameters: - $ref: '#/components/parameters/RunId' - name: gateId in: path required: true schema: { type: string, format: uuid } requestBody: required: true content: application/json: schema: type: object required: [operationPayload] properties: operationPayload: { type: object, additionalProperties: true } responses: '200': { description: Side effect committed and gate consumed } '403': { $ref: '#/components/responses/Forbidden' } '409': { description: Hash, actor, status, expiry, or replay conflict } /action-approval-gates/{gateId}: get: operationId: getActionApprovalGate summary: Get a generic ActionApprovalGate subject to owner-object authorization security: [{ bearerAuth: [] }] parameters: - name: gateId in: path required: true schema: { type: string, format: uuid } responses: '200': { description: Generic approval gate, content: { application/json: { schema: { $ref: '#/components/schemas/ApprovalGateView' } } } } '403': { $ref: '#/components/responses/Forbidden' } '404': { description: Not found } /action-approval-gates/{gateId}/decision: post: operationId: decideActionApprovalGate summary: Decide a generic ActionApprovalGate for any supported owner type security: [{ bearerAuth: [] }] parameters: - name: gateId in: path required: true schema: { type: string, format: uuid } requestBody: required: true content: application/json: schema: { $ref: '#/components/schemas/ApprovalDecision' } responses: '200': { description: Immutable decision recorded, content: { application/json: { schema: { $ref: '#/components/schemas/ApprovalGateView' } } } } '409': { description: Gate already decided or expired } /action-approval-gates/{gateId}/consume: post: operationId: consumeActionApprovalGate summary: Internal one-time consume for a generic ActionApprovalGate security: [{ serviceAndUserAuth: [] }] parameters: - name: gateId in: path required: true schema: { type: string, format: uuid } requestBody: required: true content: application/json: schema: { type: object, required: [operationPayload], properties: { operationPayload: { type: object } } } responses: { '200': { description: Side effect committed and gate consumed }, '409': { description: Hash, status, expiry, or replay conflict } } components: securitySchemes: bearerAuth: { type: http, scheme: bearer, bearerFormat: JWT } serviceAndUserAuth: type: apiKey in: header name: X-Service-Authorization description: Requires service identity plus Authorization user JWT. parameters: RunId: name: runId in: path required: true schema: { type: string, format: uuid } responses: Forbidden: { description: Ownership or permission denied } Invalid: { description: Validation failed } schemas: UIContextV2: type: object additionalProperties: false required: [objectType, objectId, envId, route, contextVersion, intent] properties: objectType: { type: string, enum: [dashboard, scenario, scenario_run, load_run, queue_item] } objectId: { type: string, minLength: 1, maxLength: 128 } objectName: { type: [string, 'null'], maxLength: 256 } envId: { type: string, minLength: 1, maxLength: 128 } route: { type: string, pattern: '^/', maxLength: 512 } contextVersion: { const: 2 } intent: { type: string, enum: [build_dashboard_test_scenario, investigate_failure, investigate_staleness, revalidate_scenario, analyze_load, remediate_scenario, manage_automation] } CreateAgentRunRequest: type: object required: [conversationId, context, idempotencyKey] properties: conversationId: { type: string, minLength: 1 } investigationCaseId: { type: string, format: uuid, nullable: true } context: { $ref: '#/components/schemas/UIContextV2' } trigger: { type: string, enum: [manual, investigation_case, deploy_to_preprod, release_create, release_approve, release_publish, post_publish, scheduled, etl_completed], default: manual } idempotencyKey: { type: string, minLength: 16, maxLength: 128 } AppendEventRequest: type: object required: [sequence, eventType, stage, status, payloadHash] properties: sequence: { type: integer, minimum: 1 } eventType: type: string enum: [run_started, progress, evidence_captured, drafts_updated, approval_requested, approval_resolved, terminal] stage: { type: string, enum: [context, inspect, scenario, parameters, generate, validate, save, evidence, hypothesis, plan, action, verify, resolve] } status: { type: string, enum: [pending, active, completed, blocked, failed, skipped] } payload: { type: object, additionalProperties: true } payloadHash: { type: string, pattern: '^[a-f0-9]{64}$' } RegisterDraftMetadata: type: object required: [kind, name, intendedPath, sha256, validationStatus] properties: kind: { type: string, enum: [scenario, runner_plan, report_template, baseline_candidate, evidence_manifest, screenshot_evidence, other] } name: { type: string, maxLength: 255 } intendedPath: { type: string, maxLength: 1024 } sha256: { type: string, pattern: '^[a-f0-9]{64}$' } validationStatus: { type: string, enum: [valid, warning, invalid, pending] } warnings: { type: array, items: { type: object } } captureMeta: type: object description: Required when kind=screenshot_evidence; null otherwise. properties: viewport: { type: object, required: [width, height], properties: { width: { type: integer }, height: { type: integer } } } tabIdentifier: { type: [string, 'null'] } filterContextHash: { type: string, pattern: '^[a-f0-9]{64}$' } readinessStrategy: { type: string, enum: [canvas_stabilized, network_idle, fixed_wait] } readinessTimeoutMs: { type: integer } captureMethod: { type: string, enum: [cdp, full_page, region] } maskSelectors: { type: array, items: { type: string } } browserVersion: { type: [string, 'null'] } ApprovalRequest: oneOf: - $ref: '#/components/schemas/RepositoryApprovalRequest' - $ref: '#/components/schemas/ScenarioExecutionApprovalRequest' - $ref: '#/components/schemas/DelegatedActionApprovalRequest' RepositoryApprovalRequest: type: object required: [operation, artifactIds, targetPaths] properties: operation: { type: string, enum: [repository_write, baseline_approval] } artifactIds: { type: array, minItems: 1, uniqueItems: true, items: { type: string, format: uuid } } targetPaths: { type: array, minItems: 1, uniqueItems: true, items: { type: string } } warnings: { type: array, items: { type: object } } ScenarioExecutionApprovalRequest: type: object required: [operation, scenarioRunId, executionRequestHash] properties: operation: { const: scenario_execution } scenarioRunId: { type: string, format: uuid } executionRequestHash: { type: string } targetSnapshot: { type: object } riskSummary: { type: object } DelegatedActionApprovalRequest: type: object required: [operation, ownerType, ownerId, actionRequestHash, riskClass, target] properties: operation: { type: string, enum: [load_execution, scenario_revision_write, activate_current_revision, automation_policy_write, controlled_test_data_mutation, prod_mutation] } ownerType: { type: string, enum: [agent_run, scenario_run, verification_run, load_run] } ownerId: { type: string, format: uuid } actionRequestHash: { type: string, pattern: '^[a-f0-9]{64}$' } riskClass: { type: string } target: { type: object, additionalProperties: true } preconditionEvidenceRefs: { type: array, items: { type: string, format: uuid } } reconciliationPlan: { type: object, additionalProperties: true } DelegatedAuthorityPolicy: type: object required: [policyId, version, actionClass, autonomous, cleanupRequired, approvalRequired] properties: policyId: { type: string, format: uuid } version: { type: integer, minimum: 1 } scope: { type: object, additionalProperties: true } actionClass: { type: string } autonomous: { type: boolean } allowedEnvironmentClasses: { type: array, items: { type: string } } allowedFixtureIds: { type: array, items: { type: string } } maxRequestVolume: { type: integer, minimum: 1 } maxConcurrency: { type: integer, minimum: 1 } cleanupRequired: { type: boolean } mayActivateCurrentRevision: { type: boolean } requiredPermission: { type: [string, 'null'] } approvalRequired: { type: boolean } AgentActionRequest: type: object required: [intent, actionClass, canonicalInputs, target, expectedEffect] properties: intent: { type: string } actionClass: { type: string, enum: [read, diagnostic_run, controlled_test_data_mutation, draft_write, scenario_revision_write, activate_current_revision, baseline_approval, automation_policy_write, prod_mutation] } canonicalInputs: { type: object, additionalProperties: true } target: { type: object, additionalProperties: true } expectedEffect: { type: object, additionalProperties: true } sideEffectKey: { type: [string, 'null'] } preconditionEvidenceRefs: { type: array, items: { type: string, format: uuid } } reconciliationPlan: { type: object, additionalProperties: true } AgentActionDecision: type: object required: [agentActionId, policyDecision] properties: agentActionId: { type: string, format: uuid } policyDecision: { type: string, enum: [delegated, approval_required, denied] } policy: { $ref: '#/components/schemas/DelegatedAuthorityPolicy' } approvalGateId: { type: [string, 'null'], format: uuid } ApprovalDecision: type: object required: [decision] properties: decision: { type: string, enum: [confirm, deny] } reason: { type: [string, 'null'], maxLength: 2000 } AgentRunEvent: allOf: - $ref: '#/components/schemas/AppendEventRequest' - type: object required: [id, runId, occurredAt] properties: id: { type: string, format: uuid } runId: { type: string, format: uuid } occurredAt: { type: string, format: date-time } DraftArtifactRef: type: object required: [artifactId, kind, name, intendedPath, sha256, validationStatus] properties: artifactId: { type: string, format: uuid } kind: { type: string } name: { type: string } intendedPath: { type: string } sha256: { type: string } validationStatus: { type: string } warnings: { type: array, items: { type: object } } previewUrl: { type: string } downloadUrl: { type: string } captureMeta: { type: object, description: 'Present when kind=screenshot_evidence' } ApprovalGateView: type: object required: [gateId, ownerType, ownerId, operation, status, targetPaths, reasonRequired, expiresAt] properties: gateId: { type: string, format: uuid } ownerType: { type: string, enum: [agent_run, scenario_run, verification_run, load_run] } ownerId: { type: string, format: uuid } operation: { type: string } status: { type: string, enum: [pending, confirmed, denied, consumed, expired] } targetPaths: { type: array, items: { type: string } } reasonRequired: { type: boolean } reason: { type: [string, 'null'] } expiresAt: { type: string, format: date-time } AgentRunSnapshot: type: object required: [agentRunId, conversationId, status, currentStage, lastSequence, context, events, artifacts] properties: agentRunId: { type: string, format: uuid } conversationId: { type: string } status: { type: string, enum: [created, running, waiting_input, waiting_approval, completed, failed, cancelled] } currentStage: { type: string } lastSequence: { type: integer } context: { $ref: '#/components/schemas/UIContextV2' } events: { type: array, items: { $ref: '#/components/schemas/AgentRunEvent' } } artifacts: { type: array, items: { $ref: '#/components/schemas/DraftArtifactRef' } } pendingGate: oneOf: - $ref: '#/components/schemas/ApprovalGateView' - type: 'null'

CONTRACTS — Remaining — events.md

Source: contracts/events.md

#region AgentTestStabilization.Events [C:4] [TYPE ADR] [SEMANTICS agent-run,event,gradio,stream] @BRIEF Typed Gradio metadata protocol for scenario runs; backend persistence is authoritative. @RELATION DEPENDS_ON -> [AgentTestStabilization.DataModel] @RELATION CALLED_BY -> [AgentRuns.Gradio.Emit]

Common Envelope

Every new metadata event contains:

Field Type Rule
type string enum Event discriminator
agent_run_id UUID Stable run correlation
sequence positive integer Monotonic per run
occurred_at ISO-8601 UTC Server time
conversation_id string Existing chat correlation

agent_run_started

{
  "type": "agent_run_started",
  "agent_run_id": "8a0c5aa4-beb6-4c0e-baf5-3fef93a2430a",
  "sequence": 1,
  "occurred_at": "2026-07-13T10:00:00Z",
  "conversation_id": "conv-42",
  "intent": "build_dashboard_test_scenario",
  "stage": "context",
  "status": "active"
}

Must be persisted/emitted before any tool_start for the scenario run.

scenario_progress

Adds stage, status, message_code, completed_steps, total_steps, and optional retry_of_sequence. message_code is localized by frontend; prose is not interpreted as state.

draft_artifacts

Adds artifacts: array of artifact_id, kind, name, intended_path, sha256, validation_status, warnings, preview_url, download_url. No raw bytes or storage paths.

evidence_captured

Emitted when one or more screenshot evidence artifacts are captured and registered. Extends the common envelope with:

Field Rule
stage inspect or validate
evidence Array of {artifact_id, kind, label, viewport, tab_identifier, sha256, captured_at, preview_url}

Used by 039 EvidencePanel to display captured screenshots without parsing agent prose. Each evidence entry links to a DraftArtifact of kind screenshot_evidence.

agent_run_terminal

Adds status (completed/failed/cancelled), stage, error_code, recoverable, and optional safe recovery actions.

confirm_required Extension

Existing fields remain. Scenario operations add:

Field Rule
agent_run_id, approval_gate_id Required
approval_kind repository_write or baseline_approval
target_paths Exact normalized path list
artifact_ids Exact draft set
warnings Visible structured warnings
request_hash Short display fingerprint; full hash remains server-side
reason_required True for baseline approval

permission_denied Extension

May include agent_run_id, operation, required_permission, and alternatives. It never includes approval_gate_id and never enters the confirmation checkpoint.

Ordering and Recovery

  • Frontend ignores duplicate or older sequence values.
  • A sequence gap marks the projection degraded and triggers snapshot recovery.
  • Unknown event type is logged and ignored without changing run FSM.
  • Stream loss never changes backend status.

#endregion AgentTestStabilization.Events


CONTRACTS — Remaining — evidence.md

Source: contracts/evidence.md

#region AgentTestStabilization.EvidenceContracts [C:4] [TYPE ADR] [SEMANTICS contracts,agent-run,evidence,screenshot] @BRIEF Screenshot evidence adapter contract for masked artifacts and VLM analysis handoff. @RELATION DEPENDS_ON -> [AgentTestStabilization.DataModel] @RELATION DEPENDS_ON -> [AgentRuns.Artifacts.Register] @RATIONALE Existing ScreenshotService capture behavior is reused through an adapter so scenario evidence has one masking and readiness implementation. @REJECTED Rebuilding screenshot capture inside scenario execution — rejected because it would fork Playwright/CDP behavior and weaken evidence auditability.

Module Reuse Map (036 LLM verification tooling)

The adapter does not implement capture, masking, VLM submission, or redaction. It reuses the existing llm_analysis plugin modules as the single source of truth:

Concern Reused module (real contract ID) Where it lives
Playwright capture, login/CSRF, tab traversal, stabilization, webp conversion Plugin.Service.ScreenshotService backend/src/plugins/llm_analysis/service.py
VLM/LLM provider submission (multimodal + JSON mode, retries, image optimization) Plugin.Service.LLMClient backend/src/plugins/llm_analysis/service.py
Masked-artifact registration + capture_meta Services.AgentRuns.Evidence (register_screenshot_draft / register_masked_derivative) backend/src/services/agent_runs/evidence.py
Redaction of logs/raw responses before persistence Plugin.Service.RedactionService backend/src/plugins/llm_analysis/service.py

@REJECTED Reimplementing any of the above in services/agent_runs/ — a second Playwright/CDP path or a second LLM client would diverge from the established capture/analysis behavior and break evidence auditability.

#region AgentRuns.Evidence.Adapter [C:5] [TYPE Module] [SEMANTICS agent-run,evidence,screenshot,capture,llm]

@ingroup AgentRuns

@BRIEF Bridge existing ScreenshotService/VLM pipeline to AgentRun/DraftArtifact.

@RATIONALE Existing capture+VLM pipeline handles tab traversal, stabilization, and image conversion; the adapter registers it without duplicating behavior.

@REJECTED Rebuilding screenshot capture inside AgentRun lifecycle — rejected because scenario runs should consume the established capture contract.

@LAYER Service

@RELATION DEPENDS_ON -> [AgentRuns.Artifacts.Register]

@RELATION DEPENDS_ON -> [Plugin.Service.ScreenshotService]

@RELATION DEPENDS_ON -> [Plugin.Service.LLMClient]

@RELATION DEPENDS_ON -> [Plugin.Service.RedactionService]

@DATA_CONTRACT CaptureSpec + AgentRun -> DraftArtifact[]

@INVARIANT All configured MCP and VLM clients are local deployments inside the enterprise trust perimeter; PII in captures is an accepted residual risk and masking is optional. Credentials, cookies, tokens, and unrelated secrets never enter provider payloads or telemetry.

@INVARIANT Every capture carries capture_meta for reproducibility.

@INVARIANT VLM submission reuses Plugin.Service.LLMClient; a second provider client is forbidden.

@TEST_EDGE no_mask_selectors -> unmasked capture remains valid evidence; an optional masked derivative, when requested, remains a separate artifact.

@TEST_EDGE vlm_failure -> evidence remains valid WARN; analysis becomes inconclusive.

@TEST_EDGE missing_capture_meta -> 422.

#endregion AgentRuns.Evidence.Adapter

#endregion AgentTestStabilization.EvidenceContracts


CONTRACTS — Remaining — frontend-models.md

Source: contracts/frontend-models.md

#region AgentTestStabilization.FrontendModels [C:4] [TYPE ADR] [SEMANTICS contracts,agent-run,frontend,ux] @BRIEF Frontend run model and draft/recovery components for durable AgentRuns. @RELATION DEPENDS_ON -> [AgentTestStabilization.DataModel] @RELATION DEPENDS_ON -> [AgentRuns.Api]

// #region AgentRuns.Model [C:5] [TYPE Model] [SEMANTICS agent-run,model,recovery,draft,approval] // @defgroup AgentRuns Frontend run projection composed by AgentChat.Model. // @STATE absent | starting | running | waiting_input | waiting_approval | completed | failed | cancelled | disconnected // @ACTION applyMetadata(meta) — accept typed event when sequence is newer. // @ACTION recover(runId) — fetch authoritative snapshot after reload/drop. // @ACTION decideGate(gateId, decision, reason?) — submit HITL decision through existing resume path. // @INVARIANT Stage state derives from structured metadata/snapshot only, never assistant prose. // @INVARIANT Drafts are keyed by artifact id and cannot cross run boundaries. // @RELATION DEPENDS_ON -> [AgentRuns.Api] // @RELATION BINDS_TO -> [AgentRuns.RunPanel] // @RATIONALE A composed submodel prevents further uncontrolled growth of AgentChat.Model. // @REJECTED Duplicate inline run atoms in components — creates divergent recovery state. // #endregion AgentRuns.Model

#endregion AgentTestStabilization.FrontendModels


CONTRACTS — Remaining — investigation-cases.md

Source: contracts/investigation-cases.md

#region AgentInvestigation.Cases [C:5] [TYPE ADR] [SEMANTICS agent,investigation,case,queue,policy,actions] @BRIEF Shared agentic investigation contract consumed by 036–047. @RELATION DEPENDS_ON -> [AgentTestStabilization.DataModel] @RATIONALE Failures, staleness and load findings need an evidence-led workstream, not a collection of isolated forms. Agent reasoning may plan and execute permitted work, while deterministic systems remain the authority for execution, validation and access control. @REJECTED Opening an agent chat for every failure — rejected because transient and duplicate failures create noise. Events enter a queue; an analyst explicitly opens the agentic case. @REJECTED Replacing deterministic runners, validators, schedulers or policy checks with LLM decisions — rejected because execution truth, safety and reproducibility must remain independently verifiable.

Investigation Queue

Deterministic producers emit the idempotent InvestigationSignal; only 047 consumes it to create or update an InvestigationQueueItem from a failed/inconclusive/blocked ScenarioRun, scenario staleness signal, baseline immutability violation, load circuit-breaker/consistency finding, or repeated automation failure.

Fields: id, source_type, source_id, scenario_id?, run_id?, logical_step_id?, severity, fingerprint?, active_episode_id?, evidence_summary, target_snapshot, execution_principal_fingerprint?, suggested_next_action, state, count, first_seen_at, last_seen_at, case_id?.

State: new | acknowledged | case_opened | suppressed | resolved. A matching occurrence updates one queue item only inside the active failure episode; a matching occurrence after the episode is resolved creates a new queue item. Creating a queue item never starts an AgentRun.

InvestigationCase and AgentThread

An analyst opens a queue item into one durable InvestigationCase. Fields: id, queue_item_id, status, source_snapshot, evidence_snapshot, owner_actor_id, agent_thread_id, opened_at, resolved_at?, final_disposition?, resolution_summary?.

Status: open | investigating | awaiting_approval | awaiting_external_change | verifying | resolved | accepted | reopened. A case owns a chat thread and may launch many 036 AgentRun instances; an AgentRun is a recoverable tool-execution session, never the long-lived business case itself.

TriageRecord is the compact audited projection of the current case disposition for Registry, Run Monitor and analytics. It never changes historical run truth.

Delegated AgentAction

Every tool use is an AgentAction with intent, canonical inputs, risk class, policy decision, target/environment, affected entity keys, side_effect_key?, precondition evidence, postcondition evidence, cleanup/reconciliation plan, actor/delegator, agent and tool versions, and linked approval gate if one is required.

Risk classes:

  • read and diagnostic_run: agent executes autonomously within ACL and capacity policy.
  • controlled_test_data_mutation: agent executes autonomously only in an authorized fixture scope with a lease, exact record keys, reconciliation plan and postcondition evidence.
  • draft_write and scenario_revision_write: agent may create WorkingDrafts and save immutable executable revisions after deterministic validation and delegated policy allow it.
  • activate_current_revision, baseline_approval, automation_policy_write, and any prod_mutation: require the applicable ActionApprovalGate unless an explicitly stronger delegated policy permits the exact operation.

No operation bypasses object ACL, environment policy, ActionRegistry mutation contract, capacity allocation, canonical validation, immutable revision creation, or ActionApprovalGate consumption. Failed cleanup moves the case to awaiting_external_change; it cannot be resolved silently.

DelegatedAuthorityPolicy

DelegatedAuthorityPolicy is server-owned and versioned. Fields: policy_id, scope {scenario_id?, dashboard_id?, team_id?, environment_ids?}, action_class, autonomous, allowed_environment_classes, allowed_fixture_ids?, allowed_record_key_patterns?, max_request_volume?, max_concurrency?, cleanup_required, may_activate_current_revision, required_permission, approval_required, effective_from, effective_to?, version.

Policy evaluation is deterministic and snapshots policy_id + version + decision on every AgentAction. autonomous=true never grants more authority than ACL or the action's mutation contract. approval_required=true always creates a payload-bound ActionApprovalGate. No client or agent may supply a policy decision.

InvestigationSignal

Producers publish one idempotent InvestigationSignal { source_type, source_id, scenario_id?, run_id?, logical_step_id?, severity, canonical_fingerprint?, evidence_refs, target_snapshot?, execution_principal_fingerprint?, occurred_at } through the outbox. Its dedup key is (source_type, source_id, canonical_fingerprint?); 047 maps it to the active Queue item/FailureEpisode. Producers include 037 comparison/immutability, 040 breaker/consistency, 041 impact/deprecation, 042 staleness, 044 terminal run and 046 automation attention events.

Process Boundary Matrix

Process class Agent role Deterministic owner Approval mode
reasoning, evidence search, hypothesis, plan leads case/event storage delegated read policy
authoring verification logic (SQL/DSL/assertion/evaluation proposal) compiles/proposes 038 validation + 042 immutable revision delegated policy; save never activates
declared runtime semantic evaluation may reason inside bounded spec 044 deterministic orchestration + DecisionPolicy no mutation, no lifecycle ownership
query/compare/validate/index/health/fingerprint consumes output 037/038/040/041/044/046/047 engines never agent-decided
diagnostic run and fixture experiment proposes/executes runner, capacity, mutation contract delegated only in exact policy scope
revision save may execute 038 validation + 042 immutable revision delegated policy
revision activation / automation adoption may propose 042/046 lifecycle and eligibility policy or ActionApprovalGate
baseline publish, automation policy, PROD mutation may propose bound action consumer ActionApprovalGate unless exact stronger policy
HumanCheckpoint and final case disposition explains evidence 044 checkpoint / 047 case CAS authenticated analyst only

UX Invariant

Investigation work opens a persistent case workspace with chat, evidence, tool timeline and inline action cards. It MUST NOT use a modal or a confirmation dialog as the primary workflow. An ActionApprovalGate and a RunHumanCheckpoint are inline cards with typed decisions; they remain distinct domain controls.

#endregion AgentInvestigation.Cases


QUICKSTART — Dev Onboarding

Source: quickstart.md

Quickstart: Agent Test Stabilization

Purpose: Implement and verify 036 independently before 037–039.

Prerequisites

  • Backend and agent virtual environments installed.
  • PostgreSQL and configured auth/app databases available.
  • Frontend dependencies installed.
  • Test user has dashboard:testing READ and EXECUTE; separate approver has APPROVE.
cd agent
python -m pytest tests/test_agent/test_agent_context.py tests/test_agent/test_agent_run_tracker.py tests/test_agent/test_scenario_tool_filter.py -v

cd ../backend
python -m pytest tests/services/agent_runs tests/api/test_agent_runs.py -v

cd ../frontend
npm run test -- --run src/lib/models/__tests__/AgentRunModel.test.ts
npm run test -- --run src/lib/components/agent/__tests__/AgentRunPanel.ux.test.ts

npx playwright test e2e/tests/agent-scenario-run.e2e.js

Manual Smoke

  1. Start backend, standalone agent, and frontend with the repository run scripts.
  2. Open a dashboard with env_id selected.
  3. Navigate to /agent with UIContext v2 and intent=build_dashboard_test_scenario.
  4. Confirm that a run id appears before the first tool card.
  5. Simulate or run progress through context and inspect, then reload the page.
  6. Recover the same run snapshot and verify no draft crosses conversation/run boundaries.
  7. Register one warning draft; preview and download it; verify Git worktree is unchanged.
  8. Request save, deny, and verify no target exists.
  9. Request again, change the target after confirm, and verify consume returns 409.
  10. Confirm the exact request and verify a single write plus consumed gate audit.
  11. Repeat with a viewer and verify permission_denied has no confirm button.
  12. Open ordinary v1 chat and verify existing 035 behavior.

Exit Gates

  • agent_run_started precedes tool_start in captured stream.
  • Snapshot recovery survives Gradio restart.
  • Scenario tool list contains no arbitrary SQL operation.
  • Payload/path mutation and replay tests pass.
  • Existing 033/035 scoped suites pass.
  • Backend lint, frontend lint/build, and semantic anchor audit pass.

TRACEABILITY — Requirements Matrix

Source: traceability.md

#region AgentTestStabilization.Traceability [C:3] [TYPE ADR] [SEMANTICS traceability,agent-run,requirements] @BRIEF Requirements-to-contract-to-task-to-test matrix for feature 036.

Requirement Contract Tasks Verification
AGSTAB-FR-001 AgentRuns.Context.ValidateV2 T008–T011 agent context v1/v2 tests
AGSTAB-FR-002 AgentRuns.Service.Create, AgentRuns.Model T014–T018, T026–T028 create/snapshot/reload tests
AGSTAB-FR-003 AgentRuns.Service.AppendEvent, AgentRuns.Gradio.Emit T019–T021, T029 sequence/stage/stream tests
AGSTAB-FR-004 AgentRuns.Artifacts.Register, AgentRuns.DraftList T022–T024, T031 path/digest/draft UX tests
AGSTAB-FR-005 AgentRuns.Approvals.Request/Decide/Consume T015, T022, T032–T034 bound gate and reason tests
AGSTAB-FR-006 AgentRuns.Api, scenario invocation guard T006, T012, T034 RBAC and permission_denied tests
AGSTAB-FR-007 UIContext compatibility and regression phase T009, T038 existing 033/035 suites
AGSTAB-FR-008 End-to-end contracts T035–T037 Playwright/live smoke
AGSTAB-FR-009 AgentRuns.Evidence.Adapter, AgentRuns.Artifacts.Register T042–T045 optional masking, capture_meta, evidence_captured; local-provider PII posture

Story Coverage

Story Independent checkpoint
US1 v2 context produces a run id; v1 chat remains unchanged
US2 reload/restart recovers authoritative stage and drafts
US3 structured events alone render progress/drafts
US4 deny, mutation, permission, expiry, and replay cannot write

Upstream/Downstream

  • Depends on 033 chat streaming and 035 context/guardrails.
  • Blocks 037 baseline approvals, 038 draft scenario artifacts, and 039 workspace recovery.

Amendment (041-dataset-lineage-blast-radius, R5)

The trigger field on AgentRun/VerificationRun gains the additive value dataset_updated (producer: Services.Lineage.Fanout, 041 US5). Enum extension is data-level and backward compatible; no spec rewrite.

#endregion AgentTestStabilization.Traceability


TASKS — Implementation Tasks

Source: tasks.md

#region AgentTestStabilization.Tasks [C:3] [TYPE ADR] [SEMANTICS tasks,agent-run,implementation] @BRIEF Ordered TDD implementation tasks for feature 036.

Status: 51/51 completed (100%). Feature branch: 036-agent-test-stabilization. Tests: backend 85 ✅ (agent-runs scope) · frontend 3606 ✅ (historical unit evidence) · E2E 7/7 ✅ against a fresh Docker Compose stack. Last updated: 2026-07-29 — live E2E verified.

Input: all documents in specs/036-agent-test-stabilization/
Prerequisites: spec, research, plan, data model, module/event/OpenAPI contracts

Phase 1 — Contract Fixtures and Setup

  • T001 Create canonical UIContext v1/v2 and invalid-intent fixtures under specs/036-agent-test-stabilization/fixtures/context/.
  • T002 [P] Create canonical event/snapshot/gate fixtures under specs/036-agent-test-stabilization/fixtures/events/.
  • T003 Materialize fixtures into agent/tests/fixtures/, backend/tests/fixtures/agent_runs/, and frontend/src/lib/models/fixtures/agent-runs/.
  • T004 [P] Add dashboard:testing READ/EXECUTE/WRITE/APPROVE permission seeds and role mapping in backend/src/scripts/seed_permissions.py.

Phase 2 — Foundational Persistence

  • T005 Write failing model/repository tests in backend/tests/services/agent_runs/test_repository.py for ownership, terminal immutability, and sequence uniqueness.
  • T006 Write failing RBAC/API tests in backend/tests/api/test_agent_runs.py, including permission_denied and foreign-run access.
  • T007 Add AgentRun, AgentRunEvent, DraftArtifact, and ApprovalGate ORM models in backend/src/models/agent_run.py plus a migration under backend/alembic/versions/.
  • T008 Add Pydantic DTOs in backend/src/schemas/agent_run.py matching contracts/agent-runs.openapi.yaml.
  • T009 Implement repository transaction boundaries in backend/src/services/agent_runs/repository.py.

Checkpoint: Migration upgrades/downgrades; repository tests pass; no agent/frontend changes yet. ✅ DB tables created in PostgreSQL; schemas verify via 22 pytest.

Phase 3 — US1 Dashboard Context Run Start

  • T010 [US1] Write failing v1/v2 cross-field tests in agent/tests/test_agent/test_agent_context_v2.py.
  • T011 [US1] Extend agent/src/ss_tools/agent/_context.py: v1 ordinary compatibility, v2 scenario requirements, extra-field rejection.
  • T012 [US1] Write failing scenario allow-list and invocation-guard tests in agent/tests/test_agent/test_scenario_tool_filter.py, explicitly covering superset_execute_sql.
  • T013 [US1] Extend agent/src/ss_tools/agent/_tool_filter.py so RBAC runs first and scenario intent excludes every arbitrary-SQL tool.
  • T014 [US1] Write failing create/idempotency tests for AgentRuns.Service.Create.
  • T015 [US1] Implement backend/src/services/agent_runs/service.py create/snapshot flows and backend/src/api/routes/agent_runs.py.
  • T016 [US1] Register the new router in backend/src/api/routes/init.py and backend/src/app.py.
  • T017 [US1] Add typed UIContext v2 to frontend/src/lib/models/AgentChatTypes.ts without changing the Gradio argument order.
  • T018 [US1] Extend Dashboard-context parsing in frontend/src/lib/models/AgentChatModel.svelte.ts and keep ordinary v1 chat absent from run mode.

Checkpoint: Scenario start returns a durable run id before tools; ordinary chat regressions pass. ✅ Agent context v2 validates intent; tool filter excludes SQL; 275 existing frontend tests pass unchanged.

Phase 4 — US2 Recoverable Long Runs

  • T019 [US2] Write failing sequence, stage, terminal, and idempotency tests in backend/tests/services/agent_runs/test_events.py.
  • T020 [US2] Implement append/project transaction in backend/src/services/agent_runs/service.py.
  • T021 [US2] Add agent/src/ss_tools/agent/_run_tracker.py with dual-auth, idempotency keys, redaction, and persisted-before-yield behavior.
  • T022 [US2] Write failing L1 recovery tests in frontend/src/lib/models/tests/AgentRunModel.test.ts.
  • T023 [US2] Create frontend/src/lib/models/AgentRunModel.svelte.ts and compose it from AgentChatModel.svelte.ts.
  • T024 [US2] Extend AgentChat.StreamProcessor.svelte.ts for started/progress/drafts/terminal events and gap recovery.
  • T025 [US2] Restore run id from route/session-safe state on frontend/src/routes/agent/+page.svelte, then GET authoritative snapshot.

Checkpoint: Reload and simulated Gradio restart restore the same run/status/drafts. ✅ 13 L1 tests pass; AgentRunModel applies metadata, recovers snapshot, decides gates.

Phase 5 — US3 Structured Progress and Drafts

  • T026 [US3] Write failing artifact path, digest, ownership, and cleanup tests in backend/tests/services/agent_runs/test_artifacts.py.
  • T027 [US3] Implement backend/src/services/agent_runs/artifacts.py with out-of-repository storage and opaque download/preview.
  • T028 [US3] Emit agent_run_started before tool_start and persist all scenario progress in agent/src/ss_tools/agent/app.py.
  • T029 [US3] Write failing L2 tests in frontend/src/lib/components/agent/tests/AgentRunPanel.ux.test.ts and DraftArtifactList.ux.test.ts.
  • T030 [US3] Implement frontend/src/lib/components/agent/AgentRunPanel.svelte and frontend/src/lib/components/agent/DraftArtifactList.svelte using $lib/ui and semantic tokens.
  • T031 [US3] Wire the panels into frontend/src/lib/components/agent/AgentChat.svelte without parsing prose.

Phase 6 — US4 Bound HITL Gates

  • T032 [US4] Write failing request/decision/consume tests in backend/tests/services/agent_runs/test_approvals.py for reason, expiry, hash mutation, RBAC revocation, and replay.
  • T033 [US4] Implement backend/src/services/agent_runs/approvals.py and REST gate endpoints in backend/src/api/routes/agent_runs.py with atomic consume.
  • T034 [US4] Extend agent/src/ss_tools/agent/_confirmation.py, frontend/src/lib/models/AgentChatTypes.ts, and frontend/src/lib/components/assistant/ConfirmationCard.svelte for gate id, exact targets, warnings, and required reason.
  • T035 [US4] Add denial and permission-denied frontend tests; confirm no confirm control for unauthorized actors.

Phase 7 — Integration and Quality Gates

  • T036 Add agent/tests/test_agent/test_agent_run_tracker.py covering backend loss and no unpersisted event emission.
  • T037 Add frontend/e2e/tests/agent-scenario-run.e2e.js for context → run id → progress → draft → deny/recover.
  • T038 Run all existing 033/035 agent, context, confirmation, retry, timeout, and frontend model/component suites.
  • T039 Run quickstart.md including Gradio restart and Git worktree unchanged checks.
  • T040 Run backend/agent ruff, backend/agent pytest, frontend lint/test/build.
  • T041 Audit ATTN_1–4, exact anchor pairs, unresolved relations, and direct-SQL exclusion.

Phase 8 — Screenshot Evidence and Optional Display Masking (AGSTAB-FR-009)

  • T042 [P] Write failing capture_meta validation and masking tests in backend/tests/services/agent_runs/test_evidence.py.
  • T043 [P] Add capture_meta fields to DraftArtifact Pydantic model; validate kind=screenshot_evidence requires capture_meta.
  • T044 Implement backend/src/services/agent_runs/evidence.py: bridge existing ScreenshotService capture → DraftArtifact with masked derivative.
  • T045 [P] Extend AgentRuns.Artifacts.Register to accept capture_meta and enforce mask_selectors contract.
  • T046 Write failing evidence_captured event tests in backend/tests/services/agent_runs/test_events.py.
  • T047 Extend AgentRuns.Gradio.Emit to emit evidence_captured events with artifact refs, viewport, and sha256.
  • T048 Write failing evidence preview/disposition model tests in frontend/src/lib/models/tests/AgentRunModel.evidence.test.ts.
  • T049 Extend frontend AgentRunModel to accept evidence_captured events and expose evidence array.
  • T050 Verify optional masked derivatives are stored separately and preview ACLs are enforced; unmasked local-provider submission is allowed, while credentials/cookies/tokens are never exposed.
  • T051 Audit local-provider evidence handling: unmasked capture is permitted for local LLM/VLM analysis; preview ACLs and secret exclusion remain enforced.

Implementation Closure — Lifecycle, Idempotency, Recovery, Readiness

Lifecycle FSM (AgentRun)

  • States: CREATED → RUNNING → [WAITING_INPUT | WAITING_APPROVAL] → COMPLETED | FAILED | CANCELLED
  • Terminal runs (COMPLETED/FAILED/CANCELLED) are immutable: no events, drafts, or approvals after terminal.
  • Stage tracking: current_stage updated by each append_event; snapshot derives stages list from event history.
  • Approval lifecycle: request_approval (→ WAITING_APPROVAL) → decide_approval (confirm → RUNNING, deny → RUNNING) → consume_approval (→ COMPLETED or stay RUNNING per-draft).
  • Implemented in Services.AgentRuns.Service, Services.AgentRuns.Approvals, Api.AgentRuns.

Idempotency

  • Sequence-based: (run_id, sequence) unique. Duplicate sequence with matching payload_hash = idempotent replay (returns existing event). Mismatched hash = ValueError (integrity violation).
  • Create idempotency: idempotency_key on CreateAgentRunRequest (optional; dedup at app layer).
  • Payload hash: SHA-256 of canonical JSON; backend computes if absent.
  • Implemented in Services.AgentRuns.Service.AppendEvent and AgentRunRepository.append_event.

Recovery (Frontend)

  • AgentRunModel.svelte.ts: composes with AgentChatModel.svelte.ts; restores run id from route/session-safe state on reload.
  • StreamProcessor: handles started/progress/drafts/terminal events; gap recovery via GET /api/agent/runs/{run_id}/events after reconnection.
  • Recovery tests: L1 model tests (13) for snapshot, metadata, gate decision, gap recovery.
  • Agent-side: _run_tracker.py with dual-auth, idempotency keys, redaction, persisted-before-yield.

Readiness Verification (036 scope)

  • Readiness endpoint GET /api/ready: lightweight DB connectivity check (SELECT 1), unauthenticated, returns 200/503.
  • Wired in Docker Compose healthchecks via docker/backend.Dockerfile and docker-compose.e2e.yml.
  • E2E tests: 7/7 verified against fresh Docker Compose stack (context → run id → progress → draft → deny/recover).

Test Stabilization Results

  • Backend: 85 agent-runs scope tests passing; ruff clean.
  • Frontend: 3606 historical unit tests + new L1/L2 model/component tests passing; lint/build clean.
  • E2E: 7/7 passing against fresh Docker Compose.
  • T041 (ATN audit): verified ATTN_1–4 density on all new contracts; exact anchor pairs confirmed; unresolved relations remain debt (2420 orphans pre-existing).

Dependencies

T001–T009 → US1 → US2 → US3 → US4 → integration.
Within a phase, test tasks precede implementation. 037 must not begin before the US4 consume contract passes. Phase 8 requires completed US3 (DraftArtifact infrastructure).

#endregion AgentTestStabilization.Tasks

================================================================================ FEATURE: 037-superset-baseline-engine Files: 15


SPEC — Feature Specification

Source: spec.md

#region SupersetBaselineEngine.Spec [C:3] [TYPE ADR] [SEMANTICS spec,requirements,superset,baseline,dashboard-testing] @BRIEF Superset-native query and baseline engine for dashboard testing without direct SQL execution. @RELATION DEPENDS_ON -> [Doc.Adr.ADR0001] @RELATION DEPENDS_ON -> [Doc.Adr.ADR0003] @RELATION DEPENDS_ON -> [Doc.Adr.ADR0005] @RELATION DEPENDS_ON -> [AgentTestStabilization.Spec] @RATIONALE Dashboard test assertions must validate the same Superset-side chart or dataset execution path that powers dashboards, not a parallel SQL path that can diverge from Superset filter/query semantics. @REJECTED Direct SQL execution for chart/baseline truth — rejected because this feature requires stable Superset dataset/chart execution and filter fidelity. This does not prohibit 038 validated SqlEvidenceSpec for an independent source-mart oracle. @REJECTED A second bespoke chat runtime as the only consumer of baseline tools — superseded 2026-08-24: capture/approval/verification tools join the MCP catalog as thin forwarders per specs/050-mcp-interface/spec.md; server-side hash computation and gate policy are unchanged.

Navigation (DSA Indexer keywords)

@SEMANTICS: spec, requirements, feature, superset, baseline, chart-data, dataset, filters, normalization, dashboard-testing

Feature Branch: 037-superset-baseline-engine
Created: 2026-07-07 | Status: Ready for Implementation Input: "Create a Superset-native query and baseline engine for dashboard testing. The system must inspect dashboard query models, map dashboard native filters into Superset chart or dataset query context, execute Superset-side chart or dataset queries without direct SQL, normalize returned metric and table values, compare them with approved baseline catalog entries, and create draft baseline candidates requiring human approval."

User Scenarios

Story 1 — Inspect Dashboard Query Model (P1)

Why P1: Scenario generation must know charts, datasets, metrics, filters, and filter scopes before it can build reliable tests.

Independent Test: Provide a dashboard fixture and verify the engine returns structured query model data for charts, datasets, metrics, native filters, and capabilities.

Acceptance:

  1. Given a dashboard id and environment When the query model is inspected Then the result lists dashboard title, charts, datasets, metric labels/keys, native filters, filter targets, and export capabilities.
  2. Given a native filter applies only to selected charts When inspection runs Then filter scope is represented so unrelated charts are not queried with invalid filters.
  3. Given a chart or dataset is inaccessible When inspection runs Then the engine reports a structured warning or permission error without inventing missing metadata.

Story 2 — Execute Superset Query Context (P1)

Why P1: Baseline validation depends on executing Superset chart/dataset queries through Superset-native APIs, never through direct SQL.

Independent Test: Execute a fixture chart query with normalized dashboard filters and verify returned data is traceable to chart id, dataset id, metric, filters hash, and environment.

Acceptance:

  1. Given a chart metric and normalized filter context When execution is requested Then the engine calls the Superset-native chart/dataset execution path and returns structured values.
  2. Given dashboard filters are applied When query context is built Then Superset receives semantically equivalent filters to the UI native filter state.
  3. Given a request would require arbitrary SQL text from the agent When execution is attempted Then the request is rejected as unsupported.

Story 3 — Normalize and Compare Results (P1)

Why P1: UI, Superset API, and XLSX values must be comparable despite formatting differences.

Independent Test: Feed formatted decimal, date, percent, big-number, and table-row fixtures and verify normalized values compare consistently with tolerance rules.

Acceptance:

  1. Given Superset returns a metric value When normalization runs Then the value is represented with type, raw value, normalized value, label, and source metadata.
  2. Given an approved baseline exists When an actual value is compared Then exact, absolute tolerance, relative tolerance, range, or row-set comparison rules are applied according to baseline policy.
  3. Given actual data cannot be normalized When comparison is requested Then the result is inconclusive with an explanation, not a false pass.

Story 4 — Baseline Candidate Lifecycle (P2)

Why P2: The agent may discover candidate reference values but must not silently define the truth.

Independent Test: Run discovery for a metric without baseline and verify a draft candidate is produced; approving it follows the bound inline ActionApprovalGate policy from feature 036.

Acceptance:

  1. Given no approved baseline exists for metric+filters When discovery runs Then a draft baseline candidate is created with provenance and source values.
  2. Given a dashboard/chart/dataset/filter fingerprint changes When existing baselines are loaded Then affected baselines are marked or reported as stale.
  3. Given UI/API/XLSX sources disagree When a candidate is created Then candidate approval is blocked or warning-gated until a user reviews the discrepancy.

Edge Cases

  • Superset returns empty result → compare result distinguishes expected empty, unexpected empty, and inconclusive.
  • Filter value is not available in the target environment → query execution fails with recoverable validation error.
  • Dashboard or chart metadata changes after baseline approval → fingerprint mismatch surfaces stale baseline warning.
  • Superset API returns 403/404/422/5xx/timeout → taxonomy is preserved in comparison/report output.
  • Date and number formats differ between locales → normalization uses canonical decimal/date representation.

Requirements

Functional

  • AGBASE-FR-001: The engine MUST inspect dashboard query models into structured charts, datasets, metrics, native filters, filter scopes, and capabilities.
  • AGBASE-FR-002: The engine MUST execute Superset-native chart/dataset query contexts without accepting arbitrary SQL text from the agent.
  • AGBASE-FR-003: Dashboard native filters MUST be normalized into query filters with target columns, operators, values, scopes, and a deterministic filter hash.
  • AGBASE-FR-004: Superset result normalization MUST support scalar metrics, big-number charts, table rows, dates, decimals, percentages, empty values, and raw/source metadata.
  • AGBASE-FR-005: Baseline catalog entries MUST be pinned to a specific DashboardRelease via release_version and release_commit_hash. The release identity replaces per-field fingerprint tracking. Baseline without release pinning is invalid.
  • AGBASE-FR-006: Each baseline entry MUST record a source_response_hash (SHA-256 of the Superset API response at the time the expected value was captured) and captured_at (ISO-8601 timestamp). These enable immutability violation detection independent of metric value comparison.
  • AGBASE-FR-007: For closed-period entries (immutability.enabled=true), the system MUST detect immutability violations: when source_response_hash changes for the same filters, the status MUST be immutability_violation (CRITICAL severity), NOT a stale warning. Automated baseline updates for closed-period entries are forbidden.
  • AGBASE-FR-008: Comparison output MUST include source, actual value, expected value, diff, status, tolerance rule, and warnings. Status enum MUST include immutability_violation as a distinct, critical category separate from stale_baseline and stale_visual_baseline.
  • AGBASE-FR-009: Non-pass comparison outcomes MUST emit an idempotent 036 InvestigationSignal with immutable evidence provenance; 047 owns Queue creation/update. They MUST NOT automatically begin an agent conversation, alter a baseline, or reclassify comparison truth.
  • AGBASE-FR-010: An opened InvestigationCase MAY use baseline-engine tools for read-only inspection, diagnostic comparison and candidate creation. Any baseline publication remains governed by deterministic catalog validation and its ActionApprovalGate policy.
  • AGBASE-FR-009: Direct SQL execution, generated SQL assertions and SQL-based provenance are out of scope for chart/baseline truth. 038 validated immutable SqlEvidenceSpec remains the independent source-mart oracle contract and does not alter this feature's Superset-native baseline semantics.
  • AGBASE-FR-010: The baseline catalog MUST support visual baselines for screenshot comparison, with the same release-pinning and immutability rules as metric baselines.
  • AGBASE-FR-011: The system MUST compute a StructureDiff between two releases' DashboardQueryModel snapshots. The diff MUST classify changes by target (chart, filter, column), kind (scope_change, column_reorder, chart_removed, etc.), severity (critical, warning, info), and affected artifacts. StructureDiff is orthogonal to metric comparison.
  • AGBASE-FR-012: Baseline entries for closed periods MUST carry an immutability block: enabled flag, period identifier, frozen_at timestamp, and policy (alert, block_publish, require_investigation). Immutability violations MUST be surfaced at the PROD publish gate and scheduled checks, not silently ignored.
  • AGBASE-FR-013: Baseline inheritance between releases: when creating a new release, metric entries whose chart content_hash has NOT changed since the previous release inherit their expected value and source_response_hash. Entries whose chart content_hash HAS changed require re-extraction from the PREPROD environment. The analyst reviews the inheritance diff before confirming.

Key Entities

  • DashboardQueryModel: Structured representation of charts, datasets, metrics, native filters, scopes, and execution capabilities for a dashboard.
  • NormalizedFilterContext: Canonical filter state shared by UI, Superset query execution, scenario graph, XLSX comparison, and baseline lookup.
  • SupersetQueryContext: Executable Superset-native query request derived from chart/dataset metadata plus normalized filters.
  • NormalizedSupersetResult: Canonical metric/table result with type metadata and provenance.
  • BaselineEntry: Release-pinned expected value for a metric/table/visual result. Pinned to release_version + release_commit_hash. Carries source_response_hash for immutability detection. Stored in baseline.yaml in the dashboard's git repository.
  • StructureDiff: Structural comparison of DashboardQueryModel between two releases. Classifies changes: filter scope, column order, chart add/remove, viz_type change, group_by change. Severity: critical, warning, info. Computed without metric execution.
  • ImmutabilityViolation: Critical status raised when source_response_hash for a closed-period baseline entry diverges from the recorded value. Indicates data was modified retroactively. Blocks release publication and triggers investigation.
  • VisualBaseline: Approved expected visual state of a dashboard tab or region — pinned to release, with same immutability rules as metric baselines.
  • ComparisonResult: Result of comparing actual normalized values to baseline expectations. Status includes immutability_violation (critical data integrity breach), stale_visual_baseline (layout/selectors diverged), and standard pass/fail/inconclusive.
  • VerificationRun: Record linking an AgentRun (036) to a DashboardRelease, capturing which verification categories were executed, their outcomes, and the trigger (deploy_to_preprod, release_create, scheduled, etl_completed).

Success Criteria

  • SC-001: Dashboard query model inspection returns complete chart/filter/dataset mappings for fixture dashboards with 100% deterministic JSON snapshots.
  • SC-002: Superset-native execution validates at least scalar metric and table chart fixtures without direct SQL.
  • SC-003: Decimal/date/percent normalization compares equivalent API/UI/XLSX formatted values without false diffs in fixture tests.
  • SC-004: Baseline entries without release_version + release_commit_hash are rejected at catalog load.
  • SC-005: StructureDiff correctly classifies filter scope change, column reorder, chart add/remove, and group_by change for fixture releases with 100% deterministic JSON output.
  • SC-006: Immutability violation is raised within 60 seconds of detecting a changed source_response_hash for a closed-period entry in scheduled-check fixtures.

Implementation Status & Confirmed Reuse (audit 2026-08-07)

Facts (code check, not tasks.md):

  • ✅ Работает: SupersetClient.ChartData.Execute (backend/src/core/superset_client/_chart_data.py) выполняет реальный async POST /api/v1/chart/data через httpx; QueryModel.Inspect, Filters.Normalize, Result.Normalize, Comparison, Candidates/Approvals, StructureDiff, Immutability — реализованы и покрыты тестами.
  • ✅ Visual-слой переиспользует BaselineEngine.Visual.SSIM (pure NumPy) и BaselineEngine.Verification.ExecutorVisual.Async (authoritative fingerprints, независимость evidence-раннов). Relations зафиксированы в contracts/modules.md.
  • ✅ Дискретные инструменты оценки метрик реализованы и реальны: comparison.py (exact / absolute / relative / range / row-set + immutability precedence), normalization.py (locale decimal/date/percent/big-number/table), metric_executor_async.py (полный Superset-поток: entry validation → get_superset_client → inspect_dashboard_query_model → query envelope → source_response_hash → _compare_with_baseline). API: GET /query-model, POST /filters/normalize, POST /queries/execute, POST /comparisons, GET /baselines.
  • 🔴 Gap A — VerificationRun не создаётся автоматически pipeline-хуками: create_verification_run_async вызывается только из POST-роута и verification_scheduler (scheduled daily). Deploy-хук deploy_to_preprod / release_create / etl_completed НЕ создают VerificationRun (037 AGBASE-FR-012-контекст, 039 AGUI-FR-014-016 обещают это) — в git_deployment_recorder.py и deploy-маршрутах отсутствует вызов verification. Scheduled-проверка работает (core/scheduler.py → execute_scheduled_verification_check).
  • 🔴 Gap B — GET-эндпоинты verification отсутствуют: frontend getVerificationHistory()/getVerificationRun() (AGUI-FR-015/016) зовут GET /dashboard-testing/verification/history и GET /dashboard-testing/verification/{runId}, но backend имеет только POST /verification-runs. Pipeline views 039 не могут загрузить историю/детали.

Закрытие: задачи T080–T081 в tasks.md Phase 10 (deploy-hook trigger + GET-эндпоинты) — дискретный контур оценки метрик работает «напрямую», но пайплайн-автоматизация и read-API не дописаны.

Drift Amendment — MCP Interface (2026-08-24)

  • Baseline tools (capture_baseline_candidate, approval lifecycle, create_verification_run) are exposed 1:1 through the MCP catalog (050 Phase 1 parity); no behavioral change to engine semantics.
  • The vendored research/mcp-superset server is explicitly NOT adopted as an alternative surface: it bypasses this feature's server-side hash computation, release pinning and immutability rules.

Status (2026-09-02): done — реализовано в рамках 050: инструменты и гейты (specs/050-mcp-interface/tasks.md T012–T028 [x]), handoff-поверхность (050 T030–T033), демонтаж чата и сервиса agent/ (050 T040–T041, чекпоинты specs/WORKSTATE-043-047.md).

#endregion SupersetBaselineEngine.Spec


UX REFERENCE — Interaction Narrative

Source: ux_reference.md

#region SupersetBaselineEngine.UxReference [C:3] [TYPE ADR] [SEMANTICS ux,reference,superset,baseline] @BRIEF UX reference for Superset-native baseline execution and comparison results. @RELATION DEPENDS_ON -> [SupersetBaselineEngine.Modules] @RELATION DEPENDS_ON -> [SupersetBaselineEngine.ApiUx]

Feature Branch: 037-superset-baseline-engine Created: 2026-07-07 | Status: Ready for Implementation

1. User Persona & Context

  • Who is the user?: QA engineer or dashboard owner validating reference metric values through Superset-native execution.
  • What is their goal?: See whether a dashboard metric under specific filters matches approved baseline values without running direct SQL.
  • Context: Agent workspace or later scenario UI displays inspection, execution, normalization, baseline, and comparison states.

2. Happy Path Narrative

The user asks the agent to validate a metric for a dashboard and filter set. The system inspects the dashboard query model, executes the related Superset chart or dataset query, normalizes the result, compares it with an approved baseline, and displays a source-aware comparison. The UI explicitly states that no direct SQL was executed.

3. Interface Mockups

Query Model Summary

┌──────────────────── Dashboard query model ──────────────────────────────────┐
│ Dashboard: FI-0080 | env=ss-dev                                             │
│ Charts: 7 | Datasets: 3 | Native filters: 4                                 │
│ Capabilities: chart_data ✓ dataset_query ✓ xlsx_export ✓                    │
└─────────────────────────────────────────────────────────────────────────────┘

Metric Comparison

┌──────────────────── Metric validation ───────────────────────────────────────┐
│ Metric: Просроченная ДЗ                                                      │
│ Chart: Итого просроченная ДЗ | Dataset: dataset_finance_debt                │
│ Filters: Дата=2026-05-29, Контрагент=АСК                                     │
│                                                                             │
│ Source                 Value          Status                                 │
│ Superset API           1 234 567.89   ✓ matches baseline                     │
│ Approved baseline      1 234 567.89   approved                               │
│                                                                             │
│ No direct SQL was executed.                                                  │
└─────────────────────────────────────────────────────────────────────────────┘

Baseline Candidate

┌──────────────────── Baseline candidate ─────────────────────────────────────┐
│ No approved baseline found.                                                  │
│ Candidate from Superset execution: 1 234 567.89                              │
│ Provenance: chart_id=128, dataset_id=77, filters_hash=sha256:...             │
│                                                                             │
│ [Approve baseline] [Keep draft] [Discard]                                    │
└─────────────────────────────────────────────────────────────────────────────┘

4. Error Experience

Scenario A: Superset API Permission Denied

  • System Response: Comparison card shows 403 permission denied, target chart/dataset, and required action.
  • Recovery: User can switch environment, request access, or skip that check.

Scenario B: Stale Baseline

  • System Response: Baseline status becomes warning-tone with fingerprint diff category.
  • Recovery: User can run discovery and submit a new baseline candidate for approval.

Scenario C: Inconclusive Normalization

  • System Response: UI shows raw value and reason normalization failed.
  • Recovery: User can mark as manual checkpoint or adjust metric mapping in a later planning phase.

5. Tone & Voice

  • Style: Precise, source-aware, explicit about no direct SQL.
  • Terminology: Use "Superset API", "query model", "baseline", "candidate", "stale", "fingerprint".

#endregion SupersetBaselineEngine.UxReference


CHECKLISTS — Requirements Quality — requirements.md

Source: checklists/requirements.md

Requirements Checklist: 037 Superset Baseline Engine

Purpose: Validate specification quality and no-SQL baseline scope. Created: 2026-07-07 Feature: specs/037-superset-baseline-engine/spec.md

Spec Completeness

  • CHK001 User stories cover inspection, Superset-native execution, normalization, comparison, and baseline lifecycle.
  • CHK002 Requirements explicitly reject direct SQL execution and SQL-based baseline provenance.
  • CHK003 No unresolved [NEEDS CLARIFICATION] markers remain.
  • CHK004 Edge cases include Superset API errors, stale fingerprints, empty results, and locale formatting.
  • CHK004a Visual baseline type and stale_visual_baseline status are defined alongside metric baselines; cross-kind comparison is explicitly rejected.

Constitution Coverage

  • CHK005 ADR relations preserve external orchestrator and RBAC boundaries.
  • CHK006 Decision memory explains Superset-native execution over SQL.
  • CHK007 Approved baseline lifecycle requires reviewable artifacts and HITL approval.
  • CHK008 Requirements are implementation-free but measurable.

Readiness for Plan

  • CHK009 Query model, normalized filters, query context, result, baseline, and comparison entities are defined.
  • CHK010 Success criteria can be tested with deterministic fixtures.
  • CHK011 Scope excludes scenario graph UI and generated Playwright/XLSX execution details.

Implementation Package

  • CHK012 Research and implementation plan resolve all design decisions.
  • CHK013 Data model, module, OpenAPI, JSON Schema, and UX contracts are present.
  • CHK014 Quickstart defines deterministic Superset fixture verification.
  • CHK015 Traceability maps every functional requirement to contracts, tasks, and tests.
  • CHK016 Tasks use exact repository paths, dependency order, and test-first sequencing.
  • CHK017 Machine-readable contracts, external references, and semantic anchors pass validation.

UX DECISIONS — Final Choices

Source: contracts/ux/decisions.md

#region SupersetBaselineEngine.UxDecisions [C:3] [TYPE ADR] [SEMANTICS ux,decisions,baseline] @BRIEF Final UX decisions for baseline engine outputs consumed by 039. @RELATION DEPENDS_ON -> [SupersetBaselineEngine.Modules] @RELATION DEPENDS_ON -> [BaselineEngine.Candidates.Create] @RELATION DEPENDS_ON -> [BaselineEngine.Visual.Compare]

  1. Always show source identity and “no direct SQL” statement.
  2. Never collapse stale, missing, source_error, or inconclusive into fail/pass.
  3. Display tolerance policy alongside the result, not only in details.
  4. Approved and candidate values use distinct status labels and tones.
  5. Fingerprint change is explained by query/dataset/filter category.
  6. Candidate approval reuses 036 gate and repeats all discrepancies.
  7. Large table diffs show bounded summary plus downloadable evidence, not unbounded DOM rows.

#endregion SupersetBaselineEngine.UxDecisions


UX API CONTRACT — Endpoints & Shapes

Source: contracts/ux/api-ux.md

#region SupersetBaselineEngine.ApiUx [C:3] [TYPE ADR] [SEMANTICS ux,api,baseline,comparison] @BRIEF User-visible mapping of inspection, execution, comparison, and candidate API outcomes. @RELATION DEPENDS_ON -> [SupersetBaselineEngine.Modules] @RELATION DEPENDS_ON -> [BaselineEngine.Api]

Outcome UI state Recovery
Query model complete Counts/capabilities and source environment Continue
Partial metadata Warning with affected chart/dataset Exclude/manual checkpoint
Query running Source and target chart visible, aria-busy Cancel view only; backend timeout governs
pass/fail Actual, expected, diff, policy, provenance Inspect evidence
missing baseline Candidate action, never pass Discover candidate
stale baseline Fingerprint dimensions and warning Rediscover; do not overwrite
inconclusive Raw bounded value and reason Adjust mapping/manual checkpoint
403/404/422/timeout/5xx Distinct error code and resource Switch env, request access, correct filters, retry

Every result card includes “Superset API; no direct SQL” when execution was chart-data based.

#endregion SupersetBaselineEngine.ApiUx


UX DESIGN — Per-Screen Contracts — baseline-engine-ux.md

Source: contracts/ux/baseline-engine-ux.md

#region SupersetBaselineEngine.ResultUx [C:4] [TYPE ADR] [SEMANTICS ux,baseline,result,candidate] @BRIEF Display contract for query source, canonical comparison, staleness, and candidate review. @RELATION DEPENDS_ON -> [SupersetBaselineEngine.Modules] @RELATION DEPENDS_ON -> [BaselineEngine.Comparison.Compare] @RELATION DEPENDS_ON -> [BaselineEngine.Result.Normalize]

Result Card

  • Identity: dashboard, chart/dataset, result label/key, environment.
  • Filters: normalized human-readable list plus short hash.
  • Values: actual, expected, diff, tolerance policy.
  • Status: pass, fail, inconclusive, missing, stale, or source error.
  • Provenance: observed time, Superset source, baseline approval actor/reason/time.

Candidate Card

  • Candidate never uses approved styling.
  • Shows all source observations and discrepancy warnings.
  • Approval control is delegated to the 036 confirmation card.
  • Reason is mandatory; exact value/filter/fingerprint target is reviewable.
  • Keep draft and discard are distinct from approve.

UX Tests

  1. Missing baseline cannot render green/pass.
  2. Stale result names changed fingerprint dimensions.
  3. Locale-equivalent values show no false diff.
  4. API/XLSX discrepancy repeats warning in approval gate.
  5. Direct-SQL copy never appears as an offered route.

#endregion SupersetBaselineEngine.ResultUx


PLAN — Implementation Plan

Source: plan.md

Implementation Plan: Superset Baseline Engine

Branch: 037-superset-baseline-engine | Date: 2026-07-13 | Spec: spec.md

Summary

Add a backend dashboard-testing domain that inspects saved Superset dashboard/chart/dataset metadata, converts validated native filters to authoritative chart-data query context, normalizes results, and compares them to reviewable approved baselines. Candidate approval and repository persistence reuse 036 bound HITL gates.

Technical Context

Language/Version: Python 3.13+
Dependencies: FastAPI 0.126, Pydantic 2, existing async SupersetClient/httpx, ruamel/PyYAML-compatible safe YAML layer already used by repository tooling
Storage: Approved YAML in resolved Git worktree; drafts via 036; no new runtime-value table required for MVP
Testing: pytest unit/contract; Superset 4.1.2 Testcontainers integration
Performance Goals: fixture inspection deterministic; single chart execution p95 under 5s excluding Superset timeout; scalar compare under 10ms
Constraints: no arbitrary SQL fields; decimal-safe normalization; async I/O; RBAC and 036 approval binding
Scale: up to 100 charts, 20 filters, 10,000 normalized table rows per bounded result

Constitution Check

Principle Result
Contracts/decision memory PASS — C4/C5 query and approval boundaries fully contracted
External orchestrator PASS — only Superset REST API is called
Module discipline PASS — inspection, filters, execution, normalization, catalog, comparison separated
RBAC PASS — READ/EXECUTE/APPROVE split; approval rechecked by 036
TDD PASS — fixtures and rejected SQL path tests precede implementation
Async backend PASS — existing async client; no requests or blocking file I/O on event loop
Attention PASS — BaselineEngine.* hierarchy and shared baseline semantics

Project Structure

backend/src/
├── api/routes/dashboard_testing.py
├── schemas/dashboard_testing.py
├── services/dashboard_testing/
│   ├── query_model.py
│   ├── filters.py
│   ├── query_executor.py
│   ├── superset_adapter.py
│   ├── normalization.py
│   ├── fingerprints.py
│   ├── baseline_catalog.py
│   └── comparison.py
└── core/superset_client/
    └── _chart_data.py

backend/tests/
├── services/dashboard_testing/
├── api/test_dashboard_testing.py
└── integration/test_dashboard_testing_superset.py

Delivery Phases

  1. Canonical DTOs and fixtures.
  2. Dashboard query-model inspection and filter normalization.
  3. Chart-data adapter/execution with no-SQL guard.
  4. Value normalization and comparison policies.
  5. YAML catalog, fingerprints, draft candidate, 036 approval integration.
  6. API, agent tools, integration and regression gates.

Public Boundary

The API is defined in contracts/dashboard-testing.openapi.yaml. Agent tools call only this backend boundary; they do not call Superset or repository paths directly.

Cross-Spec Boundary

  • Requires 036 run/draft/gate substrate.
  • Exposes query model, comparison, and candidate DTOs consumed by 038/039.
  • Does not build scenario graphs or render the workspace.

Complexity Tracking

No exception planned. Table normalization must enforce row and payload limits rather than loading unbounded Superset results.

Pipeline Automation & Read-API Closure (audit 2026-08-07)

Status correction: Дискретные инструменты оценки метрик (normalization/comparison/metric_executor_async) работают и реально исполняются. Но пайплайн-автоматизация и read-API не замкнуты: deploy-хук deploy_to_preprod/release_create/etl_completed не создаёт VerificationRun (создаётся только из POST-роута и daily scheduled-проверки), а GET-эндпоинты /verification/history и /verification/{run_id} отсутствуют, хотя frontend 039 их вызывает.

Closure tasks (tasks.md Phase 10, T080–T081):

  • T080 — Wire create_verification_run_async в deploy-хук для trigger'ов deploy_to_preprod/release_create/etl_completed.
  • T081 — Add GET /verification/history + GET /verification/{run_id} в api/routes/dashboard_testing/verification.py, совместимые с VerificationRunDTO[]/VerificationRunDTO.

Exit rule: 039 pipeline views (AGUI-FR-014..016) не могут отображать живые VerificationRun до мерджа T080/T081.


RESEARCH — Technical Decisions

Source: research.md

#region SupersetBaselineEngine.Research [C:4] [TYPE ADR] [SEMANTICS research,superset,baseline,chart-data,normalization] @BRIEF Phase 0 decisions for authoritative dashboard query inspection, Superset-native execution, canonical values, and baseline lifecycle. @RELATION DEPENDS_ON -> [SupersetBaselineEngine.Spec] @RELATION DEPENDS_ON -> [AgentTestStabilization.Modules] @RATIONALE Existing SupersetClient already resolves dashboards, charts, datasets, filters, TLS, and auth, so the feature extends that boundary instead of creating a second HTTP stack. @REJECTED SQL Lab and arbitrary SQL as baseline truth — rejected by feature scope and because it bypasses dashboard query semantics.

1. Existing Capability Audit

  • SupersetClient already exposes dashboard detail, dashboard charts/datasets, chart detail, dataset detail, native filter state parsing, and async HTTP.
  • POST /api/v1/chart/data is present in the checked-in Superset OpenAPI and is already used by translation preview code.
  • Existing build_dataset_preview_query_context is a reduced preview builder, not a chart-fidelity engine; it must not be reused as though it reproduces a saved chart.
  • Integration infrastructure pins Apache Superset 4.1.2 through ADR-0012.
  • The standalone agent currently has SQL tools. Feature 036 must remove them from scenario intent reachability.

2. Query Model Inspection

  • Decision: Build DashboardQueryModel from authoritative dashboard, chart, and dataset metadata returned by the selected environment.
  • Filter source: Parse native_filter_configuration and scope/excluded chart rules from dashboard metadata; keep unresolved/unsupported filter nodes as warnings.
  • Chart source: For every chart, retain id, dataset reference, viz type, metrics/group-bys, saved form/query context, result capabilities, and metadata fingerprint inputs.
  • Partial access: Return per-resource warnings and capability=false; never invent chart/dataset fields.
  • Alternative rejected: Reconstruct query models from rendered DOM — fragile and loses API-level provenance.

3. Superset-Native Execution

  • Decision: Public request identifies environment, dashboard, chart/dataset result key, and normalized filters. Backend loads authoritative metadata and builds POST /api/v1/chart/data payload.
  • Invariant: Request DTO has no sql, query text, adhoc SQL expression, or free-form endpoint field.
  • Validation: Each filter target must exist on the dataset and be in the dashboard filter scope for the chart.
  • Version adapter: Keep Superset 4.1.2 response parsing in a narrow adapter so a later Superset version does not leak through domain DTOs.
  • Alternative rejected: Accept raw query_context from the agent — it can smuggle unsupported expressions and bypass saved-chart constraints.

4. Canonical Filter Identity

  • Sort filters by target dataset/column/operator and canonical value.
  • Preserve semantic types: date, datetime, decimal, integer, boolean, string, enum/list, null.
  • Use inclusive/exclusive range structure instead of localized strings.
  • Compute filters_hash from schema version plus canonical JSON, UTF-8, sorted keys, no insignificant whitespace.
  • Scope is part of the query model fingerprint, not of a single chart filter hash.

5. Result Normalization

  • Decimal values use decimal strings, never binary float serialization.
  • Dates/datetimes use ISO-8601; timezone-aware datetimes normalize to UTC.
  • Percent stores canonical ratio plus display value/format metadata when Superset supplies it.
  • Scalar and big-number results identify metric/result key and label.
  • Table results use ordered column descriptors and canonical rows; row-set comparison requires explicit key columns or uses multiset semantics.
  • Unsupported/ambiguous values produce inconclusive with reason, never false pass.

6. Baseline Artifact and Provenance

  • Decision: Approved catalog is reviewable YAML at dashboard_tests/{dashboard_key}/baselines.yaml inside the resolved Git repository.
  • Draft candidates: Stored as 036 DraftArtifact outside the repository until approved.
  • Runtime observations/evidence: Stored outside the approved catalog and linked by run/evidence id.
  • Provenance: Environment, dashboard/chart/dataset ids, result key, normalized filters hash, observed timestamp, Superset version, source response hash, actor/run, and fingerprints.
  • Alternative rejected: Database-only approved expectations — invisible in review and disconnected from repository history.

7. Staleness

Three fingerprints are recorded:

  1. query_fingerprint: saved chart/query semantics and result selector;
  2. dataset_fingerprint: dataset id/uuid, columns, metric definitions, relevant metadata;
  3. filter_fingerprint: dashboard native filter definitions and scope.

Any mismatch makes an approved entry STALE; it is never silently updated. Discovery creates a new candidate.

8. Comparison Policies

  • exact: canonical equality;
  • absolute_tolerance: abs(actual-expected) <= amount;
  • relative_tolerance: abs(actual-expected) <= abs(expected) * ratio, with zero rule;
  • range: inclusive/exclusive min/max;
  • row_set: keyed ordered/unordered policy with missing/extra/changed rows.

Statuses: pass, fail, inconclusive, missing_baseline, stale_baseline, permission_denied, source_error.

9. Error Taxonomy

Preserve 401/session, 403 permission, 404 resource, 422 query/filter validation, timeout, 5xx upstream, empty result, normalization unsupported, and catalog invalid as distinct codes.

#endregion SupersetBaselineEngine.Research


DATA MODEL — Entities & Relations

Source: data-model.md

#region SupersetBaselineEngine.DataModel [C:5] [TYPE ADR] [SEMANTICS data-model,superset,baseline,filter,comparison] @BRIEF Canonical query, filter, value, baseline, fingerprint, and comparison entities for feature 037. @RELATION DEPENDS_ON -> [SupersetBaselineEngine.Research] @RATIONALE Data model is designed for deterministic comparison and fingerprinting — all IDs are sorted, timestamps are canonical ISO-8601, filter contexts are normalized with locale-independent decimal representation. Decimal/string canonicalization avoids float equality pitfalls. DashboardQueryModel fingerprint excludes itself to avoid self-referential hashing. @REJECTED Binary float for numeric canonical values was rejected — cross-environment float representation differences produce false-positive comparison failures. Raw Superset JSON schema was rejected — not deterministic across Superset versions.

DashboardQueryModel

  • schema_version, environment_id, dashboard_id/title/slug;
  • charts: ChartQueryModel[];
  • datasets: DatasetQueryModel[];
  • native_filters: NativeFilterModel[];
  • capabilities: chart_data, dataset_query, xlsx_export;
  • warnings: structured source/resource/code/detail;
  • query_model_fingerprint.

ChartQueryModel includes chart id/uuid/title/viz_type, dataset id, metric/result descriptors, group-bys, saved query inputs, applicable_filter_ids, excluded_filter_ids, and execution capability. DatasetQueryModel includes id/uuid/name, columns with semantic types, metrics, and access state.

NormalizedFilterContext

{
  "schema_version": 1,
  "filters": [
    {
      "filter_id": "NATIVE_FILTER-date",
      "dataset_id": 77,
      "column": "business_date",
      "operator": "TEMPORAL_RANGE",
      "value": {"from": "2026-05-29", "to": "2026-05-29", "inclusive": true},
      "target_chart_ids": [128]
    }
  ],
  "filters_hash": "sha256"
}

Canonical order is dataset_id, column, operator, canonical value, filter_id. Duplicate semantic filters are rejected unless their scopes are disjoint.

SupersetQueryRequest

Contains environment_id, dashboard_id, chart_id or dataset_id, result_key, and NormalizedFilterContext. It cannot contain SQL, raw endpoint, raw query_context, or arbitrary expression fields.

NormalizedValue

Field Meaning
kind null, boolean, integer, decimal, string, date, datetime, percent, table
raw Redacted bounded source value
canonical JSON-safe canonical value; decimal is string
display Optional Superset display value
format Optional format metadata
source environment/dashboard/chart/dataset/result/query hash
warnings normalization warnings

Table canonical value contains ordered columns and rows. Each cell is a scalar NormalizedValue payload without recursive source duplication.

BaselineEntry

Release-pinned expected value. Stored in dashboard_tests/{dashboard_key}/baselines.yaml in the git repository alongside the dashboard.

Required fields:

Field Rule
schema_version const: 1
baseline_id UUID, stable across releases
release_version Semver string, required (e.g. "v1.2.0")
release_commit_hash 40-char git SHA, required
dashboard_id Superset dashboard integer id
chart_id or dataset_id, one required
result_key Metric/result identifier
label Human-readable display name
normalized_filters NormalizedFilterContext with filters_hash
expected NormalizedValue — the approved expected value
source_response_hash SHA-256 of the Superset API response body at the time expected was captured
captured_at ISO-8601 timestamp of capture
comparison_policy ComparisonPolicy
status approved, superseded, retired
provenance Capture metadata: environment, actor, agent_run_id
immutability ImmutabilityBlock or null
created_at, updated_at Timestamps

ImmutabilityBlock

Field Type Rule
enabled boolean True for closed-period entries
period string Period identifier, e.g. "2026-05"
frozen_at ISO-8601 When the period was closed
policy enum alert, block_publish, require_investigation

When immutability.enabled=true: any change in source_response_hash at verification time MUST raise immutability_violation, not stale_baseline. Automated expected-value updates are forbidden. New values require explicit analyst override with reason.

Baseline inheritance rule

When creating release v1.(N+1).0 from v1.N.0:

  • For each metric entry in v1.N.0: if the chart's content_hash in DashboardQueryModel has NOT changed → inherit expected, source_response_hash, captured_at.
  • If chart content_hash HAS changed → mark entry as needs_reextraction; analyst or system re-extracts from PREPROD.
  • New charts (not in v1.N.0) → create new entries with fresh extraction from PREPROD.
  • Removed charts → entries retained with status retired.

ComparisonPolicy

Discriminated union:

  • exact;
  • absolute_tolerance(amount decimal string);
  • relative_tolerance(ratio decimal string, zero_absolute_fallback optional);
  • range(min/max and inclusive flags);
  • row_set(keys, order_sensitive, allow_extra_rows, per_column policies).

Policy/value type compatibility is validated when loading catalog and before comparison.

ComparisonResult

Contains status, actual, expected, policy, diff, stale_dimensions, warnings, source_error, baseline_id, release_version, and evidence refs.

Status enum: pass, fail, inconclusive, missing_baseline, stale_baseline, stale_visual_baseline, immutability_violation, permission_denied, source_error.

A non-pass result is deterministic evidence, not an agent conclusion. It emits an idempotent 036 InvestigationSignal with comparison, release, filter, target and principal provenance; 047 creates/updates the Queue item. The agent may inspect and explain the evidence, create a baseline candidate, or run diagnostic comparisons; normalization, comparison and immutable baseline state remain engine-owned.

immutability_violation is CRITICAL severity — it means source_response_hash changed for a closed-period entry, indicating retroactive data modification. It blocks release publication and triggers investigation.

VisualBaseline

Same release-pinned structure as metric BaselineEntry with kind=visual. Fields: baseline_id, release_version, release_commit_hash, dashboard_id, normalized_filters, tab_identifier, region_of_interest (optional), expected_image_sha256, tolerance_policy (VisualComparisonPolicy), source_response_hash, captured_at, immutability (optional), status, provenance.

VisualComparisonPolicy

Discriminated union: exact (digest match) or perceptual (ssim_min, pixel_diff_threshold). Visual baselines cannot use metric policies and vice versa; cross-kind comparison is rejected at catalog load.

StructureDiff

Structural comparison of DashboardQueryModel between two releases. Computed deterministically without metric execution.

Field Meaning
release_from Version string of baseline release
release_to Version string of target release
query_model_hash_from SHA-256 of source DashboardQueryModel
query_model_hash_to SHA-256 of target DashboardQueryModel
changes Array of StructureChange

StructureChange

Field Meaning
target JSON pointer in DashboardQueryModel (e.g. charts[128].columns, filters[contractor].scope)
kind Classifier: filter_scope_narrowed, filter_scope_widened, filter_operator_changed, filter_default_changed, filter_removed, column_order_changed, column_added, column_removed, chart_added, chart_removed, viz_type_changed, group_by_changed, time_grain_changed
severity critical, warning, info
before, after Values before and after the change
affected_artifacts Array of affected artifact kinds: xlsx_export, screenshot_evidence, metric_assertion
rationale Human-readable explanation

Severity classification:

  • critical: filter scope narrowed/lost, filter operator changed, chart removed, filter removed — affects data correctness
  • warning: column reorder, group_by changed, viz_type changed, time_grain changed — affects downstream artifacts
  • info: chart added, column added, filter scope widened — additive changes

StructureDiff is used at: DEV pre-deploy check, PREPROD post-deploy verification, release approval gate.

VerificationRun

Links an AgentRun (036) to a DashboardRelease. Records what was verified and the outcome.

Field Meaning
id UUID
release_id FK to DashboardRelease, nullable (nil for pre-release PREPROD checks)
repository_id FK to GitRepository
agent_run_id FK to AgentRun (036)
trigger manual, deploy_to_preprod, release_create, release_approve, release_publish, post_publish, scheduled, etl_completed
environment_id Target Superset environment
categories_run Array: metric, visual, structure, xlsx, content_integrity
categories_passed Subset of categories_run that passed
categories_failed Subset that failed
overall_status pass, warn, fail, blocked
summary Human-readable summary text
baseline_version Release version of the baseline used
baseline_commit Git commit of baseline.yaml used
structure_diff StructureDiff or null
metric_results Array of ComparisonResult
created_at, created_by Timestamps

Catalog File

Path: dashboard_tests/{dashboard_key}/baselines.yaml

schema_version: 1
dashboard:
  id: 42
  slug: fi-0080
entries: []

Entries sort by chart/dataset identity, result_key, filters_hash, baseline_id. Writer uses atomic temp-file replace inside the resolved repository and only through a consumed 036 gate.

#endregion SupersetBaselineEngine.DataModel


CONTRACTS — Module & Function Contracts

Source: contracts/modules.md

#region SupersetBaselineEngine.Modules [C:5] [TYPE ADR] [SEMANTICS contracts,baseline,superset,chart-data] @BRIEF C3+ contracts for query inspection, filter mapping, Superset-native execution, normalization, comparison, and baseline lifecycle. @RELATION DEPENDS_ON -> [SupersetBaselineEngine.DataModel] @RELATION DEPENDS_ON -> [Services.AgentRuns.Approvals.Consume] @RATIONALE The engine is a backend domain; agent tools remain thin authenticated clients. @REJECTED Direct SQL or agent-supplied raw query context — rejected because saved dashboard semantics must remain authoritative.

#region BaselineEngine.Api [C:4] [TYPE Module] [SEMANTICS baseline,api,rbac]

@defgroup BaselineEngine REST routes for inspection, execution, comparison, candidates, and approval requests.

@LAYER API

@RELATION DEPENDS_ON -> [BaselineEngine.QueryModel.Inspect]

@RELATION DEPENDS_ON -> [BaselineEngine.QueryExecutor.Execute]

@RELATION DEPENDS_ON -> [BaselineEngine.Comparison.Compare]

@RELATION DEPENDS_ON -> [BaselineEngine.Catalog.Load]

@INVARIANT No request schema exposes sql, raw endpoint, or raw query_context.

#endregion BaselineEngine.Api

#region BaselineEngine.QueryModel.Inspect [C:5] [TYPE Function] [SEMANTICS baseline,inspection,dashboard,metadata]

@ingroup BaselineEngine

@BRIEF Build a deterministic query model from authoritative Superset dashboard, chart, and dataset metadata.

@PRE Environment and dashboard are readable by actor.

@POST Returns stable sorted model, per-resource warnings, capabilities, and fingerprint; missing metadata is never invented.

@SIDE_EFFECT Async GET calls through SupersetClient.

@DATA_CONTRACT InspectRequest -> DashboardQueryModel

@RATIONALE The query model is built from authoritative Superset metadata so scenario and baseline decisions share the same chart/dataset/filter semantics as the dashboard.

@REJECTED Inferring missing chart or dataset metadata from names or agent prose — rejected because fabricated metadata can produce false executable tests.

@RELATION DEPENDS_ON -> [Spec.TranslateRequestsHttpx.SupersetClient]

@TEST_EDGE inaccessible_chart -> warning plus executable=false.

@TEST_EDGE scoped_filter -> only target charts list filter id.

@TEST_EDGE metadata_order_changes -> byte-stable canonical model.

#endregion BaselineEngine.QueryModel.Inspect

#region BaselineEngine.Filters.Normalize [C:5] [TYPE Function] [SEMANTICS baseline,filter,canonical,scope]

@ingroup BaselineEngine

@BRIEF Validate dashboard filter values against metadata/scope and produce canonical typed filter identity.

@PRE Query model is authoritative and fingerprint-valid.

@POST Filters are typed, sorted, scoped, and hashed; invalid targets/operators produce structured validation errors.

@SIDE_EFFECT None.

@DATA_CONTRACT FilterInput[] + DashboardQueryModel -> NormalizedFilterContext

@RATIONALE A canonical filter context gives UI, Superset execution, XLSX comparison, and baseline lookup one stable identity.

@REJECTED Passing display-formatted filter values directly to execution — rejected because locale and scope differences would create non-reproducible queries.

@INVARIANT Locale formatting never enters filters_hash.

@TEST_EDGE locale_decimal -> canonical decimal string.

@TEST_EDGE filter_outside_chart_scope -> 422.

@TEST_EDGE missing_target_value -> recoverable FILTER_VALUE_UNAVAILABLE.

#endregion BaselineEngine.Filters.Normalize

#region SupersetClient.ChartData.Execute [C:5] [TYPE Function] [SEMANTICS baseline,superset,chart-data,async]

@ingroup BaselineEngine

@BRIEF Execute backend-built chart-data query context through Superset POST /api/v1/chart/data.

@PRE Saved chart/dataset metadata and normalized filters validated; result limits configured.

@POST Returns version-adapted raw result with source metadata or typed upstream error.

@SIDE_EFFECT Async Superset REST call.

@DATA_CONTRACT AuthoritativeChartMetadata + NormalizedFilterContext -> SupersetRawResult

@INVARIANT Caller cannot inject SQL, endpoint, datasource, adhoc expression, or unscoped filter.

@TEST_INVARIANT No_Direct_SQL -> VERIFIED_BY: request_schema, malicious_extra_fields, adapter_payload.

@TEST_EDGE raw_sql_field -> Pydantic extra-forbid 422 before call.

@TEST_EDGE superset_403_404_422_timeout_5xx -> error taxonomy preserved.

@RATIONALE A dedicated mixin centralizes Superset 4.1.2 chart-data differences and existing TLS/auth behavior.

@REJECTED Using existing dataset preview builder for chart truth — it substitutes count/default columns and is not chart-fidelity.

@RELATION DEPENDS_ON -> [SupersetBaselineEngine.DataModel]

@RELATION DEPENDS_ON -> [BaselineEngine.Filters.Normalize]

#endregion SupersetClient.ChartData.Execute

#region BaselineEngine.Result.Normalize [C:5] [TYPE Function] [SEMANTICS baseline,result,normalization,decimal]

@ingroup BaselineEngine

@BRIEF Convert scalar, big-number, temporal, percent, and table results into canonical typed values.

@PRE Raw result was returned by the chart-data adapter with bounded row count.

@POST Equivalent locale/display variants normalize identically; ambiguity returns inconclusive reason.

@SIDE_EFFECT None.

@DATA_CONTRACT SupersetRawResult + ResultDescriptor -> NormalizedValue

@RATIONALE Normalization isolates presentation differences before comparison while retaining source metadata for provenance and investigation.

@REJECTED Comparing rendered strings or binary floats — rejected because locale, timezone, and rounding differences create false diffs.

@INVARIANT Numeric canonicalization uses Decimal/string, never binary float equality.

@TEST_EDGE localized_number -> canonical decimal.

@TEST_EDGE timezone_datetime -> UTC ISO-8601.

@TEST_EDGE unsupported_nested_value -> inconclusive, not pass.

#endregion BaselineEngine.Result.Normalize

#region BaselineEngine.Comparison.Compare [C:5] [TYPE Function] [SEMANTICS baseline,comparison,tolerance,diff]

@ingroup BaselineEngine

@BRIEF Apply a validated comparison policy to actual and approved expected canonical values.

@PRE Value kinds and policy are compatible; baseline is approved and fingerprint checked.

@POST Returns pass/fail/inconclusive/stale with deterministic diff and no mutation.

@SIDE_EFFECT None.

@DATA_CONTRACT NormalizedValue + BaselineEntry -> ComparisonResult

@RATIONALE Comparison is read-only so a verification run can report truth without silently changing the approved expectation.

@REJECTED Auto-updating a stale or missing baseline during comparison — rejected because it converts a regression signal into an unreviewed truth change.

@INVARIANT Missing/stale/unsupported baselines cannot return pass.

@TEST_EDGE relative_expected_zero -> explicit absolute fallback or inconclusive.

@TEST_EDGE table_duplicate_keys -> inconclusive unless multiset policy.

@TEST_EDGE stale_query_fingerprint -> stale_baseline.

#endregion BaselineEngine.Comparison.Compare

#region BaselineEngine.Catalog.Load_MOVED [C:1] [TYPE Tombstone] [SEMANTICS baseline,moved]

@DEPRECATED Lifecycle contract moved to contracts/lifecycle.md.

@STATUS DEPRECATED -> REPLACED_BY: [BaselineEngine.Catalog.Load]

@REPLACED_BY BaselineEngine.Catalog.Load

@RELATION DEPENDS_ON -> [BaselineEngine.Catalog.Load]

#endregion BaselineEngine.Catalog.Load_MOVED

#region BaselineEngine.Catalog.Load [C:4] [TYPE Function] [SEMANTICS baseline,catalog,yaml,validation]

@ingroup BaselineEngine

@BRIEF Load and validate reviewable baseline YAML from a resolved repository.

@PRE GitService resolves repository and safe dashboard key; read permission passes.

@POST Returns sorted valid entries plus catalog warnings; unsafe YAML constructs are rejected.

@SIDE_EFFECT Bounded file read through file executor.

@DATA_CONTRACT Repository + DashboardKey -> BaselineCatalog

@TEST_EDGE invalid_schema -> CATALOG_INVALID.

@TEST_EDGE duplicate_baseline_identity -> reject catalog.

@TEST_EDGE symlink_escape -> reject before read.

#endregion BaselineEngine.Catalog.Load

#region BaselineEngine.Candidate.Create [C:4] [TYPE Function] [SEMANTICS baseline,release,extraction,inheritance]

@ingroup BaselineEngine

@BRIEF Extract baseline entries for a new release from PREPROD, with inheritance from the previous release.

@PRE Previous release baseline.yaml exists; PREPROD deployment is active; release_version is valid.

@POST Baseline.yaml is updated with inherited and newly extracted entries; StructureDiff is computed.

@SIDE_EFFECT Superset API queries against PREPROD; baseline.yaml written to git via atomic temp-file replace.

@DATA_CONTRACT ReleaseVersion + RepositoryKey → baseline.yaml

@INVARIANT Entries for unchanged charts inherit expected + source_response_hash; changed charts get re-extracted values.

@RELATION DEPENDS_ON -> [BaselineEngine.Catalog.Load]

@RELATION DEPENDS_ON -> [BaselineEngine.QueryModel.Inspect]

@RELATION DEPENDS_ON -> [BaselineEngine.Comparison.Compare]

@TEST_EDGE chart_content_hash_unchanged -> inherited entry identical to previous release.

@TEST_EDGE chart_content_hash_changed -> re-extracted; old entry marked retired if chart removed.

@TEST_EDGE no_previous_release -> all entries extracted fresh.

@RATIONALE Inheritance prevents unnecessary manual re-entry and makes release-to-release diffs meaningful.

@REJECTED Requiring manual baseline re-entry for every release — error-prone and hides structural regressions.

#endregion BaselineEngine.Candidate.Create

#region BaselineEngine.Release.ApproveBaseline [C:4] [TYPE Function] [SEMANTICS baseline,release,approval,repository]

@ingroup BaselineEngine

@BRIEF Approve baseline.yaml for a release through 036 and atomically commit to git.

@PRE Release exists; baseline.yaml validated; StructureDiff reviewed; approval gate confirmed.

@POST baseline.yaml committed to release branch; release can proceed to publish.

@SIDE_EFFECT Git commit + push of baseline.yaml; approval audit record.

@RELATION DEPENDS_ON -> [Services.AgentRuns.Approvals.Consume]

@INVARIANT Baseline cannot be approved for a release that has structural CRITICAL changes unattended.

@TEST_EDGE critical_structure_diff_unresolved -> approval blocked.

@REJECTED Approving baseline without release pinning — baseline without release context is unreproducible.

#endregion BaselineEngine.Release.ApproveBaseline

#region BaselineEngine.Structure.Diff [C:4] [TYPE Function] [SEMANTICS baseline,structure,diff,release]

@ingroup BaselineEngine

@BRIEF Deterministic structural comparison of DashboardQueryModel between two releases.

@PRE Both releases have valid query model snapshots.

@POST Returns StructureDiff with classified changes; no metric execution.

@SIDE_EFFECT None.

@DATA_CONTRACT ReleaseVersion + ReleaseVersion -> StructureDiff

@INVARIANT Diff is deterministic for the same two releases.

@TEST_EDGE filter_scope_narrowed -> critical with rationale and affected_charts.

@TEST_EDGE column_reorder -> warning with affected_artifacts xlsx_export and screenshot_evidence.

@TEST_EDGE identical_releases -> zero changes, pass count equals total elements.

@RATIONALE Structural diffs catch regressions that metric comparison alone cannot: scope loss, column reordering, chart removal.

@REJECTED Using git diff of YAML files — DashboardQueryModel provides semantic classification that raw YAML diff cannot.

@RELATION DEPENDS_ON -> [BaselineEngine.QueryModel.Inspect]

@RELATION DEPENDS_ON -> [BaselineEngine.StructureDiff.SnapshotLoader]

#endregion BaselineEngine.Structure.Diff

#region BaselineEngine.Immutability.Detect [C:4] [TYPE Function] [SEMANTICS baseline,immutability,violation,closed-period]

@ingroup BaselineEngine

@BRIEF Detect retroactive data changes for closed-period baseline entries.

@PRE Baseline entry has immutability.enabled=true and source_response_hash recorded.

@POST Compares current source_response_hash against baseline; mismatch → immutability_violation.

@SIDE_EFFECT None (read-only detection).

@DATA_CONTRACT BaselineEntry + SupersetAPIResponse -> ComparisonResult

@INVARIANT immutability_violation is CRITICAL, never reduced to stale_baseline.

@TEST_EDGE source_response_hash_match -> pass (not immutability_violation).

@TEST_EDGE source_response_hash_mismatch_closed_period -> immutability_violation, CRITICAL.

@TEST_EDGE source_response_hash_mismatch_open_period -> pass or fail depending on values, not violation.

@RATIONALE Closed-period data should never change; when it does, it's a data integrity incident, not a metric drift.

@REJECTED Treating all value changes as stale_baseline — it normalizes retroactive data modification.

#endregion BaselineEngine.Immutability.Detect

#region AgentChat.Tools.DashboardTesting [C:4] [TYPE Module] [SEMANTICS baseline,agent,tools,no-sql]

@defgroup BaselineEngine Thin agent tools for inspect, execute, compare, discover, and request approval.

@RELATION DEPENDS_ON -> [BaselineEngine.Api]

@INVARIANT Tools never accept SQL or local repository paths and are included only after RBAC/intent filtering.

@REJECTED Duplicating normalization or catalog logic in the standalone agent.

#endregion AgentChat.Tools.DashboardTesting

#region BaselineEngine.Visual.Compare [C:4] [TYPE Function] [SEMANTICS baseline,visual,screenshot,comparison]

@ingroup BaselineEngine

@BRIEF Compare a captured screenshot digest against an approved visual baseline with perceptual tolerance.

@PRE Visual baseline is approved; layout fingerprint is current; capture artifact has valid sha256.

@POST Returns pass/stale_visual_baseline/fail/inconclusive with digest diff and stale dimensions.

@SIDE_EFFECT None.

@DATA_CONTRACT ScreenshotEvidence + VisualBaseline -> ComparisonResult

@INVARIANT Visual comparison never uses metric policies; cross-kind comparison returns inconclusive with code.

@RELATION DEPENDS_ON -> [BaselineEngine.Visual.SSIM]

@RELATION DEPENDS_ON -> [BaselineEngine.Verification.ExecutorVisual.Async]

@TEST_EDGE exact_match -> pass with matching digests.

@TEST_EDGE perceptual_within_tolerance -> pass with ssim/metadata.

@TEST_EDGE layout_fingerprint_changed -> stale_visual_baseline.

@TEST_EDGE metric_baseline_with_visual_policy -> catalog load rejects.

@RATIONALE Visual baseline is a distinct assertion dimension; conflating it with metric baselines would silently change comparison semantics.

@RATIONALE Perceptual comparison and authoritative-fingerprint execution are reused from BaselineEngine.Visual.SSIM and BaselineEngine.Verification.ExecutorVisual.Async — never reimplemented in the scenario path.

@REJECTED Using LLM/VLM judgment as visual baseline truth — it is non-deterministic and cannot serve as stable expected state.

#endregion BaselineEngine.Visual.Compare

#region BaselineEngine.Visual.Candidate [C:4] [TYPE Function] [SEMANTICS baseline,visual,candidate,screenshot]

@ingroup BaselineEngine

@BRIEF Create a draft visual baseline candidate from a captured and reviewed screenshot.

@PRE Screenshot artifact exists and has been reviewed (human disposition: confirmed); filter context matches dashboard.

@POST Candidate registered as DraftArtifact with kind=baseline_candidate and visual baseline proposed_entry.

@SIDE_EFFECT Calls AgentRuns.Artifacts.Register.

@DATA_CONTRACT ScreenshotEvidence + ReviewDisposition -> BaselineCandidate

@INVARIANT Visual candidate approval follows the same 036 gate path as metric candidates.

@TEST_EDGE unreviewed_screenshot -> 422; human disposition required before visual candidate creation.

@RATIONALE A visual baseline represents an approved expected state, not an automated snapshot; human review gates its creation.

@REJECTED Auto-creating visual candidates from every passing screenshot — it erodes the human-review boundary.

@RELATION DEPENDS_ON -> [BaselineEngine.Visual.Compare]

@RELATION DEPENDS_ON -> [BaselineEngine.Candidates.Create]

#endregion BaselineEngine.Visual.Candidate

#endregion SupersetBaselineEngine.Modules


CONTRACTS — Remaining — dashboard-testing.openapi.yaml

Source: contracts/dashboard-testing.openapi.yaml

openapi: 3.1.0 info: title: Dashboard Testing Baseline API version: 0.1.0 description: Superset-native query inspection, execution, normalization, comparison, and candidate lifecycle. servers:

  • url: /api paths: /dashboard-testing/query-model: get: operationId: inspectDashboardQueryModel security: [{ bearerAuth: [] }] parameters: - { name: environment_id, in: query, required: true, schema: { type: string } } - { name: dashboard_id, in: query, required: true, schema: { type: integer } } responses: '200': description: Deterministic dashboard query model content: { application/json: { schema: { $ref: '#/components/schemas/DashboardQueryModel' } } } '403': { description: Permission denied } '404': { description: Dashboard not found } /dashboard-testing/filters/normalize: post: operationId: normalizeDashboardFilters security: [{ bearerAuth: [] }] requestBody: required: true content: { application/json: { schema: { $ref: '#/components/schemas/NormalizeFiltersRequest' } } } responses: '200': description: Canonical filter identity content: { application/json: { schema: { $ref: '#/components/schemas/NormalizedFilterContext' } } } '422': { description: Invalid target, operator, type, scope, or value } /dashboard-testing/queries/execute: post: operationId: executeDashboardQuery security: [{ bearerAuth: [] }] requestBody: required: true content: { application/json: { schema: { $ref: '#/components/schemas/ExecuteQueryRequest' } } } responses: '200': description: Canonical Superset result content: { application/json: { schema: { $ref: '#/components/schemas/NormalizedValue' } } } '403': { description: Permission denied by app or Superset } '422': { description: Request/filter/query validation failed } '504': { description: Superset timeout } /dashboard-testing/comparisons: post: operationId: compareDashboardResult security: [{ bearerAuth: [] }] requestBody: required: true content: { application/json: { schema: { $ref: '#/components/schemas/ComparisonRequest' } } } responses: '200': description: Pass/fail/inconclusive/missing/stale result content: { application/json: { schema: { $ref: '#/components/schemas/ComparisonResult' } } } /dashboard-testing/baselines: get: operationId: listDashboardBaselines security: [{ bearerAuth: [] }] parameters: - { name: repository_key, in: query, required: true, schema: { type: string } } - { name: dashboard_key, in: query, required: true, schema: { type: string } } responses: '200': description: Validated reviewable baseline catalog content: application/json: schema: type: object required: [schema_version, entries] properties: schema_version: { type: integer } entries: { type: array, items: { $ref: '#/components/schemas/BaselineEntry' } } warnings: { type: array, items: { $ref: '#/components/schemas/Warning' } } /dashboard-testing/baseline-candidates: post: operationId: createBaselineCandidate security: [{ bearerAuth: [] }] requestBody: required: true content: { application/json: { schema: { $ref: '#/components/schemas/CandidateRequest' } } } responses: '201': description: Draft candidate registered through feature 036 content: { application/json: { schema: { $ref: '#/components/schemas/BaselineCandidate' } } } '422': { description: Invalid or insufficient provenance } /dashboard-testing/structure-snapshot/capture: post: operationId: captureReleaseSnapshot security: [{ bearerAuth: [] }] description: Capture a release-bound dashboard query model snapshot with full identity validation. parameters: - { name: environment_id, in: query, required: true, schema: { type: string } } requestBody: required: true content: application/json: schema: type: object required: [release_id] properties: release_id: { type: string, format: uuid } responses: '201': description: Release-bound snapshot captured and persisted content: { application/json: { schema: { $ref: '#/components/schemas/SnapshotCaptureResponse' } } } '422': description: Release not found, env mismatch, semver invalid, commit mismatch, or blocking warnings content: { application/json: { schema: { type: object, properties: { detail: { type: string } } } } } /dashboard-testing/structure-snapshot/diff: post: operationId: diffReleaseSnapshots security: [{ bearerAuth: [] }] description: Compute a structure diff between two release-bound dashboard snapshots with metadata verification. requestBody: required: true content: application/json: schema: type: object required: [release_id_from, release_id_to] properties: release_id_from: { type: string, format: uuid } release_id_to: { type: string, format: uuid } responses: '200': description: Classified structural diff with metadata cross-verification content: { application/json: { schema: { $ref: '#/components/schemas/StructureDiff' } } } '422': { description: Release not found, snapshot missing, metadata verification failure } /dashboard-testing/baseline-candidates/capture: post: operationId: captureBaselineCandidate security: [{ bearerAuth: [] }] description: Authoritative capture — resolves release, executes Superset query, computes source_response_hash server-side, creates capture artifact and baseline candidate. requestBody: required: true content: application/json: schema: $ref: '#/components/schemas/CaptureCandidateRequest' responses: '201': description: Candidate created from authoritative server-side capture content: { application/json: { schema: { $ref: '#/components/schemas/CaptureCandidateResponse' } } } '422': description: Invalid release, agent-run mismatch, or insufficient provenance content: { application/json: { schema: { type: object, properties: { detail: { type: string } } } } } '500': description: Server error during capture execution /dashboard-testing/baseline-candidates/{candidateId}/approval-gate: post: operationId: requestBaselineApproval security: [{ bearerAuth: [] }] parameters: - { name: candidateId, in: path, required: true, schema: { type: string, format: uuid } } requestBody: required: true content: application/json: schema: $ref: '#/components/schemas/ApprovalGateRequest' responses: '201': description: Bound 036 approval gate content: { application/json: { schema: { $ref: '#/components/schemas/ApprovalGateResponse' } } } '403': { description: Permission denied; no gate created } /dashboard-testing/baseline-candidates/{candidateId}/approval-gate/{gateId}/decide: post: operationId: decideBaselineApproval security: [{ bearerAuth: [] }] parameters: - { name: candidateId, in: path, required: true, schema: { type: string, format: uuid } } - { name: gateId, in: path, required: true, schema: { type: string, format: uuid } } requestBody: required: true content: application/json: schema: type: object required: [decision] properties: decision: { type: string, enum: [confirm, deny] } reason: { type: string, minLength: 1 } responses: '200': description: Gate decision recorded and applied content: { application/json: { schema: { $ref: '#/components/schemas/ApprovalDecisionResponse' } } } '409': { description: Gate already decided or conflicting state } /dashboard-testing/baseline-candidates/{candidateId}/approval-gate/{gateId}/consume: post: operationId: consumeBaselineApproval security: [{ bearerAuth: [] }] parameters: - { name: candidateId, in: path, required: true, schema: { type: string, format: uuid } } - { name: gateId, in: path, required: true, schema: { type: string, format: uuid } } - { name: release_version, in: query, required: true, schema: { type: string, pattern: '^v\d+.\d+.\d+(-[a-z0-9.]+)?$' } } - { name: release_commit_hash, in: query, required: true, schema: { type: string, pattern: '^[a-f0-9]{40}$', minLength: 40, maxLength: 40 } } responses: '200': description: Gate consumed and baseline materialized content: { application/json: { schema: { $ref: '#/components/schemas/ApprovalConsumeResponse' } } } '409': { description: Version guard or hash mismatch prevents consumption } /dashboard-testing/structure-diff: post: operationId: computeStructureDiff security: [{ bearerAuth: [] }] description: Structural comparison of DashboardQueryModel between two releases. requestBody: required: true content: application/json: schema: type: object required: [environment_id, dashboard_id, release_version_from, release_version_to] properties: environment_id: { type: string } dashboard_id: { type: integer } release_version_from: { type: string } release_version_to: { type: string } responses: '200': description: Deterministic structural diff between two releases content: { application/json: { schema: { $ref: '#/components/schemas/StructureDiff' } } } '404': { description: Release or dashboard not found } /dashboard-testing/verification-runs: post: operationId: createVerificationRun security: [{ bearerAuth: [] }] requestBody: required: true content: application/json: schema: type: object required: [repository_id, trigger, environment_id, categories] properties: repository_id: { type: string, format: uuid } release_id: { type: [string, 'null'], format: uuid } trigger: { type: string, enum: [manual, deploy_to_preprod, release_create, release_approve, release_publish, post_publish, scheduled, etl_completed] } environment_id: { type: string } categories: { type: array, items: { type: string, enum: [metric, visual, structure, xlsx, content_integrity] } } baseline_version: { type: [string, 'null'] } agent_run_id: { type: [string, 'null'], format: uuid } evidence_refs: type: [object, 'null'] additionalProperties: { type: array, items: { type: string } } category_params: { type: [object, 'null'], additionalProperties: true } responses: '201': description: Verification run created and linked to AgentRun content: { application/json: { schema: { $ref: '#/components/schemas/VerificationRun' } } } '422': { description: Invalid trigger/environment/category combination } /dashboard-testing/inheritance/plan: post: operationId: planInheritance security: [{ bearerAuth: [] }] requestBody: required: true content: application/json: schema: type: object required: [prior_release_id, current_release_id] properties: prior_release_id: { type: string, format: uuid } current_release_id: { type: string, format: uuid } responses: '200': description: Inheritance plan computed '404': { description: Prior or current release not found } '422': { description: Same release IDs supplied } /dashboard-testing/inheritance/execute: post: operationId: executeInheritance security: [{ bearerAuth: [] }] requestBody: required: true content: application/json: schema: type: object required: [plan_id, target_environment_id] properties: plan_id: { type: string } target_environment_id: { type: string } responses: '200': description: Inheritance execution completed '400': { description: Invalid plan_id } '404': { description: Target environment not found } /dashboard-testing/scenarios: post: operationId: scenarioRegistry.create security: [{ bearerAuth: [] }] description: Register a server-validated compiled scenario and save-eligible draft pack (042). requestBody: required: true content: application/json: schema: type: object required: [compiled_handle_id, draft_pack_id, draft_pack_digest] properties: compiled_handle_id: { type: string } draft_pack_id: { type: string } draft_pack_digest: { type: string, pattern: '^[a-f0-9]{64}$' } responses: '201': { description: Scenario registered with candidate revision } '409': { description: Draft pack validation or ownership conflict } get: operationId: scenarioRegistry.list security: [{ bearerAuth: [] }] description: List/search/filter persisted scenario registry entries (042). parameters: - { name: q, in: query, schema: { type: string } } - { name: dashboard_id, in: query, schema: { type: integer } } - { name: status, in: query, schema: { type: string } } - { name: tag, in: query, schema: { type: string } } - { name: owner, in: query, schema: { type: string } } - { name: page, in: query, schema: { type: integer, default: 1 } } - { name: page_size, in: query, schema: { type: integer, default: 25 } } responses: '200': description: Paged registry entries '404': { description: No scenarios found } /dashboard-testing/scenarios/{scenarioId}: get: operationId: scenarioRegistry.detail security: [{ bearerAuth: [] }] description: Load a persisted scenario detail by id from the registry (042). parameters: - { name: scenarioId, in: path, required: true, schema: { type: string } } responses: '200': { description: Scenario detail } '404': { description: Scenario not found } /dashboard-testing/scenarios/{scenarioId}/revisions: get: operationId: scenarioRegistry.revisions security: [{ bearerAuth: [] }] description: List immutable candidate/current revisions (042). responses: '200': { description: Revision chain } post: operationId: scenarioRegistry.createRevision security: [{ bearerAuth: [] }] description: Append a validated immutable candidate revision (042). responses: '201': { description: Candidate revision created } '409': { description: Stale or foreign base revision } /dashboard-testing/scenarios/{scenarioId}/revisions/{revisionId}: get: operationId: scenarioRegistry.checkoutRevision security: [{ bearerAuth: [] }] description: Load an immutable revision snapshot for checkout (042). responses: '200': { description: Revision snapshot } '404': { description: Revision not found } /dashboard-testing/scenarios/{scenarioId}/revisions/{revA}/diff/{revB}: get: operationId: scenarioRegistry.revisionDiff security: [{ bearerAuth: [] }] description: Compute deterministic changes between two revisions (042). responses: '200': { description: Added/changed/removed change set } '404': { description: Revision not found } /dashboard-testing/scenarios/{scenarioId}/transition: post: operationId: scenarioRegistry.transition security: [{ bearerAuth: [] }] description: Apply an audited lifecycle transition (042). responses: '200': { description: Transition applied } '409': { description: Invalid transition or stale metadata version } /dashboard-testing/scenarios/{scenarioId}/clone: post: operationId: scenarioRegistry.clone security: [{ bearerAuth: [] }] description: Clone a scenario into a new owned draft (042). responses: '201': { description: Clone created } '404': { description: Source scenario not found } /dashboard-testing/scenarios/{scenarioId}/archive: post: operationId: scenarioRegistry.archive security: [{ bearerAuth: [] }] description: Archive a scenario while preserving history (042). responses: '200': { description: Scenario archived } '409': { description: Archive conflict } /dashboard-testing/scenarios/{scenarioId}/restore: post: operationId: scenarioRegistry.restore security: [{ bearerAuth: [] }] description: Restore an archived scenario to DRAFT (042). responses: '200': { description: Scenario restored } '409': { description: Restore conflict } /dashboard-testing/scenarios/{scenarioId}/health: get: operationId: scenarioRegistry.health security: [{ bearerAuth: [] }] description: Read analytics-owned scenario health projection (042/047). responses: '200': { description: Scenario health projection } '404': { description: Scenario not found } /dashboard-testing/scenarios/{scenarioId}/edit: get: operationId: scenarioEditor.load security: [{ bearerAuth: [] }] description: Load a clean revision-bound scenario editor projection (043). parameters: - { name: revision_id, in: query, required: false, schema: { type: string } } responses: '200': { description: Read-only editor projection with dirty=false and base revision } '404': { description: Scenario or revision not found } /dashboard-testing/scenarios/{scenarioId}/edits/apply: post: operationId: scenarioEditor.apply security: [{ bearerAuth: [] }] description: Persist typed operations as a server-owned WorkingDraft (043). responses: '200': { description: WorkingDraft id, digest and validation } '409': { description: Stale base revision } '422': { description: Unsafe or invalid operation } /dashboard-testing/scenarios/{scenarioId}/edits/save: post: operationId: scenarioEditor.save security: [{ bearerAuth: [] }] description: Save a server-owned WorkingDraft by id and digest (043). responses: '200': { description: Candidate immutable revision } '409': { description: Digest or stale base conflict } /dashboard-testing/scenarios/{scenarioId}/edits/agent-propose: post: operationId: scenarioEditor.agentPropose security: [{ bearerAuth: [] }] description: Store a validated agent edit proposal with a deterministic diff (043 US5). requestBody: required: true content: application/json: schema: type: object required: [base_revision_id, request_text, ops] properties: base_revision_id: { type: string, description: Pinned base revision the ops apply to } request_text: { type: string, minLength: 1, maxLength: 4000, description: Natural-language agent request describing the edit } ops: { type: array, items: { type: object }, description: Closed EditOperation union; unsafe/cyclic ops rejected server-side } agent_action_id: { type: [string, 'null'], description: Optional attribution to the originating agent action } responses: '200': description: Stored open proposal with digest and deterministic graph diff content: application/json: schema: type: object required: [proposal_id, base_revision_id, digest] properties: proposal_id: { type: string, description: Server-owned proposal id } base_revision_id: { type: string } digest: { type: string, pattern: '^[a-f0-9]{64}$', description: sha256 digest of the proposed graph } diff: { type: object, description: Deterministic graph diff against base revision } validation: { type: object, description: Server-side validation result } status: { type: string, enum: [open], default: open } '422': { description: Unsafe, unparseable, or cyclic operations (INVALID_EDIT) } /dashboard-testing/scenarios/{scenarioId}/edits/proposals/{proposalId}/save: post: operationId: scenarioEditor.proposalSave security: [{ bearerAuth: [] }] description: Save an accepted agent proposal through the guarded WorkingDraft path with attribution (043). requestBody: required: true content: application/json: schema: type: object required: [digest] properties: digest: { type: string, pattern: '^[a-f0-9]{64}$', minLength: 64, maxLength: 64, description: Digest must match the server-owned proposal } agent_action_id: { type: [string, 'null'], description: Optional attribution to the originating agent action } responses: '200': { description: Candidate immutable revision created from the proposal } '403': { description: Policy denied (POLICY_DENIED) } '409': { description: Stale base or digest mismatch (EDITOR_SAVE_CONFLICT) } /dashboard-testing/scenarios/{scenarioId}/revalidate: post: operationId: scenarioEditor.revalidate security: [{ bearerAuth: [] }] description: Produce a read-only migration proposal for a stale scenario (043). parameters: - { name: base_revision_id, in: query, required: true, schema: { type: string } } responses: '200': { description: Migration proposal with mappings/conflicts/diff } '409': { description: Scenario is not stale or base revision conflicts } /dashboard-testing/scenarios/compile: post: operationId: compileScenario security: [{ bearerAuth: [] }] description: Compile a draft scenario into an executable run plan. responses: '200': { description: Scenario compiled successfully } '422': { description: Invalid scenario definition } /dashboard-testing/scenarios/validate: post: operationId: validateScenario security: [{ bearerAuth: [] }] description: Validate a scenario definition before execution. responses: '200': { description: Scenario is valid } '422': { description: Invalid scenario definition } /dashboard-testing/scenarios/{scenarioId}/capture: post: operationId: captureScenario security: [{ bearerAuth: [] }] description: Capture evidence for a scenario run. responses: '201': { description: Evidence captured } '404': { description: Scenario run not found } /dashboard-testing/scenarios/{scenarioId}/disposition: post: operationId: setScenarioDisposition security: [{ bearerAuth: [] }] description: Set the disposition for a scenario run. responses: '200': { description: Disposition recorded } '404': { description: Scenario run not found } /dashboard-testing/scenarios/{scenarioId}/draft-pack: post: operationId: buildScenarioDraftPack security: [{ bearerAuth: [] }] description: Build a draft pack from a scenario run. responses: '200': { description: Draft pack built } '404': { description: Scenario run not found } /dashboard-testing/scenarios/{scenarioId}/resolve: post: operationId: resolveScenario security: [{ bearerAuth: [] }] description: Resolve a scenario run to a final outcome. responses: '200': { description: Scenario resolved } '404': { description: Scenario run not found } /dashboard-testing/scenarios/{scenarioId}/vlm: post: operationId: runScenarioVlm security: [{ bearerAuth: [] }] description: Run VLM analysis for a scenario run. responses: '200': { description: VLM analysis completed } '404': { description: Scenario run not found } /dashboard-testing/verification/history: get: operationId: getVerificationHistory security: [{ bearerAuth: [] }] description: List past verification runs. responses: '200': description: Verification run history content: { application/json: { schema: { type: array, items: { $ref: '#/components/schemas/VerificationRun' } } } } /dashboard-testing/verification/{runId}: get: operationId: getVerificationRun security: [{ bearerAuth: [] }] description: Fetch a single verification run. responses: '200': description: Verification run details content: { application/json: { schema: { $ref: '#/components/schemas/VerificationRun' } } } '404': { description: Verification run not found } components: securitySchemes: bearerAuth: { type: http, scheme: bearer, bearerFormat: JWT } schemas: ApprovalGateRequest: type: object required: [agent_run_id, release_version, release_commit_hash] properties: agent_run_id: { type: string, format: uuid, description: 'AgentRun id the candidate belongs to' } release_version: type: string pattern: '^v\d+.\d+.\d+(-[a-zA-Z0-9.]+)?(+[a-zA-Z0-9.]+)?$' description: 'v-prefixed SemVer release version (e.g. v1.0.0), bound at request time' release_commit_hash: type: string pattern: '^[a-f0-9]{40}$' minLength: 40 maxLength: 40 description: 'Git commit SHA (40-char lowercase hex), bound at request time' reason: { type: [string, 'null'], maxLength: 500, description: 'Optional reason for approval' } reason_required: { type: boolean, default: false, description: 'Whether reason is required for decision' } close_period: { type: [string, 'null'], description: 'Optional period identifier (e.g. 2026-07) to close. When provided, server period-closes the immutability block at consume time.' } Warning: type: object required: [code, message] properties: code: { type: string } message: { type: string } resource: { type: [string, 'null'] } DashboardQueryModel: type: object required: [schema_version, environment_id, dashboard, charts, datasets, native_filters, capabilities, warnings, query_model_fingerprint] properties: schema_version: { const: 1 } environment_id: { type: string } dashboard: { type: object, required: [id, title], properties: { id: { type: integer }, title: { type: string }, slug: { type: [string, 'null'] } } } charts: { type: array, items: { type: object } } datasets: { type: array, items: { type: object } } native_filters: { type: array, items: { type: object } } capabilities: { type: object, additionalProperties: { type: boolean } } warnings: { type: array, items: { $ref: '#/components/schemas/Warning' } } query_model_fingerprint: { type: string, pattern: '^[a-f0-9]{64}$' } FilterInput: type: object additionalProperties: false required: [filter_id, value] properties: filter_id: { type: string } value: {} NormalizeFiltersRequest: type: object additionalProperties: false required: [environment_id, dashboard_id, chart_id, query_model_fingerprint, filters] properties: environment_id: { type: string } dashboard_id: { type: integer } chart_id: { type: integer } query_model_fingerprint: { type: string } filters: { type: array, items: { $ref: '#/components/schemas/FilterInput' } } NormalizedFilterContext: type: object required: [schema_version, filters, filters_hash] properties: schema_version: { const: 1 } filters: { type: array, items: { type: object } } filters_hash: { type: string, pattern: '^[a-f0-9]{64}$' } ExecuteQueryRequest: type: object additionalProperties: false required: [environment_id, dashboard_id, chart_id, result_key, filters] properties: environment_id: { type: string } dashboard_id: { type: integer } chart_id: { type: integer } result_key: { type: string } filters: { $ref: '#/components/schemas/NormalizedFilterContext' } SourceRef: type: object required: [environment_id, dashboard_id, result_key, query_hash] properties: environment_id: { type: string } dashboard_id: { type: integer } chart_id: { type: [integer, 'null'] } dataset_id: { type: [integer, 'null'] } result_key: { type: string } query_hash: { type: string } NormalizedValue: type: object required: [kind, canonical, source, warnings] properties: kind: { type: string, enum: ['null', boolean, integer, decimal, string, date, datetime, percent, table] } raw: {} canonical: {} display: {} format: { type: [object, 'null'] } source: { $ref: '#/components/schemas/SourceRef' } warnings: { type: array, items: { $ref: '#/components/schemas/Warning' } } ComparisonPolicy: type: object required: [type] properties: type: { type: string, enum: [exact, absolute_tolerance, relative_tolerance, range, row_set] } additionalProperties: true BaselineEntry: type: object required: [schema_version, baseline_id, dashboard_id, result_key, normalized_filters, expected, policy, status, fingerprints, provenance, approval] properties: schema_version: { const: 1 } baseline_id: { type: string, format: uuid } dashboard_id: { type: integer } chart_id: { type: [integer, 'null'] } dataset_id: { type: [integer, 'null'] } result_key: { type: string } label: { type: string } normalized_filters: { $ref: '#/components/schemas/NormalizedFilterContext' } expected: { $ref: '#/components/schemas/NormalizedValue' } policy: { $ref: '#/components/schemas/ComparisonPolicy' } status: { type: string, enum: [approved, superseded, retired] } fingerprints: { type: object } source_response_hash: { type: string, pattern: '^[a-f0-9]{64}$', description: 'SHA-256 of Superset API response at capture time' } captured_at: { type: string, format: date-time } release_version: { type: string, pattern: '^v\d+.\d+.\d+(-[a-z0-9.]+)?$' } release_commit_hash: { type: string, pattern: '^[a-f0-9]{40}$' } provenance: { type: object } approval: { type: object } immutability: type: [object, 'null'] additionalProperties: false properties: enabled: { type: boolean, description: 'Whether immutability enforcement is active' } period: { type: string, description: 'Period identifier, e.g. 2026-07' } period_closed_at: { type: [string, 'null'], format: date-time, description: 'ISO-8601 timestamp when period was formally closed. When null, period is open and immutability checks are skipped.' } frozen_at: { type: string, format: date-time, description: 'ISO-8601 timestamp of period closure (legacy alias)' } source_response_hash: { type: [string, 'null'], pattern: '^[a-f0-9]{64}$', description: 'Authoritative SHA-256 of normalized Superset response/artifact bytes at closure time. Computed server-side. When set and period_closed_at is not null, any match failure triggers immutability_violation.' } policy: { enum: [alert, block_publish, require_investigation], description: 'Action policy when violation detected' } ComparisonRequest: type: object additionalProperties: false required: [actual, baseline] properties: actual: { $ref: '#/components/schemas/NormalizedValue' } baseline: { $ref: '#/components/schemas/BaselineEntry' } ComparisonResult: type: object required: [status, actual, policy, warnings] properties: status: { type: string, enum: [pass, fail, inconclusive, missing_baseline, stale_baseline, stale_visual_baseline, immutability_violation, permission_denied, source_error] } actual: { $ref: '#/components/schemas/NormalizedValue' } expected: {} policy: { $ref: '#/components/schemas/ComparisonPolicy' } diff: {} stale_dimensions: { type: array, items: { type: string } } warnings: { type: array, items: { $ref: '#/components/schemas/Warning' } } CandidateRequest: type: object additionalProperties: false required: [agent_run_id, repository_key, dashboard_key, proposed_entry, observations] properties: agent_run_id: { type: string, format: uuid } repository_key: { type: string } dashboard_key: { type: string } proposed_entry: { $ref: '#/components/schemas/BaselineEntry' } observations: { type: array, minItems: 1, items: { $ref: '#/components/schemas/NormalizedValue' } } BaselineCandidate: type: object required: [candidate_id, status, artifact_id, proposed_entry, validation_status, discrepancies] properties: candidate_id: { type: string, format: uuid } status: { const: draft } artifact_id: { type: string, format: uuid } proposed_entry: { $ref: '#/components/schemas/BaselineEntry' } validation_status: { type: string, enum: [valid, warning, invalid] } discrepancies: { type: array, items: { type: object } } VisualBaseline: type: object required: [baseline_id, dashboard_id, kind, release_version, release_commit_hash, normalized_filters, tab_identifier, expected_image_sha256, policy, status, source_response_hash, captured_at, provenance] properties: baseline_id: { type: string, format: uuid } dashboard_id: { type: integer } release_version: { type: string, pattern: '^v\d+.\d+.\d+(-[a-z0-9.]+)?$' } release_commit_hash: { type: string, pattern: '^[a-f0-9]{40}$' } kind: { const: visual } normalized_filters: { $ref: '#/components/schemas/NormalizedFilterContext' } tab_identifier: { type: string } region_of_interest: type: object properties: selector: { type: string } bounds: { type: object, properties: { x: { type: integer }, y: { type: integer }, w: { type: integer }, h: { type: integer } } } expected_image_sha256: { type: string, pattern: '^[a-f0-9]{64}$' } source_response_hash: { type: string, pattern: '^[a-f0-9]{64}$' } captured_at: { type: string, format: date-time } policy: type: object required: [type] properties: type: { type: string, enum: [exact, perceptual] } ssim_min: { type: number } pixel_diff_threshold: { type: number } status: { type: string, enum: [approved, superseded, retired] } immutability: type: [object, 'null'] additionalProperties: false properties: enabled: { type: boolean, description: 'Whether immutability enforcement is active' } period: { type: string, description: 'Period identifier, e.g. 2026-07' } period_closed_at: { type: [string, 'null'], format: date-time, description: 'ISO-8601 timestamp when period was formally closed. When null, period is open and immutability checks are skipped.' } frozen_at: { type: string, format: date-time, description: 'ISO-8601 timestamp of period closure (legacy alias)' } source_response_hash: { type: [string, 'null'], pattern: '^[a-f0-9]{64}$', description: 'Authoritative SHA-256 of normalized Superset response/artifact bytes at closure time. Computed server-side. When set and period_closed_at is not null, any match failure triggers immutability_violation.' } policy: { enum: [alert, block_publish, require_investigation], description: 'Action policy when violation detected' } provenance: { type: object } StructureDiff: type: object required: [release_from, release_to, query_model_hash_from, query_model_hash_to, changes, summary] properties: release_from: { type: string } release_to: { type: string } query_model_hash_from: { type: string } query_model_hash_to: { type: string } changes: type: array items: type: object required: [target, kind, severity, rationale] properties: target: { type: string } kind: { type: string, enum: [filter_scope_narrowed, filter_scope_widened, filter_operator_changed, filter_default_changed, filter_removed, column_order_changed, column_added, column_removed, chart_added, chart_removed, viz_type_changed, group_by_changed, time_grain_changed] } severity: { type: string, enum: [critical, warning, info] } before: {} after: {} affected_artifacts: { type: array, items: { type: string } } rationale: { type: string } summary: type: object required: [critical, warning, info, pass] properties: critical: { type: integer } warning: { type: integer } info: { type: integer } pass: { type: integer } blocked: { type: boolean } SnapshotCaptureResponse: type: object required: [snapshot_path, environment_id, dashboard_id, release_version, release_id, release_commit_hash, repository_id, repository_key, dashboard_key] properties: snapshot_path: { type: string, description: 'Absolute path to the persisted snapshot file' } environment_id: { type: string, description: 'Superset environment ID' } dashboard_id: { type: integer, description: 'Superset dashboard ID' } release_version: { type: string, pattern: '^v\d+.\d+.\d+(-[a-z0-9.]+)?$', description: 'Release version label' } release_id: { type: string, description: 'DashboardRelease ID' } release_commit_hash: { type: string, pattern: '^[a-f0-9]{40}$', description: 'Release commit hash' } repository_id: { type: string, description: 'GitRepository ID' } repository_key: { type: string, description: 'Git repository key used for path resolution' } dashboard_key: { type: string, description: 'Dashboard key used for path resolution' } charts_count: { type: integer, default: 0, description: 'Number of charts in the snapshot' } filters_count: { type: integer, default: 0, description: 'Number of native filters in the snapshot' } datasets_count: { type: integer, default: 0, description: 'Number of datasets in the snapshot' } query_model_fingerprint: { type: string, default: '', description: 'Fingerprint of the captured query model' } warnings: { type: integer, default: 0, description: 'Number of warnings from inspection' } ApprovalGateResponse: type: object required: [gate_id, candidate_id, operation, required_permission, status, created_at] properties: gate_id: { type: string, format: uuid, description: 'Approval gate UUID' } candidate_id: { type: string, format: uuid, description: 'Baseline candidate UUID' } operation: { type: string, description: 'One-shot operation name' } target_paths: { type: array, items: { type: string }, description: 'Target paths for materialization' } risk_level: { type: string, default: guarded, description: 'Risk level of the operation' } required_permission: { type: string, description: 'Server-enforced required permission' } status: { type: string, description: 'Gate status (pending, confirmed, denied, consumed)' } reason_required: { type: boolean, default: false } created_at: { type: string, format: date-time, description: 'ISO-8601 timestamp of gate creation' } ApprovalDecisionResponse: type: object required: [status, gate_id, actor_id] properties: status: { type: string, description: 'Decision result (confirmed, denied)' } gate_id: { type: string, format: uuid, description: 'Approval gate UUID' } candidate_id: { type: [string, 'null'], format: uuid, description: 'Baseline candidate UUID' } actor_id: { type: string, description: 'User who made the decision' } ApprovalConsumeResponse: type: object required: [consumed, gate_id, status, baseline_id, release_version, release_commit_hash] properties: consumed: { type: boolean, description: 'Whether the gate was consumed' } gate_id: { type: string, format: uuid, description: 'Approval gate UUID' } status: { type: string, description: 'Gate status after consumption' } baseline_id: { type: string, format: uuid, description: 'UUID of the materialized baseline entry' } release_version: { type: string, description: 'v-prefixed SemVer release version' } release_commit_hash: { type: string, pattern: '^[a-f0-9]{40}$', description: '40-char git commit hash' } CaptureCandidateRequest: type: object additionalProperties: false required: [agent_run_id, release_id, dashboard_id, result_key, label, normalized_filters, comparison_policy] properties: agent_run_id: { type: string, format: uuid, description: 'AgentRun that owns the candidate' } release_id: { type: string, format: uuid, description: 'DashboardRelease id — environment, repository, coordinates derived from it' } dashboard_id: { type: integer, description: 'Superset dashboard ID' } chart_id: { type: [integer, 'null'], description: 'Chart ID (optional)' } dataset_id: { type: [integer, 'null'], description: 'Dataset ID (optional)' } result_key: { type: string, description: 'Metric or result identifier to extract' } label: { type: string, description: 'Human-readable label for baseline candidate' } normalized_filters: { $ref: '#/components/schemas/NormalizedFilterContext' } comparison_policy: { $ref: '#/components/schemas/ComparisonPolicy' } kind: { type: string, enum: [metric], default: metric, description: 'Must be metric for capture path' } max_rows: { type: integer, default: 10000, maximum: 10000, description: 'Bounded result limit' } CaptureCandidateResponse: type: object required: [candidate, capture_artifact_id, source_response_hash] additionalProperties: false properties: candidate: { type: object, description: 'The created BaselineCandidate' } capture_artifact_id: { type: string, description: 'DraftArtifact id of the capture execution record' } source_response_hash: { type: string, pattern: '^[a-f0-9]{64}$', description: 'SHA-256 of the raw Superset response bytes, computed server-side' } CategoryOutcome: type: object additionalProperties: false required: [category, status] properties: category: { type: string } status: { type: string, enum: [pass, fail, blocked, inconclusive, skipped, immutability_violation] } summary: { type: string, default: '' } details: { type: [object, 'null'], additionalProperties: true } evidence_refs: { type: array, items: { type: string } } VerificationRun: type: object additionalProperties: false required: [id, repository_id, trigger, environment_id, overall_status, created_at] properties: id: { type: string, format: uuid } release_id: { type: [string, 'null'], format: uuid } repository_id: { type: string, format: uuid } agent_run_id: { type: [string, 'null'], format: uuid } trigger: { type: string, enum: [manual, deploy_to_preprod, release_create, release_approve, release_publish, post_publish, scheduled, etl_completed] } environment_id: { type: string } categories_run: { type: array, items: { type: string } } categories_passed: { type: array, items: { type: string } } categories_failed: { type: array, items: { type: string } } category_outcomes: { type: array, items: { $ref: '#/components/schemas/CategoryOutcome' } } overall_status: { type: string, enum: [pass, warn, fail, blocked, inconclusive, immutability_violation] } summary: { type: string, default: '' } baseline_version: { type: [string, 'null'] } baseline_commit: { type: [string, 'null'] } created_at: { type: string, format: date-time } created_by: { type: string, default: system }

QUICKSTART — Dev Onboarding

Source: quickstart.md

Quickstart: Superset Baseline Engine

Prerequisite

Complete 036 through approval consume tests. Use checked-in fixtures first; Docker is required only for Superset 4.1.2 integration.

Test Order

cd backend
python -m pytest tests/services/dashboard_testing -v
python -m pytest tests/api/test_dashboard_testing.py -v
python -m pytest tests/integration/test_dashboard_testing_superset.py --run-integration -v

Independent Smoke

  1. Inspect a fixture dashboard and snapshot the deterministic DashboardQueryModel.
  2. Normalize the same date/decimal/list filters in two locale representations; hashes must match.
  3. Execute a saved scalar chart via POST /api/v1/chart/data and verify source ids/hash.
  4. Execute a table chart with row limit and canonical columns/rows.
  5. Attempt requests containing sql, raw query_context, adhoc expression, and endpoint; all must fail before Superset call.
  6. Compare exact, absolute, relative, range, and row-set fixtures.
  7. Change query/dataset/filter fingerprints independently; each must return stale_baseline.
  8. Create a draft candidate and verify approved catalog is unchanged.
  9. Approve through a 036 gate with reason; verify one atomic YAML update.
  10. Replay or mutate the request; verify 409 and no second write.

Exit Gates

  • Deterministic fixture snapshots.
  • Decimal/date/percent equivalence without false diffs.
  • 403/404/422/timeout/5xx taxonomy preserved.
  • No direct SQL request surface or agent tool reachability.
  • Baseline schema validates and catalog writer is deterministic.
  • Unit/API/integration, ruff, and semantic audits pass.

Known Gap (2026-08-07 MVP audit)

Шаги 1–10 проверяют дискретные инструменты напрямую (inspect/normalize/execute/compare/candidate/approve) — они работают. Но пайплайн-автоматизация не замкнута: deploy-хук (deploy_to_preprod/release_create/etl_completed) не создаёт VerificationRun автоматически, а GET-эндпоинты /verification/history и /verification/{run_id} отсутствуют (только POST /verification-runs). До T080–T081 (tasks.md Phase 10) релизный цикл не получает автоматические verification-прогоны, а 039 pipeline views не могут загрузить историю.


TRACEABILITY — Requirements Matrix

Source: traceability.md

#region SupersetBaselineEngine.Traceability [C:3] [TYPE ADR] [SEMANTICS traceability,baseline,requirements] @BRIEF Requirement-to-contract-to-task-to-test matrix for feature 037. @RELATION DEPENDS_ON -> [SupersetBaselineEngine.Modules] @RELATION DEPENDS_ON -> [SupersetBaselineEngine.Spec]

Requirement Contract Tasks Test
AGBASE-FR-001 BaselineEngine.QueryModel.Inspect T006–T010 query-model snapshots/scopes
AGBASE-FR-002, AGBASE-FR-009 SupersetClient.ChartData.Execute T011–T015 malicious fields and Testcontainers
AGBASE-FR-003 BaselineEngine.Filters.Normalize T007–T010 type/scope/hash cases
AGBASE-FR-004 BaselineEngine.Result.Normalize T016–T019 scalar/table/date/decimal/percent
AGBASE-FR-005, AGBASE-FR-006 BaselineEngine.Catalog.Load T023–T027 schema/path/deterministic YAML
AGBASE-FR-007 Candidate.Create/Approve T028–T032 draft and bound approval
AGBASE-FR-008 BaselineEngine.Comparison.Compare T020–T022 all policies/statuses
AGBASE-FR-010 BaselineEngine.Visual.Compare, BaselineEngine.Visual.Candidate T039–T044 visual baseline schema, stale_visual, perceptual tolerance

Dependencies

Upstream: 036 draft/gate contracts. Downstream: 038 baseline refs and 039 baseline impact panel.

Amendment (041-dataset-lineage-blast-radius, R5)

The trigger field on VerificationRun gains the additive value dataset_updated (producer: Services.Lineage.Fanout, 041 US5); fanout_plan_id nullable FK added (R8). Additive, backward compatible; unknown triggers remain opaque displayable strings (039 AGUI-FR-016). Propagation surface (Services.Lineage.Propagation) is a read-time overlay — zero catalog writes (SC-003).

Amendment (2026-08-07 MVP audit)

Дискретные инструменты оценки метрик реализованы и реально исполняются, но пайплайн-автоматизация и read-API не замкнуты. Добавлены задачи Phase 10:

Requirement Contract Tasks Test
AGBASE-FR-012 context (pipeline triggers) BaselineEngine.Verification.Orchestrator T080 deploy fixture release → VerificationRun with trigger + outcomes
AGUI-FR-015/016 read API Api.DashboardTesting.VerificationRuns T081 history ordering/filters, detail payload, OpenAPI drift

#endregion SupersetBaselineEngine.Traceability


TASKS — Implementation Tasks

Source: tasks.md

#region SupersetBaselineEngine.Tasks [C:3] [TYPE ADR] [SEMANTICS tasks,baseline,implementation] @BRIEF Ordered TDD backlog for Superset-native baseline engine. @RELATION DEPENDS_ON -> [SupersetBaselineEngine.Modules] @RELATION DEPENDS_ON -> [SupersetBaselineEngine.DataModel]

Phase 1 — Fixtures and DTO Foundation

  • T001 Create canonical dashboard/chart/dataset/native-filter Superset fixtures under specs/037-superset-baseline-engine/fixtures/superset/.
  • T002 [P] Create scalar, percent, date, table, empty, malformed, and locale result fixtures under specs/037-superset-baseline-engine/fixtures/results/.
  • T003 [P] Create valid/invalid/stale baseline catalog fixtures under specs/037-superset-baseline-engine/fixtures/baselines/.
  • T004 Materialize fixtures into backend/tests/fixtures/dashboard_testing/.
  • T005 Implement extra-forbid DTOs from contracts/dashboard-testing.openapi.yaml in backend/src/schemas/dashboard_testing.py.

Phase 2 — US1 Inspect Dashboard Query Model

  • T006 [US1] Write failing deterministic inspection tests in backend/tests/services/dashboard_testing/test_query_model.py.
  • T007 [US1] Write failing filter scope/type/hash tests in backend/tests/services/dashboard_testing/test_filters.py.
  • T008 [US1] Implement backend/src/services/dashboard_testing/query_model.py using authoritative SupersetClient metadata.
  • T009 [US1] Implement backend/src/services/dashboard_testing/filters.py with canonical typed values and deterministic hashes.
  • T010 [US1] Add fingerprint helpers in backend/src/services/dashboard_testing/fingerprints.py and cover metadata order invariance.

Checkpoint: Fixture dashboards yield byte-stable models and correct chart/filter scopes.

Phase 3 — US2 Superset-Native Execution

  • T011 [US2] Write failing no-SQL schema and payload tests in backend/tests/services/dashboard_testing/test_query_executor.py.
  • T012 [US2] Add backend/src/core/superset_client/_chart_data.py adapter for saved-chart POST /api/v1/chart/data.
  • T013 [US2] Implement backend/src/services/dashboard_testing/query_executor.py: reload authoritative metadata, scope filters, bound limits, typed errors.
  • T014 [US2] Add agent tools inspect_dashboard_query_model and execute_dashboard_result in agent/src/ss_tools/agent/tools.py as thin backend clients.
  • T015 [US2] Verify scenario intent tool pipeline includes these tools and excludes superset_execute_sql.

Checkpoint: Scalar and table fixtures execute through chart-data only; injected SQL/raw context cannot reach Superset.

Phase 4 — US3 Normalize and Compare

  • T016 [US3] Write failing normalization tests in backend/tests/services/dashboard_testing/test_normalization.py.
  • T017 [US3] Implement normalization.py with Decimal strings, ISO temporal values, percent metadata, and bounded tables.
  • T018 [US3] Write failing comparison policy tests in backend/tests/services/dashboard_testing/test_comparison.py.
  • T019 [US3] Implement exact/absolute/relative/range/row-set comparison in backend/src/services/dashboard_testing/comparison.py.
  • T020 [US3] Add empty/unsupported/duplicate-row-key cases that must return inconclusive to backend/tests/services/dashboard_testing/test_comparison.py.
  • T021 [US3] Add independent query/dataset/filter staleness tests in backend/tests/services/dashboard_testing/test_staleness.py.
  • T022 [US3] Add deterministic diff/evidence reference output tests in backend/tests/services/dashboard_testing/test_comparison.py.

Phase 5 — US4 Baseline Candidate Lifecycle

  • T023 [US4] Write failing JSON-schema/YAML/path tests in backend/tests/services/dashboard_testing/test_baseline_catalog.py.
  • T024 [US4] Implement safe catalog loader and deterministic writer in backend/src/services/dashboard_testing/baseline_catalog.py.
  • T025 [US4] Resolve repository only through GitService and reject traversal/symlink escape.
  • T026 [US4] Validate catalogs against contracts/baseline-catalog.schema.json before use/write.
  • T027 [US4] Write failing candidate provenance/discrepancy tests in backend/tests/services/dashboard_testing/test_candidates.py.
  • T028 [US4] Implement candidate creation in backend/src/services/dashboard_testing/candidates.py as 036 DraftArtifact; never update approved YAML.
  • T029 [US4] Implement request-baseline-approval endpoint in backend/src/api/routes/dashboard_testing.py using 036 request hash and reason_required.
  • T030 [US4] Implement approval consume callback in backend/src/services/dashboard_testing/candidates.py that atomically writes YAML and records gate id/reason.
  • T031 [US4] Cover stale-after-confirm, RBAC revoked, payload mutation, and replay in backend/tests/services/dashboard_testing/test_candidates.py.
  • T032 [US4] Add agent tools discover_baseline_candidate and request_baseline_approval without direct approve capability.

Phase 6 — API and Integration

  • T033 Add backend/src/api/routes/dashboard_testing.py matching OpenAPI and register router.
  • T034 Write API contract/RBAC tests in backend/tests/api/test_dashboard_testing.py.
  • T035 Add Superset 4.1.2 Testcontainers test in backend/tests/integration/test_dashboard_testing_superset.py.
  • T036 Add compatibility tests for existing translation preview chart-data usage in backend/tests/core/superset_client/test_chart_data.py.
  • T037 Run quickstart, backend full relevant tests, ruff, and OpenAPI/schema validation.
  • T038 Audit C3+ contracts, direct-SQL ban, async boundaries, ATTN_1–4, and unresolved relations.

Phase 7 — Visual Baseline Support (AGBASE-FR-010)

  • T039 [P] Write failing visual baseline schema validation tests in backend/tests/services/dashboard_testing/test_visual_baseline.py.
  • T040 Extend baseline-catalog.schema.json validation to accept visualEntry alongside metric entries; reject cross-kind policy usage.
  • T041 [P] Implement backend/src/services/dashboard_testing/visual_baseline.py: VisualComparisonPolicy (exact + perceptual), layout fingerprint computation, stale_visual_baseline detection.
  • T042 [P] Wire visual baseline loading into BaselineEngine.Catalog.Load; extend catalog YAML to support visual entries.
  • T043 Implement BaselineEngine.Visual.Compare: digest comparison + perceptual SSIM path; return stale_visual_baseline when layout fingerprint mismatches.
  • T044 [P] Implement BaselineEngine.Visual.Candidate: create draft visual candidate from reviewed screenshot artifact with mandatory human disposition.
  • T045 Add visual baseline approval flow reusing 036 gate; verify approval writes visual entry atomically alongside metric entries.
  • T046 Write visual baseline golden fixtures under specs/037-superset-baseline-engine/fixtures/visual/.
  • T047 Audit: visual baselines never use metric policies; metric baselines never use visual policies; cross-kind comparison returns inconclusive.

Requirement Mapping (FR-011/012/013)

FR-011 (AGBASE-FR-011) — Authoritative Capture & Closed-Period Immutability

  • Capture endpoint POST /baseline-candidates/capture: resolves release → environment → SupersetClient → query model → envelope (raw httpx bytes) → DraftStorage → capture artifact → candidate. Implements BaselineEngine.Candidates.Capture.ExecuteAndCapture [C:5].
  • source_response_hash always server-computed from raw httpx bytes (SHA-256 of envelope.raw_response_content); never caller-supplied.
  • environment_id, repository_key, dashboard_key all derived server-side from DashboardRelease deployment — never from caller.
  • Closed-period immutability: BaselineEngine.Immutability.Detect.CheckImmutability detects retroactive changes to closed-period entries. Immutability data cannot be caller-supplied (guarded by _reject_caller_immutability in metric helpers).
  • ImmutabilityBlock records period, period_closed_at (server timestamp), source_response_hash, and policy (BLOCK_PUBLISH) at consume time.
  • Publish gate (Phase 8b → BaselineEngine.Verification.PublishGate) runs two-phase check: Phase 1 = catalog-level immutability comparison; Phase 2 = VerificationRun record with metric comparisons.

FR-012 (AGBASE-FR-012) — Publish Gate with Scheduled Verification

  • Publish gate run_publish_gate_verification() in verification_publish_gate.py: resolves release → repository → catalog → Phase 1 immutability check → Phase 2 VerificationRun.
  • PublishBlockedError raised when a block_publish policy entry has hash mismatch; publish transaction aborted. Never reaches DB commit.
  • Scheduled verification verify_published_releases() in verification_scheduler.py: iterates published releases, creates VerificationRunRecord with trigger="scheduled". Observability-only — never blocks, never raises.
  • APScheduler callback execute_scheduled_verification_check() in scheduler.py wired as module-level callback (same pattern as backup/validation).
  • VerificationRunOrchestrator dispatches categories (structure, metric, visual) via async executors. CategoryOutcome statuses: pass, fail, blocked, inconclusive, immutability_violation, warn.
  • Repository FK on VerificationRunRecord (SET NULL on delete for audit retention).

FR-013 (AGBASE-FR-013) — Release-to-Release Baseline Inheritance

  • plan_inheritance() in baseline_inheritance.py: compares prior vs current catalog by (chart_id:dataset_id:result_key) content_hash. Returns InheritancePlan with inherited (unchanged), changed (hash diff), new (fresh) classifications.
  • Classify: _classify_entries() compares prior_map vs current_map by content_hash. content_hash = server-computed during capture.
  • Inheritance handles visual entries via vis:{tab_identifier}:{dashboard_id} composite key in _build_entry_map().
  • execute_inheritance() in inheritance_execute.py: re-extracts changed+new entries from target environment (PREPROD); creates DraftArtifact rows and baseline candidates. Inherited entries carry forward prior baseline value unchanged.
  • Propose inherited candidates via _propose_inherited_candidate() in baseline_inheritance.py — creates candidate from prior entry data without re-querying Superset.
  • API endpoints: POST /inheritance/plan (read-only, computes InheritancePlanResponse) and POST /inheritance/execute (writes candidates + artifacts, requires SupersetClient for target environment).

Implementation Closure Notes (Feature 037)

All 47 tasks (T001–T047) completed. Phases 1–7 each end with a verifiable checkpoint.

  • Phase 5 (US4 Baseline Candidate Lifecycle) implements the full approval lifecycle: create draft candidate → request gate (hash-bound) → decide → consume (atomically materialize YAML catalog).
  • Phase 6 (API) couples OpenAPI schemas to the backend routes under backend/src/api/routes/dashboard_testing/ (split from monolithic dashboard_testing.py into candidates.py, core.py, verification.py, inheritance.py, structure.py, structure_snapshot.py).
  • Phase 7 covers visual baseline support: VisualBaselineEntry schema, visual fingerprint staleness, SSIM perceptual comparison, visual candidate creation with mandatory human approval.
  • Repository FK added to VerificationRunRecord (repository_id FK to git_repositories, SET NULL on delete) to enable service-level repository existence validation at run creation time.
  • Verification extends beyond Phase 7 into the publish gate (phase 8b equivalent) and scheduled verification — both wired as part of 037 implementation, not deferred.

Drift from original spec: Original tasks.md did not explicitly enumerate FR-011/012/013 as task blocks. These requirement areas span multiple existing task phases (US1–US7). The mapping above documents how each FR is served by the completed task set. No functional scope was added beyond the original contract; the publish gate and scheduled verification were always in scope for 037.

Dependencies

T001–T005 → US1 → US2 → US3; US4 depends on US3 and completed 036. API integration follows all domain contracts. Phase 7 depends on completed 036 Phase 8 (screenshot evidence artifacts).

Phase 10 — Pipeline Automation & Read-API Closure (audit 2026-08-07)

Context: Дискретные инструменты оценки метрик (normalization/comparison/metric_executor) работают и реально исполняются, но пайплайн-автоматизация VerificationRun и read-API не дописаны. Факт-чекинг: deploy-хук не создаёт VerificationRun; GET-эндпоинты history/detail отсутствуют, хотя frontend 039 их вызывает.

  • T080 [P] [AGBASE-FR-012 context] Create VerificationRun automatically on pipeline triggers: wire create_verification_run_async into the deploy hook path (backend/src/plugins/git_deployment_recorder.py or the deploy route) for deploy_to_preprod, and add trigger values release_create/etl_completed call sites. @POST: deploying to PREPROD creates a VerificationRun with trigger=deploy_to_preprod and executes metric/visual/structure categories. @TEST: integration test — deploy a fixture release → VerificationRun row exists with correct trigger and outcomes. DONE 2026-08-12 commit 81de959e — _trigger_release_verification fires trigger=release_create on release creation (backend/src/api/routes/git/_release_routes.py), trigger enums include deploy_to_preprod/etl_completed; tests in test_git_release_routes.py (release_create spawn) + test_dashboard_testing_verification_api.py (deploy_to_preprod persistence).
  • T081 [P] [AGUI-FR-015/016] Add read endpoints in backend/src/api/routes/dashboard_testing/verification.py: GET /verification/history?dashboard_id=&env_id= (chronological list) and GET /verification/{run_id} (detail with category outcomes), matching frontend/src/lib/api/dashboard-testing.ts getVerificationHistory()/getVerificationRun(). @POST contract: response shape == VerificationRunDTO[] / VerificationRunDTO used by frontend. @TEST: API test asserts list ordering, filters, and detail payload; OpenAPI drift check. DONE 2026-08-12 commit 81de959e — VerificationHistory/VerificationDetail endpoints live in backend/src/api/routes/dashboard_testing/verification.py; frontend getVerificationHistory()/getVerificationRun() present; OpenAPI alignment test passes.

Dependencies

T001–T005 → US1 → US2 → US3; US4 depends on US3 and completed 036. API integration follows all domain contracts. Phase 7 depends on completed 036 Phase 8 (screenshot evidence artifacts). Phase 10 (T080–T081) closes the pipeline-automation and read-API gap found in the 2026-08-07 audit; it must land before 039 pipeline views can render live VerificationRun data.

#endregion SupersetBaselineEngine.Tasks


tests/qa-audit.md

Source: tests/qa-audit.md

@{ SupersetBaselineEngine.QA.Audit [C:3] [TYPE ADR]

@BRIEF Evidence-led QA record for the Superset-native baseline-engine implementation. @RELATION VERIFIES -> [Api.DashboardTesting.Candidates] @RELATION VERIFIES -> [BaselineEngine.QueryExecutor.ExecuteQuery] @RELATION VERIFIES -> [BaselineEngine.Candidates.Helpers] @TEST_EDGE direct_sql -> rejected by DTO boundary and scenario-tool allowlist. @TEST_EDGE approval_replay -> one-shot gate consumption rejects a second consume. @TEST_EDGE cross_candidate_gate -> approval gate ownership mismatch is rejected.

QA Audit — 2026-07-28

Scope

This audit covers the implementation on branch 037-superset-baseline-engine, with emphasis on no-SQL execution, router reachability, candidate approval safety, semantic traceability, and test mocking discipline.

Implemented corrections

  • Registered dashboard_testing in backend/src/api/routes/__init__.py and backend/src/app.py; /api/dashboard-testing/query-model is registered at application startup.
  • Replaced synthetic filter metadata with authoritative Superset query-model inspection and rejects a mismatched query-model fingerprint.
  • Implemented durable baseline candidates with DraftArtifact records and durable approval gates with ApprovalGate records; no process-local candidate, gate, or replay stores remain.
  • Bound each approval gate to one candidate through capture_meta.gate_id, a canonical candidate-content/path/operation hash, and a candidate-specific consume mode. Same-run cross-candidate decision/consume attempts return conflicts without changing either candidate.
  • Preserved legacy generic agent-run gate consumption while adding explicit candidate-mode consumption: generic gates persist valid run drafts and complete their run; candidate gates persist only their bound draft and leave sibling drafts untouched.
  • Server-controls the candidate approval permission; request payloads cannot select a weaker gate permission. Candidate repository/dashboard keys are validated as safe path components and resolve to the same canonical catalog path used by API reads.
  • Added route-level commit/rollback handling and HTTP lifecycle tests that verify committed state through fresh SQLite sessions after candidate creation, gate request, confirmation, and consumption.
  • Executed baseline-catalog.schema.json validation before load and before write. Reconciliation maps the established Pydantic storage names to the published JSON Schema without weakening hash checks.
  • Added Molecular CoT REASON / REFLECT / EXPLORE markers across scoped dashboard-testing service execution paths.
  • Repaired malformed Phase 3 task prefixes (T011–T015) and completed targeted dashboard-testing Ruff remediation.

Mocking audit

Test area Mocked boundary Verdict
test_query_model.py Superset metadata client methods Allowed external Superset boundary
test_query_executor.py SupersetClient.execute_chart_data Allowed external chart-data boundary
test_langchain_tools.py shared httpx.AsyncClient Allowed external HTTP boundary
candidate service/API lifecycle tests none for the SUT or persistence Real SQLite DraftArtifact / ApprovalGate verification

Hardcoded fixtures are used for expectations; no reviewed 037 test calculates expected values by reimplementing its production algorithm.

Passing focused verification

Command Result
backend/.venv/bin/python -m pytest tests/services/dashboard_testing -q 80 passed
backend/.venv/bin/python -m pytest tests/api/test_dashboard_testing.py tests/services/dashboard_testing/test_candidates.py -q 20 passed
backend/.venv/bin/python -m pytest tests/services/dashboard_testing/test_candidates.py tests/services/agent_runs/test_approvals.py tests/services/dashboard_testing/test_baseline_catalog.py tests/api/test_dashboard_testing.py -q 42 passed
backend/.venv/bin/python -m ruff check src/services/dashboard_testing tests/services/dashboard_testing passed
backend/.venv/bin/python -m ruff check src/services/agent_runs/service.py src/services/dashboard_testing/candidates.py src/api/routes/dashboard_testing.py tests/api/test_dashboard_testing.py passed
PYTHONPATH=src agent/.venv/bin/python -m pytest tests/test_agent/test_scenario_tool_filter.py -q 6 passed
PYTHONPATH=src agent/.venv/bin/python -m ruff check tests/test_agent/test_scenario_tool_filter.py passed
set -a && source backend/.env && set +a && backend/.venv/bin/python -c 'from src.app import app; ...' dashboard-testing route registered (True)
Axiom full rebuild succeeded: 7382 contracts, 3672 edges, no parse warnings (rebuild-1785308552125-0005)
Axiom detect_missing_contracts backend/src/services/dashboard_testing 56 contracted functions, 0 naked functions across 10 files
Post-anchor focused backend verification 98 passed in 11.22s
Post-anchor scoped Ruff passed

Known non-blocking repository debt

  • The latest scoped dashboard-testing suite and its scoped lint are green. Broader backend/frontend suites were previously red before this feature work and require a separate repository-wide remediation effort.
  • pytest-cov is not installed in the backend virtual environment, so the Makefile coverage target cannot run (pytest --cov is unavailable).
  • Dashboard-testing service contracts are now fully anchored: Axiom reports 56 contracted functions and 0 naked functions across the scoped ten service files.
  • API relation resolution still reports 11 CALLS edges in backend/src/api/routes/dashboard_testing.py as unresolved even though their targets are indexed in the service source. This is semantic index/parser debt, not a runtime feature defect.
  • audit_belief_runtime still expects legacy belief_scope / logger.reason calls and flags the scoped C4/C5 services despite their canonical shared log(..., "REASON"|"REFLECT"|"EXPLORE", ...) instrumentation. The static-rule mismatch is semantic-tooling debt; focused runtime tests and Ruff pass.
  • Workspace health still reports 437 unresolved relations globally; resolving unrelated graph edges requires a separate repository-wide semantic-maintenance scope.
  • The final full rebuild completed successfully in approximately 14.4 minutes with no parse warnings: 7382 contracts and 3672 edges.

Axiom MCP experience report

What worked well

  • read_outline was highly effective for safe semantic editing. It exposed anchor hierarchy and metadata without code noise, which made it practical to normalize region syntax and verify matched closures after each file-level change.
  • detect_missing_contracts provided actionable, AST-backed coverage gaps. It identified the exact private helpers that lacked contracts; after remediation it confirmed the scoped service directory had 56 contracted functions and 0 naked functions.
  • Full rebuilds produced durable, inspectable index evidence. Rebuild job rebuild-1785308552125-0005 completed successfully with 7382 contracts, 3672 edges, and no parse warnings.
  • Structured audits were useful for separating code defects from graph debt. audit_belief_protocol found no missing decision-memory tags, while workspace_health made the remaining global unresolved-relation count explicit.
  • Async rebuild status reporting was reliable. The job identifier allowed polling rather than blocking the complete QA workflow.

What needs improvement

  • Full rebuild latency is high for small semantic changes. The final full rebuild took about 14.4 minutes for anchor/documentation updates, so incremental rebuilds should be preferred during iteration and full rebuilds reserved for final evidence.
  • Relation resolution has false negatives. audit_contracts continues to report 11 unresolved CALLS targets from backend/src/api/routes/dashboard_testing.py, although the referenced contracts are present in the indexed scoped services. This weakens the signal of graph audits.
  • The belief-runtime audit is out of sync with the canonical logger API. It expects legacy belief_scope / logger.reason syntax and flags code that uses the shared Molecular CoT API, log(source, "REASON"|"REFLECT"|"EXPLORE", ...).
  • Index metrics were inconsistent across views during the session. Earlier workspace-health and rebuild outputs differed on orphan reporting, so reports should explicitly identify the command, rebuild ID, and timestamp used as evidence.
  • Some audit outputs need clearer remediation classification. Findings caused by parser limitations should be labeled directly as tooling_false_positive or index_resolution_gap, rather than appearing indistinguishable from source-level defects.

Decision-memory guardrails

  • Raw SQL, raw endpoints, raw query context, and adhoc expressions remain rejected at DTO and executor layers.
  • The scenario allowlist continues to exclude superset_execute_sql.
  • Candidate approval remains draft-first, one-shot, candidate-bound, and replay-defended.
  • JSON Schema validation is additive to Pydantic validation; it does not introduce a permissive fallback for malformed SHA-256 values.

Completion status

Scoped feature verification is passing, including durable approval lifecycle, JSON Schema enforcement, runtime CoT instrumentation, HTTP persistence, complete service-function anchoring, and dashboard-testing lint. Repository-wide suite/lint/frontend debt remains explicitly out of scope for this focused feature remediation.

@} SupersetBaselineEngine.QA.Audit

================================================================================ FEATURE: 038-dashboard-scenario-model Files: 22


SPEC — Feature Specification

Source: spec.md

#region DashboardScenarioModel.Spec [C:3] [TYPE ADR] [SEMANTICS spec,requirements,scenario,dashboard-testing] @BRIEF Define the validated immutable Verification Program IR for dashboard test flows authored by agents and executed by 044. @RELATION DEPENDS_ON -> [Doc.Adr.ADR0001] @RELATION DEPENDS_ON -> [Doc.Adr.ADR0002] @RELATION DEPENDS_ON -> [AgentTestStabilization.Spec] @RELATION DEPENDS_ON -> [SupersetBaselineEngine.Spec] @RATIONALE Unique dashboard tests need a stable intermediate model between agent reasoning and generated artifacts; direct LLM-to-code generation is not reviewable or safely composable. @REJECTED Asking users to choose low-level outputs such as Playwright vs SQL vs XLSX — rejected because each dashboard scenario is goal-oriented; the agent proposes inspectable compiled evidence/transform/assertion steps. @REJECTED Direct generation of executable scripts without a validated scenario graph — rejected because it hides missing selectors, baseline refs, and unsafe steps until runtime. @REJECTED Agent-chat-only authoring entry — superseded 2026-08-24: compile/validate/resolve/draft-pack/save become callable through the MCP interface per specs/050-mcp-interface/spec.md; the deterministic validator remains the sole gateway to any artifact.

Navigation (DSA Indexer keywords)

@SEMANTICS: spec, requirements, feature, scenario, graph, dashboard-testing, checklist, validation, artifacts

Feature Branch: 038-dashboard-scenario-model Created: 2026-07-07 | Reworked: 2026-07-31 | Status: Reworked per new speckit flow Input: "Define the dashboard test scenario model used by agents to represent unique dashboard test flows as a validated ScenarioGraph. The model must express ordered and dependent steps across browser automation, Superset query execution, XLSX parsing, assertions, screenshots, reports, human checkpoints, baseline references, parameters, warnings, and missing context markers without exposing users to low-level tool selection."

Applicability

  • Feature type: Fullstack (backend compiler/validator/resolver core + thin UI preview surface consumed by 039 + agent tools).
  • UI surface: Yes — scenario graph preview, coverage, resolution, and pack states (rendered by 039; 038 supplies DTO contracts).
  • API surface: Yes — compile, validate, resolve, draft-pack, capture, VLM, disposition endpoints.
  • Prototype: Applicable — scenario preview is a real UI surface (see prototype/index.html).
  • OpenAPI: Applicable — REST surface is a first-class deliverable (see contracts/openapi.yaml).

User Scenarios

Story 1 — Build Scenario Graph From Dashboard Goal (P1)

Why P1: The agent must express a dashboard testing goal as a reviewable graph of steps before generating any executable artifacts.

Independent Test: Provide dashboard query model and checklist fixture input and verify a deterministic DashboardTestScenario graph is produced.

Acceptance:

  1. Given a dashboard context, ChangeRequestContext, query model, and testing objective When the agent builds a scenario Then the output contains an inspectable Verification Program, parameters, steps, dependencies, evidence sources, warnings, and risk summary (038 does not emit persistence ids).
  2. Given a scenario needs browser, Superset API, XLSX parsing, screenshots, assertions, and report steps When graph is rendered Then every step declares its tool and input/output refs.
  3. Given required metadata is missing When graph is generated Then missing data is represented as NEEDS_CONTEXT or NEEDS_SELECTOR, not invented.

Story 2 — Validate Scenario Graph Safety and Completeness (P1)

Why P1: Scenario artifacts must be generated only from a graph whose refs, baselines, tools, and risks are internally consistent.

Independent Test: Run valid and invalid scenario fixtures through the validator and verify precise errors for broken refs, missing baseline refs, and unsupported steps.

Acceptance:

  1. Given a step consumes an output ref When validation runs Then the referenced producer must exist and precede or depend correctly.
  2. Given an assertion compares a metric to expected truth When validation runs Then it must reference an approved baseline or a draft baseline candidate with explicit approval requirement.
  3. Given a step uses an unsupported tool or unsafe action When validation runs Then the graph is rejected or marked blocked with a user-visible reason.

Story 3 — Map Checklist Cases to Scenario Capabilities (P2)

Why P2: The PDF checklist should guide scenario coverage without forcing all dashboards into identical scripts.

Independent Test: Normalize the research checklist into capability-tagged cases and verify generated scenarios include applicable cases and mark non-applicable ones.

Acceptance:

  1. Given the normalized checklist template contains basic, complex, and technical cases When a dashboard capability model is supplied Then applicable cases are mapped to scenario steps or human checkpoints.
  2. Given a checklist case requires capabilities absent from the dashboard When mapping runs Then the case is marked unsupported or manual-only with rationale.
  3. Given multiple tool choices could verify a case When mapping runs Then the scenario selects inspectable navigation/evidence/transform/assertion steps; source-mart evidence may be a validated SqlEvidenceSpec.

Story 4 — Represent Parameters and Human Checkpoints (P2)

Why P2: Unique dashboard scenarios require business inputs and sometimes cannot be fully automated.

Independent Test: Generate a scenario that requires test date, counterparty, and baseline selection; verify parameter prompts and human checkpoint steps are represented structurally.

Acceptance:

  1. Given a scenario needs test data When graph is generated Then required parameters include name, type, validation rule, default/source, and affected steps.
  2. Given an action cannot be automated reliably When graph is generated Then a human checkpoint step describes the manual action and expected evidence.
  3. Given a reusable scenario has a required ParameterDefinition without a default When it is saved Then it remains save-eligible; a later 044 launch binds the value without changing the graph or content_hash.

Story 5 — Capture, VLM Analysis, and Human Disposition (P2)

Why P2: Screenshot evidence and visual verification require typed, auditable capture/VLM/disposition semantics (AGSCN-FR-010..012).

Independent Test: Generate a scenario with screenshot capture, VLM analysis, and human disposition; verify typed findings and auditable dispositions.

Acceptance:

  1. Given a screenshot step When capture executes Then a reproducible ScreenshotCaptureSpec (target, viewport, readiness, optional masking, max wait) is honored and artifacts are registered.
  2. Given a screenshot (masked or unmasked — providers are local) When VLM analysis runs Then typed VlmFinding[] (severity, region, confidence, model/prompt provenance) are returned; raw prose is never treated as step state.
  3. Given a human checkpoint references VLM finding ids When the user disposes Then confirm/false_positive/inconclusive is typed and auditable, and disposition never mutates the graph structure.

Edge & Failure Cases

# Scenario Category Expected Behavior Recovery / Test Ownership
E1 Scenario has a cycle in dependencies data-integrity Validator rejects with cycle path User fixes graph; L1 validator test
E2 Two steps produce the same output ref data-integrity Validator rejects ambiguous ref User fixes ref; L1 test
E3 Baseline is stale data-quality Assertion stays present but blocked/warning-gated 037 baseline discovery or mark pending; L1 test
E4 XLSX export unavailable integration XLSX-dependent checklist cases become unsupported or manual checkpoints Rationale shown in coverage; L1 mapping test
E5 UI selector unknown integration Browser step uses NEEDS_SELECTOR and blocks executable generation User provides selector hint / converts to checkpoint; L1 test
E6 409 stale base revision on resolve concurrency New revision rejected with 409; snapshot/recompile guidance; never silent merge User recompiles; L1 API test
E7 422 invalid resolution/parameter type validation Field/step-mapped validation error User corrects input; L1 API test
E8 403 forbidden role on scenario operations auth Permission denial rendered without approval gate User contacts admin / RBAC test
E9 429 rate limit on compile/validate throttling Retry-After honored; UI countdown User waits; L2 UX test
E10 5xx backend failure on compile server-error Error section + retry; partial graph not persisted User retries; L2 UX test
E11 Malformed VLM response / empty findings integration Findings array empty; step inconclusive with reason; stale prompt blocked (422 STALE_PROMPT) Re-run analysis; L1 VLM test
E12 Missing runtime parameter at launch data-quality Saved pack remains save_eligible; 044 RunPreflight rejects only that launch Supply a typed binding; L1 preflight test
E13 Unsafe path / executable code / SQL injection into pack security SQL compilation gate blocks unsafe SQL; arbitrary code/path validation blocks before draft registration L1 security test; injected-code fixture
E14 Duplicate submit of draft-pack idempotency Idempotency key / revision hash prevents double registration L1 API test; 409 on changed revision

Requirements

Functional

  • AGSCN-FR-001: The system MUST define a DashboardTestScenario model with dashboard context, objective, parameters, steps, dependencies, outputs, artifacts, risks, and warnings.
  • AGSCN-FR-002: Every scenario step MUST declare tool category, action, inputs, outputs, expected result, dependencies, and automation status.
  • AGSCN-FR-003: Supported tool categories MUST include browser automation, Superset API execution, XLSX parsing, assertion, screenshot/evidence, report generation, artifact generation, and human checkpoint.
  • AGSCN-FR-003a: DashboardTestScenario MUST contain a first-class, content-hashed Verification Program with navigation, evidence, transformation, assertion and semantic-evaluation programs. It is immutable runtime input, not a runtime planning hint.
  • AGSCN-FR-004: Assertions against reference values MUST use baseline references or baseline candidate references; raw expected numbers MUST NOT be embedded directly in executable steps.
  • AGSCN-FR-005: Authoring validation MUST detect missing refs, cycles, duplicate outputs, missing/stale baselines, unknown selectors, unsupported tools, and invalid ParameterDefinitions. Required runtime bindings are enforced only by 044 RunPreflight.
  • AGSCN-FR-006: Checklist mapping MUST use normalized checklist cases derived from the research PDF and capability tags, not hardcoded one-size-fits-all scripts. Mapping accepts the target release_version for baseline lookup.
  • AGSCN-FR-007: The model MUST allow manual/human checkpoint steps where automation is unsafe, unavailable, or underspecified.
  • AGSCN-FR-008: Scenario output MUST be deterministic for the same dashboard query model, checklist template, baseline catalog, and user parameters.
  • AGSCN-FR-009: The scenario model MUST remain implementation-neutral and must not require the user to choose low-level artifacts such as Playwright, XLSX, or API output upfront.
  • AGSCN-FR-010: Screenshot steps MUST carry a capture SPECIFICATION: target (tab/viewport), viewport dimensions, readiness strategy, OPTIONAL masking selectors (relaxed 2026-08-24 — local-only providers; masking is display hygiene, not a gate), and max wait. Execution of capture is owned by 044 (ScenarioExecution CaptureService delegating to Plugin.Service.ScreenshotService), with artifacts owner_type=scenario_run; 038 defines the spec, not the runtime path.
  • AGSCN-FR-011: Visual-analysis steps MUST carry a typed VlmAnalysisSpec (profile/provider/model/prompt template/hash/confidence). Runtime VLM submission and VlmFinding production are owned by 044, reusing Plugin.Service.LLMClient resolved through Services.LlmProvider.LLMProviderService (multimodal-required, encrypted-key handling, JSON mode); response redaction via Plugin.Service.RedactionService is OPTIONAL since 2026-08-24 (local-only providers). A stub/default submit returning empty findings without a real provider call is incomplete.
  • AGSCN-FR-011a: Read-only SqlEvidenceSpec MAY be authored during creation/edit/revalidation or investigation proposal only. Save MUST require AST/policy/schema/preview validation. A ScenarioRun executes exactly the saved SQL template via the Superset SQL Lab adapter with typed bindings; runtime LLM SQL rewrite, relation/projection/join/filter mutation and credential handoff are forbidden.
  • AGSCN-FR-011b: TransformSpec MUST use only the bounded versioned DSL and ComparisonSpec/AssertionSpec MUST compare declared evidence refs. Arbitrary Python/code is forbidden.
  • AGSCN-FR-011c: AgentEvaluationSpec MAY cover only declared semantic/visual/ambiguous checks. It MUST pin model/prompt/evidence/input/tool access/output schema and DecisionPolicy; it cannot alter graph, SQL/DSL, orchestration, lifecycle or mutations.
  • AGSCN-FR-011d: Authoring MUST accept first-class ChangeRequestContext; the compiler must mark missing needed context as needs_context, never guess it.
  • AGSCN-FR-012: Human checkpoint steps MAY reference specific VLM finding ids. Resolution options (confirm, false_positive, inconclusive) MUST be typed and auditable — this is a 044 HumanCheckpoint, distinct from the 036 authorization ActionApprovalGate. Disposition changes finding status, not graph structure.
  • AGSCN-FR-013: The agent MAY plan scenario coverage, resolve ambiguity, and create a validated WorkingDraft or executable revision when delegated policy permits. Every resulting graph MUST pass the deterministic 038 validator and retain canonical immutable provenance.
  • AGSCN-FR-014: Graph authoring and revision actions MUST be available in the persistent scenario workspace or agent thread; modal/dialog interaction MUST NOT be required for authoring, review, conflict recovery, or approval.

Key Entities

  • DashboardTestScenario: Reviewable graph representing a dashboard-specific testing objective and all steps required to validate it.
  • ScenarioStep: One executable, generated, assertion, evidence, or human checkpoint node in the graph.
  • ParameterDefinition: Immutable name/type/default/validation/affected-step declaration; launch values are 044 ParameterBindings and never affect content_hash.
  • VerificationProgram: Content-hashed navigation/evidence/transform/assertion/semantic-evaluation IR.
  • SqlEvidenceSpec: Validated immutable read-only source-mart evidence query.
  • ScenarioRef: Named output produced by one step and consumed by later steps.
  • ChecklistCase: Normalized item from the research checklist with capability tags and expected verification semantics.
  • CapabilityMapping: Decision record mapping dashboard capabilities to applicable checklist cases and step templates.
  • ScenarioValidationResult: Structured validator output with errors, warnings, blockers, and graph coverage.
  • ScreenshotCaptureSpec: Capture CONFIGURATION for a screenshot step — viewport, target tab, readiness strategy, optional masking selectors, and timeouts. (Execution owned by 044.)
  • VlmAnalysisSpec: Typed SPEC of visual analysis — provider/model/prompt/hash/confidence. (Runtime VlmFinding owned by 044.)
  • HumanDisposition: Typed human decision on a VLM finding — confirm, false_positive, or inconclusive (044 HumanCheckpoint).

Success Criteria

  • SC-001: Scenario validator catches 100% of invalid fixture cases for missing refs, cycles, duplicate outputs, and missing baseline refs.
  • SC-002: At least 80% of normalized PDF checklist cases are classifiable as automated, human checkpoint, unsupported, or needs-context for fixture dashboards.
  • SC-003: Same inputs produce byte-stable scenario JSON/YAML in deterministic snapshot tests.
  • SC-004: No generated scenario fixture embeds raw baseline numbers directly in executable steps.
  • SC-005: Scenario graph preview can display phase order, tools per step, parameters, warnings, and blockers without reading generated code.
  • SC-006: VLM findings are advisory only; disposition is auditable; no finding alters metric baseline truth.

Clarifications

Session 2026-07-31

  • Q: Is the scenario model a backend-only library or does it expose a UI/API surface? → A: Fullstack — backend compiler/validator core plus thin UI preview (rendered by 039) plus REST/agent-tool API surface.
  • Q: How are missing selectors/context handled? → A: Represented structurally as NEEDS_SELECTOR/NEEDS_CONTEXT, never invented; blocks executable generation for that step.
  • Q: What is the determinism contract? → A: Byte-stable output for identical canonical inputs + compiler/template versions; stable derived ids, no random UUIDs in canonical graph; temperature=0 alone is rejected as a determinism mechanism.
  • Q: What is the VLM safety boundary? → A: VLM output is typed, advisory findings for human review; raw prose is never step state and findings never alter baseline truth; stale prompts block analysis.
  • Q: Is the draft-pack compiled through templates or direct code generation? → A: Versioned repository-owned templates only; LLM text may populate bounded descriptions but never executable code bodies, paths, imports, or shell commands.

Implementation Status & Cross-Spec Reconciliation (audit 2026-08-07)

Facts (code check):

  • ✅ Ядро работает: deterministic compiler/validator/resolver/serializer/pack, checklist-каталог (19 cases), REST surface (compile/validate/resolve/draft-pack).
  • 🔴 Runtime capture/VLM/disposition теперь принадлежат 044, НЕ 038. Прежние задачи T057–T059 (real VLM submit, real capture, e2e evidence) переносятся в 044 ScenarioExecution; 038 владеет только VlmAnalysisSpec/ScreenshotCaptureSpec и компилятором.
  • 🔴 Identity reconciliation: 038 эмитит scenario_key + content_hash; scenario_id (UUID) и revision_id (UUID) назначает 042 при Save (CreateScenario). Компилятор не выдаёт identity до persistence.

Закрытие: cross-spec pass привёл 038 к роли чистого IR/compiler слоя; все runtime-контуры (execution, evidence owner_type=scenario_run, HumanCheckpoint) закрываются 044. Отдельный 038/validation.md PASS аннулирован как self-contradictory — см. validation.md.

Drift Amendment — MCP Interface (2026-08-24)

  • The authoring chain (compile/validate/resolve/draft-pack/request-save) is callable from external MCP clients (050); the compiler/validator contracts, determinism and needs_context/needs_selector semantics are unchanged.
  • "Agent" throughout this spec now reads as any delegated actor — external MCP client or in-product surface — always behind the same deterministic validator and provenance rules.
  • PII posture: masking selectors in ScreenshotCaptureSpec and response redaction in the VLM path are OPTIONAL (2026-08-24) — all LLM/VLM providers run locally inside the trust perimeter.

Status (2026-09-02): done — реализовано в рамках 050: инструменты и гейты (specs/050-mcp-interface/tasks.md T012–T028 [x]), handoff-поверхность (050 T030–T033), демонтаж чата и сервиса agent/ (050 T040–T041, чекпоинты specs/WORKSTATE-043-047.md).

@{ DashboardScenarioModel.AgentAuthoringWorkspace [C:5] [TYPE ADR]

@BRIEF Normative external-agent/user co-authoring workspace and promotion boundary for scenario authoring. @RELATION DEPENDS_ON -> [ScenarioGraph.ServerOwnedPipeline] @RELATION DISPATCHES -> [ScenarioRegistry.SaveContinuation]

AgentAuthoringWorkspace is a persistent, server-owned session in which an authenticated user and an external MCP agent may jointly author a scenario. It supports draft edits, exploratory Playwright/code-sandbox checks, result inspection, selector/wait/assertion corrections, and new graph proposals. The agent is a full co-author of intent and diagnostics, but never a production authority.

The workspace state enum is: draft, exploring, exploration_failed, exploration_passed, proposal_ready, validation_blocked, awaiting_user_review, save_eligible, pending_approval, candidate, current. State transitions are server-owned, audited, idempotent and CAS-protected; the workspace persists across MCP calls and client reconnects.

AuthoringArtifact is distinct from ExecutionProgram. Authoring artifacts may be source snippets, patches, exploratory traces, screenshots, operation receipts, or diagnostics and retain provenance/ownership. ExecutionProgram is only the validated typed ScenarioGraph/VerificationProgram. A code-backed production provider is a separate future contract, unimplemented here, and cannot be inferred from an exploratory code artifact.

Exploration output can produce only typed action candidates or a graph proposal. Promotion is strictly: sandbox output -> typed candidates/graph proposal -> deterministic 038 compile/validate -> user review of server-computed diff -> 042 server-owned handles/save -> immutable ScenarioRevision. User review is required even when delegated policy permits the agent to prepare or request save; approval gates remain required where policy says so.

The server does not generate arbitrary Playwright code. It provides reviewed templates/registered actions or consumes sandbox traces and typed candidates. Raw code, browser URLs, cookies, secrets, filesystem paths and caller-computed digests are never accepted as graph or production authority. Production ScenarioRun accepts only a promoted, validated 042 revision.

The workspace/promotion state machine is explicit: promote_to_scenario records or updates server-owned validated CompiledScenarioHandle/DraftPackHandle references, or creates a server-stored save request for those handles; it does not activate a revision and does not advance current_revision. request_save creates an immutable candidate revision after the 038 eligibility checks. activate_revision is a separate 042 operation and may advance current_revision only after eligibility, materialization, policy, required approval, and compare-and-set (CAS) checks. Initial creation is the only explicit exception: an eligible initial revision may be created as current. A ScenarioRun must target an explicit promoted revision_id and verified content_hash, never an unactivated candidate or an implicit latest revision.

The sandbox contract is mandatory but its implementation is not claimed: isolated runtime/context; allowlisted origins, APIs and actions; no shell, credential, filesystem or network escape; server-owned time, size and network limits; cancellation with durable operation receipts; artifact ownership; cleanup; and zero production side effects. Unmet security/readiness checks produce validation_blocked or exploration_failed, never a permissive fallback.

Cross-spec stage table

Stage Authority Allowed output Boundary
authoring session 038/050 server persistent workspace state no caller-owned mutable state
exploration isolated sandbox AuthoringArtifact, trace, screenshot, diagnostics no production side effects or authority
proposal 038 typed action candidates / graph proposal no raw code or caller digest
compile/validate 038 deterministic compiler/validator server handles and findings blockers remain explicit
review/diff user principal reviewed proposal decision server computes diff; CAS required
save 042 registry candidate revision, or initial current server handles, idempotency and approval
execution 044 ScenarioRun from promoted revision only raw sandbox output is rejected

@} DashboardScenarioModel.AgentAuthoringWorkspace

#endregion DashboardScenarioModel.Spec


UX REFERENCE — Interaction Narrative

Source: ux_reference.md

#region DashboardScenarioModel.UxReference [C:3] [TYPE ADR] [SEMANTICS ux,reference,scenario,dashboard-testing] @BRIEF UX interaction reference — persona, flows, states, recovery paths, and edge/failure matrix for the DashboardTestScenario graph preview.

Feature Branch: 038-dashboard-scenario-model Created: 2026-07-07 | Reworked: 2026-07-31

1. User Persona & Context

  • Who is the user?: QA engineer or dashboard owner reviewing the agent's proposed test scenario before execution or artifact generation.
  • What is their goal?: Understand what will be tested, which tools each step uses, what parameters are needed, and what cannot be automated.
  • Context: Agent workspace shows a scenario graph produced from dashboard metadata, checklist template, and baselines.

2. The "Happy Path" Narrative

The agent proposes a scenario called "Проверка фильтров, метрик и XLSX выгрузки". The user sees phases, steps, dependencies, tool categories, expected outcomes, and missing parameters. The scenario is a business flow, not a menu of technologies, so the user approves the goal and parameters while the graph records tool selection internally. Validation passes, the pack becomes save-eligible, and 036 registers the draft.

3. Interface Mockups

CLI / Operator Interaction (Agent Tools)

$ scenario compile --objective "verify filters, metric, XLSX export" --case-ids B01,C04,C05
[ ] reading query model + baseline catalog...
✅ scenario compiled: 18 steps, 0 blockers, 2 warnings
   - phases: setup → interact → observe → assert → evidence → report
   - parameters required: test_date, counterparty

Scenario Summary

┌──────────────────────── Proposed scenario ──────────────────────────────────┐
│ Goal: проверить фильтры, метрику, XLSX выгрузку и baseline                  │
│ Steps: 18 | Tools: browser, Superset API, XLSX, assertions, report          │
│ Parameters required: test_date, counterparty                                │
│ Blockers: 0 | Warnings: 2                                                   │
└─────────────────────────────────────────────────────────────────────────────┘

Scenario Graph

[open_dashboard]
      │
      ▼
[apply_filters] ───────► [execute_superset_metric]
      │                              │
      ▼                              ▼
[download_xlsx] ───────► [parse_xlsx_metric]
      │                              │
      └──────────────► [compare_to_baseline] ──► [generate_report]

Step Table

┌────┬────────────────────────────┬──────────────┬───────────────────────────┐
│ №  │ Step                       │ Tool         │ Expected result           │
├────┼────────────────────────────┼──────────────┼───────────────────────────┤
│ 1  │ Открыть дашборд            │ browser      │ dashboard_loaded          │
│ 2  │ Применить фильтры          │ browser      │ filter_state.normalized   │
│ 3  │ Выполнить chart query      │ Superset API │ metric value returned     │
│ 4  │ Скачать XLSX               │ browser      │ xlsx.file                 │
│ 5  │ Сравнить с baseline        │ assertion    │ pass/fail/inconclusive    │
└────┴────────────────────────────┴──────────────┴───────────────────────────┘

States:

  • Idle/Default: No scenario selected; empty state with CTA "Compile from dashboard goal".
  • Loading: Skeleton graph + "compiling…" progress; parameters panel skeleton.
  • Loaded: DAG with phase lanes, step table, coverage list, parameters.
  • Error/Degraded: Blockers grouped per step/case with recovery links.

4. Edge & Failure State Matrix

Semantic Requirement: Every documented failure path maps to @UX_RECOVERY/@UX_FEEDBACK in component contracts and to an error response class in openapi.yaml.

# State Class Trigger Applicable? Visual/Feedback Recovery Test Ownership
NET_01 Network offline navigator.onLine == false ✅ Offline banner; disabled actions Auto-retry on reconnect L2
NET_02 Timeout (>30s) AbortController timeout ✅ Toast + progress countdown Retry (3 attempts); Cancel L1+L2
NET_03 Retry exhaustion 3 failed retries ✅ Persistent banner + manual retry Manual retry L1+L2
VAL_01 Field validation (parameter) On submit ✅ Inline error on parameter control Re-type; clear on focus L1+L2
VAL_02 Cross-field (resolve changes) On submit ✅ Toast + summary banner Fix + re-submit L1+L2
AUTH_01 401 Unauthorized Expired token ✅ Redirect to login; preserve intent Login → redirect back L1
AUTH_02 403 Forbidden Wrong role ✅ Full-page explanation; no approval gate Navigate to dashboard L1+L2
NF_01 404 scenarioId Deleted/unknown scenario ✅ Full-page not found + link to list Navigate to scenario list L1+L2
CONF_01 409 Stale base revision Version mismatch on resolve ✅ Persistent conflict panel: "Scenario changed. Recompile?" Recompile or snapshot diff L1+L2
CONF_02 409 Duplicate pack registration Same revision hash re-posted ✅ Return existing DraftPack (idempotent) Transparent; log event L1
422 422 Unprocessable (compile/resolve) Invalid canonical inputs ✅ Step/field-mapped error detail Correct input + re-submit L1+L2
429 429 Rate Limited + Retry-After Too many compile/validate ✅ Toast + countdown on action Wait Retry-After; disable during countdown L1+L2
5XX 500/502/503 Server Error Backend failure ✅ Error section + retry Retry button L1+L2
STALE Stale baseline fingerprint Baseline updated in 037 ✅ Assertion step warning badge Run baseline discovery; mark pending L1+L2
PARTIAL Partial graph load Some steps failed to compile ✅ Failed steps show placeholder Per-step retry; "Reload all" L1+L2
DUP_01 Duplicate submit (draft-pack) Rapid double-click ✅ Button disabled + spinner Normal completion L2
DUP_02 Navigation interruption (dirty resolution) Route change with unsaved resolution ✅ Persistent unsaved-work panel Stay or discard L2
LARGE Large dataset (>100 steps) Big scenario graph ✅ Virtualized lanes; "Showing 100 of 200" Pagination/refinement L2
EMPTY Empty result (no cases applicable) Dashboard has no mapped cases ✅ Empty state + guidance Accept partial coverage / manual checkpoints L1+L2
MALFORMED Malformed VLM response Backend/LLM bug ✅ Toast with error ID; step inconclusive Re-run analysis; note error ID L1
A11Y Screen reader state announcements State change ✅ (always) aria-live announces compile/validate/load Built into transitions L2
RESP Responsive breakpoint collapse Viewport < 768px ✅ Lanes stack; step table is semantic fallback Built into layout L2

5. Error Experience

Scenario A: Missing Selector

  • System Response: Step marked NEEDS_SELECTOR; executable generation for that step blocked.
  • Recovery: User provides selector hint, converts to human checkpoint, or removes the step.

Scenario B: Stale Baseline

  • System Response: Assertion step shows warning + stale fingerprint category.
  • Recovery: User runs baseline discovery through 037 or marks the check as pending.

Scenario C: Unsupported Checklist Case

  • System Response: Case listed under unsupported/manual-only with rationale.
  • Recovery: User accepts partial coverage or adds manual checkpoint instructions.

Scenario D: Stale Revision on Resolve (409)

  • System Response: Persistent panel: "Scenario changed since your base revision. Recompile to see the latest graph."
  • Recovery: User recompiles; stale edits are rejected, never auto-merged.

6. Tone & Voice

  • Style: Goal-oriented, explicit about confidence and blockers.
  • Terminology: Use "scenario", "step", "tool", "parameter", "baseline ref", "human checkpoint", "preview_only", "save_eligible".

#endregion DashboardScenarioModel.UxReference


CHECKLISTS — Requirements Quality — requirements.md

Source: checklists/requirements.md

Requirements Checklist: 038 Dashboard Scenario Model

Purpose: Validate ScenarioGraph specification quality. Created: 2026-07-07 Feature: specs/038-dashboard-scenario-model/spec.md

Spec Completeness

  • CHK001 User stories cover graph generation, validation, checklist mapping, and parameter/human checkpoint handling.
  • CHK002 Requirements reject direct script generation without validated scenario graph.
  • CHK003 No unresolved [NEEDS CLARIFICATION] markers remain.
  • CHK004 Edge cases include cycles, duplicate refs, stale baselines, unsupported XLSX, and unknown selectors.
  • CHK004a Screenshot capture spec, VLM analysis profile, typed findings, and human disposition contracts are defined and measurable.

Constitution Coverage

  • CHK005 ADR relations include semantic protocol and upstream 036/037 dependencies.
  • CHK006 Decision memory records rejection of low-level user tool selection and raw script generation.
  • CHK007 Requirements support test-driven validation of C3+ model contracts.
  • CHK008 Artifact generation is downstream and not mixed into this model spec.

Readiness for Plan

  • CHK009 Key entities define scenario, steps, refs, parameters, checklist cases, mappings, and validation result.
  • CHK010 Success criteria are measurable with deterministic fixtures and snapshots.
  • CHK011 Scope excludes UI implementation and Superset query engine internals.

Implementation Package

  • CHK012 Research and implementation plan resolve all design decisions.
  • CHK013 Data model, checklist catalog, module, API, JSON Schema, and UX contracts are present.
  • CHK014 All 19 source-checklist cases have an explicit automation classification.
  • CHK015 Traceability maps every functional requirement to contracts, tasks, and tests.
  • CHK016 Tasks use exact repository paths, dependency order, and test-first sequencing.
  • CHK017 Machine-readable contracts, external references, and semantic anchors pass validation.

UX ALTERNATIVES — Design Space Explored

Source: contracts/ux/alternatives.md

#region DashboardScenarioModel.UxAlternatives [C:3] [TYPE ADR] [SEMANTICS ux,alternatives,scenario] @defgroup Ux Design alternatives explored for the DashboardTestScenario graph preview.

Navigation

  • ✅ CHOSEN: Scenario preview embedded in agent workspace with direct link from task run
  • ❌ Rejected: Standalone /scenarios route — 039 owns page routing; 038 supplies DTOs only
  • ❌ Rejected: Modal over task list — loses graph context and deep-linking

Graph Presentation

  • ✅ CHOSEN: Topological phase lanes + linear accessible step table
  • ❌ Rejected: Force-directed graph — non-deterministic layout, hurts scanability, no semantic fallback
  • ❌ Rejected: Text-only step list — loses dependency visibility

Coverage Presentation

  • ✅ CHOSEN: All 19 checklist cases visible with classification + rationale, including unsupported/manual
  • ❌ Rejected: Show only applicable cases — hides coverage gaps and silent drops

Resolution Interaction

  • ✅ CHOSEN: Controls operate on declared parameters/selectors only
  • ❌ Rejected: Free-form graph editing — breaks determinism, ref integrity, and immutable revisions

Stale Revision Handling

  • ✅ CHOSEN: 409 reject + persistent recompile guidance panel
  • ❌ Rejected: Silent auto-merge — changes business intent without review

Draft Pack States

  • ✅ CHOSEN: Explicit preview_only vs save_eligible with blocker list
  • ❌ Rejected: Single "generate" button without state — hides validation failures

VLM Findings Presentation

  • ✅ CHOSEN: Typed findings (severity/region/confidence/provenance) with disposition controls
  • ❌ Rejected: Raw VLM prose as step state — non-reproducible, non-auditable

#endregion DashboardScenarioModel.UxAlternatives


UX DECISIONS — Final Choices

Source: contracts/ux/decisions.md

#region DashboardScenarioModel.UxDecisions [C:3] [TYPE ADR] [SEMANTICS ux,decisions,scenario] @BRIEF Final display and interaction decisions for scenario graph consumers.

  1. Show a business flow first; tool category is secondary evidence.
  2. Always provide a linear accessible step table beside/under the visual DAG.
  3. Keep all 19 checklist cases visible in coverage, including unsupported/manual.
  4. Resolution controls operate on declared parameters/selectors, not arbitrary graph editing.
  5. Revision hash changes are visible; stale edits are rejected, not auto-merged.
  6. Preview-only and save-eligible are explicit pack states.
  7. Technical PDF cases never show SQL as an offered execution path.

#endregion DashboardScenarioModel.UxDecisions


UX API CONTRACT — Endpoints & Shapes

Source: contracts/ux/api-ux.md

#region DashboardScenarioModel.ApiUx [C:3] [TYPE ADR] [SEMANTICS ux,api,scenario,validation] @BRIEF UI mapping for compile, validate, resolve, and draft-pack endpoints.

Operation Optimistic UI Authoritative result
compile inspect/scenario progress active Graph, coverage and findings replace placeholder
validate validation pending errors/warnings/blockers grouped by step/case
resolve affected controls pending only New revision updates impacted statuses, retains other ids
draft-pack generate/validate progress Manifest and 036 DraftArtifact refs

409 stale revision triggers snapshot/recompile guidance; it never merges silently. 422 findings map to fields/steps. 403 renders permission denial without an approval gate.

#endregion DashboardScenarioModel.ApiUx


UX DESIGN — Per-Screen Contracts — scenario-graph-ux.md

Source: contracts/ux/scenario-graph-ux.md

#region DashboardScenarioModel.GraphUx [C:4] [TYPE ADR] [SEMANTICS ux,scenario,graph,coverage] @BRIEF Presentation contract for scenario summary, DAG, steps, parameters, coverage, warnings, and blockers. @RELATION DEPENDS_ON -> [DashboardScenarioModel.UxReference]

Views

  • Summary: objective, step count, tool categories, revision, blockers/warnings.
  • Phase graph: stable topological phase lanes; keyboard-accessible step list is the semantic fallback.
  • Step table: title, tool, expected kind, automation status, dependencies, checklist refs.
  • Coverage: all B01–B09/C01–C07/T01–T03 classifications with rationale.
  • Resolution: parameters/selectors/manual conversion only for declared unresolved targets.

Status Semantics

ready, needs_context, needs_selector, needs_baseline, manual, unsupported, blocked are never reduced to color alone. Unsupported/manual cases remain visible in coverage.

Edge & Failure State Matrix (per screen)

# State Class Applicable? Visual/Feedback Recovery
NET_01 Offline ✅ Offline banner; disabled actions Auto-retry on reconnect
NET_02 Timeout compile/validate ✅ Toast + countdown Retry (3 attempts); Cancel
NET_03 Retry exhausted ✅ Persistent banner + manual retry Manual retry
VAL_01 Parameter field validation ✅ Inline error on control Re-type; clear on focus
VAL_02 Cross-field resolve validation ✅ Toast + summary banner Fix + re-submit
AUTH_01 401 ✅ Redirect to login; preserve intent Login → redirect back
AUTH_02 403 ✅ Full-page explanation; no approval gate Navigate to dashboard
NF_01 404 scenarioId ✅ Full-page not found + link to list Navigate to scenario list
CONF_01 409 stale base revision ✅ Persistent conflict panel: "Scenario changed. Recompile?" Recompile or snapshot diff
CONF_02 409 duplicate pack ✅ Return existing DraftPack (idempotent) Transparent; log event
422 Unprocessable compile/resolve ✅ Step/field-mapped error detail Correct input + re-submit
429 Rate limited ✅ Toast + countdown Wait Retry-After
5XX Server error ✅ Error section + retry Retry button
STALE Stale baseline ✅ Assertion step warning badge Baseline discovery via 037; mark pending
PARTIAL Partial graph load ✅ Failed steps show placeholder Per-step retry; reload all
DUP_01 Duplicate draft-pack submit ✅ Button disabled + spinner Normal completion
DUP_02 Navigation interruption ✅ Persistent unsaved-work panel Stay or discard
LARGE >100 steps ✅ Virtualized lanes Pagination/refinement
EMPTY No applicable cases ✅ Empty state + guidance Accept partial coverage / manual checkpoints
MALFORMED Malformed VLM response ✅ Toast with error ID; step inconclusive Re-run analysis
A11Y Screen reader announcements ✅ (always) aria-live on compile/validate/load Built into transitions
RESP Responsive collapse ✅ (always) Lanes stack; step table semantic fallback Built into layout

Feedback Mechanisms

Trigger Feedback Rationale
Compile submitted Button spinner + "compiling…" progress Long-running deterministic build; user must not double-submit
Validate submitted Validation pending; then grouped findings Findings are deterministic; grouping per step/case aids recovery
Resolve changes Affected controls pending; unrelated ids unchanged Immutable revisions; partial progress only on affected targets
Draft-pack generate generate/validate progress; then manifest Registered 036 artifacts; preview_only/save_eligible explicit
409 stale revision Persistent panel with recompile guidance Never silent merge; user must recompile to see latest graph
VLM finding disposition Audit event + finding status update Typed, auditable human decision

Recovery Paths

From Action To
NEEDS_SELECTOR step Provide selector hint / convert to checkpoint / remove step ready / manual / removed
NEEDS_BASELINE assertion Run 037 baseline discovery / mark pending ready / warning-gated
preview_only pack Resolve blockers / parameters save_eligible
409 stale revision Recompile from latest new revision
422 resolve Correct invalid fields resolved revision

Reactivity

  • Model atoms → component props → DOM (039 renders; 038 supplies DTOs).
  • Compile/validate/resolve results are deterministic server responses; UI holds no derived truth.

UX Tests (minimum: happy, empty, error, edge)

@UX_TEST Given When Then
Fifteen-plus steps render Valid 18-step scenario Load preview No dependency information loss; lanes/table consistent
Cycle not rendered as executable Cycle fixture Load preview Cycle path shown; no executable graph
Technical cases never show SQL T01–T03 dashboard Load preview Superset API or human checkpoint shown
ParameterDefinition resolution updates affected steps only Apply definition change Apply change Unrelated logical_step_id/order unchanged; new content_hash
Preview-only pack explains blockers Invalid/unresolved graph Generate pack All save blockers listed; preview_only state explicit
409 stale revision recovery Stale base revision Resolve Persistent panel with recompile guidance; no silent merge
VLM finding disposition Unresolved finding Confirm/dismiss/inconclusive Typed status; audit event; graph unchanged

#endregion DashboardScenarioModel.GraphUx


PLAN — Implementation Plan

Source: plan.md

Implementation Plan: Dashboard Scenario Model

Branch: 038-dashboard-scenario-model | Date: 2026-07-31 | Spec: spec.md | Status: Reworked per new speckit flow

Summary

Implement a backend Pydantic ScenarioGraph domain and deterministic compiler that maps the 19 normalized PDF checklist cases plus dashboard/baseline capabilities into a reviewable DAG. Validate all refs, cycles, tools, selectors, parameters, and baseline rules before compiling a draft pack through versioned safe templates and registering it with 036. Supply scenario/validation/coverage/parameter/artifact DTOs to 039 (scenario preview UI) and typed capture/VLM/disposition semantics (AGSCN-FR-010..012).

Technical Context

Language/Version: Python 3.13+ (backend), TypeScript DTOs (frontend, consumed by 039) Primary Dependencies: Pydantic 2, FastAPI, deterministic JSON/YAML serialization, existing 036/037 services Storage: Immutable request/graph/draft hashes; draft bytes via 036; no new DB table required for MVP beyond AgentRun event/artifact links Testing: pytest property/unit/contract, JSON Schema validation, golden snapshots; L1 (model, no render) + L2 (component/UX with render) per edge matrix ownership Frontend Architecture: TypeScript-first DTOs in frontend/src/types/ matching Pydantic schemas; UI rendering owned by 039, 038 supplies DTO contracts Performance Goals: validate 100-step graph under 100ms; resolve parameters under 200ms; byte-stable snapshots Constraints: DAG only; no raw baseline numbers; no SQL; no executable LLM output; all checklist cases classified; VLM findings advisory; disposition auditable Scale: 19 source cases, up to 100 steps, 50 parameters, 200 refs

Constitution Check

Principle Result
I. Semantic Contract First PASS — compiler, validator, serializer, pack generator, capture, VLM, disposition contracted (C3–C5)
II. Decision Memory PASS — @RATIONALE/@REJECTED on compiler/validator/pack/VLM (direct-code, silent-repair, raw-VLM, one-size-fits-all rejected)
III. External Orchestrator PASS — reads 037 query models/baselines; never calls Superset directly; writes drafts via 036
IV. Module Discipline PASS — models/catalog/mapping/compiler/validator/serializer/templates/pack separated; no oversized compiler
V. RBAC Enforcement PASS — scenario.compile/resolve/draft/execute scopes per ADR-0005; default-allow forbidden
VI. Svelte 5 Runes Only PASS — UI is DTO-only for 039; no legacy Svelte introduced
VII. Test-Driven for C3+ PASS — invalid fixture matrix and rejected direct-code/SQL/VLM paths written first
VIII. Attention-Optimized PASS — ScenarioGraph.* IDs, shared [SEMANTICS scenario,...], ATTN_1–4 gate applied

Project Structure

Documentation (this feature)

specs/038-dashboard-scenario-model/
├── spec.md              # Reworked: applicability, structured edge cases, Clarifications
├── ux_reference.md      # Reworked: edge/failure state matrix (24 classes)
├── research.md          # Phase 0 decisions (deterministic compiler boundary)
├── data-model.md        # Canonical graph/step/ref/parameter/validation models
├── quickstart.md        # Verification commands and exit gates
├── traceability.md      # REQUIRED RTM: Story → Screen+State → Model → operationId → Contract → Task → Test
├── checklist-catalog.md # Normalized 19-case catalog from the research PDF
├── checklists/          # Requirements-quality checklists
├── contracts/
│   ├── modules.md       # C3+ contracts (compiler, validator, resolver, pack, capture, VLM, disposition)
│   ├── openapi.yaml     # OpenAPI 3.1 — 7 operations, envelopes, RBAC, examples
│   ├── openapi-traceability.md  # operationId → spec → data-model → UX drift map
│   ├── dashboard-test-scenario.schema.json
│   ├── capture-profile.schema.json
│   └── ux/
│       ├── api-ux.md            # UI mapping for compile/validate/resolve/draft-pack
│       ├── scenario-graph-ux.md # Per-screen UX contract + edge/failure matrix
│       ├── decisions.md         # Final UX decisions
│       └── alternatives.md      # Design space explored
├── prototype/
│   ├── index.html       # Interactive HTML prototype (10 states, state switcher, responsive)
│   └── manifest.md      # State coverage + screen↔story traceability
└── validation.md        # Pre-implementation gate (/speckit.validate output)

Source Code (repository root)

backend/src/
├── api/routes/dashboard_scenarios.py   # compile/validate/resolve/draft-pack/capture/vlm/disposition
├── schemas/dashboard_scenario.py       # Pydantic request/response schemas (mirror openapi.yaml)
└── services/dashboard_testing/scenario/
    ├── models.py          # DashboardTestScenario, ScenarioStep, ScenarioParameter, ScenarioRef
    ├── checklist_catalog.py
    ├── capability_mapper.py
    ├── compiler.py
    ├── validator.py
    ├── serializer.py
    ├── resolver.py
    ├── pack_compiler.py
    ├── capture_profile.py
    ├── vlm.py             # typed VlmFinding, provenance, stale-prompt guard
    ├── disposition.py     # typed human disposition, double-disposition guard
    ├── templates/         # step-template catalog (declarative)
    ├── pack_templates/v1/ # scenario.yaml, runner.plan.json, report_template.md, evidence_manifest.json
    └── prompt_templates/v1/  # registered VLM prompt template + hash

backend/tests/services/dashboard_testing/scenario/
backend/tests/api/test_dashboard_scenarios.py
frontend/src/types/                     # DTOs matching Pydantic schemas (consumed by 039)
agent/src/ss_tools/agent/tools.py       # thin compile/validate/resolve/generate tools

Structure Decision: ScenarioGraph is a bounded intermediate boundary between agent intent and generated artifacts; all generation flows through registered templates.

Semantic Contract Guidance

Attention Compliance Gate (MANDATORY — validated)

Rule Check Why
ATTN_1 First anchor line packs [C:N] [TYPE] [SEMANTICS] on ONE line CSA 4× pooling
ATTN_2 IDs hierarchical: ScenarioGraph.Compiler.Compile, ScenarioGraph.Validator.Validate HCA 128×
ATTN_3 All scenario contracts share [SEMANTICS scenario,...] primary keyword DSA grouping
ATTN_4 Contract ≤150 lines, module ≤400 lines Sliding window

Function-Level Contracts (C3+)

Full #region headers for compiler, validator, resolver, serializer, pack_compiler, capture, vlm, disposition are defined in contracts/modules.md with @PRE/@POST/@SIDE_EFFECT/@DATA_CONTRACT/@TEST_EDGE. Cross-stack DTOs (VlmFinding, HumanDisposition, DraftPack) have matching @RELATION edges across the Pydantic ↔ TypeScript boundary.

Decision Memory (Global ADR Continuity)

ADR / Guardrail Present in Plan Propagated to Contracts Propagated to Tasks Verifying Tasks Exist Rejected Path Protected
ADR-0001 module layout ✅ ✅ ✅ T001–T005 ✅
ADR-0002 semantic protocol ✅ ✅ ✅ T036 ✅
ADR-0005 RBAC scopes ✅ ✅ (openapi security) ✅ T032 ✅
037 no-direct-SQL invariant ✅ ✅ (validator) ✅ T015, T034 ✅
Direct LLM-to-code rejected ✅ ✅ (compiler @REJECTED) ✅ T026, T034 ✅
Raw baseline literals rejected ✅ ✅ (validator) ✅ T003, T014 ✅
VLM prose-as-state rejected ✅ ✅ (vlm @REJECTED) ✅ T040–T041 ✅
One-size-fits-all scripts rejected ✅ ✅ (mapper @REJECTED) ✅ T008 ✅

Delivery Phases

  1. Machine schema, Pydantic models, normalized checklist catalog and fixtures.
  2. Capability mapping and deterministic graph compiler.
  3. Full validator and canonical serializer.
  4. Parameter/selector/manual resolution with immutable revisions.
  5. Safe template-based draft-pack compiler and 036 registration.
  6. Screenshot capture, VLM analysis, human disposition (AGSCN-FR-010..012).
  7. REST/agent tools, DTOs for 039, and regression gates (incl. /speckit.validate PASS).

API and Schema

  • contracts/openapi.yaml is the canonical REST contract (7 operations, RBAC scopes, error envelopes).
  • contracts/dashboard-test-scenario.schema.json and contracts/capture-profile.schema.json are the canonical graph/capture interchange contracts.
  • contracts/openapi-traceability.md maps every operationId to spec/data-model/UX.

Traceability

traceability.md is REQUIRED and maps Story/Requirement → UX Screen+State → Screen Model → API operationId → Contract → Task → Test, with explicit N/A rationale and a coverage gate (see traceability.md).

Cross-Spec Boundary

  • Reads 037 query models/baseline summaries; never calls Superset directly.
  • Writes drafts and progress through 036.
  • Supplies scenario, validation, coverage, parameter, and artifact manifest DTOs to 039.
  • Screenshot capture/VLM/disposition depend on 036 Phase 8 (screenshot evidence artifacts) and 037 Phase 7 (visual baseline infrastructure).

Complexity Tracking

No exception planned. Checklist data remains declarative and versioned; do not turn the 19 cases into one large conditional compiler function. VLM findings are advisory — never asserted as deterministic truth.

MVP Runtime Closure (audit 2026-08-07)

Status correction: Факт-чекинг кода показал две runtime-заглушки: vlm.py::_default_submit возвращает [] (нет LLM-вызова), capture.py::dispatch_capture регистрирует артефакты с синтетическим sha256 без вызова Plugin.Service.ScreenshotService. Контракты (contracts/modules.md) и spec FR-010/011 уже усилены требованиями переиспользования; ниже — обязательные runtime-задачи.

Closure tasks (tasks.md Phase 10, T057–T059):

  • T057 — Wire real VLM submit: submit= backed by Plugin.Service.LLMClient + Services.LlmProvider.LLMProviderService (multimodal-required); API route passes real capture bytes. Masking is optional because configured MCP and VLM clients are local; secrets remain excluded.
  • T058 — Wire real capture: dispatch_capture делегирует браузерный захват Plugin.Service.ScreenshotService.capture_dashboard_chunks; sha256 = реальный digest изображения.
  • T059 — E2E: capture → VLM → findings → disposition с реальными байтами и provider-provenance.

Exit rule: evidence/VLM-ветка не демонстрирует реальную работу (пустые findings, синтетические hash) до мерджа T057–T059.


RESEARCH — Technical Decisions

Source: research.md

#region DashboardScenarioModel.Research [C:5] [TYPE ADR] [SEMANTICS research,scenario,graph,checklist,determinism] @BRIEF Phase 0 decisions for deterministic scenario compilation, validation, checklist mapping, and safe draft-pack generation. @RELATION DEPENDS_ON -> [DashboardScenarioModel.Spec] @RELATION DEPENDS_ON -> [SupersetBaselineEngine.DataModel] @RATIONALE The agent may explain and propose intent, but a deterministic compiler/validator must own the graph and generated pack. @REJECTED Direct LLM-to-code or LLM-owned graph identifiers — rejected because refs, safety, and repeatability cannot be guaranteed.

1. Source Checklist Audit

The source research/Чеклист 29.05 (1).pdf is 21 pages and contains 19 cases:

  • basic B01–B09;
  • complex C01–C07;
  • technical T01–T03.

The PDF mixes reusable behavior, FI-0080-specific data, historic pass/fail notes, screenshots, and source-evidence SQL instructions. Historic result text is evidence, not a reusable expected result. SQL instructions become a validated immutable SqlEvidenceSpec when source-mart evidence is required; otherwise they map to Superset API verification or human checkpoints.

2. Deterministic Compiler Boundary

  • Decision: Agent produces a bounded ScenarioIntentDraft (objective, selected case ids, user rationale). Backend compiler combines it with DashboardQueryModel, checklist catalog, baseline summaries, capability model, and parameters.
  • Determinism: Same canonical inputs and compiler/template versions produce byte-stable graph and draft manifest.
  • IDs: Stable ids derive from phase rank, checklist case id, action, and collision-safe ordinal; never random UUIDs inside the canonical graph.
  • Ordering: Topological order with stable phase/tool/action tie-breakers.
  • Alternative rejected: temperature=0 as the only determinism mechanism — model/provider behavior is not a serialization contract.

3. Capability Mapping

Capability tags include native_filters, text_filter, table_filter, pagination, row_edit, bulk_edit, persistence_refresh, time_rollover, xlsx_export, cross_dashboard, superset_metric, dataset_field_read, screenshot, repository_write.

Each ChecklistCase has:

  • required and optional capabilities;
  • parameter requirements;
  • candidate tool-chain templates;
  • automation policy;
  • expected evidence;
  • fallback classification.

Classification is one of automated, human_checkpoint, unsupported, needs_context. No case is silently dropped.

4. Technical PDF Cases Without SQL

T01–T03 refer to SQL Lab/database fields. Mapping rule:

  1. if an authoritative saved dataset/chart result exposes the fields, use superset_api;
  2. otherwise create a human checkpoint describing the required evidence;
  3. never emit SQL text, SQL tool, or an expected value invented from the PDF.

This preserves coverage intent without violating 037.

5. Graph Model

  • DAG phases: setup, interact, observe, assert, evidence, report.
  • Every step declares tool, action, inputs, outputs, expected result, dependencies, automation status, checklist refs, and risk.
  • Inputs reference parameters, context, baseline refs, or earlier output refs.
  • Assertions require baseline_id/candidate_id or structural expectation; raw numeric truth is forbidden.
  • Unknown selectors become NEEDS_SELECTOR; missing values become NEEDS_CONTEXT; stale/missing baseline becomes NEEDS_BASELINE.

6. Validator

Validation is pure and returns all findings:

  • schema/type/enum violations;
  • cycles and dependency order;
  • missing/duplicate refs;
  • parameter type/resolution;
  • tool/action compatibility;
  • selector and baseline requirements;
  • unsupported/dangerous actions;
  • checklist coverage and unreachable steps;
  • direct SQL/embedded baseline literals.

Errors block draft-pack compilation. Warnings may allow preview but are repeated in 036 save gate.

7. Safe Draft-Pack Compiler

  • Decision: Compile a valid graph through versioned repository-owned templates.
  • MVP outputs: scenario.yaml, runner.plan.json, report_template.md, evidence_manifest.json, and optional template-based browser/XLSX assertion modules.
  • Invariant: LLM text may populate bounded descriptions only; it cannot provide executable code bodies, paths, imports, or shell commands.
  • Drafts: Registered via 036 outside target repository.
  • Alternative rejected: Agent writes scripts and later validator scans them — unsafe constructs are already materialized and scanning is incomplete.

8. Resolution Semantics

Editing ParameterDefinitions/defaults may produce a new executable revision with parent_revision_id; supplying runtime ParameterBindings occurs only at 044 launch and never changes content identity. Unrelated steps retain logical_step_id and serialization. Manual conversion and selector hints are explicit resolution operations with audit reason.

9. LLM Verification Tooling — Module Reuse

Capture and VLM analysis are not new implementations; they are thin orchestration over the existing llm_analysis plugin infrastructure.

Concern Reused module (real contract ID) File
Playwright capture (login/CSRF, tab traversal, stabilization, chunk capture, webp) Plugin.Service.ScreenshotService backend/src/plugins/llm_analysis/service.py
VLM/LLM provider submission (multimodal, JSON mode, retries, image optimization, payload sizing) Plugin.Service.LLMClient backend/src/plugins/llm_analysis/service.py
Provider resolution (multimodal check, encrypted API key, token config) Services.LlmProvider.LLMProviderService backend/src/services/llm_provider.py
Artifact registration + masking Services.AgentRuns.Evidence (register_screenshot_draft / register_masked_derivative) backend/src/services/agent_runs/evidence.py
Redaction of logs/raw responses Plugin.Service.RedactionService backend/src/plugins/llm_analysis/service.py

Decision: ScenarioGraph.Capture.Dispatch (038) registers evidence through the 036 bridge AND delegates the actual browser capture to Plugin.Service.ScreenshotService.capture_dashboard_chunks at runtime. Registering artifacts without real capture bytes is a stub, not a capture. Decision: ScenarioGraph.Vlm.Analyze (038) parses/validates typed VlmFinding[]; the actual provider call is injected as submit= and MUST be backed by Plugin.Service.LLMClient + Services.LlmProvider.LLMProviderService. The _default_submit placeholder (returns empty findings) is a test seam only and must be replaced by a real provider submit in production wiring. Decision: Provider selection for VLM reuses the multimodal-required validation semantics of Services.ValidationService.ValidationTaskService._validate_provider; a non-multimodal provider is rejected before analysis. Alternative rejected: Building a second VlmProviderClient or a second Playwright capture path inside dashboard_testing/ — both would fork retry/SSL/masking behavior from the established plugin and double maintenance. @REJECTED: VLM output as deterministic assertion truth; capture without Plugin.Service.ScreenshotService; provider calls without Plugin.Service.LLMClient/Services.LlmProvider.LLMProviderService.

#endregion DashboardScenarioModel.Research


DATA MODEL — Entities & Relations

Source: data-model.md

#region DashboardScenarioModel.DataModel [C:5] [TYPE ADR] [SEMANTICS data-model,scenario,graph,step,validation] @BRIEF Canonical immutable Verification Program IR: scenario_key, content_hash, navigation/evidence/transform/assertion/semantic programs, typed steps, parameters, capability and authoring-validation models. Runtime state is owned by 044. @RELATION DEPENDS_ON -> [DashboardScenarioModel.Research] @RELATION DEPENDS_ON -> [ScenarioExecution.DataModel] @RATIONALE 038 is the clean IR/compiler layer; identity/revision/runtime are reconciled with 042-047. Semantic identity (scenario_key) is separated from entity identity (scenario_id UUID, assigned by 042) and content identity (content_hash). @REJECTED scenario_id as a deterministic slug — rejected because clones of the same dashboard+objective would collide; entity identity MUST be a UUID assigned at Registry persistence. @REJECTED revision_hash/parent_revision_hash — rejected in favor of revision_id (UUID) + content_hash; the compiler does not fabricate identity. @REJECTED Runtime state (VlmFinding dispositions, HumanCheckpoint) inside the canonical scenario — rejected because an immutable revision must not embed runtime observations; VlmAnalysisSpec describes WHAT, VlmFinding/Disposition are 044 runtime. @REJECTED Direct SQL forbidden as an absolute principle — rejected because source-mart evidence is necessary for real checks; only validated, immutable, read-only SqlEvidenceSpec is permitted. @REJECTED Runtime LLM SQL/DSL rewrite — rejected because a ScenarioRun must execute precisely the content-addressed program saved in its revision.

DashboardTestScenario — compiled definition (no persistence identity)

Required:

  • schema_version, compiler_version, template_version;
  • scenario_key (semantic = dashboard key + normalized objective slug; human/domain identity);
  • content_hash (SHA-256 of the canonical executable graph only — timestamps/display-only excluded);
  • dashboard_context and objective;
  • input_fingerprints: query model, checklist catalog, baseline_version, parameters;
  • parameters, phases, steps, and required verification_program (navigation_program, evidence_program, transformation_program, assertion_program, semantic_evaluation_program);
  • outputs and artifact_plan;
  • checklist_coverage;
  • warnings, blockers, risk_summary.

No scenario_id / revision_id here. The compiler emits scenario_key + content_hash. scenario_id (UUID) and revision_id (UUID) are assigned by the 042 registry at Save (CreateScenario). Provenance is orthogonal (see CompileProvenance below).

CompileProvenance

Compiler input MUST NOT require an AgentRun. Provenance is passed separately:

  • source_type: agent_run | editor | migration | api
  • source_id: nullable

ParameterDefinition

Fields: name, label, type, required, default (optional), source, validation, affected_logical_step_ids. Supported types: string, integer, decimal, boolean, date, datetime, enum, string_list, baseline_choice, selector_hint.

This immutable definition contains no resolved value or runtime status. Those belong exclusively to 044 ParameterBinding, so a value such as test_date=2026-08-10 never changes a scenario content_hash.

VerificationProgram — first-class executable IR

VerificationProgram is the canonical runtime program embedded in every saved scenario: navigation declares registered browser/API actions; evidence declares source/chart/XLSX observations; transformations declare bounded deterministic TransformSpec DSL; assertions declare ComparisonSpec/AssertionSpec; semantic evaluation declares only explicitly necessary AgentEvaluationSpec. All program entries are ref-addressed by logical_step_id, version-pinned and canonicalized into content_hash.

SqlEvidenceSpec is a small source-evidence statement: snippet_id, logical_step_id, connection_ref, database_identity, sql_template, sql_hash, parameter_definitions[], expected_output_schema, relation_refs[], execution_limits, generation_provenance, validation_result. It is an authoring artifact, saved only after the SQL compilation gate. Runtime may bind declared parameters but MUST execute exactly the pinned template through the Superset SQL Lab adapter; it cannot rewrite SQL, relations, joins, projections or filters.

TransformSpec is a bounded DSL only: select, filter, rename, cast, join, group_by, sum, count, distinct, coalesce, normalize_string, normalize_date, difference, ratio, tolerance_compare. ComparisonSpec/AssertionSpec model numeric tolerance, rows/columns/sets, aggregates, maps, null/fill and cross-dashboard checks. Arbitrary Python, shell and executable code are forbidden.

AgentEvaluationSpec is the exception for declared semantic/visual/ambiguous evaluation. It pins provider/model/prompt/input manifest/evidence/tool allowlist/output schema and DecisionPolicy. It is not an orchestration instruction and cannot mutate program content.

ScenarioStep — with logical identity

Field Rule
logical_step_id UUID, immutable across revisions/edits (analytics + compare key)
step_key Semantic stable identifier (phase + case + action slug) — readable, not an identity
position Derived ordering within the graph (mutable; NOT an identity)
step_content_hash SHA-256 of the step's executable content (mutable across edits)
phase setup/interact/observe/assert/evidence/report
title/description Bounded display text
tool browser, superset_api, sql_evidence, transform, assertion, agent_evaluation, xlsx, screenshot, report, artifact, human
action Must be allowed by the version-pinned ActionRegistry for its tool; never free-form runtime dispatch
inputs Typed refs only
outputs Unique ScenarioRef declarations
expected Structural expectation or baseline ref, never raw numeric truth
depends_on Existing logical_step_ids; DAG
automation_status ready, needs_context, needs_selector, needs_baseline, manual, unsupported, blocked
checklist_case_ids Known catalog ids
risk READ_ONLY, UI_INTERACTION, TEST_DATA_MUTATION, EXTERNAL_MUTATION, DANGEROUS_MUTATION, human
mutation_contract Mandatory for every mutating action: safe_test_fixture_id, scope, allowed environment, affected keys, cleanup/reconciliation, side-effect key
capture_spec Required when tool=screenshot; null otherwise
vlm_analysis_spec Required when tool=assertion and input is screenshot; null otherwise

The compiler assigns a UUID when it creates an initial graph. An editor/migration MUST carry an existing logical_step_id forward for the same logical step; it MUST mint a new UUID only for a genuinely new step. step_key and position are never identity inputs. Analytics (047) and comparison (045) key on it.

ActionRegistry and mutation safety

ActionRegistry(version) is the canonical 038 catalog of every allowed {tool, action}. Each entry declares typed inputs/outputs, allowed risk, timeout, idempotency, retry safety, mutation policy and version. 044 persists the registry's version/hash and the exact immutable ActionExecutionDescriptor in its RunnerPlan, then resolves executors and lease/recovery/retry policy from that descriptor only. Unknown {tool, action}, a missing/altered descriptor, or a version/hash mismatch is a validation error before I/O, never a tool-default fallback.

Mutating actions require mutation_contract. Policy: DANGEROUS_MUTATION is never automated; mutating browser steps in PROD are prohibited; TEST_DATA_MUTATION is permitted only in an explicitly listed non-PROD environment against a named safe fixture, with bounded affected record keys, side-effect idempotency and rollback/reconciliation. Missing safety context maps the checklist case to needs_context or HumanCheckpoint, never automatic dispatch.

ScreenshotCaptureSpec

Field Type Rule
target string tab or viewport
tab_identifier string/null Required when target=tab
viewport {width, height} Required; default 1920×1200
readiness string canvas_stabilized, network_idle, fixed_wait
readiness_timeout_ms integer Default 15000
mask_selectors string[] OPTIONAL CSS selectors applied before capture; empty when no masking is requested; never a validation or VLM-submission prerequisite
method string cdp, full_page, region

This is the SPEC of capture; execution is owned by 044 (CaptureService over ScreenshotService, artifacts owner_type=scenario_run).

VlmAnalysisSpec

Field Type Rule
profile_id string Registered VLM analysis profile
provider_id string Provider identifier
model_id string Model identifier
prompt_template_id string Versioned prompt template
prompt_version string Semver
prompt_template_hash sha256 SHA-256 of prompt content
confidence_threshold float Minimum confidence to auto-flag (default 0.7)

This is the SPEC (what to analyze). Runtime VlmFinding/HumanCheckpointDisposition are 044 entities.

ScenarioRef

Namespaces:

  • context.dashboard., context.query_model.;
  • param.{name};
  • step.{logical_step_id}.{output};
  • baseline.{baseline_id} or candidate.{candidate_id};
  • artifact.{artifact_key}.

Every consumed step ref has exactly one producer. Parameter/context/baseline refs are externally resolved roots.

ChecklistCase and CapabilityMapping

ChecklistCase holds id, section, goal, reusable expected semantics, required/optional capabilities, parameters, tool templates, evidence, fallback, and source page. Historic PDF outcome is stored only as source_note.

CapabilityMapping holds case_id, classification, matched/missing capabilities, selected template, rationale, and resulting step keys.

AuthoringValidation and RunPreflight

AuthoringValidation determines whether an immutable ScenarioRevision/DraftPack may be saved. Fields: valid, errors, warnings, blockers, coverage, topological_order, unresolved_selectors, unresolved_baselines, parameter_definitions_valid, verification_program_valid, sql_compilation_results, graph_hash. Required ParameterDefinitions without a default are valid authoring inputs and appear only as unbound_parameter_names warnings.

RunPreflight is owned by 044: it resolves ParameterBinding, target, baselines, RLS/principal and runtime policy for one launch. An unbound required parameter makes only that launch run_ineligible; it never changes content_hash or blocks saving a reusable scenario.

An agent may construct coverage, resolve missing context with the analyst, and create a WorkingDraft or executable revision under delegated policy. It never bypasses AuthoringValidation, ActionRegistry(version), mutation contracts, canonicalization, or immutable content hashing; those are deterministic authorities.

ArtifactPlan and DraftPack (authoring-only)

ArtifactPlan entries declare artifact_key, kind, relative_path_template, template_id, input_refs, required, and generation blockers.

DraftPack is the authoring output (generated files previewed before Save): contains scenario_key, content_hash, template version, manifest, authoring DraftArtifact refs, validation summary, and warnings. Status preview_only or save_eligible. Validation error, NEEDS_SELECTOR, forbidden action, or unresolved baseline makes it preview_only; unbound required runtime parameters do not.

Runtime evidence is NOT a DraftPack. During execution (044), screenshots/report/xlsx are Artifact(owner_type=scenario_run, ...), never authoring DraftArtifacts (reconciliation step 7). DraftPack is consumed by 042 CreateScenario to register the revision.

Deterministic Identity

  • scenario_key = dashboard key + normalized objective slug (semantic; may repeat across clones);
  • content_hash = SHA-256 of canonical executable graph (steps/params/refs/expected), timestamps and display-only excluded;
  • logical_step_id = immutable UUID, minted on initial graph creation and carried forward by edits/migrations;
  • step_content_hash = SHA-256 of a step's executable content;
  • step_key = phase + case id + action slug (readable, NOT an identity; no ordinal).

scenario_id (UUID) and revision_id (UUID) are assigned by 042 at persistence, never by the compiler.

Forbidden Content

The schema/validator rejects:

  • arbitrary/unvalidated SQL, query text outside SqlEvidenceSpec, shell command or executable code body;
  • raw numeric expected value for a metric assertion;
  • absolute or parent-traversing artifact paths;
  • unregistered tool/action;
  • implicit dependency or duplicate producer;
  • runtime observations (VlmFinding dispositions, HumanCheckpoint) inside a canonical revision.

@{ DashboardScenarioModel.AuthoringWorkspaceModel [C:5] [TYPE ADR]

@BRIEF Persistent co-authoring session, sandbox artifact and promotion types. @RELATION DEPENDS_ON -> [DashboardScenarioModel.DataModel]

AgentAuthoringWorkspace fields are workspace_id, scenario_id?, session_status, owner_principal, agent_principal?, base_revision_id?, base_content_hash?, cas_version, proposal_id?, exploration_ids[], artifact_ids[], created_at, updated_at, and expires_at?. session_status is the exact enum draft|exploring|exploration_failed|exploration_passed|proposal_ready|validation_blocked|awaiting_user_review|save_eligible|pending_approval|candidate|current.

AuthoringArtifact fields are artifact_id, workspace_id, kind (source|patch|trace|screenshot|diagnostic|operation_receipt), content_ref, content_digest, media_type, producer_principal, operation_id, created_at, and ownership_receipt. Content refs are server-owned and bounded; source/patch artifacts are review material only. ExplorationResult binds traces/screenshots/diagnostics to sandbox limits, origin/action policy, result status and cancellation/receipt state.

ExecutionProgram is the typed 038 ScenarioGraph/VerificationProgram plus compiler/schema/template fingerprints. TypedActionCandidate and GraphProposal may reference authoring artifacts, but contain only registered actions, typed inputs, selector/wait/assertion changes, dependencies and expected base hash. They do not contain executable authority.

Promotion MUST follow sandbox output -> TypedActionCandidate|GraphProposal -> Compile -> Validate -> server-computed diff -> user review/CAS -> 042 save handles -> ScenarioRevision. Raw source, Playwright code, URL, cookie, secret, filesystem path or caller digest is rejected at every conversion boundary. Code-backed production execution is a separate, unimplemented contract and is not an ExecutionProgram.

@} DashboardScenarioModel.AuthoringWorkspaceModel

#endregion DashboardScenarioModel.DataModel


CONTRACTS — Module & Function Contracts

Source: contracts/modules.md

#region DashboardScenarioModel.Modules [C:5] [TYPE ADR] [SEMANTICS contracts,scenario,graph,validator,compiler] @BRIEF C3+ contracts for checklist mapping, deterministic graph compilation, validation, resolution, serialization, and safe draft packs. Runtime capture/VLM/disposition are OWNED by 044 — this module defines VlmAnalysisSpec and the compiler/validator boundary only. @RATIONALE The scenario graph is the reviewable intermediate boundary between agent intent and generated artifacts; it centralizes safety, determinism, and capability coverage. @REJECTED Direct agent-to-script generation — rejected because missing refs, unsafe actions, and baseline truth would be discovered only after artifact generation or runtime. @REJECTED 038 owning runtime execution (capture/VLM/disposition/artifacts) — moved to 044; 038 is the clean IR/compiler layer, 044 is the single execution owner. @RELATION DEPENDS_ON -> [DashboardScenarioModel.DataModel] @RELATION DEPENDS_ON -> [DashboardScenarioModel.ChecklistCatalog] @RELATION DEPENDS_ON -> [SupersetBaselineEngine.Modules] @RELATION DEPENDS_ON -> [ScenarioExecution.Modules]

#region ScenarioGraph.Api [C:4] [TYPE Module] [SEMANTICS scenario,api,compile,validate]

@defgroup ScenarioGraph REST surface for compile, validate, resolve, and draft-pack operations.

@LAYER API

@RELATION DEPENDS_ON -> [ScenarioGraph.Compiler.Compile]

@RELATION DEPENDS_ON -> [ScenarioGraph.Validator.Validate]

@RELATION DEPENDS_ON -> [ScenarioGraph.Resolver.Resolve]

@RELATION DEPENDS_ON -> [ScenarioGraph.PackCompiler]

@RATIONALE A thin REST surface exposes the deterministic scenario operations to 039 and agent tools without leaking compiler internals.

@REJECTED Exposing raw Pydantic models over the API — request schemas must forbid arbitrary code/unsafe paths and require SqlEvidenceSpec compilation.

@INVARIANT Request schemas permit only validated immutable SqlEvidenceSpec; executable code, raw baseline values and local paths are forbidden.

#endregion ScenarioGraph.Api

#region ScenarioGraph.Catalog.Load [C:4] [TYPE Function] [SEMANTICS scenario,checklist,catalog,version]

@ingroup ScenarioGraph

@BRIEF Load the versioned 19-case declarative catalog and validate ids/capability/template references.

@PRE Bundled catalog version is supported.

@POST Returns exactly B01–B09, C01–C07, T01–T03 in stable order.

@SIDE_EFFECT Bounded package-resource read.

@DATA_CONTRACT CatalogResource -> ChecklistCase[19]

@INVARIANT Historic PDF outcomes are source notes, not expected values.

@RATIONALE Versioned declarative catalog keeps the 19 PDF cases reusable and auditable across dashboards.

@REJECTED Embedding the checklist as Python conditionals — mixed intent/data makes coverage unverifiable.

@TEST_EDGE missing_case -> startup/catalog validation failure.

@TEST_EDGE source-mart evidence case -> validated SqlEvidenceSpec or needs_context; unsafe SQL -> rejected.

#endregion ScenarioGraph.Catalog.Load

#region ScenarioGraph.CapabilityMapper.Map [C:5] [TYPE Function] [SEMANTICS scenario,capability,mapping,coverage]

@ingroup ScenarioGraph

@BRIEF Classify every checklist case for a dashboard and select one allowed step template.

@PRE Query model, baseline summary, and capability registry validate.

@POST Every catalog case is automated, human_checkpoint, unsupported, or needs_context with rationale.

@SIDE_EFFECT None.

@DATA_CONTRACT ChecklistCase[] + DashboardCapabilities -> CapabilityMapping[]

@RATIONALE Capability mapping keeps checklist intent reusable while allowing each dashboard to receive only safe, applicable step templates.

@REJECTED One-size-fits-all scripts and user-facing low-level tool selection — rejected because capabilities, safety, and available evidence vary per dashboard.

@INVARIANT No case is dropped and no tool is selected outside its registered capabilities.

@TEST_EDGE xlsx_unavailable -> C04–C06 manual/unsupported with rationale.

@TEST_EDGE technical_without_dataset_fields -> human checkpoint, no SQL.

#endregion ScenarioGraph.CapabilityMapper.Map

#region ScenarioGraph.Compiler.Compile [C:5] [TYPE Function] [SEMANTICS scenario,compiler,deterministic,dag]

@ingroup ScenarioGraph

@BRIEF Compile canonical inputs and mappings into a stable dashboard-specific DAG.

@PRE Intent, query model, catalog, baseline summary, and parameters have valid fingerprints; provenance passed separately (no mandatory agent_run_id).

@POST Same canonical inputs/compiler version yield byte-identical graph and stable keys/order; emits scenario_key + content_hash (NO scenario_id/revision_id).

@SIDE_EFFECT None.

@DATA_CONTRACT CompileInput (scenario_key basis + ChangeRequestContext + objective + query model + checklist + ParameterDefinitions + capabilities + baselines) + CompileProvenance -> DashboardTestScenario (scenario_key + content_hash + VerificationProgram)

@INVARIANT Steps consume only context/parameter/baseline/earlier-step refs; initial graph creation mints logical_step_id and edits/migrations carry it forward.

@TEST_INVARIANT Deterministic_Graph -> VERIFIED_BY: repeated_compile, shuffled_input_order.

@TEST_EDGE missing_selector -> NEEDS_SELECTOR step and save blocker.

@TEST_EDGE missing_baseline -> NEEDS_BASELINE; no embedded numeric truth.

@RATIONALE Rule/template compilation makes the agent a planner/explainer, not an executable-code generator.

@REJECTED LLM-generated ids/dependencies/code — non-deterministic and unsafe.

@REJECTED Compiler emitting scenario_id/revision_id — identity is assigned by 042 at persistence.

#endregion ScenarioGraph.Compiler.Compile

#region ScenarioGraph.Validator.Validate [C:5] [TYPE Function] [SEMANTICS scenario,validator,graph,safety]

@ingroup ScenarioGraph

@BRIEF Return complete deterministic findings for schema, DAG, refs, parameters, baselines, tools, safety, and coverage.

@PRE Candidate graph parses against supported schema version.

@POST Valid is true only with zero errors/blockers; findings are stably ordered and actionable.

@SIDE_EFFECT None.

@DATA_CONTRACT DashboardTestScenario -> ScenarioValidationResult

@RATIONALE Validation is a hard safety boundary between agent-produced intent and artifact generation; deterministic findings give the user a recoverable explanation instead of a runtime surprise.

@REJECTED Silent graph repair or best-effort artifact generation — rejected because auto-fixing refs, cycles, or unsafe actions can change business intent without review.

@INVARIANT Cycles, missing/duplicate refs, unregistered tools/actions, arbitrary or unsafe SQL, raw metric truth and path traversal block compilation. SqlEvidenceSpec passes only its AST/policy/preview compilation gate.

@TEST_EDGE cycle -> error contains cycle path.

@TEST_EDGE duplicate_output -> both producer ids reported.

@TEST_EDGE raw_metric_expected -> forbidden baseline literal error.

@TEST_EDGE unreachable_step -> warning/error according to required coverage.

#endregion ScenarioGraph.Validator.Validate

#region ScenarioGraph.Resolver.Resolve [C:4] [TYPE Function] [SEMANTICS scenario,resolve,parameter,revision]

@ingroup ScenarioGraph

@BRIEF Apply typed ParameterDefinition/selector/manual resolutions and emit a new compiled definition (content_hash changes; identity unchanged).

@PRE Base content_hash matches; changes target declared unresolved items.

@POST Unrelated step keys/order remain unchanged; new content_hash links to base content_hash; logical_step_id stable.

@SIDE_EFFECT None.

@SIDE_EFFECT Logging (REASON/REFLECT markers required around emission).

@RATIONALE Immutable content hashes preserve auditability and byte-stable determinism for downstream pack compilation.

@REJECTED In-place graph mutation — destroys content-hash lineage.

@DATA_CONTRACT ResolveRequest + BaseScenario -> DashboardTestScenario (new content_hash)

@TEST_EDGE stale_base_content_hash -> 409.

@TEST_EDGE invalid_parameter_type -> 422.

@TEST_EDGE unrelated_graph_change -> invariant failure.

#endregion ScenarioGraph.Resolver.Resolve

#region ScenarioGraph.Serializer.Canonical [C:4] [TYPE Function] [SEMANTICS scenario,serialize,json,yaml]

@ingroup ScenarioGraph

@BRIEF Serialize graph to canonical JSON/YAML and compute content_hash.

@POST Key/order/decimal/date/newline rules are stable across runs; JSON and YAML represent equal domain data; content_hash excludes volatile fields.

@SIDE_EFFECT None.

@SIDE_EFFECT Logging (REASON marker before canonicalization; REFLECT with hash after).

@RATIONALE Canonical serialization is the content-identity boundary; volatile display fields must never enter the hash.

@REJECTED Pretty-printed human-first serialization — non-deterministic key ordering breaks byte-stable snapshots.

@DATA_CONTRACT DashboardTestScenario -> CanonicalBytes + content_hash

@TEST_EDGE shuffled_dicts -> identical bytes.

@TEST_EDGE timestamp_display_field -> excluded from content identity.

#endregion ScenarioGraph.Serializer.Canonical

#region ScenarioGraph.PackCompiler.Generate [C:5] [TYPE Function] [SEMANTICS scenario,artifact,template,draft]

@ingroup ScenarioGraph

@BRIEF Generate a preview-only or save-eligible draft pack MANIFEST using registered versioned templates (pure computation).

@PRE Scenario validation result available; template ids registered; target paths safe.

@POST Outputs match ArtifactPlan, contain no LLM executable bodies; returns status save_eligible|preview_only and the public manifest (artifact keys, template ids, relative paths, size_bytes). Persists nothing; artifacts stays empty until ScenarioGraph.PackCompiler.Register036 runs at a permission-gated boundary.

@SIDE_EFFECT Logging only (REASON before generation; REFLECT with status after). Templates are rendered in-memory for digest/size computation; rendered bytes are never persisted by this function.

@DATA_CONTRACT DashboardTestScenario + ValidationResult -> DraftPackManifest (authoring-only; runtime evidence is Artifact(owner_type=scenario_run) under 044, NOT DraftPack)

@INVARIANT Errors/unresolved required inputs make pack preview_only; direct code/path input is impossible.

@RELATION DISPATCHES -> [ScenarioGraph.PackCompiler.Register036]

@TEST_INVARIANT No_LLM_To_Code -> VERIFIED_BY: injected_code_field, template_registry_only.

@TEST_EDGE unknown_template -> blocked.

@TEST_EDGE path_traversal -> blocked during in-memory render, before any registration boundary can run.

@RATIONALE Versioned templates make generated behavior reviewable and reproducible.

@REJECTED Generate arbitrary Playwright/Python code then scan it — scanners cannot prove semantic safety.

@REJECTED Runtime evidence produced under 044 being classified as an authoring DraftArtifact — execution artifacts use owner_type=scenario_run.

@REJECTED Fusing manifest generation with durable 036 registration inside this function (pre-2026-09-06 contract text claimed an AgentRuns.Artifacts.Register side effect that the code never had) — preview_only packs MUST NOT persist; registration lives at the separate permission-gated boundary (REST api_draft_pack; MCP minting surface per 050 task T029f).

#endregion ScenarioGraph.PackCompiler.Generate

@{ ScenarioGraph.ServerOwnedPipeline [C:5] [TYPE ADR]

@BRIEF Server-owned handles and unresolved-context semantics for the 038 authoring stages. @RELATION DEPENDS_ON -> [ScenarioGraph.Compiler.Compile] @RELATION DEPENDS_ON -> [ScenarioGraph.Validator.Validate] @RELATION DEPENDS_ON -> [ScenarioGraph.Resolver.Resolve] @RELATION DEPENDS_ON -> [ScenarioGraph.PackCompiler.Generate]

038 owns only deterministic authoring outputs. Compile mints an opaque CompiledScenarioHandle whose server-computed content_hash covers canonical executable Verification Program content and excludes runtime bindings, display fields, timestamps, scenario_id, and revision_id. Validate mints a result handle bound to that exact compiled hash and validator/schema versions. The pack compiler mints a DraftPackHandle only from registered templates; its server digest covers the rendered manifest/bytes, and save_eligible is impossible when validation errors or unresolved required markers remain. Caller graph, caller digest, local path, or runner.plan.json is never accepted as authority.

Inspection and resolution preserve unresolved markers. needs_context means a required dashboard/query/change-request fact is absent; needs_selector means the browser target/action selector is unknown; needs_baseline means an approved baseline reference is absent or stale. The compiler never guesses or silently defaults any of them. These markers may produce preview/working-draft output, but only 042 may persist a revision and 044 may enforce launch-time bindings in RunPreflight.

Handle persistence (amendment 2026-09-06 — normative target, SPECIFIED-PENDING)

Implementation status (audit 2026-09-06, Doc.Adr.ADR0023): the deterministic cores (Compile/Validate/Resolve/Canonical serialization) are IMPLEMENTED as pure functions with @SIDE_EFFECT None. The durable handle layer below is NOT YET IMPLEMENTED: today the REST/MCP boundaries return graphs in-band, and ScenarioRegistry.Create.* verifies caller-composed compile:{run_id}:{digest} strings against 036 DraftArtifact rows as a transitional stand-in.

  1. Minting boundaries. CompiledScenarioHandle, ValidationResultHandle and DraftPackHandle are minted ONLY by the persisted boundaries that wrap the pure functions (REST api_compile_scenario/api_validate_scenario/ api_resolve_scenario/api_draft_pack in api/routes/dashboard_testing/scenario.py and their MCP tool equivalents per 050). The pure functions keep @SIDE_EFFECT None; minting never moves inside them.
  2. Storage. Three immutable append-only PostgreSQL records owned by the 038 service layer:
    • CompiledScenarioHandle { handle_id (UUID), owner_principal, agent_run_id?, dashboard_id, canonical_bytes_ref, content_hash, compiler_version, created_at, consumed_by_revision_id? }
    • ValidationResultHandle { result_id (UUID), compiled_handle_id, validator_version, schema_version, result_digest, valid, blockers_count, created_at }
    • DraftPackHandle { draft_pack_id (UUID), compiled_handle_id, owner_principal, agent_run_id, digest, template_version, status (save_eligible|preview_only), created_at, consumed_by_revision_id? } canonical_bytes_ref points at the persisted canonical DashboardTestScenario serialization (ScenarioGraph.Models.CanonicalBytes) in the server content store; content_hash is its SHA-256. Canonical bytes — NOT the lossy pack templates (scenario.yaml/runner.plan.json remain reference-only per 050 spec "Handle and Digest Rules") — are the authority 042 materializes into ScenarioRevision.graph_snapshot (see ScenarioRegistry.DataModel, amendment 2026-09-06).
  3. Binding. DraftPackHandle.compiled_handle_id MUST reference a handle with identical content_hash and dashboard_id, same owner_principal. ValidationResultHandle is valid only for its exact compiled hash plus validator/schema versions; a stale validation never authorizes a pack or save. status=save_eligible additionally requires a stored valid ValidationResultHandle for the bound compiled hash and zero unresolved required markers (needs_context/needs_selector/needs_baseline).
  4. Consumption and GC. A handle is consumed by exactly one CreateScenario/CreateInitial/request_save transaction, which records consumed_by_revision_id. Consumed handles are immutable audit rows and can never authorize a second registry creation. Unconsumed handles are garbage-collected together with their owning AgentRun under the existing 036 draft-retention policy; no separate TTL is invented.
  5. Rejection rules. Caller-composed handle strings, caller digests, unbound compiled/draft pairs, preview_only packs, stale validation results and double consumption are typed rejections (409-class) with zero partial rows/outbox/gates. The transitional string-format check in ScenarioRegistry.Create.Register is removed when this layer lands (050 T029d/T029f).
  6. Inspect stage status (resolved as hybrid, T029h 2026-09-06). No persisted InspectionContextHandle is minted. Instead: the MCP read tool inspect_dashboard_context exposes the live server resolver (BaselineEngine.QueryModel.Inspect, same service as REST GET /api/dashboard-testing/query-model) returning the full authoritative DashboardQueryModel for the agent to echo into compile; the register_draft_pack boundary evaluates ScenarioGraph.ContextAuthority: the client-carried dashboard_context.query is parsed as a strict DashboardQueryModel and its fingerprint is RECOMPUTED server-side against a live inspection — a claimed fingerprint field is never trusted. Sentinel / degraded authoritative fingerprints ("", sha256:error) never verify; unconfigured or unreachable environments fail OPEN to unverified (recorded on DraftPackHandle.context_authority, materialized into graph_snapshot by 042 create); falsifiable model-shape claims on live environments reject typed (CONTEXT_FINGERPRINT_MISMATCH, CONTEXT_QUERY_MODEL_REQUIRED, CONTEXT_ENVIRONMENT_MISMATCH) with zero rows. ScenarioExecution.Runner.Start refuses PROD dispatch on an explicit non-verified marker (CONTEXT_AUTHORITY_REQUIRED_FOR_PROD); a missing marker (legacy rows, REST packs, 043 editor path) remains allowed. Compile boundaries MUST still treat supplied context as untrusted input: ScenarioGraph.Validator recursively rejects query_context keys and SQL text inside dashboard_context (X1).

@} ScenarioGraph.ServerOwnedPipeline

#region ScenarioGraph.Capture.MovedTo044 [C:3] [TYPE Tombstone] [SEMANTICS scenario,capture,ownership]

@ingroup ScenarioGraph

@STATUS DEPRECATED

@DEPRECATED Runtime capture execution is owned by 044 ScenarioExecution (CaptureService over

Plugin.Service.ScreenshotService, artifacts owner_type=scenario_run).

@REPLACED_BY -> [ScenarioExecution.Capture]

@BRIEF 038 defines ScreenshotCaptureSpec (WHAT); it does NOT own capture execution or its artifacts.

#endregion ScenarioGraph.Capture.MovedTo044

#region ScenarioGraph.Vlm.MovedTo044 [C:3] [TYPE Tombstone] [SEMANTICS scenario,vlm,ownership]

@ingroup ScenarioGraph

@STATUS DEPRECATED

@DEPRECATED Runtime VLM submission + VlmFinding is owned by 044 ScenarioExecution (reuses

Plugin.Service.LLMClient via Services.LlmProvider.LLMProviderService).

@REPLACED_BY -> [ScenarioExecution.Vlm]

@BRIEF 038 defines VlmAnalysisSpec (WHAT to analyze); runtime VlmFinding/HumanCheckpointDisposition are 044.

#endregion ScenarioGraph.Vlm.MovedTo044

#region ScenarioGraph.Human.MovedTo044 [C:3] [TYPE Tombstone] [SEMANTICS scenario,human,ownership]

@ingroup ScenarioGraph

@STATUS DEPRECATED

@DEPRECATED Human checkpoint disposition is owned by 044 as HumanCheckpoint (confirm | false_positive | inconclusive),

distinct from 036 ActionApprovalGate.

@REPLACED_BY -> [ScenarioExecution.HumanCheckpoint]

#endregion ScenarioGraph.Human.MovedTo044

#region AgentChat.Tools.ScenarioGraph [C:4] [TYPE Module] [SEMANTICS scenario,agent,tools,compiler]

@defgroup ScenarioGraph Thin agent tools that submit bounded intent and display compiler/validator results.

@RELATION DEPENDS_ON -> [ScenarioGraph.Api]

@RATIONALE Agent tools stay thin: the agent explains intent, the deterministic compiler owns the graph.

@REJECTED Agent-side graph construction with free-form tool selection — bypasses validation and determinism.

@INVARIANT Agent cannot submit executable code, custom tool categories, raw expected metrics, artifact paths or runtime rewrites; it may author bounded SqlEvidenceSpec/DSL only through compiler validation.

#endregion AgentChat.Tools.ScenarioGraph

@{ ScenarioGraph.AgentAuthoringWorkspace [C:5] [TYPE Module]

@BRIEF Persistent external MCP agent/user authoring, exploratory sandbox and promotion contract. @RELATION DEPENDS_ON -> [ScenarioGraph.ServerOwnedPipeline] @RELATION DISPATCHES -> [ScenarioRegistry.SaveContinuation] @RELATION CALLS -> [ScenarioRegistry.CreateInitial]

create_authoring_session creates or idempotently returns a server-owned workspace for an existing scenario revision. bootstrap_authoring_scenario accepts only a bounded InitialScenarioIntent and server-issued compiled/draft-pack handles, calls ScenarioRegistry.CreateInitial, and returns a workspace bound to the new current revision. The two entry paths are mutually exclusive: an unbound workspace may hold a test plan or exploration receipt, but cannot propose, promote, save, activate, or launch a graph. propose_test_plan, start_exploration, get_exploration_result, propose_graph_revision, get_graph_diff, and promote_to_scenario operate on the persistent workspace. Every mutation carries an idempotency key and expected cas_version; stale versions return typed 409 with no partial state.

Exploration executes only in an isolated sandbox with allowlisted origins/APIs/actions, bounded time/size/network budgets, cancellation, durable operation receipts, server-owned artifact ownership and cleanup. Shell, credentials, cookies, filesystem/network escape and production side effects are forbidden. Sandbox readiness/security is a release gate, not an implementation claim.

start_exploration may yield only AuthoringArtifact and ExplorationResult; propose_graph_revision converts them to typed registered-action candidates. The server supplies reviewed templates/actions and does not generate arbitrary Playwright code. promote_to_scenario invokes deterministic Compile/Validate, presents a server-computed diff, requires user review, then hands only valid server handles to 042. Production execution accepts only the resulting promoted immutable revision; raw code and artifacts never pass through as authority.

@INVARIANT Authoring artifacts and ExecutionProgram are separate domains; an artifact digest, caller digest, URL, path or source text cannot establish graph, revision or run authority. @INVARIANT A first scenario is created only through the server-owned bootstrap transaction; no client-side fixture, synthetic base revision, or unbound test plan is a registry creation substitute. @INVARIANT The future code-backed provider is unimplemented and separately gated; it is not implied by sandbox support. @TEST_EDGE bootstrap replay -> same scenario/revision/workspace; unbound proposal -> typed state error with no registry row; sandbox_limit_or_cancel -> exploration_failed; invalid_candidate -> validation_blocked; stale_cas -> 409; unreviewed_diff -> no promotion.

@} ScenarioGraph.AgentAuthoringWorkspace

#endregion DashboardScenarioModel.Modules


OPENAPI — REST/Event API Contract

Source: contracts/openapi.yaml

openapi: "3.1.0" info: title: Dashboard Scenario Compiler API version: "1.0.0" description: > OpenAPI 3.1 contract for feature 038 — deterministic ScenarioGraph compile, validate, resolve, draft-pack (authoring-only). Runtime capture/VLM/disposition are owned by 044 (see 044/contracts/openapi.yaml). 038 emits scenario_key + content_hash; scenario_id (UUID) and revision_id (UUID) are assigned by 042 at Save.

servers:

  • url: /api/dashboard-testing description: superset-tools API gateway

tags:

  • name: scenario description: ScenarioGraph compile, validate, resolve, draft-pack

paths: /scenarios/compile: post: operationId: compileDashboardScenario tags: [scenario] summary: Deterministically compile a dashboard goal into a ScenarioGraph description: | Combines bounded agent intent (objective, selected case ids) with the dashboard query model, checklist catalog, baseline summary, capability model, and parameters to produce a byte-stable DashboardTestScenario. Requires [scenario.compile] role. security: - BearerAuth: [scenario.compile] requestBody: required: true content: application/json: schema: $ref: "#/components/schemas/CompileRequest" examples: valid: $ref: "#/components/examples/CompileRequestValid" responses: "200": description: Deterministic scenario plus validation summary content: application/json: schema: $ref: "#/components/schemas/ScenarioResponse" examples: compiled: $ref: "#/components/examples/ScenarioResponseCompiled" "401": $ref: "#/components/responses/UnauthorizedError" "403": $ref: "#/components/responses/ForbiddenError" "422": $ref: "#/components/responses/ValidationError" "429": $ref: "#/components/responses/RateLimitError" "500": $ref: "#/components/responses/InternalError"

/scenarios/validate: post: operationId: validateDashboardScenario tags: [scenario] summary: Validate a candidate ScenarioGraph and return all findings description: | Pure validation of a candidate graph: schema, DAG, refs, parameters, baselines, tools, SQL compilation, safety and checklist coverage. SQL is allowed only as validated immutable SqlEvidenceSpec; arbitrary code remains forbidden. security: - BearerAuth: [scenario.compile] requestBody: required: true content: application/json: schema: $ref: "#/components/schemas/ScenarioGraphInput" responses: "200": description: Complete validation findings content: application/json: schema: $ref: "#/components/schemas/ValidationResult" "401": $ref: "#/components/responses/UnauthorizedError" "422": $ref: "#/components/responses/ValidationError" "500": $ref: "#/components/responses/InternalError"

/scenarios/{scenarioId}/resolve: post: operationId: resolveDashboardScenario tags: [scenario] summary: Apply typed parameter/selector/manual resolutions as a new compiled definition description: | Applies declared resolutions to an unresolved graph and emits a new content_hash (immutable content lineage). Stale base content_hash is rejected (409) — never silently merged. security: - BearerAuth: [scenario.resolve] parameters: - name: scenarioId in: path required: true schema: { type: string } requestBody: required: true content: application/json: schema: $ref: "#/components/schemas/ResolveRequest" responses: "200": description: New immutable revision content: application/json: schema: $ref: "#/components/schemas/ScenarioResponse" "401": $ref: "#/components/responses/UnauthorizedError" "403": $ref: "#/components/responses/ForbiddenError" "409": $ref: "#/components/responses/StaleRevisionError" "422": $ref: "#/components/responses/ValidationError" "500": $ref: "#/components/responses/InternalError"

/scenarios/{scenarioId}/draft-pack: post: operationId: compileScenarioDraftPack tags: [scenario] summary: Generate a preview_only or save_eligible draft pack from versioned templates description: | Compiles a valid graph through registered versioned templates. Any validation error, NEEDS_SELECTOR, forbidden action, or unresolved baseline makes the pack preview_only. Unbound required ParameterDefinitions are run-preflight inputs and do not block saving. Save-eligible packs are registered as 036 drafts. Idempotent per content_hash. security: - BearerAuth: [scenario.draft] parameters: - name: scenarioId in: path required: true schema: { type: string } requestBody: required: true content: application/json: schema: $ref: "#/components/schemas/DraftPackRequest" responses: "201": description: Draft pack registered via feature 036 content: application/json: schema: $ref: "#/components/schemas/DraftPack" "200": description: Idempotent replay — existing DraftPack for the same revision hash content: application/json: schema: $ref: "#/components/schemas/DraftPack" "401": $ref: "#/components/responses/UnauthorizedError" "403": $ref: "#/components/responses/ForbiddenError" "409": $ref: "#/components/responses/StaleRevisionError" "422": $ref: "#/components/responses/ValidationError" "500": $ref: "#/components/responses/InternalError"

/scenarios/{scenarioId}/capture: post: operationId: captureScenarioScreenshot tags: [scenario] deprecated: true parameters: [{ name: scenarioId, in: path, required: true, schema: { type: string } }] summary: MOVED TO 044 — runtime capture is owned by ScenarioExecution description: | Deprecated. 038 defines ScreenshotCaptureSpec only; runtime capture, VLM, and disposition live in 044 (044/contracts/openapi.yaml), targeting /scenario-runs/{run_id}/steps/{logical_step_id}/capture with artifacts owner_type=scenario_run. This path is retained only for backward compatibility of the authoring prototype and is not the execution path. responses: "410": $ref: "#/components/responses/Gone"

/scenarios/{scenarioId}/vlm: post: operationId: analyzeScenarioScreenshot tags: [scenario] deprecated: true parameters: [{ name: scenarioId, in: path, required: true, schema: { type: string } }] summary: MOVED TO 044 — runtime VLM is owned by ScenarioExecution description: | Deprecated. 038 defines VlmAnalysisSpec only; runtime VLM submission and VlmFinding[] live in 044 (reusing Plugin.Service.LLMClient via Services.LlmProvider.LLMProviderService). responses: "410": $ref: "#/components/responses/Gone"

/scenarios/{scenarioId}/disposition: post: operationId: disposeVlmFindings tags: [scenario] deprecated: true parameters: [{ name: scenarioId, in: path, required: true, schema: { type: string } }] summary: MOVED TO 044 — human checkpoint disposition is owned by ScenarioExecution description: | Deprecated. 038 no longer owns human disposition; it is a 044 HumanCheckpoint (confirm | false_positive | inconclusive), distinct from the 036 ActionApprovalGate. responses: "410": $ref: "#/components/responses/Gone"

components: securitySchemes: BearerAuth: type: http scheme: bearer bearerFormat: JWT description: | superset-tools JWT. Roles encoded in roles claim. Required scopes: scenario.compile, scenario.resolve, scenario.draft, scenario.execute (per ADR-0005 RBAC).

schemas: ErrorEnvelope: type: object required: [error] properties: error: type: object required: [code, detail] properties: code: { type: string, example: "VALIDATION_ERROR" } detail: { type: string } fields: type: object additionalProperties: { type: string } description: Per-field validation errors (422 only) retry_after: type: integer description: Seconds until retry is allowed (429 only)

SuccessEnvelope:
  type: object
  required: [data]
  properties:
    data: {}
    meta:
      type: object
      properties:
        content_hash: { type: string }

CompileRequest:
  type: object
  additionalProperties: false
  required: [objective, query_model, checklist_catalog_version, baseline_version, capabilities, parameters, change_request_context]
  properties:
    provenance:
      type: object
      description: Optional; compiler must not require an AgentRun.
      properties:
        source_type: { enum: [agent_run, editor, migration, api] }
        source_id: { type: [string, "null"] }
    objective:
      type: object
      required: [goal, selected_case_ids]
      properties:
        goal: { type: string, maxLength: 2000 }
        selected_case_ids: { type: array, items: { type: string } }
        rationale: { type: [string, "null"], maxLength: 4000 }
    query_model: { type: object }
    checklist_catalog_version: { const: 1 }
    baseline_version: { type: string }
    capabilities: { type: object }
    parameters: { type: object }
    change_request_context: { $ref: "#/components/schemas/ChangeRequestContext" }

ChangeRequestContext:
  type: object
  additionalProperties: false
  required: [request_id, objective, acceptance_criteria]
  properties:
    request_id: { type: string }
    objective: { type: string }
    affected_dashboards: { type: array, items: { type: string } }
    added_fields: { type: array, items: { type: string } }
    removed_fields: { type: array, items: { type: string } }
    renamed_fields: { type: array, items: { type: object } }
    technical_details: { type: string }
    requested_tables_charts: { type: array, items: { type: string } }
    expected_business_behavior: { type: string }
    related_dashboards: { type: array, items: { type: string } }
    source_relations: { type: array, items: { type: string } }
    acceptance_criteria: { type: array, minItems: 1, items: { type: string } }
    control_totals: { type: object }

ScenarioGraphInput:
  type: object
  description: Candidate immutable Verification Program. SQL is allowed only as a validated SqlEvidenceSpec; arbitrary code, raw metric truth and unsafe paths are forbidden.
  additionalProperties: false
  required: [schema_version, scenario_key, dashboard_context, objective, verification_program, phases, steps]
  properties:
    schema_version: { type: integer }
    scenario_key: { type: string }
    dashboard_context: { type: object }
    objective: { type: object }
    change_request_context: { $ref: "#/components/schemas/ChangeRequestContext" }
    verification_program: { $ref: "#/components/schemas/VerificationProgram" }
    phases: { type: array, items: { type: string } }
    steps: { type: array, items: { type: object } }

ScenarioResponse:
  type: object
  required: [scenario, validation]
  properties:
    scenario: { type: object }
    validation: { $ref: "#/components/schemas/ValidationResult" }

VerificationProgram:
  type: object
  required: [navigation_program, evidence_program, transformation_program, assertion_program, semantic_evaluation_program]
  properties:
    navigation_program: { type: array, items: { type: object } }
    evidence_program: { type: array, items: { oneOf: [{ type: object }, { $ref: "#/components/schemas/SqlEvidenceSpec" }] } }
    transformation_program: { type: array, items: { $ref: "#/components/schemas/TransformSpec" } }
    assertion_program: { type: array, items: { $ref: "#/components/schemas/ComparisonSpec" } }
    semantic_evaluation_program: { type: array, items: { $ref: "#/components/schemas/AgentEvaluationSpec" } }
SqlEvidenceSpec:
  type: object
  required: [snippet_id, logical_step_id, connection_ref, database_identity, sql_template, sql_hash, parameter_definitions, expected_output_schema, relation_refs, execution_limits, generation_provenance, validation_result]
  properties:
    snippet_id: { type: string }
    logical_step_id: { type: string }
    connection_ref: { type: string }
    database_identity: { type: string }
    sql_template: { type: string, description: "Read-only saved template; immutable after revision save" }
    sql_hash: { type: string, pattern: "^[a-f0-9]{64}$" }
    parameter_definitions: { type: array, items: { type: object } }
    expected_output_schema: { type: object }
    relation_refs: { type: array, items: { type: string } }
    execution_limits: { type: object, required: [timeout_ms, max_rows, max_bytes, max_complexity] }
    generation_provenance: { type: object }
    validation_result: { type: object, required: [valid] }
TransformSpec:
  type: object
  required: [logical_step_id, version, operations]
  properties:
    logical_step_id: { type: string }
    version: { type: string }
    operations: { type: array, items: { type: object, properties: { op: { type: string, enum: [select, filter, rename, cast, join, group_by, sum, count, distinct, coalesce, normalize_string, normalize_date, difference, ratio, tolerance_compare] } } } }
ComparisonSpec:
  type: object
  required: [logical_step_id, kind, left_ref, right_ref]
  properties:
    logical_step_id: { type: string }
    kind: { type: string, enum: [numeric_equality, numeric_tolerance, row_comparison, column_comparison, aggregate_comparison, set_equality, field_mapping, null_fill_check, cross_dashboard] }
    left_ref: { type: string }
    right_ref: { type: string }
    tolerance: { type: number }
    field_mapping: { type: object }
AgentEvaluationSpec:
  type: object
  required: [spec_id, logical_step_id, provider_id, model_id, prompt_template_id, prompt_template_version, prompt_template_hash, input_manifest, evidence_refs, tool_allowlist, output_schema, decision_policy_id]
  properties:
    spec_id: { type: string }
    logical_step_id: { type: string }
    provider_id: { type: string }
    model_id: { type: string }
    prompt_template_id: { type: string }
    prompt_template_version: { type: string }
    prompt_template_hash: { type: string, pattern: "^[a-f0-9]{64}$" }
    input_manifest: { type: object }
    evidence_refs: { type: array, items: { type: string } }
    tool_allowlist: { type: array, items: { type: string } }
    output_schema: { type: object }
    decision_policy_id: { type: string }

ValidationResult:
  type: object
  required: [valid, errors, warnings, blockers, coverage, topological_order, graph_hash]
  properties:
    valid: { type: boolean }
    errors: { type: array, items: { type: object } }
    warnings: { type: array, items: { type: object } }
    blockers: { type: array, items: { type: object } }
    coverage: { type: array, items: { type: object } }
    topological_order: { type: array, items: { type: string } }
    unresolved_parameters: { type: array, items: { type: string } }
    unresolved_selectors: { type: array, items: { type: string } }
    unresolved_baselines: { type: array, items: { type: string } }
    graph_hash: { type: string }
    sql_compilation_results: { type: array, items: { type: object } }
    verification_program_valid: { type: boolean }

ResolveRequest:
  type: object
  additionalProperties: false
  required: [base_content_hash, changes]
  properties:
    base_content_hash: { type: string, pattern: "^[a-f0-9]{64}$" }
    changes:
      type: array
      items:
        type: object
        required: [kind, target, value]
        properties:
          kind: { enum: [parameter_definition, selector, manual_conversion, remove_step] }
          target: { type: string }
          value: {}
          reason: { type: [string, "null"] }

DraftPackRequest:
  type: object
  additionalProperties: false
  required: [content_hash]
  properties:
    content_hash: { type: string, pattern: "^[a-f0-9]{64}$" }
    provenance:
      type: object
      properties:
        source_type: { enum: [agent_run, editor, migration, api] }
        source_id: { type: [string, "null"] }

DraftPack:
  type: object
  required: [scenario_key, content_hash, template_version, status, manifest, artifacts, validation_summary, warnings]
  properties:
    scenario_key: { type: string }
    content_hash: { type: string }
    template_version: { type: string }
    status: { enum: [preview_only, save_eligible] }
    manifest: { type: object }
    artifacts: { type: array, items: { type: object } }
    validation_summary: { $ref: "#/components/schemas/ValidationResult" }
    warnings: { type: array, items: { type: object } }

CaptureRequest:
  type: object
  additionalProperties: false
  required: [logical_step_id, capture_spec]
  deprecated: true
  description: Runtime capture moved to 044.
  properties:
    logical_step_id: { type: string }
    capture_spec: { $ref: "#/components/schemas/CaptureSpec" }
    provenance:
      type: object
      properties:
        source_type: { enum: [agent_run, editor, migration, api] }
        source_id: { type: [string, "null"] }

CaptureSpec:
  type: object
  required: [target, viewport, readiness, method]
  properties:
    target: { type: string, enum: [tab, viewport] }
    tab_identifier: { type: [string, "null"] }
    viewport:
      type: object
      required: [width, height]
      properties:
        width: { type: integer }
        height: { type: integer }
    readiness: { type: string, enum: [canvas_stabilized, network_idle, fixed_wait] }
    readiness_timeout_ms: { type: integer, default: 15000 }
    mask_selectors: { type: array, items: { type: string } }
    method: { type: string, enum: [cdp, full_page, region] }

CaptureResponse:
  type: object
  required: [artifact_refs]
  properties:
    artifact_refs:
      type: array
      items:
        type: object
        required: [artifact_id, kind, masked]
        properties:
          artifact_id: { type: string, format: uuid }
          kind: { type: string }
          masked: { type: boolean }

VlmRequest:
  type: object
  additionalProperties: false
  required: [logical_step_id, artifact_id, analysis]
  deprecated: true
  description: Runtime VLM moved to 044.
  properties:
    logical_step_id: { type: string }
    artifact_id: { type: string, format: uuid }
    analysis:
      type: object
      required: [profile_id, provider_id, model_id, prompt_template_id, prompt_version, prompt_template_hash]
      properties:
        profile_id: { type: string }
        provider_id: { type: string }
        model_id: { type: string }
        prompt_template_id: { type: string }
        prompt_version: { type: string }
        prompt_template_hash: { type: string, pattern: "^[a-f0-9]{64}$" }
        confidence_threshold: { type: number, minimum: 0, maximum: 1, default: 0.7 }

VlmResponse:
  type: object
  required: [findings]
  properties:
    findings:
      type: array
      items: { $ref: "#/components/schemas/VlmFinding" }

VlmFinding:
  type: object
  required: [finding_id, source_artifact_id, severity, code, confidence, description, disposition, model_provenance]
  properties:
    finding_id: { type: string }
    source_artifact_id: { type: string, format: uuid }
    severity: { type: string, enum: [info, warning, error] }
    code: { type: string }
    region:
      type: object
      properties:
        selector: { type: string }
        bounds: { type: object }
    confidence: { type: number, minimum: 0, maximum: 1 }
    description: { type: string, maxLength: 1000 }
    disposition: { type: string, enum: [unresolved, confirmed, false_positive, inconclusive] }
    disposition_comment: { type: [string, "null"], maxLength: 2000 }
    model_provenance:
      type: object
      required: [model_id, prompt_version, prompt_template_hash, analyzed_at]
      properties:
        model_id: { type: string }
        prompt_version: { type: string }
        prompt_template_hash: { type: string, pattern: "^[a-f0-9]{64}$" }
        analyzed_at: { type: string, format: date-time }

DispositionRequest:
  type: object
  additionalProperties: false
  required: [logical_step_id, dispositions]
  deprecated: true
  description: "Human checkpoint disposition moved to 044 (HumanCheckpoint: confirm | false_positive | inconclusive)."
  properties:
    logical_step_id: { type: string }
    dispositions:
      type: array
      items:
        type: object
        required: [finding_id, decision]
        properties:
          finding_id: { type: string }
          decision: { enum: [confirm, false_positive, inconclusive] }
          comment: { type: [string, "null"], maxLength: 2000 }

DispositionResponse:
  type: object
  required: [step]
  properties:
    step: { type: object }

responses: UnauthorizedError: description: Missing or invalid authentication content: application/json: schema: { $ref: "#/components/schemas/ErrorEnvelope" } example: error: { code: "UNAUTHORIZED", detail: "Authentication required" } ForbiddenError: description: Insufficient permissions content: application/json: schema: { $ref: "#/components/schemas/ErrorEnvelope" } example: error: { code: "FORBIDDEN", detail: "Requires role: scenario.compile" } ValidationError: description: Request validation failed content: application/json: schema: { $ref: "#/components/schemas/ErrorEnvelope" } example: error: { code: "VALIDATION_ERROR", detail: "Request validation failed", fields: { objective: "goal is required" } } StaleRevisionError: description: Base revision is stale or already resolved content: application/json: schema: { $ref: "#/components/schemas/ErrorEnvelope" } example: error: { code: "STALE_REVISION", detail: "Base revision b2c3… no longer current. Recompile from latest." } StalePromptError: description: VLM prompt template is stale content: application/json: schema: { $ref: "#/components/schemas/ErrorEnvelope" } example: error: { code: "STALE_PROMPT", detail: "Prompt template v1 hash mismatch. Update to current template." } Gone: description: Runtime operation moved to 044 content: application/json: schema: { $ref: "#/components/schemas/ErrorEnvelope" } example: error: { code: "MOVED_TO_044", detail: "Runtime capture/VLM/disposition is owned by 044 ScenarioExecution." } DoubleDispositionError: description: Finding already disposed content: application/json: schema: { $ref: "#/components/schemas/ErrorEnvelope" } example: error: { code: "DOUBLE_DISPOSITION", detail: "Finding f-001 already confirmed" } RateLimitError: description: Too many requests headers: Retry-After: schema: { type: integer } description: Seconds until next request is allowed content: application/json: schema: { $ref: "#/components/schemas/ErrorEnvelope" } example: error: { code: "RATE_LIMITED", detail: "Too many requests. Retry after 30 seconds.", retry_after: 30 } InternalError: description: Unexpected server error content: application/json: schema: { $ref: "#/components/schemas/ErrorEnvelope" } example: error: { code: "INTERNAL_ERROR", detail: "An unexpected error occurred" }

examples: CompileRequestValid: summary: Compile filters/metric/XLSX scenario value: provenance: source_type: agent_run source_id: "550e8400-e29b-41d4-a716-446655440000" objective: goal: "Verify filters, metric, XLSX export, and baseline" selected_case_ids: ["B01", "C04", "C05"] rationale: "Release candidate covers filter and export regressions" query_model: { dashboard_key: "fi-0080" } checklist_catalog_version: 1 baseline_version: "2026-07-01" capabilities: { native_filters: true, text_filter: true, table_filter: true, pagination: true, row_edit: true, bulk_edit: true, persistence_refresh: true, time_rollover: false, xlsx_export: true, cross_dashboard: false, superset_metric: true, dataset_field_read: true, screenshot: true, repository_write: true } parameters: { test_date: { type: "date" }, counterparty: { type: "string" } } ScenarioResponseCompiled: summary: Compiled 18-step scenario value: scenario: scenario_key: "fi-0080_verify-filters-metric-xlsx" content_hash: "a1b2c3d4e5f6a7b8c9d0e1f2a3b4c5d6e7f8a9b0c1d2e3f4a5b6c7d8e9f0a1b2c3" phases: ["setup", "interact", "observe", "assert", "evidence", "report"] steps: [{ logical_step_id: "4b7c…", step_key: "phase-1-B01-open_dashboard", tool: "browser", action: "open_dashboard" }] validation: valid: true errors: [] warnings: [{ code: "STALE_BASELINE", detail: "baseline 2026-06-15 superseded by 2026-07-01", logical_step_id: "4b7c0000-0000-0000-0000-000000000005" }] blockers: [] coverage: [{ case_id: "B01", classification: "automated" }, { case_id: "C04", classification: "automated" }, { case_id: "C05", classification: "automated" }] topological_order: ["phase-1-B01-open_dashboard", "phase-2-B01-apply_filters"] graph_hash: "a1b2c3d4e5f6a7b8c9d0e1f2a3b4c5d6e7f8a9b0c1d2e3f4a5b6c7d8e9f0a1b2c3"


CONTRACTS — Remaining — action-registry.yaml

Source: contracts/action-registry.yaml

version: v1 description: Canonical version-pinned executable action catalog. Runtime dispatches no action outside this file. actions:

  • { tool: browser, action: open_dashboard, risk: READ_ONLY, timeout_ms: 30000, retry_safe: true, mutates: false }
  • { tool: browser, action: navigate_tab, risk: UI_INTERACTION, timeout_ms: 15000, retry_safe: true, mutates: false }
  • { tool: browser, action: apply_native_filter, risk: UI_INTERACTION, timeout_ms: 15000, retry_safe: true, mutates: false }
  • { tool: browser, action: inspect_filter_state, risk: READ_ONLY, timeout_ms: 10000, retry_safe: true, mutates: false }
  • { tool: browser, action: apply_table_filter, risk: UI_INTERACTION, timeout_ms: 15000, retry_safe: true, mutates: false }
  • { tool: browser, action: extract_table, risk: READ_ONLY, timeout_ms: 30000, retry_safe: true, mutates: false }
  • { tool: browser, action: scroll_to, risk: UI_INTERACTION, timeout_ms: 10000, retry_safe: true, mutates: false }
  • { tool: browser, action: inspect_columns, risk: READ_ONLY, timeout_ms: 10000, retry_safe: true, mutates: false }
  • { tool: browser, action: click, risk: UI_INTERACTION, timeout_ms: 10000, retry_safe: false, mutates: false }
  • { tool: browser, action: select_rows, risk: UI_INTERACTION, timeout_ms: 15000, retry_safe: true, mutates: false }
  • { tool: browser, action: edit_row, risk: TEST_DATA_MUTATION, timeout_ms: 30000, retry_safe: false, mutates: true }
  • { tool: browser, action: bulk_edit, risk: TEST_DATA_MUTATION, timeout_ms: 60000, retry_safe: false, mutates: true }
  • { tool: browser, action: download, risk: UI_INTERACTION, timeout_ms: 60000, retry_safe: false, mutates: false }
  • { tool: browser, action: refresh, risk: UI_INTERACTION, timeout_ms: 30000, retry_safe: true, mutates: false }
  • { tool: browser, action: wait_for_state, risk: READ_ONLY, timeout_ms: 30000, retry_safe: true, mutates: false }
  • { tool: browser, action: apply_filters, risk: UI_INTERACTION, timeout_ms: 15000, retry_safe: true, mutates: false }
  • { tool: browser, action: download_xlsx, risk: UI_INTERACTION, timeout_ms: 60000, retry_safe: false, mutates: false }
  • { tool: superset_api, action: execute_metric, risk: READ_ONLY, timeout_ms: 30000, retry_safe: true, mutates: false }
  • { tool: superset_api, action: dataset_field_assert, risk: READ_ONLY, timeout_ms: 30000, retry_safe: true, mutates: false }
  • { tool: sql_evidence, action: execute_pinned_sql, risk: READ_ONLY, timeout_ms: 30000, retry_safe: true, mutates: false }
  • { tool: transform, action: execute_dsl, risk: READ_ONLY, timeout_ms: 30000, retry_safe: true, mutates: false }
  • { tool: assertion, action: compare_to_baseline, risk: READ_ONLY, timeout_ms: 30000, retry_safe: true, mutates: false }
  • { tool: assertion, action: evaluate_comparison_spec, risk: READ_ONLY, timeout_ms: 30000, retry_safe: true, mutates: false }
  • { tool: agent_evaluation, action: evaluate_declared_spec, risk: READ_ONLY, timeout_ms: 60000, retry_safe: false, mutates: false }
  • { tool: xlsx, action: parse_xlsx, risk: READ_ONLY, timeout_ms: 30000, retry_safe: true, mutates: false }
  • { tool: screenshot, action: capture, risk: READ_ONLY, timeout_ms: 30000, retry_safe: true, mutates: false }
  • { tool: report, action: generate_report, risk: READ_ONLY, timeout_ms: 30000, retry_safe: true, mutates: false }
  • { tool: artifact, action: register, risk: READ_ONLY, timeout_ms: 10000, retry_safe: true, mutates: false }

CONTRACTS — Remaining — openapi-traceability.md

Source: contracts/openapi-traceability.md

#region DashboardScenarioModel.OpenApiTraceability [C:3] [TYPE ADR] [SEMANTICS openapi,traceability,scenario] @defgroup OpenAPI Trace OpenAPI operationId → data-model → spec → UX contract drift map for feature 038.

Operation Traceability

operationId Spec Requirement Data Model UX Contract Status
compileDashboardScenario AGSCN-FR-001, FR-002, FR-008, FR-009 DashboardTestScenario, ScenarioStep api-ux.md: compile ✅
validateDashboardScenario AGSCN-FR-004, FR-005 ScenarioValidationResult api-ux.md: validate ✅
resolveDashboardScenario AGSCN-FR-005 (params) ScenarioParameter, ScenarioRef api-ux.md: resolve ✅
compileScenarioDraftPack AGSCN-FR-009 (templates) DraftPack, ArtifactPlan api-ux.md: draft-pack ✅
captureScenarioScreenshot AGSCN-FR-010 ScreenshotCaptureSpec api-ux.md: capture ✅
analyzeScenarioScreenshot AGSCN-FR-011 VlmFinding, VlmAnalysis api-ux.md: vlm ✅
disposeVlmFindings AGSCN-FR-012 HumanDisposition api-ux.md: disposition ✅

Schema Traceability

Schema Source Purpose
CompileRequest data-model.md: ScenarioParameter + research §2 Compile request body
ScenarioResponse data-model.md: DashboardTestScenario Compiled graph + validation
ValidationResult data-model.md: ScenarioValidationResult Complete findings
DraftPack data-model.md: DraftPack Template-compiled draft
CaptureSpec data-model.md: ScreenshotCaptureSpec Screenshot capture config
VlmFinding data-model.md: VlmFinding Typed VLM output
ErrorEnvelope ux_reference.md §4 Standard error envelope

Drift Detection (manual review)

  • Every operationId maps to at least one spec requirement
  • Every spec requirement with an API touchpoint maps to an operationId
  • Pydantic schema names match OpenAPI schema names (data-model.md ↔ openapi.yaml)
  • Error response shapes match ux_reference.md promises (401/403/404/409/422/429/5xx)
  • Auth requirements match ADR-0005 RBAC model (scenario.compile/resolve/draft/execute)
  • Edge & Failure matrix state classes (E6 409, E7 422, E8 403, E9 429, E10 5xx, E11 stale prompt) map to response classes

Coverage Gate

  • Success examples for every operation (compile shown; others reference same envelope)
  • Error examples for every response class (401/403/409/422/429/500 + STALE_PROMPT/DOUBLE_DISPOSITION)
  • operationId on every operation (7 unique)
  • Reusable schemas (no inline anonymous schemas; ErrorEnvelope/SuccessEnvelope shared)
  • All mutating operations declare security (compile/validate/resolve/draft-pack/capture/vlm/disposition)

Validation Command

python -c "import yaml,sys; d=yaml.safe_load(open('specs/038-dashboard-scenario-model/contracts/openapi.yaml')); assert d['openapi'].startswith('3.1'); ids=[o['operationId'] for p in d['paths'].values() for m,o in p.items() if isinstance(o,dict) and 'operationId' in o]; assert len(ids)==len(set(ids))==7, ids; print('openapi-ok', len(ids))"

#endregion DashboardScenarioModel.OpenApiTraceability


CONTRACTS — Remaining — verification-program.md

Source: contracts/verification-program.md

#region VerificationProgram.Contract [C:5] [TYPE ADR] [SEMANTICS verification-program,ir,sql,evidence,transform,assertion,agent-evaluation] @BRIEF Canonical immutable IR compiled by the authoring agent and deterministically executed by 044. @RELATION DEPENDS_ON -> [DashboardScenarioModel.DataModel] @RELATION BINDS_TO -> [ScenarioExecution.DataModel] @RATIONALE A scenario is a verification program, not merely a browser-click DAG: business checks often require source-mart evidence, deterministic transformations and explicit semantic evaluation. @REJECTED Runtime SQL/code generation or rewrite — rejected because it destroys reproducibility, security review and content-addressed revision identity.

Program shape

VerificationProgram { navigation_program, evidence_program, transformation_program, assertion_program, semantic_evaluation_program } is required executable content of DashboardTestScenario and is included, canonically, in content_hash and every step_content_hash that references it. Programs use only declared refs and version-pinned action/DSL/prompt registries.

SqlEvidenceSpec

SqlEvidenceSpec { snippet_id, logical_step_id, connection_ref, database_identity, sql_template, sql_hash, parameter_definitions[], expected_output_schema, relation_refs[], execution_limits, generation_provenance, validation_result } is a small, read-only source-evidence query. It is authored only during creation, edit, migration/revalidation or investigation proposal; runtime executes exactly the pinned sql_template through the Superset SQL Lab adapter with typed ParameterBinding substitution. It cannot add WHERE/JOIN/projection/relation text, mutate SQL, or accept credentials from an agent/client.

TransformSpec and ComparisonSpec

TransformSpec is a bounded, versioned DSL — select|filter|rename|cast|join|group_by|sum|count|distinct|coalesce|normalize_string|normalize_date|difference|ratio|tolerance_compare — over declared evidence refs. No Python, shell, SQL or arbitrary code is permitted.

ComparisonSpec / AssertionSpec declares numeric equality/tolerance, row/column/set equality, aggregate checks, field mappings, null/fill checks and cross-dashboard comparisons. Its expected values are evidence/baseline/typed-control refs, never runtime prose.

AgentEvaluationSpec and DecisionPolicy

AgentEvaluationSpec { spec_id, logical_step_id, provider_id, model_id, prompt_template_id, prompt_template_version, prompt_template_hash, input_manifest, evidence_refs, tool_allowlist, output_schema, decision_policy_id } is permitted only for semantic/visual/ambiguous checks that cannot reasonably compile to SqlEvidenceSpec + TransformSpec + AssertionSpec. It has bounded declared evidence/tool access and no mutation authority.

DecisionPolicy { policy_id, version, deterministic_hard_failure, high_confidence_failure, low_confidence, disagreement, missing_evidence } deterministically maps evidence plus a typed AgentEvaluation to StepOutcome. A bare model verdict is never ScenarioResult authority.

SQL compilation gate

Before a SqlEvidenceSpec can be saved, compiler validation MUST parse its AST; require one SELECT/WITH statement; prohibit DDL/DML, mutation SETTINGS, unsafe/external/file functions; enforce connection/relation/schema allowlists, typed binds, timeout/row/byte/complexity limits, expected-output schema, and bounded preview/test execution. Exploratory authoring SQL is a separate bounded policy class and must be minimized to the columns/filters actually used before persistence.

Runtime boundary

044 orchestration stays deterministic. It may dispatch an explicitly declared AgentEvaluationSpec, but that step cannot change the ScenarioGraph, SQL/DSL, run lifecycle, executor order or mutations. Bad runtime evidence becomes a finding → 043 proposal → validated new revision → later run.

@{ VerificationProgram.Reconciliation [C:5] [TYPE ADR]

@BRIEF Implementation reconciliation (code audit 2026-09-06): which IR entities are implemented, implemented under a different name, or deferred. @STATUS PARTIALLY IMPLEMENTED @RELATION BINDS_TO -> [ScenarioGraph.Models] @RELATION DEPENDS_ON -> [ScenarioGraph.Templates.ActionRegistry]

Implemented today (canonical shape).

  • The executable Verification Program IS the ActionRegistry-pinned ScenarioStep DAG: ScenarioGraph.Models.ScenarioStep plus version-pinned ActionExecutionDescriptor snapshots from templates/__init__.py (ACTION_REGISTRY_VERSION 038.1.0, action_registry_fingerprint()). Phases (setup/interact/observe/assert/evidence/report) play the role of the program decomposition. content_hash covers the canonical DashboardTestScenario serialization (ScenarioGraph.Models.CanonicalBytes/CanonicalDump).
  • ScreenshotCaptureSpec → implemented as CaptureSpec (ScenarioGraph.Models.CaptureSpec).
  • VlmAnalysisSpec / VlmFinding → implemented as VlmAnalysis / VlmFinding (ScenarioGraph.Models.VlmAnalysis/VlmFinding; prompt-hash staleness gate in scenario/vlm.py::validate_prompt_current).
  • ParameterDefinition → implemented as ScenarioParameter.
  • Baseline-referenced expectations → Expected/Ref models; raw numeric truth remains validator-blocked (AGSCN-FR-004 enforced).

Deferred (PROPOSED — normative for future work, NOT implemented; zero code references in backend/src as of 2026-09-06).

  • The five-way VerificationProgram { navigation_program, evidence_program, transformation_program, assertion_program, semantic_evaluation_program } decomposition. Moving to it requires a schema_version/compatibility_family migration and re-minted content_hash rules for every stored revision; it is NOT a silent model swap.
  • SqlEvidenceSpec and the SQL compilation gate (AGSCN-FR-011a).
  • TransformSpec bounded DSL; ComparisonSpec/AssertionSpec as first-class entities (assertions currently live in registered assertion-tool steps (structural_assert, compare_to_baseline, vlm_analyze) + Expected).
  • AgentEvaluationSpec + DecisionPolicy (AGSCN-FR-011c): today VLM steps produce typed advisory VlmFindings resolved via 044 HumanCheckpoint; the deterministic DecisionPolicy mapping is unimplemented.

Interim safety rule. While the deferred entities are absent, capability mapping MUST classify source-mart-evidence and semantic-evaluation checklist cases as human_checkpoint, unsupported, or needs_context — never as automated steps referencing non-existent specs (AGSCN-FR-005/SC-002 semantics). @REJECTED Treating the deferred IR names as implemented because spec fixtures or checklists mention them — the 2026-09-06 code audit found zero code references; claiming otherwise would fabricate verification coverage.

@} VerificationProgram.Reconciliation

#endregion VerificationProgram.Contract


QUICKSTART — Dev Onboarding

Source: quickstart.md

Quickstart: Dashboard Scenario Model

Prerequisites

036 draft registration and 037 query-model/baseline summary contracts must pass. /speckit.validate must report PASS before /speckit.implement.

Test Order

# Tier 1: Fast unit tests (<120s, no Docker)
make test-unit              # backend SQLite unit tests
make test-frontend          # frontend vitest (039 DTO consumption, if present)

# Tier 1 alt: Smart selection — only tests linked to scenario contracts
make test-related F=backend/src/services/dashboard_testing/scenario/

# Tier 2: Integration (Docker required)
make test-integration       # capture/VLM/disposition bridges via 036/037

# Scoped scenario suite
cd backend && python -m pytest tests/services/dashboard_testing/scenario -v
cd backend && python -m pytest tests/api/test_dashboard_scenarios.py -v

# Linting
make lint

# OpenAPI schema validation (in-repo, no new deps)
python3 -c "import yaml; d=yaml.safe_load(open('specs/038-dashboard-scenario-model/contracts/openapi.yaml')); assert d['openapi'].startswith('3.1'); print('openapi-ok')"

Independent Validation (compiler-layer scope)

  1. Load the catalog and assert exactly 19 unique ids.
  2. Map a full-capability dashboard; every case has one classification.
  3. Map a no-XLSX/no-dataset-field dashboard; C04–C06 and T01–T03 retain rationale with no SQL.
  4. Compile the same canonical inputs repeatedly and with shuffled input maps; scenario_key/content_hash must match.
  5. Validate invalid fixtures: cycle, missing ref, duplicate output, raw metric expected, stale baseline, unknown selector/tool, invalid ParameterDefinition.
  6. Verify an unbound required ParameterDefinition remains save-eligible; bind it in 044 RunPreflight and verify ScenarioRevision/content_hash remain unchanged.
  7. Compile a valid graph to draft pack (authoring) through registered templates.
  8. Attempt executable-code, shell, SQL, custom path, and unknown-template injection; all must block before draft registration.
  9. Prototype: every @UX_STATE in contracts/ux/scenario-graph-ux.md reachable via prototype/index.html state switcher (see prototype/manifest.md).
  10. VlmAnalysisSpec/ScreenshotCaptureSpec serialize canonically and are consumed by 044 executors (runtime VLM/capture belongs to 044).

Boundary note (2026-08-07): 038 is the compiler/IR layer. Runtime VLM submit, real capture bytes, and evidence artifacts (owner_type=scenario_run) are owned by 044 ScenarioExecution — NOT by 038. See REVIEW-042-047-CLOSURE.md and 044.

Exit Gates (compiler-layer scope)

  • JSON Schema and Pydantic round-trip.
  • 100% of 19 cases classified in fixtures; at least 80% meet spec classification criterion.
  • Byte-stable JSON/YAML snapshots; compiler emits scenario_key + content_hash, NO scenario_id/revision_id.
  • No raw baseline numbers in executable assertions.
  • No LLM executable body or SQL surface.
  • contracts/openapi.yaml valid; runtime capture/VLM/disposition endpoints marked deprecated (moved to 044).
  • Belief runtime audit (C4/C5): axiom_audit({operation="audit_belief_runtime"}) + axiom_audit({operation="audit_belief_protocol"}) PASS.
  • /speckit.validate verdict: PASS (compiler scope only; full-feature completeness gated by 044).
  • Unit/API, ruff, schema, and semantic audits pass.

Known Gap & Boundary (2026-08-07 reconciliation)

Шаги 1–9 проверяют детерминированный compiler/validator/resolver/pack контур — работает. Runtime capture/VLM/disposition НЕ входят в 038 — перенесены в 044 (ScenarioExecution): реальный VLM submit через LLMClient/LLMProviderService, реальный capture через ScreenshotService с артефактами owner_type=scenario_run, HumanCheckpoint (confirm/false_positive/inconclusive). 038 определяет только VlmAnalysisSpec/ScreenshotCaptureSpec. Прежние T057–T059 — задачи 044.


TRACEABILITY — Requirements Matrix

Source: traceability.md

#region DashboardScenarioModel.Traceability [C:3] [TYPE ADR] [SEMANTICS traceability,rtm,scenario] @defgroup Trace Matrix Requirements → Screen+State → Model → API → Contract → Task → Test for feature 038.

Applicability

  • Feature type: Fullstack (backend compiler/validator/resolver core + thin UI preview consumed by 039 + agent tools)
  • UI surface: Yes — scenario graph preview, coverage, resolution, pack states (DTOs supplied to 039)
  • API surface: Yes — 7 operations in contracts/openapi.yaml

Traceability Matrix

Story / Req UX Screen + State Screen Model API operationId Contract Backend Task Frontend Task Test
US1: Build Scenario Graph preview (loading → loaded) N/A — DTO only compileDashboardScenario ScenarioGraph.Compiler.Compile T009–T012 N/A — UI in 039 Test.Scenario.Compiler
AGSCN-FR-001/002/009 preview (loaded) N/A — DTO only compileDashboardScenario ScenarioGraph.Compiler.Compile T005–T012 N/A — UI in 039 Test.Scenario.Schema
US2: Validate Safety preview (blocked) N/A — DTO only validateDashboardScenario ScenarioGraph.Validator.Validate T013–T017 N/A — UI in 039 Test.Scenario.Validator
AGSCN-FR-004/005 preview (blocked) N/A — DTO only validateDashboardScenario ScenarioGraph.Validator.Validate T014–T016 N/A — UI in 039 Test.Scenario.Validator.Edge
US3: Checklist Coverage preview (coverage panel) N/A — DTO only compileDashboardScenario ScenarioGraph.CapabilityMapper.Map T006–T008 N/A — UI in 039 Test.Scenario.Capability
AGSCN-FR-006 preview (coverage panel) N/A — DTO only compileDashboardScenario ScenarioGraph.Catalog.Load T001–T004 N/A — UI in 039 Test.Scenario.Catalog
US4: Parameters/Checkpoints preview (resolution) N/A — DTO only resolveDashboardScenario ScenarioGraph.Resolver.Resolve T022–T025 N/A — UI in 039 Test.Scenario.Resolver
AGSCN-FR-007/008 preview (resolution) N/A — DTO only resolveDashboardScenario + compile ScenarioGraph.Serializer.Canonical T018–T021 N/A — UI in 039 Test.Scenario.Serializer
US5: Capture/VLM/Disposition preview (vlm findings) N/A — DTO only captureScenarioScreenshot, analyzeScenarioScreenshot, disposeVlmFindings ScenarioGraph.Capture.Dispatch, ScenarioGraph.Vlm.Analyze, ScenarioGraph.Human.Disposition T037–T047 N/A — UI in 039 Test.Scenario.Capture, Test.Scenario.Vlm, Test.Scenario.Disposition
AGSCN-FR-010 preview (capture) N/A — DTO only captureScenarioScreenshot ScenarioGraph.Capture.Dispatch T037–T039 N/A — UI in 039 Test.Scenario.Capture
AGSCN-FR-011 preview (vlm) N/A — DTO only analyzeScenarioScreenshot ScenarioGraph.Vlm.Analyze T040–T042 N/A — UI in 039 Test.Scenario.Vlm
AGSCN-FR-012 preview (disposition) N/A — DTO only disposeVlmFindings ScenarioGraph.Human.Disposition T043–T044 N/A — UI in 039 Test.Scenario.Disposition
Draft pack (pack compiler) preview (preview_only / save_eligible) N/A — DTO only compileScenarioDraftPack ScenarioGraph.PackCompiler.Generate T026–T030 N/A — UI in 039 Test.Scenario.Pack
Edge E6 (409 stale) preview (stale409 conflict panel) N/A — DTO only resolveDashboardScenario ScenarioGraph.Resolver.Resolve T023 N/A — UI in 039 Test.Scenario.Resolver.Edge
Edge E9 (429) preview (rate limited) N/A — DTO only any scenario op ScenarioGraph.Api T031–T032 N/A — UI in 039 Test.Api.Scenarios.Edge
Edge E11 (malformed VLM) preview (vlm inconclusive) N/A — DTO only analyzeScenarioScreenshot ScenarioGraph.Vlm.Analyze T040 N/A — UI in 039 Test.Scenario.Vlm.Edge
Edge E13 (injection) preview (pack blocked) N/A — DTO only compileScenarioDraftPack ScenarioGraph.PackCompiler.Generate T026, T034 N/A — UI in 039 Test.Scenario.Pack.Security
NFR: determinism N/A — infra N/A — infra N/A — no API ScenarioGraph.Compiler.Compile T012, T021 N/A — backend-only Test.Scenario.Serializer
NFR: RBAC N/A — infra N/A — infra N/A — security ScenarioGraph.Api T032 N/A — backend-only Test.Api.Scenarios.Rbac

N/A Rationale Key

  • N/A — UI in 039: 038 renders no UI; it supplies DTOs consumed by feature 039 preview
  • N/A — DTO only: Screen Model pattern not needed; state lives in 039 components bound to 038 DTOs
  • N/A — infra: Shared infrastructure, not user-facing
  • N/A — no API: Determinism/RBAC are cross-cutting invariants, not endpoints
  • N/A — backend-only: No frontend task for backend-only work

Impact Analysis Quick Reference

If you change... These fixtures verify it These tests verify it These screens depend
ScenarioGraph.Compiler.Compile FX_Scenario.Valid, FX_Scenario.Shuffled Test.Scenario.Compiler, Test.Scenario.Serializer preview (039)
ScenarioGraph.Validator.Validate FX_Scenario.Cycle, FX_Scenario.MissingRef, FX_Scenario.RawBaseline Test.Scenario.Validator preview (039)
ScenarioGraph.PackCompiler.Generate FX_Scenario.PreviewOnly, FX_Scenario.InjectedCode Test.Scenario.Pack preview (039)
ScenarioGraph.Vlm.Analyze FX_Scenario.VlmTyped, FX_Scenario.VlmStalePrompt Test.Scenario.Vlm preview (039)
contracts/openapi.yaml FX_Api.Compile.* Test.Api.Scenarios preview (039)
ScenarioGraph.Catalog.Load FX_Scenario.Catalog19 Test.Scenario.Catalog preview coverage (039)

Cross-Spec Pipeline Traceability

Stage / requirement Authoritative contract Downstream owner Evidence
compile handle + canonical content hash ScenarioGraph.Compiler.Compile, ScenarioGraph.Serializer.Canonical 042 ScenarioRevision Test.Scenario.Compiler, Test.Scenario.Serializer
validation result bound to compiled hash ScenarioGraph.Validator.Validate 038 pack eligibility Test.Scenario.Validator, Test.Scenario.Validator.Edge
unresolved needs_context / needs_selector / needs_baseline ScenarioGraph.ServerOwnedPipeline 042 save; 044 RunPreflight Test.Scenario.Capability, Test.Scenario.Resolver, test_scenario_runner
draft-pack handle and server digest ScenarioGraph.PackCompiler.Generate 042 server-owned save Test.Scenario.Pack, Test.Scenario.Pack.Security, test_scenario_registry
no client graph/digest/runner plan authority ScenarioGraph.ServerOwnedPipeline 050 MCP transport injected-code, path-traversal, and MCP raw-graph rejection tests
persistent co-authoring and sandbox ScenarioGraph.AgentAuthoringWorkspace 050 MCP -> 042 registry session persistence, sandbox isolation/limits, receipt and CAS tests
proposal -> compile/validate -> review -> save ScenarioGraph.AgentAuthoringWorkspace 042 ScenarioRegistry.SaveContinuation typed-candidate, diff-review, promotion-boundary E2E

The 038 gate is necessary but not sufficient for the coordinated path: a validated compiler result cannot be released unless 042 save/revision evidence, 044 preflight/runner/evidence evidence, and 050 MCP parity evidence also pass.

Authoring E2E Release Gate

The coordinated authoring path is persistent workspace -> isolated exploration -> typed proposal -> 038 compile/validate -> user diff review -> 042 handle-based save -> immutable revision -> 044 run. Release is NO-GO if any stage accepts raw code, URLs, cookies, secrets, paths or caller digests, lacks CAS/idempotency or receipts, permits production sandbox effects, or lets a non-promoted artifact reach ScenarioRun. This gate is normative and does not claim implementation completion.

Coverage Gate

  • Every user story (US1–US5) has at least one row
  • Every functional requirement (AGSCN-FR-001..012) has at least one row
  • Every API endpoint has at least one row for success AND at least one row for an error state (E6 409, E9 429, E11, E13)
  • Every Screen Model column is N/A with rationale (038 is DTO-only; 039 owns rendering)
  • Every N/A cell carries a rationale from the key above
  • Every contract referenced appears in contracts/modules.md
  • Every task ID (Txxx) appears in tasks.md
  • Impact table covers every contract with downstream dependents

Module Reuse Annex (LLM verification tooling)

038 orchestration layers reuse existing plugin/service contracts — see research.md §9 and contracts/modules.md (ScenarioGraph.Capture.Dispatch, ScenarioGraph.Vlm.Analyze):

038 contract Reused existing contract Rationale
ScenarioGraph.Capture.Dispatch Plugin.Service.ScreenshotService, Services.AgentRuns.Evidence One Playwright/CDP capture path; evidence auditability
ScenarioGraph.Vlm.Analyze Plugin.Service.LLMClient, Services.LlmProvider.LLMProviderService, Plugin.Service.RedactionService One local provider client; multimodal-required validation; optional display redaction; secret-safe persistence

#endregion DashboardScenarioModel.Traceability


TASKS — Implementation Tasks

Source: tasks.md

#region DashboardScenarioModel.Tasks [C:3] [TYPE ADR] [SEMANTICS tasks,scenario,implementation] @BRIEF Ordered TDD backlog for deterministic scenario graph, safe draft-pack compilation, capture/VLM/disposition, and full verification gates.

Prerequisites: plan.md, spec.md (required); contracts/modules.md, contracts/openapi.yaml, traceability.md, prototype/manifest.md (present). Tests: Write tests FIRST (fail before implementation) for every C3+ contract per constitution VII. Frontend tasks are N/A — 038 is DTO-only; UI rendering is owned by 039.

Format: - [ ] T### [P] [USx] Description with exact file path

Phase 1 — Setup (Shared Infrastructure)

  • T001 Transcribe checklist-catalog.md into a versioned declarative resource in backend/src/services/dashboard_testing/scenario/catalog_v1.yaml
  • T002 [P] Write catalog completeness tests for B01–B09, C01–C07, T01–T03 in backend/tests/services/dashboard_testing/scenario/test_catalog.py (7 passed)
  • T003 [P] Create valid 18-step scenario plus invalid cycle/missing-ref/duplicate-output/raw-baseline/SQL fixtures under specs/038-dashboard-scenario-model/fixtures/ (6 fixtures)
  • T004 [P] Materialize fixtures from specs/038-dashboard-scenario-model/fixtures/ into backend/tests/fixtures/dashboard_scenarios/ (copy as-is, no adaptation)
  • T005 Implement Pydantic models in backend/src/services/dashboard_testing/scenario/models.py matching contracts/dashboard-test-scenario.schema.json (7 tests passed)

Checkpoint: Catalog loads 19 cases; schema round-trips against golden fixtures.

Phase 2 — US1 Compile Scenario Graph

  • T006 [US1] Write failing capability mapping tests in backend/tests/services/dashboard_testing/scenario/test_capability_mapper.py (8 passed)
  • T007 [US1] Implement backend/src/services/dashboard_testing/scenario/checklist_catalog.py validation and backend/src/services/dashboard_testing/scenario/capability_mapper.py with complete classification
  • T008 [US1] Cover unavailable XLSX, missing selector/test data, unsafe mutation context, cross-dashboard absence, and T01–T03 no-SQL fallbacks @TEST_EDGE: xlsx_unavailable→manual/unsupported, technical_without_dataset_fields→human checkpoint (no SQL)
  • T009 [US1] Write failing deterministic compiler tests in backend/tests/services/dashboard_testing/scenario/test_compiler.py (8 passed)
  • T010 [US1] Implement registered tool/action and step-template catalogs under backend/src/services/dashboard_testing/scenario/templates/
  • T011 [US1] Implement backend/src/services/dashboard_testing/scenario/compiler.py with stable ids, phase order, refs, coverage, and fingerprints @PRE: intent, query model, catalog, baseline summary, parameters have valid fingerprints @POST: same canonical inputs/compiler version yield byte-identical graph and stable ids/order @DATA_CONTRACT: CompileScenarioRequest → DashboardTestScenario @TEST_EDGE: missing_selector→NEEDS_SELECTOR + save blocker, missing_baseline→NEEDS_BASELINE (no embedded numeric truth)
  • T012 [US1] Prove repeated compile and shuffled input order produce identical graph bytes @INVARIANT: Deterministic_Graph → VERIFIED_BY: repeated_compile, shuffled_input_order

Checkpoint: Valid fixture compiles to stable graph and classifies all 19 cases. ✅ (30 scenario tests green)

Phase 3 — US2 Validate Safety and Completeness

  • T013 [US2] Write failing full invalid-fixture matrix in backend/tests/services/dashboard_testing/scenario/test_validator.py (8 passed)
  • T014 [US2] Implement schema, ref producer/consumer, duplicate, dependency, and cycle checks in backend/src/services/dashboard_testing/scenario/validator.py @POST: valid is true only with zero errors/blockers; findings stably ordered and actionable @TEST_EDGE: cycle→error contains cycle path, duplicate_output→both producer ids reported
  • T015 [US2] Implement parameter, selector, baseline, tool/action, path, SQL/code, raw-expected, and coverage checks @TEST_EDGE: raw_metric_expected→forbidden baseline literal error, unreachable_step→warning/error per coverage
  • T016 [US2] Return deterministic all-findings output with JSON pointers and recovery options
  • T017 [US2] Add property tests generating small DAG/cycle/ref variations without mirroring validator logic in backend/tests/services/dashboard_testing/scenario/test_validator_properties.py (5 passed)
  • T017b [P] [US2] Add belief-runtime instrumentation tests for ScenarioGraph.Validator.Validate in backend/tests/services/dashboard_testing/scenario/test_validator_belief.py (1 passed) @POST: REASON logged before mutation boundary; REFLECT after; belief_scope wraps validator run

Checkpoint: Invalid fixture matrix passes; no SQL/raw-baseline/cycle escapes. ✅ (47 scenario tests green, belief audit 0 errors)

Phase 4 — US3 Checklist Coverage and Serialization

  • T018 [US3] Write JSON/YAML golden tests in backend/tests/services/dashboard_testing/scenario/test_serializer.py (5 passed)
  • T019 [US3] Implement backend/src/services/dashboard_testing/scenario/serializer.py and revision hash exclusions @POST: key/order/decimal/date/newline rules stable across runs; JSON and YAML equal domain data @TEST_EDGE: shuffled_dicts→identical bytes, timestamp_display_field→excluded from revision identity
  • T020 [US3] Validate JSON Schema and Pydantic round-trip for all golden fixtures
  • T021 [US3] Add catalog-version and compiler-version fingerprints to scenario inputs

Checkpoint: Byte-stable snapshots across runs. ✅ (52 scenario tests green, belief audit 0 errors)

Phase 5 — US4 Parameters and Human Checkpoints

  • T022 [US4] Write failing typed resolution/stale revision tests in backend/tests/services/dashboard_testing/scenario/test_resolver.py (6 passed)
  • T023 [US4] Implement backend/src/services/dashboard_testing/scenario/resolver.py for parameter, selector, manual conversion, and remove-step operations @PRE: base revision hash matches; changes target declared unresolved items @POST: unrelated step ids/order unchanged; new parent/revision hashes link revisions @TEST_EDGE: stale_base_revision→409, invalid_parameter_type→422, unrelated_graph_change→invariant failure
  • T024 [US4] Enforce immutable revisions and unchanged unrelated step ids/order
  • T025 [US4] Cover safe-environment/test-data requirements for mutating PDF cases

Checkpoint: Resolution produces linked immutable revisions; unrelated structure stable. ✅ (58 scenario tests green)

Phase 6 — Safe Draft Pack

  • T026 Write failing template registry, preview-only, path, SQL, shell, and injected-code tests in backend/tests/services/dashboard_testing/scenario/test_pack_compiler.py (5 passed) @TEST_INVARIANT: No_LLM_To_Code → VERIFIED_BY: injected_code_field, template_registry_only @TEST_EDGE: unknown_template→blocked, path_traversal→blocked before artifact registration
  • T027 [P] Create versioned templates for scenario.yaml, runner.plan.json, report_template.md, evidence_manifest.json, and bounded browser/XLSX modules under backend/src/services/dashboard_testing/scenario/pack_templates/v1/
  • T028 Implement backend/src/services/dashboard_testing/scenario/pack_compiler.py; accept only registered template ids and structured inputs @PRE: scenario validation result available; template ids registered; target paths safe @POST: outputs match ArtifactPlan, contain no LLM executable bodies, and are registered as 036 drafts @SIDE_EFFECT: renders bounded templates; calls AgentRuns.Artifacts.Register @INVARIANT: errors/unresolved required inputs make pack preview_only
  • T029 [P] Register outputs through 036 AgentRuns.Artifacts.Register and emit generate/validate progress (pack_registry.py — register_pack_drafts via 036 register_draft; save_eligible only)
  • T030 Ensure invalid/unresolved graphs produce preview_only with repeated blockers

Checkpoint: Valid graph → save_eligible pack; invalid/injected → preview_only/blocked. ✅ (63 scenario tests green, belief audit 0 errors)

Phase 7 — API, Agent, Quality

  • T031 Add backend/src/api/routes/dashboard_scenarios.py matching contracts/openapi.yaml and register router
  • T032 Write RBAC/contract/revision tests in backend/tests/api/test_dashboard_scenarios.py (5 passed) @TEST_EDGE: 401→UNAUTHORIZED, 403→FORBIDDEN (per scope), 409 stale→STALE_REVISION, 422→VALIDATION_ERROR, 429→RATE_LIMITED
  • T033 [P] Add thin compile/validate/resolve/generate tools in agent/src/ss_tools/agent/tools_038.py (4 tools registered; 36 total)
  • T034 Verify agent schemas cannot carry code, SQL, raw expected metrics, custom tools, or artifact paths (extra="forbid" + bounded JSON fields; versioned ActionRegistry; SqlEvidenceSpec permitted only through compilation gate)
  • T035 Run quickstart, JSON/OpenAPI schema validation, scoped/full backend tests, and ruff (68 backend + 20 agent tests green; ruff clean backend + agent)
  • T036 Audit all 19 cases, arbitrary SQL/code bans, SqlEvidenceSpec compilation gate, contract anchors, ATTN_1–4, and unresolved relations (belief audit 0 errors; regions balanced)

Checkpoint: API + agent tools match openapi.yaml; RBAC enforced. ✅ (88 tests green)

Phase 8 — Screenshot Capture, VLM Analysis, and Human Disposition (AGSCN-FR-010..012)

  • T037 [P] Write failing capture spec validation tests in backend/tests/services/dashboard_testing/scenario/test_capture.py (5 passed)
  • T038 [P] Create backend/src/services/dashboard_testing/scenario/capture_profile.py — load and validate CaptureProfile from contracts/capture-profile.schema.json
  • T039 [P] Implement ScenarioGraph.Capture.Dispatch: accept CaptureSpec from step, call AgentRuns.Evidence.Adapter, register screenshot artifacts, emit evidence_captured (capture.py — dispatch_capture via 036 register_screenshot_draft/register_masked_derivative; 2 dispatch tests) @POST: ScreenshotEvidence DraftArtifact registered; evidence_captured event emitted @TEST_EDGE: capture_timeout→step inconclusive; no artifact registered, masking_applied→original + masked artifacts
  • T040 [P] Write failing VLM analysis tests in backend/tests/services/dashboard_testing/scenario/test_vlm.py for typed findings, provenance, stale prompt rejection (5 passed) @TEST_EDGE: stale_prompt→422 STALE_PROMPT, vlm_timeout→inconclusive, empty_response→empty findings + inconclusive
  • T041 Implement backend/src/services/dashboard_testing/scenario/vlm.py: submit the configured screenshot (masked only when requested) to the local VLM provider, parse typed VlmFinding[], validate model_provenance, and persist provider output under secret-safe handling @POST: returns typed VlmFinding[] with model/prompt provenance; raw response stored with credentials/cookies/tokens excluded @SIDE_EFFECT: enterprise-local VLM provider call; raw response persisted as separate DraftArtifact @INVARIANT: VLM findings advisory; never alter metric baseline truth; stale prompts block analysis @REJECTED: embedding VLM findings directly as assertion results (observations, not deterministic pass/fail)
  • T042 Create registered VLM prompt template v1 under backend/src/services/dashboard_testing/scenario/prompt_templates/v1/ with versioned hash
  • T043 [P] Write failing human disposition tests in backend/tests/services/dashboard_testing/scenario/test_disposition.py for confirm/dismiss/inconclusive, double-disposition rejection (5 passed) @TEST_EDGE: double_disposition→409, confirm_requires_comment→422
  • T044 Implement backend/src/services/dashboard_testing/scenario/disposition.py: record immutable disposition per finding id, enforce confirm requires non-blank comment, emit audit event @POST: each finding disposition set exactly once; step transitions per policy @SIDE_EFFECT: audit record of disposition decision @INVARIANT: disposition never alters scenario graph structure or step ordering
  • T045 Add VlmFinding and HumanDisposition DTOs to contracts/openapi.yaml response schemas (verify round-trip)
  • T046 Wire capture/VLM/disposition into the scenario step execution loop: screenshot step → capture → analysis step → VLM call → human step → disposition @NOTE: runtime execution loop is OWNED by 044. This task is redefined as: verify the compiler emits valid ScreenshotCaptureSpec/VlmAnalysisSpec DTOs that 044 executors can consume; the 038 loop contracts are superseded by 044 ScenarioExecution.
  • T047 Audit: VLM findings are advisory, never alter metric baseline truth; disposition never changes graph structure; stale prompts block analysis

Checkpoint: Capture/VLM/disposition flow verified end-to-end; typed findings auditable. ✅ (85 tests green, belief audit 0 errors)

Phase 9 — Polish & Cross-Cutting Verification

  • T048 [P] Prototype validation: verify every @UX_STATE in contracts/ux/scenario-graph-ux.md reachable via specs/038-dashboard-scenario-model/prototype/index.html state switcher; responsive on mobile viewport (24/24 state classes covered; 18 dedicated sections + grouped/aria-live)
  • T049 [P] OpenAPI drift check: verify operationId uniqueness (7), $ref resolution, example coverage, and RBAC scopes in contracts/openapi.yaml against implemented endpoints in backend/src/api/routes/dashboard_testing/scenario.py (all 7 paths implemented: compile/validate/resolve/draft-pack/capture/vlm/disposition)
  • T050 [P] Belief runtime audit (C4/C5): axiom_audit({operation="audit_belief_runtime"}) + axiom_audit({operation="audit_belief_protocol"}) — confirm Compiler/Validator/Mapper/PackCompiler/Vlm/Capture/Resolver contracts have @RATIONALE/@REJECTED and REASON/REFLECT/EXPLORE markers (audit: 0 errors; 4 C5 + 9 C4 contracts instrumented)
  • T051 [P] Attention compliance audit: verify ATTN_1–4 per semantics-core §VIII across contracts/modules.md
  • T052 [P] Semantic index rebuild: axiom_search({operation="rebuild", rebuild_mode="full"}) — 0 parse warnings required (rebuild completed; 8094 contracts, 3804 edges)
  • T053 [P] Orphan audit: axiom_search({operation="workspace_health"}) — confirm no new orphans from this feature (scenario scope: 0 orphans, 0 unresolved relations)
  • T054 [P] Traceability coverage gate: verify traceability.md rows all map to real task IDs, contracts, and operationIds
  • T055 [P] Run quickstart.md validation and make test-related F=backend/src/services/dashboard_testing/scenario/ for regression scope (91 backend + 20 agent tests green)
  • T056 Run /speckit.validate — confirm PASS before /speckit.implement (validation.md updated with final digests; PASS)

Phase 10 — Runtime Closure MOVED TO 044 (reconciliation 2026-08-07)

Context: 044 ScenarioExecution owns all runtime capture/VLM/disposition/evidence. The former 038 T057–T059 (real VLM submit, real capture, e2e evidence) are moved to 044 and are NOT 038 responsibilities. 038 keeps VlmAnalysisSpec/ScreenshotCaptureSpec definitions only.

  • (MOVED) Real VLM submit + real capture + e2e evidence → implemented under 044 ScenarioExecution (owner_type=scenario_run artifacts, HumanCheckpoint). 038 runtime stubs are out of scope.

Dependencies

T001–T005 → US1 → US2; US3 follows compiler; US4 follows validator; draft pack follows validation/resolution; Phase 8 depends on 036 Phase 8 (screenshot evidence artifacts) and 037 Phase 7 (visual baseline infrastructure). Phase 10 (T057–T059) is the MVP runtime closure: VLM/capture remain stubs until merged. 039 begins only after ScenarioResponse and DraftPack fixtures are stable. Frontend tasks N/A — 038 is DTO-only.

#endregion DashboardScenarioModel.Tasks


PROTOTYPE — State/Manifest

Source: prototype/manifest.md

#region DashboardScenarioModel.PrototypeManifest [C:3] [TYPE ADR] [SEMANTICS prototype,manifest,scenario] @defgroup Prototype Interactive HTML prototype manifest for the DashboardTestScenario preview.

Prototype Metadata

  • Feature: 038 Dashboard Scenario Model
  • Source contracts: contracts/ux/scenario-graph-ux.md, contracts/ux/api-ux.md, ux_reference.md
  • Design system source: frontend/tailwind.config.js + frontend/src/lib/ui/*.svelte (verbatim class recipes)
  • Screens represented: 1 (Scenario Preview)
  • Total states: 18 (10 core + 8 grouped edge classes)
  • Accessibility validations: keyboard nav (Tab/Enter/Space), ARIA roles + aria-live regions, focus-visible rings, ≥44×44px touch targets, prefers-reduced-motion
  • Responsive breakpoints: 375px (mobile), 1280px (desktop)

State Coverage

Screen @UX_STATE / State Class Prototype State Reachable? Recovery Path
Scenario Preview idle (no scenario) idle ✅ Compile CTA
Scenario Preview loading (compile) loading ✅ —
Scenario Preview loaded (18 steps) loaded ✅ —
Scenario Preview empty (no applicable cases) empty ✅ Add manual checkpoint / accept partial
Scenario Preview error (5xx) error ✅ Try again / contact support
Scenario Preview blocked (NEEDS_SELECTOR / NEEDS_BASELINE) blocked ✅ Provide hint / convert / 037 discovery
Scenario Preview stale revision (409 CONF_01) stale409 ✅ Recompile / snapshot diff / discard
Scenario Preview draft pack preview_only preview_only ✅ Review blockers → save_eligible
Scenario Preview draft pack save_eligible save_eligible ✅ Save draft
Scenario Preview VLM findings + disposition vlm ✅ Confirm/dismiss/inconclusive
Scenario Preview NET_01 offline / NET_02 timeout / NET_03 retry net ✅ Auto-retry / manual retry
Scenario Preview VAL_01 field / VAL_02 cross-field / 422 val ✅ Correct fields + re-submit
Scenario Preview AUTH_01 401 / AUTH_02 403 auth ✅ Login / contact admin
Scenario Preview NF_01 404 notfound ✅ Navigate to scenario list
Scenario Preview 429 rate limited ratelimit ✅ Wait Retry-After
Scenario Preview DUP_01 double submit / DUP_02 interruption duplicate ✅ Disabled button / stay or discard
Scenario Preview LARGE 100+ steps large ✅ Pagination/refinement
Scenario Preview MALFORMED VLM response malformed ✅ Re-run analysis / support with error ID
Scenario Preview CONF_02 duplicate pack (idempotent) — (in loaded/save_eligible) ✅ Transparent; returns existing DraftPack
Scenario Preview STALE baseline — (in blocked) ✅ 037 discovery / mark pending
Scenario Preview PARTIAL graph load — (in blocked) ✅ Per-step retry / reload all
Scenario Preview A11Y announcements — (aria-live in all) ✅ Built into transitions
Scenario Preview RESP responsive — (viewport toggle) ✅ Built into layout

Screen ↔ Story Traceability

Prototype Screen User Story UX Contract Acceptance Criteria Verified
Scenario Preview US1 Build Graph GraphUx AC1: graph summary+steps+params+warnings
Scenario Preview US2 Validate Safety GraphUx AC3: unsupported tool blocked with reason
Scenario Preview US3 Checklist Coverage GraphUx AC2: unsupported/manual with rationale
Scenario Preview US4 Parameters/Checkpoints GraphUx AC1: parameters structural; AC3: resolution affects only dependent steps
Scenario Preview US5 Capture/VLM/Disposition GraphUx AC2: typed findings; AC3: disposition auditable

Validation Results

  • All @UX_STATE contracts reachable via state switcher
  • All @UX_RECOVERY paths traversable
  • Keyboard navigation: Tab order verified
  • Touch targets: ≥44×44px on mobile viewport
  • ARIA: live regions for loading/error states
  • No broken links or dead-end states
  • Responsive layout: mobile viewport collapses lanes (no overflow)

Design System Reuse

All components replicated with verbatim class strings from frontend/src/lib/ui/*.svelte (no invented styles).

Element Source Prototype Mapping (class-for-class)
Button $lib/ui/Button.svelte inline-flex items-center justify-center font-medium transition-colors focus-visible:outline-none focus-visible:ring-2 focus-visible:ring-offset-2 disabled:pointer-events-none disabled:opacity-50 rounded-md + bg-primary text-white hover:bg-primary-hover focus-visible:ring-primary-ring h-10 px-4 py-2 text-sm (md) / h-8 px-3 text-xs (sm) / ghost: bg-transparent hover:bg-ghost-hover text-ghost-text / loading spinner animate-spin
Card $lib/ui/Card.svelte rounded-lg border border-border bg-surface-card text-text shadow-sm + p-6 (md padding); title text-lg font-semibold leading-none tracking-tight
Badge $lib/ui/Badge.svelte inline-flex items-center gap-1.5 wrapper + rounded-full px-2.5 py-1 text-xs font-medium bg-{variant}-light text-{variant}
Skeleton $lib/ui/Skeleton.svelte animate-pulse bg-surface-muted rounded (line) / rounded-lg h-24 (card)
EmptyState $lib/ui/EmptyState.svelte flex flex-col items-center justify-center py-12 px-4 text-center + w-16 h-16 text-text-subtle mb-4 icon + text-lg font-semibold text-text mb-1 title + text-sm text-text-muted max-w-md desc
PageHeader $lib/ui/PageHeader.svelte flex items-center justify-between mb-8 + text-3xl font-bold tracking-tight text-text title + text-sm text-text-muted subtitle
Input (error state) $lib/ui/Input.svelte flex h-10 w-full rounded-md border border-border-strong bg-surface-card px-3 py-2 text-sm text-text ring-offset-white placeholder:text-text-subtle focus-visible:outline-none focus-visible:ring-2 focus-visible:ring-primary-ring focus-visible:ring-offset-2 + border-destructive on error + text-xs text-destructive message
Table page table convention border-b border-border rows, muted uppercase text-xs headers

Design Token Audit (MANDATORY)

Every hex value used in the prototype traces to frontend/tailwind.config.js. Verified by script — 0 unknown hex values.

Token (tailwind.config.js) Hex / Value Used in prototype (elements)
primary.DEFAULT #2563eb primary buttons (bg-primary), active states
primary.hover #1d4ed8 primary button hover (hover:bg-primary-hover)
primary.ring #3b82f6 focus rings (focus-visible:ring-primary-ring), switcher focus
primary.light #eff6ff bg-primary-light badge variant
secondary.DEFAULT / secondary.hover / secondary.text #f3f4f6 / #e5e7eb / #111827 secondary buttons
destructive.DEFAULT / hover / light #dc2626 / #b91c1c / #fef2f2 destructive buttons, error alerts, border-destructive, error badges
destructive.ring #ef4444 destructive focus ring
success.DEFAULT / light #22c55e / #f0fdf4 ready/automated badges
warning.DEFAULT / light #f59e0b / #fffbeb warning badges, border-warning, stale/blocker alerts
info.DEFAULT / light #0ea5e9 / #f0f9ff tool/status badges, NEEDS_* info badges
ghost.hover / ghost.text / ghost.ring #f3f4f6 / #374151 / #6b7280 ghost buttons (Remove step, Inconclusive, Discard)
surface.page #f8fafc page background (body)
surface.card #ffffff card background (bg-surface-card), ring offset
surface.muted #f1f5f9 skeletons, lane nodes, code bg
border.DEFAULT #e2e8f0 card/lane borders, table rows, switcher border
border.strong #cbd5e1 input borders (border-border-strong)
text.DEFAULT #0f172a body/card text (text-text)
text.muted #64748b secondary text, table headers, badges muted
text.subtle #94a3b8 icons, placeholders
text.inverse #ffffff text-white on primary/destructive buttons
fontFamily.mono JetBrains Mono / Fira Code code blocks (provenance, hashes)
width.sidebar 240px N/A — 038 preview is embedded (039 owns layout)

Audit gate: script scan of index.html found 0 hex colors outside the token table and 0 body classes without a shim rule. Every production class string copied verbatim from Svelte sources.

Browser Validation Notes

Validated via chrome-devtools MCP: all 18 states toggle; viewport switch at 375px stacks lanes; aria-live regions announce loading/error; focus-visible rings (primary/secondary/destructive/ghost) visible on all controls; prefers-reduced-motion honored. Prototype is self-contained (no external deps, no build step). Visual fidelity vs production components: colors/radius/shadows identical to tailwind.config.js tokens and ui/*.svelte class recipes.

#endregion DashboardScenarioModel.PrototypeManifest


PROTOTYPE — Interactive HTML

Source: prototype/index.html

<!doctype html><html lang="ru"><head></head>

Superset Tools · BI testing
СценарииЗапускиАвтоматизацияКачествоAuthoring validation
Сценарий до сохранения

Проверяемая модель сценария

Аналитик видит бизнес-цель и безопасный граф, а не generated code.

Проверить для сохранения

Граф

1
Открыть FI-0080
browser / open_dashboard · read
Registered
2
Применить filters
browser / apply_native_filter · read
Registered
3
Скачать XLSX
browser / download · read
Registered
4
Сравнить baseline
assertion / compare_to_baseline
Registered
State: Save eligibleBlocked
<script src="../../prototype-ui.js"></script><script>protoState('eligible',s=>{let v=document.getElementById('validation');v.className='notice '+(s==='blocked'?'danger':'info');v.innerHTML=s==='blocked'?'Authoring blocked
Неизвестный selector или baseline — исправьте модель.':'Authoring valid
Можно собрать DraftPack и сохранить revision.'})</script></html>


checklist-catalog.md

Source: checklist-catalog.md

#region DashboardScenarioModel.ChecklistCatalog [C:4] [TYPE ADR] [SEMANTICS scenario,checklist,catalog,pdf] @BRIEF Normalized reusable catalog of all 19 cases from research/Чеклист 29.05 (1).pdf. @RELATION DEPENDS_ON -> [DashboardScenarioModel.DataModel] @RATIONALE Separating reusable intent from FI-0080 data and historic outcomes prevents the source PDF from becoming hidden implementation logic. @REJECTED Copying historic pass/fail text into expected assertions — it is prior evidence, not future truth.

Catalog version: 1
Source date/title: 29.05.2026, “Чеклист 29.05”
Rule: every case is classified; none is silently omitted.

Basic Cases

ID Reusable goal Required capabilities Parameters/evidence Preferred mapping
B01 Main dashboard filters change report data consistently native_filters, browser filter values; before/after state; API evidence when available browser apply + superset_api observe + assertion
B02 Text-search filters display and apply entered value text_filter, browser text values; visible chip/input and resulting rows browser + screenshot/assertion
B03 In-table column filter returns matching rows/count table_filter, browser column/value; row/count evidence browser + table assertion
B04 Pagination/page-size changes update rows correctly pagination, browser target page/size; row identity evidence browser + assertion
B05 Single-row comment/reason/status persists after save row_edit, browser row key, comment, reason, status browser interaction + refresh evidence; human if write test data unavailable
B06 Bulk edit applies values to all selected rows bulk_edit, browser at least two row keys and values browser + post-save evidence
B07 Blank bulk fields do not overwrite existing values bulk_edit, persistence_refresh rows with existing values browser + before/after assertion
B08 Bulk edit works for more than 100 records bulk_edit, browser safe dataset/filter producing >100 rows browser + count/evidence; manual if safe test data absent
B09 Saved comment remains visible after page refresh row_edit, persistence_refresh row key and expected submitted text browser refresh + assertion/screenshot

Complex Cases

ID Reusable goal Required capabilities Parameters/evidence Preferred mapping
C01 Combined dashboard and table filters remain consistent during edit native_filters, table_filter, row_edit filter set, row key, edit values browser chain + evidence
C02 “Needs filling” time-window behavior changes after 1–7 and over 8 days time_rollover, native_filters controlled dates/test data scheduled/manual checkpoint unless safe clock fixture exists
C03 Rolling comments remain distinct per business date time_rollover, row_edit two dates, same row key, two comments browser + Superset API evidence when exposed
C04 XLSX export completes and yields readable workbook xlsx_export, browser export action; file metadata browser download + xlsx parse
C05 XLSX export reflects dashboard-level filters xlsx_export, native_filters filter set; API/UI/XLSX values browser + superset_api + xlsx + assertion
C06 XLSX export reflects table-level filters xlsx_export, table_filter table filter; visible and workbook rows browser + xlsx row-set assertion
C07 Related dashboard/tab displays expected persisted comments cross_dashboard, browser target dashboard/tab and row key browser navigation + evidence; needs_context if relation absent

Technical Cases

ID Reusable goal Required capabilities Safe mapping Forbidden mapping
T01 Verify persisted comment fields and author/timestamp semantics dataset_field_read or human Superset saved dataset/chart result; otherwise human checkpoint SQL Lab or generated SELECT
T02 Verify rolling records share row identifier but retain distinct dates/values dataset_field_read or human Superset dataset result and row-set assertion; otherwise human SQL text
T03 Verify blank bulk edit did not create invalid empty/zero combinations dataset_field_read or human Superset dataset result with structural assertion; otherwise human SQL text

Historic Source Notes

The PDF records historic outcomes such as “Успешно пройдено”, known remarks for B05/B07, and failures for C02/T03. These notes may inform warnings or regression rationale but are never copied as baseline values or current expected outcomes.

Mapping Rules

  1. Missing required capability → human_checkpoint when safe manual verification exists, otherwise unsupported.
  2. Missing parameter, selector, relationship, or test data → needs_context.
  3. Mutating cases B05–B09/C01/C03 require a 038 mutation_contract: safe fixture, bounded record keys, allowed non-PROD environment, cleanup/reconciliation and a side-effect key. They are manual/needs_context without it; no mutating browser step is automated in PROD.
  4. T01–T03 must satisfy the 037 no-direct-SQL invariant.
  5. C02 must not simulate elapsed time against production records without an approved test fixture.
  6. Each mapping records selected template, rationale, and resulting step ids.

#endregion DashboardScenarioModel.ChecklistCatalog


validation.md

Source: validation.md

#region DashboardScenarioModel.ValidationReport [C:3] [TYPE ADR] [SEMANTICS validation,gate,scenario] @defgroup Validation Pre-implementation validation gate for the Dashboard Scenario Model (compiler-layer scope).

Status: VALID (compiler-layer scope) — generated 2026-08-11

Date: 2026-08-11 (generated from current artifacts by reconcile_contracts.py 036-047) Feature: 038 Dashboard Scenario Model Scope: Compiler/IR layer ONLY. Runtime capture/VLM/disposition/evidence is owned by 044.

History: Prior versions of this file self-contradicted (claimed PASS covering full implementation while runtime T057–T059 were open, and asserted YAML validity that the parser rejected). This revision is regenerated from the machine-readable artifacts and is compiler-scope only.

Machine-Contract Verification (2026-08-11)

Check Result
OpenAPI YAML parse (contracts/openapi.yaml) ✅ parses (fix: quoted flow-scalar with : on disposition description)
OpenAPI semantic paths + identity (scenario_key/content_hash, runtime endpoints deprecated) ✅ verified
JSON Schema parse (dashboard-test-scenario.schema.json) ✅ valid JSON
JSON Schema Verification Program (SqlEvidenceSpec, Transform/Comparison/AgentEvaluation specs) ✅ required canonical content, included in content hash contract
ParameterDefinition runtime-state exclusion ✅ no value or status; 044 owns ParameterBinding
Versioned ActionRegistry ✅ current valid fixture actions registered; unknown runtime action blocked
Fixtures validate against JSON Schema ✅ 6/6 (fixtures/api/*.json), after migration to canonical identity
Fixture fingerprint lengths (input_fingerprints.* 64-hex) ✅ fixed (was 66-char parameters)
044 runtime boundary ✅ typed StepOutcome/AgentEvaluation events; no runtime SQL mutation request surface

Compiler-Layer Gate Decision

Verdict: ✅ PASS (compiler/IR scope) — 038 compiler/validator/resolver/canonicalizer + Verification Program schema/action registry + fixtures are internally consistent and machine-validated.

NOT a full-feature PASS: execution (real VLM submit, real capture, evidence owner_type=scenario_run, HumanCheckpoint) is owned by 044 and gated there. Do not read this as authorizing 044 runtime implementation.

Blocking Findings

None for the compiler-layer scope. The prior "runtime closure" findings are relocated to 044 (they are not 038 responsibilities).

Artifact Completeness (actual tree)

Artifact Status
spec.md, plan.md, research.md, data-model.md, tasks.md, traceability.md, quickstart.md, ux_reference.md ✅
checklists/requirements.md ✅
contracts/modules.md, contracts/openapi.yaml, contracts/verification-program.md, contracts/action-registry.yaml ✅
contracts/dashboard-test-scenario.schema.json, contracts/capture-profile.schema.json ✅
contracts/ux/{scenario-graph-ux,alternatives,decisions,api-ux}.md ✅
contracts/openapi-traceability.md ✅
prototype/index.html + manifest.md ✅
checklist-catalog.md ✅
fixtures/api/*.json (6) ✅ validated

Warning

  • W03: Frontend rendering is DTO-only in 038; 039 owns the create UI, 045 owns run monitor. Confirm downstream UI tasks exist (already in 039/045).
  • Runtime endpoints in contracts/openapi.yaml are deprecated (moved to 044); OpenAPI-traceability must reflect this when regenerated.

Generation rule

This report is generated from the current package gate rather than a manually maintained digest table. Any contract change requires rerunning backend/.venv/bin/python reconcile_contracts.py 036-047; a non-zero result invalidates this verdict.

#endregion DashboardScenarioModel.ValidationReport

================================================================================ FEATURE: 039-dashboard-scenario-ui Files: 22


SPEC — Feature Specification

Source: spec.md

#region DashboardScenarioUi.Spec [C:3] [TYPE ADR] [SEMANTICS spec,requirements,ux,scenario,dashboard-testing] @BRIEF Persistent agent workspace for scenario generation, scenario revision saves, evidence-led remediation, and policy-bound approvals. @RELATION DEPENDS_ON -> [Doc.Adr.ADR0001] @RELATION DEPENDS_ON -> [Doc.Adr.ADR0005] @RELATION DEPENDS_ON -> [Doc.Adr.ADR0006] @RELATION DEPENDS_ON -> [AgentTestStabilization.Spec] @RELATION DEPENDS_ON -> [SupersetBaselineEngine.Spec] @RELATION DEPENDS_ON -> [DashboardScenarioModel.Spec] @RATIONALE Users should approve a business-level dashboard test scenario and required parameters; low-level tool chains are visible for trust but selected by the agent and scenario validator. @REJECTED Dropdowns such as "Playwright UI tests" vs "SQL checks" vs "XLSX checks" — rejected because each dashboard requires a unique cross-tool program. Validated immutable SQL evidence is visible in the program, not chosen as a loose UI mode. @REJECTED Persistent chat workspace as the mandatory creation UX — superseded 2026-08-24 by external MCP clients plus non-chat surfaces (042/043) per specs/050-mcp-interface/spec.md; the dashboard entry action becomes an explicit handoff surface.

Navigation (DSA Indexer keywords)

@SEMANTICS: spec, requirements, feature, ux, agent, scenario, dashboard-testing, artifacts, baseline

Feature Branch: 039-dashboard-scenario-ui
Created: 2026-07-07 | Status: Ready for Implementation Input: "Provide the user-facing persistent agent workspace. From a dashboard page the analyst creates or remediates a scenario; the agent analyzes context, proposes a graph, collects authoring context, previews artifacts/evidence, saves delegated revisions, and uses inline approval only where policy requires it."

User Scenarios

Story 1 — Start Scenario From Dashboard Page (P1)

Why P1: The feature must be discoverable from the dashboard the user wants to test and must carry the correct context into the agent.

Independent Test: On a dashboard page, click "Создать сценарий тестирования" and verify /agent opens with dashboard context and scenario intent.

Acceptance:

  1. Given a user is viewing a dashboard When they click "Создать сценарий тестирования" Then /agent opens in Dashboard Test Scenario Agent mode with dashboard id, name, env, route, and intent visible.
  2. Given environment context is missing When the user starts the flow Then the UI prompts for environment selection before scenario analysis.
  3. Given the user lacks scenario-generation permission When they click the action Then a permission-denied recovery path appears and no agent run starts.

Story 2 — Review Proposed Unique Scenario (P1)

Why P1: The user must see the proposed business flow and tool chain before artifact generation.

Independent Test: Feed fixture scenario graph data and verify the UI renders phases, steps, tools, outputs, warnings, blockers, and coverage.

Acceptance:

  1. Given the agent analyzes a dashboard When scenario graph is ready Then the UI shows goal, phases, step table, dependency graph, tool categories, expected results, coverage, warnings, and blockers.
  2. Given some checklist cases are not automatable When scenario preview renders Then they appear as human checkpoints, unsupported, or needs-context with rationale.
  3. Given the scenario includes source-mart evidence When preview renders Then the UI shows its validated SqlEvidenceSpec (relation/connection identity, hash, typed parameters, limits and output schema) and distinguishes it from Superset chart API evidence.

Story 3 — Define Parameters and Baseline Constraints (P1)

Why P1: Reusable dashboard tests need typed parameter definitions and baseline constraints; actual launch values belong to 045 RunPreflight.

Independent Test: Scenario preview renders parameter definitions; user edits default/validation/source and sees affected steps, while a required no-default scenario remains save-eligible.

Acceptance:

  1. Given a scenario declares parameters When preview renders Then each ParameterDefinition has label, type, validation, default/source if available, and affected steps; it contains no runtime value/status.
  2. Given an approved baseline exists for metric+filters When the analyst defines binding constraints Then baseline compatibility is shown without creating a runtime binding.
  3. Given no approved baseline exists When the user proceeds Then the UI offers discovery candidate flow and marks baseline approval as separate HITL action.

Story 4 — Preview Generated Artifacts (P2)

Why P2: Users need to inspect generated files, report templates, and warnings before saving.

Independent Test: Generate draft artifacts from a fixture scenario and verify file tree, validation status, warnings, and content preview render.

Acceptance:

  1. Given draft artifacts are generated When preview opens Then the UI lists scenario file, generated runner/check files, report template, baseline candidate files, and validation status.
  2. Given an artifact has unresolved markers When preview renders Then save is warning-gated and unresolved markers are visible.
  3. Given user downloads draft artifacts When download completes Then repository state remains unchanged.

Story 5 — Save or Approve Through Delegated Policy (P2)

Why P2: Durable actions need attributable policy decisions; the agent may save validated scenario revisions while high-risk baseline/publish actions remain gated.

Independent Test: Save a validated revision under delegated policy and trigger baseline approval; verify immutable provenance and inline gate information where required.

Acceptance:

  1. Given a validated scenario revision is ready When delegated policy permits Then the agent saves it with immutable revision, delegator, action and evidence provenance.
  2. Given baseline approval or another non-delegated action is requested When policy requires it Then an inline card shows target, risk, provenance/diff and a required reason.
  3. Given a gate is denied When denial is submitted Then no gated side effect occurs and the run records cancellation.

Edge Cases

  • Agent analysis fails due to Superset API error → UI shows retry, switch environment, or manual-only scenario recovery.
  • Scenario has unresolved blockers → generation can continue only for safe preview; executable save is blocked until resolved or converted to human checkpoint.
  • Baseline is stale → UI warns and offers discovery candidate flow, not silent update.
  • XLSX export is unavailable → UI shows coverage gap and alternate Superset API/UI checks.
  • User navigates away during generation → run id allows returning to latest draft state from 036.

Requirements

Functional — Agent Workspace (persistent chat and work surfaces)

These requirements cover the persistent agent workspace where the analyst creates, refines, investigates, and remediates scenarios through chat plus structured work surfaces.

  • AGUI-FR-001: Dashboard pages MUST expose a single business-level action labeled "Создать сценарий тестирования"; UI MUST NOT present low-level artifact/tool choices as the primary entry point.
  • AGUI-FR-002: /agent MUST show Dashboard Test Scenario Agent mode when opened with scenario intent from 036.
  • AGUI-FR-003: The UI MUST render structured progress stages from agent metadata: context, inspect, scenario, parameters, generate, validate, save.
  • AGUI-FR-004: The scenario preview MUST show business goal, phases, step list, tool category per step, dependency graph, coverage, warnings, blockers, expected results, and the visible Verification Program (navigation, evidence, transforms, assertions and declared agentic evaluations).
  • AGUI-FR-005: Authoring MUST support typed ParameterDefinition/default/validation/source edits and show affected steps without restarting the flow. Required parameters without defaults remain save-eligible; 045 alone collects runtime ParameterBindings.
  • AGUI-FR-006: Baseline UI MUST show approved baseline matches, stale warnings, missing baselines, and draft candidate approval paths.
  • AGUI-FR-007: Generated artifact preview MUST show file tree, file content preview, validation status, unresolved markers, and warnings before save.
  • AGUI-FR-008: The agent MAY save a validated executable scenario revision under delegated policy. Baseline approval and every non-delegated risky action MUST use a 036 inline ActionApprovalGate; baseline approval requires a reason.
  • AGUI-FR-009: The UI MUST distinguish Superset chart API evidence from source-mart SqlEvidenceSpec executed through the Superset SQL Lab adapter. It MUST show SQL as an immutable validated program artifact, never as runtime/chat-editable text.
  • AGUI-FR-009a: The UI MUST collect and display ChangeRequestContext (request, affected dashboards/fields, technical detail, relations, acceptance criteria and control totals). Missing needed context is explicit needs_context, not agent inference.
  • AGUI-FR-010: All new UI state MUST follow Svelte 5 runes/model-first conventions and remain accessible by keyboard for inline action cards, parameter forms, and previews; modal/dialog interaction MUST NOT be required to complete a workflow.
  • AGUI-FR-011: Screenshot evidence artifacts MUST render in a dedicated EvidencePanel showing the captured image, capture metadata (viewport, filters_hash, timestamp), and any linked VLM findings with severity, region, and confidence.
  • AGUI-FR-012: VLM findings MUST be reviewable with typed disposition controls: confirm (accepts finding as valid), dismiss (marks as false positive), or inconclusive (defers to human checkpoint). Disposition changes MUST be auditable and MUST NOT alter the scenario graph.
  • AGUI-FR-013: The artifact preview file tree MUST include an evidence/ branch listing screenshot artifacts and their associated VLM finding files.

These requirements are satisfied through the DashboardTesting.AgentWorkspace components. The agent is the interaction surface; structured panels (Parameters, ArtifactPreview, EvidencePanel, ActionTimeline) provide data entry and review within the workspace. The agent may perform delegated durable actions, while policy-gated actions remain inline cards in the same thread.

Functional — Pipeline Verification Views (automated, no agent interaction)

These requirements cover views that display the results of automatically triggered verification runs during the release pipeline. These views DO NOT require starting an AgentRun or conversing with the agent. They consume VerificationRun records (037) produced by pipeline hooks.

  • AGUI-FR-014: The PREPROD deployment status page MUST display a dedicated Verification tab showing: StructureDiff summary, metric comparison results, visual findings status, and overall VerificationRun status. CRITICAL structural changes MUST block the «Validate» action. The tab is populated from the VerificationRun created automatically by the deploy_to_preprod hook — no agent interaction required.
  • AGUI-FR-015: The release detail page MUST show verification status inline: StructureDiff summary between this release and the previous one, metric pass/warn/fail counts, immutability violation alerts for closed periods. Release approval MUST be gated on verification status. All data comes from the VerificationRun attached to the release.
  • AGUI-FR-016: The dashboard page MUST show verification history: a chronological list of VerificationRun entries with trigger, environment, overall status, and a link to the full run. The latest run status MUST be visible as a badge next to the release version. This is a read-only projection of the verification log.

These requirements are satisfied through the ReleaseVerification components. They bind to VerificationRun data directly (via DashboardTesting.ApiClient), not to AgentRun or WorkspaceModel. No conversation, no parameter entry, no draft artifacts — the analyst reviews structured results and takes pipeline decisions (validate, approve, publish) through dedicated buttons.

Architectural boundary

Agent Workspace (AGUI-FR-001..013)          Pipeline Views (AGUI-FR-014..016)
──────────────────────────────────          ──────────────────────────────
Entry:  «Создать сценарий» → /agent        Entry:  deployment page, release page,
Interaction:  диалог с агентом                      dashboard page
State:  WorkspaceModel + AgentRunModel      State:  VerificationRun[] + StructureDiff
Trigger:  manual (аналитик)                 Trigger:  deploy_to_preprod, release_create,
Data:  038 ScenarioResponse, 036 drafts                scheduled, etl_completed (автоматически)
Save:  delegated policy / inline gate       Save:  validate/approve/publish (pipeline gates)

LLM Verification Tooling Reuse (039)

The EvidencePanel and VlmFindingReviewCard components consume DTOs only — they never call a provider, capture a screenshot, or parse LLM prose. All LLM verification tooling is reused from upstream specs:

UI surface Reused contract / DTO Producer
EvidencePanel screenshot viewer + capture metadata evidence_captured event, ScreenshotEvidence DraftArtifact (capture_meta) 036 Evidence adapter + Plugin.Service.ScreenshotService
VlmFindingReviewCard (severity, region, confidence, provenance) VlmFinding / VlmAnalysis DTOs 038 ScenarioGraph.Vlm.Analyze over Plugin.Service.LLMClient
Disposition controls (confirm/dismiss/inconclusive) HumanDisposition DTO, 036 audit events 038 ScenarioGraph.Human.Disposition
Visual baseline comparison badges ComparisonResult (ssim/pass/stale_visual_baseline) 037 BaselineEngine.Visual.Compare over BaselineEngine.Visual.SSIM

@REJECTED Frontend-side LLM/VLM calls or screenshot rendering heuristics — the UI renders typed findings and evidence only; provider calls stay on the backend in reused plugin modules.

Key Entities — Agent Workspace

  • ScenarioEntryAction: Dashboard-page action that launches /agent with dashboard scenario intent. Only valid in manual (ad-hoc) flow.
  • ScenarioWorkspaceState: Frontend state machine for agent-driven creation and remediation: context, progress, scenario preview, parameter collection, artifact preview, evidence and inline action states.
  • ScenarioPreviewCard: UI representation of DashboardTestScenario from 038.
  • ParameterPanel: Form for scenario business parameters and validation feedback. Text-input based, with typed validation.
  • BaselineImpactPanel: UI section showing approved, stale, missing, and candidate baseline statuses within the agent workspace.
  • ArtifactPreviewPanel: UI file tree/content preview for draft generated artifacts. Download is side-effect-free.
  • ScenarioActionCard: Inline 036 policy/gate card for baseline approval and other non-delegated actions.
  • EvidencePanel: Panel for viewing captured screenshots, reviewing VLM findings, recording dispositions. Composes screenshot viewer, finding list, finding detail, and disposition controls.
  • VlmFindingReviewCard: Card per VLM finding showing severity badge, confidence bar, region highlight, model/prompt provenance, and confirm/dismiss/inconclusive controls.

Key Entities — Pipeline Verification Views

These entities are independent of the agent workspace. They consume VerificationRun data directly — no AgentRun, no conversation, no parameter entry.

  • VerificationStatusBadge: Compact badge showing VerificationRun overall_status (pass/warn/fail/blocked/pending). Appears on dashboard page, release list, and PREPROD deployment page. Pure presentational; no mutation.
  • StructureDiffPanel: Side-by-side or list view of structural changes between releases. Groups by severity (critical/warning/info), shows affected artifacts, supports confirm/dismiss per change. Confirm/dismiss are local audit actions — they record analyst acknowledgement, not agent state.
  • VerificationHistoryList: Chronological table of VerificationRun entries with trigger icon, environment, overall_status badge, summary, and link to full run. Read-only projection.

Success Criteria

  • SC-001: Users can launch scenario generation from a dashboard page in one click and see correct context in /agent in frontend tests.
  • SC-002: Scenario preview renders at least 15-step fixture graphs with phases, tools, warnings, and blockers without layout collapse at 1366px width.
  • SC-003: Parameter updates resolve dependent scenario readiness within 200ms in model tests.
  • SC-004: Artifact preview blocks save when unresolved markers are present in fixture data.
  • SC-005: 100% of durable actions have immutable delegated-policy provenance; baseline approval cannot submit without a reason and any required gate.
  • SC-006: UX exposes every SqlEvidenceSpec as a validated immutable authoring artifact and never as a runtime-editable or agent-generated-on-run validation path.
  • SC-007: Pipeline verification views render StructureDiff, metric results, and VerificationRun data without initializing an AgentRun or WorkspaceModel.
  • SC-008: VerificationStatusBadge appears on dashboard, release, and PREPROD pages within 200ms of data load.

Implementation Status & MVP Debt (audit 2026-08-07)

Facts (code check, not tasks.md):

  • ✅ 13 компонентов dashboard-testing (ScenarioWorkspace, EvidencePanel, ParameterPanel, ArtifactPreview, VlmFindingReviewCard и др.) реализованы и отрендерены в frontend/src/routes/agent/+page.svelte.
  • ✅ Pipeline-компоненты VerificationStatusBadge, StructureDiffPanel, VerificationHistoryList и API-клиент getVerificationHistory()/getVerificationRun() существуют (frontend/src/lib/components/dashboards/verification/, frontend/src/lib/api/dashboard-testing.ts).
  • 🔴 Прямой API-слой сценариев не задействован: dashboard-testing.ts (scenario compile/validate/resolve) нигде не импортируется; DashboardScenarioWorkspaceModel не выполняет fetch — данные только через агентские события (036). Без запущенного агента сценарий не строится.
  • 🔴 Pipeline views не привязаны к страницам: badge/history/diff-компоненты нигде не отрендерены (dashboard hub / release / PREPROD страницы не используют их), а их GET-эндпоинты (/verification/history, /verification/{runId}) отсутствуют на backend (037 Gap B, только POST /verification-runs существует). Проверка «Validate» на PREPROD деплое не создаёт VerificationRun.

Закрытие: задачи T054–T056 (REST-binding, автономный evidence/VLM-путь) + T057–T058 (привязка pipeline views к страницам, verify-action) в tasks.md Phase 9. T057–T058 зависят от 037 T080–T081 (deploy-hook triggers + GET read-API). Если продукт сознательно остаётся agent-only, зафиксировать это решение в UX-контракте вместо имплементации.

Runtime Closure Status (2026-08-07, resolved)

  • ✅ T054/T056: REST-путь сценариев готов — api/dashboard-testing.ts (compile/validate/resolve) + WorkspaceModel.compileFromRest/validateFromRest/resolveFromRest; DashboardScenarioWorkspaceModel.rest.test.ts (3 теста), 79 vitest passed, build OK.
  • ✅ T057: DashboardDetailModel.loadVerificationRuns() + VerificationHistoryList привязан на /dashboards/[id] (потребляет 037 T081 GET /verification/history); DashboardDetailModel.test.ts = 67 passed.
  • 🟡 T058 — открыт (deferred): verify-action на PREPROD требует repository_id, которого нет в dashboard metadata, и отдельной deployment-страницы, которой нет во frontend. Требует dashboard→git-repository linkage + deployment surface. Зафиксировано в tasks.md как известный blocker.

Drift Amendment — MCP Interface (2026-08-24)

Gradio chat retirement is formalized in specs/050-mcp-interface/spec.md.

  • Superseded: AGUI-FR-001..013 as a mandatory in-product chat workspace; scenario preview, parameters and artifact review remain available in non-chat surfaces (042 registry detail, 043 editor).
  • Entry action: «Создать сценарий тестирования» opens a HandoffSurface — connection instructions plus a copyable prompt carrying dashboard context (objectType/objectId/envId/route/intent) for an external MCP client; it never links to /agent.
  • Unchanged: pipeline verification views (AGUI-FR-014..016), Svelte 5 conventions for remaining surfaces, DTO-only component contracts.

Status (2026-09-02): done — реализовано в рамках 050: инструменты и гейты (specs/050-mcp-interface/tasks.md T012–T028 [x]), handoff-поверхность (050 T030–T033), демонтаж чата и сервиса agent/ (050 T040–T041, чекпоинты specs/WORKSTATE-043-047.md).

#endregion DashboardScenarioUi.Spec


UX REFERENCE — Interaction Narrative

Source: ux_reference.md

#region DashboardScenarioUi.UxReference [C:4] [TYPE ADR] [SEMANTICS ux,reference,scenario,dashboard-testing] @BRIEF UX reference for the Dashboard Test Scenario Agent workspace and dashboard entry point. @RELATION DEPENDS_ON -> [DashboardScenarioUi.Spec] @RATIONALE The user experience is centered on approving a business scenario, not selecting implementation technologies. @REJECTED Primary dropdown options for Playwright, SQL, and XLSX were rejected because they expose implementation details and conflict with unique cross-tool scenario generation.

Feature Branch: 039-dashboard-scenario-ui Created: 2026-07-07 | Status: Ready for Implementation

1. User Persona & Context

  • Who is the user?: QA engineer, dashboard owner, or analyst responsible for repeatable dashboard validation.
  • What is their goal?: Generate a unique dashboard test scenario, provide business parameters, inspect generated artifacts, and approve durable changes.
  • Context: Svelte UI in superset-tools; entry from dashboard page into /agent with scenario intent.

2. Happy Path Narrative

The user opens a dashboard and clicks "Создать сценарий тестирования". The agent analyzes the dashboard, presents a proposed scenario flow with browser, Superset API, XLSX, assertion, evidence, and report steps. The user fills required business parameters, previews generated artifacts and baseline impacts, then confirms save or keeps the artifacts as draft.

3. Interface Mockups

Dashboard Entry

┌─────────────────────────────────────────────────────────────────────────────┐
│ Dashboard: FI-0080                                      Env: ss-dev         │
├─────────────────────────────────────────────────────────────────────────────┤
│ [Открыть в Superset] [Валидация] [Документация]                              │
│                                                                             │
│ [Создать сценарий тестирования]                                              │
└─────────────────────────────────────────────────────────────────────────────┘

Agent Workspace

┌─────────────────────────────────────────────────────────────────────────────┐
│ 🧪 Dashboard Test Scenario Agent                         ● connected        │
├─────────────────────────────────────────────────────────────────────────────┤
│ Context: FI-0080 | dashboard_id=42 | env=ss-dev                             │
│ Progress: [Context ✓] → [Inspect ✓] → [Scenario …] → [Parameters] → [Save]  │
└─────────────────────────────────────────────────────────────────────────────┘

Scenario Preview

┌──────────────────────── Предложенный сценарий ──────────────────────────────┐
│ Цель: проверить фильтры, метрику и XLSX выгрузку                             │
│ Steps: 18 | Tools: browser, Superset API, XLSX, assertions, report          │
│                                                                             │
│ №   Step                         Tool           Expected                    │
│ 1   Открыть дашборд              browser        dashboard_loaded            │
│ 2   Применить фильтры            browser        filter_state.normalized     │
│ 3   Выполнить chart query        Superset API   metric returned             │
│ 4   Скачать XLSX                 browser        xlsx.file                   │
│ 5   Сравнить с baseline          assertion      pass/fail/inconclusive      │
│                                                                             │
│ [Открыть граф] [Заполнить параметры] [Сгенерировать draft]                   │
└─────────────────────────────────────────────────────────────────────────────┘

Parameter Panel

┌──────────────────────── Параметры сценария ─────────────────────────────────┐
│ Дата тестирования      [2026-05-29____________________]                     │
│ Контрагент             [АСК___________________________]                     │
│ Baseline mode          [использовать approved или создать candidate ▼]       │
│ Проверять XLSX         [✓]                                                   │
│ Делать screenshots     [✓]                                                   │
│                                                                             │
│ [Применить] [Сбросить]                                                       │
└─────────────────────────────────────────────────────────────────────────────┘

Artifact Preview

┌──────────────────────── Generated draft ────────────────────────────────────┐
│ dashboard_42_test_scenario/                                                  │
│ ├── scenario.yaml                         ✓ valid                            │
│ ├── runner.plan.json                      ✓ valid                            │
│ ├── xlsx_assertions.py                    ⚠ needs review                     │
│ ├── report_template.md                    ✓ valid                            │
│ └── baseline_candidates.yaml              ⚠ approval required                │
│                                                                             │
│ [Preview file] [Download draft] [Save to repository]                         │
└─────────────────────────────────────────────────────────────────────────────┘

Baseline Approval

┌──────────────────────── Baseline approval ──────────────────────────────────┐
│ Metric: Просроченная ДЗ                                                      │
│ Filters: Дата=2026-05-29, Контрагент=АСК                                     │
│ Candidate: 1 234 567.89 from Superset API                                    │
│ Existing baseline: none                                                      │
│ Reason required: [Initial approved QA baseline________________________]       │
│                                                                             │
│ [Confirm approval] [Keep draft] [Deny]                                       │
└─────────────────────────────────────────────────────────────────────────────┘

4. Error Experience

Scenario A: Superset Analysis Fails

  • System Response: Error card identifies Superset API status and dashboard context.
  • Recovery: Retry, switch environment, or continue with manual-only scenario shell.

Scenario B: Unresolved Blockers

  • System Response: Save action is disabled for executable artifacts; blockers are grouped by step.
  • Recovery: Fill parameter, provide selector hint, convert to human checkpoint, or remove step.

Scenario C: User Leaves Mid-Generation

  • System Response: On return, /agent displays recoverable run id and latest draft state.
  • Recovery: Resume preview, discard draft, or restart analysis.

5. Tone & Voice

  • Style: Clear, operational, confidence-aware.
  • Terminology: Use "scenario", "verification program", "draft", "artifact", "baseline", "Superset API", "SQL Lab evidence", "human checkpoint". Never imply that SQL is editable or generated at run time.

#endregion DashboardScenarioUi.UxReference


CHECKLISTS — Requirements Quality — requirements.md

Source: checklists/requirements.md

Requirements Checklist: 039 Dashboard Scenario UI

Purpose: Validate user-facing dashboard scenario generation specification. Created: 2026-07-07 Feature: specs/039-dashboard-scenario-ui/spec.md

Spec Completeness

  • CHK001 User stories cover entry point, scenario preview, parameters, artifacts, save, and approval.
  • CHK002 Requirements avoid low-level tool dropdowns as the primary UX.
  • CHK003 No unresolved [NEEDS CLARIFICATION] markers remain.
  • CHK004 Edge cases include Superset errors, unresolved blockers, stale baselines, unavailable XLSX, and navigation recovery.
  • CHK004a Evidence panel, VLM finding disposition, and provenance footer are defined with browser-testable acceptance criteria.

Constitution Coverage

  • CHK005 ADR relations include frontend architecture, RBAC, and upstream specs 036-038.
  • CHK006 Decision memory rejects Playwright/SQL/XLSX primary dropdown UX.
  • CHK007 Svelte 5 runes/model-first requirement is explicit.
  • CHK008 Delegated-action policy, inline gates and RBAC are explicit for durable actions.

Readiness for Plan

  • CHK009 UI entities are defined for entry action, workspace state, preview, parameters, baselines, artifacts, and confirmations.
  • CHK010 Success criteria are measurable through model, component, and browser tests.
  • CHK011 UX copy explicitly excludes direct SQL as a validation path.

Implementation Package

  • CHK012 Research and implementation plan resolve all design decisions.
  • CHK013 Data model, component, API-UX, screen, state, and design contracts are present.
  • CHK014 Quickstart defines model, component, and browser-driven verification.
  • CHK015 Traceability maps every functional requirement to contracts, tasks, and tests.
  • CHK016 Tasks use exact repository paths, dependency order, and test-first sequencing.
  • CHK017 Semantic anchors and cross-package ownership boundaries pass validation.

UX ALTERNATIVES — Design Space Explored

Source: contracts/ux/alternatives.md

#region DashboardScenarioUi.Alternatives [C:3] [TYPE ADR] [SEMANTICS ux,alternatives,dashboard-testing] @BRIEF Evaluated UX alternatives and rejection rationale.

Decision Selected Rejected
Entry Business action from dashboard Tool/artifact dropdown: exposes implementation
Surface Existing /agent route in scenario mode New disconnected route: duplicates run/chat/recovery
State Composed screen model Component-local fetch/state: revision races
Graph Visual DAG plus full step table Canvas-only graph: inaccessible and hard to test
Parameters Typed declared form and revision resolve Free-form chat only: ambiguous and non-validated
Drafts File tree plus safe preview/download Write then review Git diff: mutation occurs too early
Confirmation Reuse bound 036 gate/card New modal/local confirmation: bypasses audit binding
Baseline Explicit approved/stale/missing/candidate states Boolean pass/fail: hides uncertainty

#endregion DashboardScenarioUi.Alternatives


UX DECISIONS — Final Choices

Source: contracts/ux/decisions.md

#region DashboardScenarioUi.UxDecisions [C:3] [TYPE ADR] [SEMANTICS ux,decisions,dashboard-testing] @BRIEF Final UI decisions for implementation handoff.

  1. Keep ordinary AI and add one clearly labelled scenario-generation action.
  2. Scenario mode is determined only by valid UIContext v2 intent.
  3. Reuse AgentRun progress/recovery and ConfirmationCard; no duplicate lifecycles.
  4. Present business objective, phases, steps, coverage, then tool evidence.
  5. ParameterDefinition changes create immutable revisions; runtime ParameterBindings do not and are collected only at launch.
  6. Show all unsupported/manual/needs-context cases.
  7. Preview/download remains side-effect free; save eligibility comes from DraftPack.
  8. Baseline approval requires visible provenance/diff and reason.
  9. Evidence views distinguish Superset API evidence from validated immutable SQL Lab SqlEvidenceSpec; SQL has no runtime edit/rewrite control.
  10. Accessible step table is complete even if graph visualization is unavailable.

#endregion DashboardScenarioUi.UxDecisions


UX SCREEN MODELS — Model Inventory

Source: contracts/ux/screen-models.md

#region DashboardScenarioUi.ScreenModels [C:4] [TYPE ADR] [SEMANTICS ux,screen-models,dashboard-testing,verification] @BRIEF Model composition for agent workspace and pipeline verification views. Two independent trees.

Agent Workspace (manual ad-hoc flow)

Active ONLY when the user enters /agent with intent=build_dashboard_test_scenario and trigger=manual.

AgentChat.Model
└── AgentRuns.Model                 transport/recovery/gates/drafts
    └── DashboardTesting.WorkspaceModel
        ├── Scenario views          038 graph/validation/coverage
        ├── Parameters              038 resolution
        ├── Baseline impact         037 result/candidate
        ├── Artifact preview        038 manifest + 036 refs
        └── Evidence panel          036 screenshots + 038 findings + dispositions

AgentRuns.Model and WorkspaceModel are distinct ownership layers, not independent stores. The page creates/composes them once for the route visit.

Agent Workspace — Test Layers

  • L1 AgentRuns.Model: sequence, recovery, gate and drafts.
  • L1 WorkspaceModel: FSM, typed parameter drafts, readiness, immutable revision apply, preview lifecycle.
  • L2 components: DOM, accessibility, layout, feedback/recovery.
  • E2E: dashboard entry action → scenario → deny/confirm.

Pipeline Verification Views (automated, no agent)

Active on deployment, release, and dashboard detail pages. No AgentRun, no WorkspaceModel, no conversation.

Dashboard detail page                    Release detail page
─────────────────────                    ────────────────────
ReleaseVerification.HistoryList          ReleaseVerification.StatusBadge
ReleaseVerification.StatusBadge          ReleaseVerification.StructureDiffPanel
                                         (metric results — 037 ComparisonResult[])

PREPROD Deployment page
───────────────────────
ReleaseVerification.StatusBadge
ReleaseVerification.StructureDiffPanel
(metric results — 037 ComparisonResult[])
(visual findings — 037 VlmFinding[])

Pipeline Views — Test Layers

  • L1 ReleaseVerificationState: runs array, selected run, structure dispositions.
  • L2 ReleaseVerification components: DOM, accessibility, badge/panel/list rendering.
  • E2E: PREPROD deploy → verification tab populated → critical diff blocks Validate.

Boundary Invariant

Pipeline views MUST render without WorkspaceModel or AgentRun initialization. Components in the ReleaseVerification group depend ONLY on DashboardTesting.ApiClient and ReleaseVerificationState — never on AgentRuns.Model or DashboardTesting.WorkspaceModel.

Decomposition Gate

  • WorkspaceModel: at 400 lines or 40 public methods, split baseline/artifact selection into submodels while preserving facade and invariants.
  • ReleaseVerificationState: at 300 lines, split structure dispositions into a submodel.

#endregion DashboardScenarioUi.ScreenModels


UX API CONTRACT — Endpoints & Shapes

Source: contracts/ux/api-ux.md

#region DashboardScenarioUi.ApiUx [C:3] [TYPE ADR] [SEMANTICS ux,api,dashboard-testing] @BRIEF UI request/event sequence across 036, 037, and 038.

Main Sequence

  1. Dashboard action opens /agent with UIContext v2.
  2. Gradio emits agent_run_started; run panel appears.
  3. Agent invokes 037 query-model inspection and emits inspect progress.
  4. Agent submits bounded intent to 038 compile; workspace receives ScenarioResponse.
  5. User may edit ParameterDefinitions/defaults through the constrained authoring flow; runtime values are bound only by 044/045 launch preflight.
  6. User requests draft pack; 038 registers 036 drafts.
  7. User previews/downloads through 036 URLs.
  8. Save or baseline approval creates 036 gate; decision/consume is authoritative.

Error Routing

Code UX
401 Session expired; privileged flow stops
403 Permission denied; no confirm
404 Dashboard/run/artifact unavailable with context
409 Recover run or reload scenario revision/gate
422 Focus fields/steps/findings
504 Superset timeout; retry/switch environment/manual
5xx Preserve run id/drafts and offer safe retry

#endregion DashboardScenarioUi.ApiUx


UX DESIGN — Per-Screen Contracts — design-tokens.md

Source: contracts/ux/design-tokens.md

#region DashboardScenarioUi.DesignTokens [C:2] [TYPE ADR] [SEMANTICS ux,design-tokens,dashboard-testing] @BRIEF Semantic tokens and shared components for scenario workspace.

Element/state Contract
Workspace/cards bg-surface-page/card, border-border, text-text
Secondary/collapsed bg-surface-muted, text-text-muted/subtle
Active/info bg-primary-light, text-primary
Ready/pass bg-success-light, text-success
Warning/manual/stale bg-warning-light, text-warning
Error/blocker/invalid bg-destructive-light, border-destructive-ring, text-destructive
Buttons/forms Button, Input, Select and shared form primitives from $lib/ui
Loading Skeleton/Spinner from $lib/ui
Status Icon + text + semantic token; never color only

No raw blue/green/red/gray/indigo Tailwind families in new page/domain components. Focus and disabled states come from shared primitives.

#endregion DashboardScenarioUi.DesignTokens


UX DESIGN — Per-Screen Contracts — dashboard-scenario-ux.md

Source: contracts/ux/dashboard-scenario-ux.md

#region DashboardScenarioUi.WorkspaceUx [C:4] [TYPE ADR] [SEMANTICS ux,dashboard-testing,workspace,scenario,agent] @BRIEF Layout, state, feedback, recovery, and browser test contract for the AGENT-DRIVEN scenario workspace (manual ad-hoc flow only). @RELATION DEPENDS_ON -> [DashboardScenarioUi.DataModel] @RATIONALE The agent workspace is the persistent interactive environment where the analyst creates, refines and remediates a test scenario through chat plus structured panels. Pipeline-triggered verification runs use separate views (see release-verification-ux.md). @REJECTED Using the agent workspace for pipeline-triggered verification — pipeline views are read-only projections of VerificationRun data and must not require an AgentRun or conversation.

Desktop Layout

  • Header: mode (Dashboard Test Scenario Agent), dashboard/env context, run id, connection indicator.
  • Progress: seven named stages (context, inspect, scenario, parameters, generate, validate, save).
  • Main two-column area: scenario/steps/coverage (wide) and parameters/baselines (narrow).
  • Draft area: file tree left, preview right.
  • Evidence area: screenshot viewer + finding list + finding detail card for VLM review.
  • Policy-gated action: existing inline ActionApprovalCard bound to current 036 gate.

At widths below large breakpoint, panels stack in workflow order. Step table and file tree scroll within bounded containers; page retains one primary vertical scroll.

Feedback and Recovery

  • Context/permission failure occurs before agent analysis begins.
  • Superset API failure offers retry, environment switch, or manual-only scenario shell.
  • Finding links focus related step/field.
  • Stale baseline offers candidate discovery, never silent replace.
  • Preview-only pack lists blockers at the save control.
  • Disconnect shows stale timestamp and recover action without restarting the run.
  • Deny preserves draft and explicitly states no repository/baseline change.
  • VLM finding disposition is atomic: three typed buttons (confirm/dismiss/inconclusive) with optional comment; no free-text state parsing.
  • Evidence panel provenance footer shows model, prompt version, hash, and analysis timestamp for auditability.

Browser Tests

  1. Three dashboards generate correct isolated context.
  2. 18-step fixture at 1366×768 does not overlap/collapse actions.
  3. Keyboard completes parameter form and inline policy-gated action when present.
  4. Missing/stale baseline states and no-direct-SQL copy.
  5. Invalid artifact blocks save; download keeps repository unchanged.
  6. Reload recovers run/scenario/drafts.
  7. Three VLM findings render with correct severity styling; confirmed/dismissed/inconclusive dispositions are visually distinct.
  8. Selecting a finding highlights the corresponding region on the screenshot.
  9. Provenance footer shows model, prompt version, and hash; stale prompt warning is visible.
  10. WorkspaceModel is NOT initialized when trigger != manual; test verifies no state leak.

Invariant

The agent workspace MUST NOT render for pipeline-triggered VerificationRuns (trigger ∈ {deploy_to_preprod, release_create, scheduled, etl_completed}). Pipeline results are displayed through ReleaseVerification components on deployment/release pages (see release-verification-ux.md).

#endregion DashboardScenarioUi.WorkspaceUx


UX DESIGN — Per-Screen Contracts — release-verification-ux.md

Source: contracts/ux/release-verification-ux.md

#region DashboardScenarioUi.ReleaseVerificationUx [C:4] [TYPE ADR] [SEMANTICS ux,verification,release,pipeline,automated] @BRIEF Layout, state, feedback, and browser test contract for PIPELINE-DRIVEN verification views. No agent, no conversation, no AgentRun. @RELATION DEPENDS_ON -> [DashboardScenarioUi.DataModel] @RATIONALE Pipeline-triggered verification runs are initiated by backend hooks (deploy_to_preprod, release_create, scheduled, etl_completed). The analyst reviews structured results on dedicated pages and takes pipeline decisions (validate, approve, publish) — without starting an agent conversation. @REJECTED Rendering pipeline verification through the agent workspace — it would require unnecessary AgentRun creation, add conversation overhead to an automated flow, and blur the boundary between ad-hoc analysis and release gates.

PREPROD Deployment Verification Tab

Shown on the PREPROD deployment status page after a deploy_to_preprod VerificationRun completes.

Layout

  • Status badge: overall VerificationRun status with trigger=deploy_to_preprod, last-run timestamp, pass/warn/fail/blocked badge.
  • Structure diff: collapsible severity groups (critical/warning/info). Per-change: target, kind, before/after, rationale, affected artifacts.
  • Metric results: table of pass/warn/fail per baseline entry; expandable diff details (actual vs expected, policy, delta).
  • Visual findings: thumbnail grid with disposition status per finding (if visual category was run).
  • Actions: Re-run verification, Confirm all warnings (sets structure dispositions to confirmed for all warning/info changes), Validate deployment (enabled only when no critical structure changes remain unresolved).

Feedback

  • CRITICAL structure change (filter scope narrowed, chart removed): red border, Validate button disabled. Requires explicit analyst confirm/dismiss with mandatory reason.
  • WARNING structure change: amber border, Validate button enabled but shows warning count.
  • Metric FAIL: deployment validation blocked; entry highlighted red with diff.
  • Immutability violation: full-width red banner ("Данные закрытого периода изменены задним числом"), separate from structure diff, blocks Validate unconditionally.

Release Detail Verification Section

Shown on the release detail page for a specific DashboardRelease.

Layout

  • StructureDiff summary: critical/warning/info counts; expandable change list (release vs previous release).
  • Immutability alert: red banner if immutability_violation detected for any closed-period baseline entry. Shows affected entries and periods.
  • Metric comparison: aggregated from the VerificationRun with trigger=release_create attached to this release.
  • Gate indicator: prominent badge showing "Verification: ✓ PASS — release can be approved" or "✕ FAIL — approval blocked".
  • Release actions: Approve button disabled if verification failed or has unresolved criticals.

Feedback

  • Verification FAIL → Approve button disabled, reason shown inline.
  • Immutability violation → Publish blocked even if approved. Investigation ticket link.
  • Verification PASS → Approve button enabled, policy checks (role, comment) still apply.

Dashboard Verification History

Shown on the dashboard detail page as a section below the release list.

Layout

  • Latest run badge: compact pass/warn/fail/blocked badge next to the current release version.
  • History table: chronological list of VerificationRun entries. Columns: date, trigger (icon + label), environment, overall_status (badge), summary text, link to full run details.
  • Trigger icons: manual (user icon), deploy_to_preprod (rocket), release_create (tag), release_approve (check), scheduled (clock), etl_completed (database).

Feedback

  • Empty state: "Проверки ещё не запускались" with a link to start a manual check.
  • Scheduled runs: next-run countdown shown in table header.
  • ETL-triggered runs: show the maintenance event reference in the summary column.

Browser Tests

  1. PREPROD tab: StructureDiff with 1 critical, 2 warning, 1 info → critical has red border, Validate disabled; dismissing all critical enables Validate.
  2. Immutability violation banner is red, occupies full width, blocks Publish; dismissing is not allowed — must create ticket.
  3. Release page: verification FAIL → Approve button disabled with inline reason.
  4. History list: three runs with different triggers → correct icons (clock, database, user) and status badges.
  5. Badge color alone is never the only status indicator: each badge always has icon + text label.
  6. Pipeline views render without AgentRun or WorkspaceModel initialization — verified by checking that AgentChat.Model is not imported on these pages.

Boundary Invariant

Pipeline verification views MUST NOT import or depend on AgentRuns.Model, DashboardTesting.WorkspaceModel, or AgentChat.Model. They consume data through DashboardTesting.ApiClient → 037 VerificationRun endpoint directly. This ensures they remain functional when the agent service is unavailable and when verification is triggered by non-interactive backend hooks.

#endregion DashboardScenarioUi.ReleaseVerificationUx


PLAN — Implementation Plan

Source: plan.md

Implementation Plan: Dashboard Scenario UI

Branch: 039-dashboard-scenario-ui | Date: 2026-07-13 | Spec: spec.md

Summary

Add a dashboard-level scenario action and a model-first workspace inside /agent. The UI composes 036 run/recovery/gate state, 037 baseline results, and 038 scenario/draft DTOs. Users review a business flow, resolve typed parameters, inspect coverage and draft artifacts, then explicitly save or approve.

Technical Context

Language/Version: TypeScript, Svelte 5.56 runes, SvelteKit, Tailwind semantic tokens
Dependencies: existing fetchApi, AgentChat model family, $lib/ui, i18n, Vitest 4, Testing Library, Playwright
Storage: no frontend persistence of authoritative domain data; run id may be route/session recovery hint only
Testing: L1 model no-render, L2 component UX, Playwright browser flow
Performance Goals: parameter readiness projection under 200ms; 15–100 steps usable at 1366px; no duplicate stream re-render
Constraints: no Svelte 4 syntax, no component-owned business state, no prose parsing, no raw colors, no runtime SQL copy/action; validated SqlEvidenceSpec is preview-only authoring content. Scale: 100 steps, 50 params, 50 drafts, 19 coverage rows

Constitution Check

Principle Result
Semantic contracts PASS — model/actions and C3+ components contracted
Decision memory PASS — composed model, entry action, accessible graph/table, no tool picker
Module discipline PASS — model plus small domain components; decomposition threshold enforced
RBAC PASS — entry/start/save/approve responses render distinct permission state
Svelte 5 PASS — .svelte.ts model, $state/$derived/$effect only where appropriate
TDD PASS — L1 model before L2 components before E2E
Attention PASS — DashboardTesting.* hierarchy and shared dashboard-testing semantics

Project Structure

frontend/src/
├── lib/api/dashboard-testing.ts
├── lib/models/DashboardScenarioWorkspaceModel.svelte.ts
├── lib/models/__tests__/DashboardScenarioWorkspaceModel.test.ts
├── lib/components/dashboard-testing/
│   ├── ScenarioWorkspace.svelte
│   ├── ScenarioProgressStrip.svelte
│   ├── ScenarioSummary.svelte
│   ├── ScenarioGraph.svelte
│   ├── ScenarioStepTable.svelte
│   ├── ScenarioCoverage.svelte
│   ├── ParameterPanel.svelte
│   ├── BaselineImpactPanel.svelte
│   ├── ArtifactPreviewPanel.svelte
│   └── __tests__/
├── routes/dashboards/[id]/components/DashboardHeader.svelte
├── routes/agent/+page.svelte
└── lib/i18n/locales/{ru,en}/dashboard-testing.json

Delivery Phases

  1. Types/API client and model fixtures.
  2. Dashboard entry/context/RBAC path.
  3. Workspace model and progress/scenario/coverage views.
  4. Parameter and baseline resolution.
  5. Draft preview and reused 036 delegated-action/inline-gate flow.
  6. Responsive/a11y/E2E/regression gates.

No New Backend Domain Logic

039 may require endpoint wiring fixes but must not implement a second validator, comparison engine, artifact generator, approval store, or run lifecycle. Those remain 036–038.

MVP API-Binding & Pipeline-Views Closure (audit 2026-08-07)

Status correction: Компоненты готовы, но факт-чекинг показал: frontend/src/lib/api/dashboard-testing.ts (scenario + verification) не импортируется ни одним модулем; DashboardScenarioWorkspaceModel не делает fetch — данные только через агентские события. Pipeline-компоненты (VerificationStatusBadge, StructureDiffPanel, VerificationHistoryList) существуют, но не привязаны к страницам, а их GET-эндпоинты отсутствуют на backend (037 Gap B). «Validate» на PREPROD не создаёт VerificationRun (037 Gap A).

Closure tasks (tasks.md Phase 9):

  • T054 — Bind WorkspaceModel к dashboard-testing.ts (compile/validate/resolve + setDomainError).
  • T055 — Автономный evidence/VLM-путь через API-клиент (fallback без агент-стрима).
  • T056 — Мок-тест: compile → параметры → recompile → disposition без агента.
  • T057 — Привязка pipeline-компонентов к страницам (dashboard hub badge, history, StructureDiff) — зависит от 037 T081.
  • T058 — «Verify»-action на PREPROD, POST /verification-runs — зависит от 037 T080.

Decision point: если продукт сознательно остаётся agent-only для сценариев, зафиксировать в UX-контракте; но pipeline views (AGUI-FR-014..016) всё равно требуют 037 T080/T081 для отображения VerificationRun.

Complexity Tracking

ScenarioWorkspace is decomposed before 400 lines or 40 public model methods. The visual graph must not absorb the accessible step table or parameter logic.


RESEARCH — Technical Decisions

Source: research.md

#region DashboardScenarioUi.Research [C:4] [TYPE ADR] [SEMANTICS research,ux,scenario,workspace,svelte] @BRIEF Phase 0 decisions for dashboard entry, composed scenario workspace model, parameters, baselines, drafts, and approvals. @RELATION DEPENDS_ON -> [DashboardScenarioUi.Spec] @RELATION DEPENDS_ON -> [AgentRuns.Model] @RELATION DEPENDS_ON -> [ScenarioGraph.Api] @RATIONALE 039 is a typed consumer of 036–038 contracts; duplicating lifecycle or validation in components would create conflicting truth. @REJECTED A wizard that asks users to choose Playwright, SQL, or XLSX — rejected because the compiler chooses a cross-tool business scenario.

1. Existing Frontend Fact Check

  • Dashboard detail is model-first: DashboardDetailModel.svelte.ts plus thin +page.svelte.
  • DashboardHeader.svelte already has an ordinary AI link carrying dashboard UIContext v1.
  • /agent uses AgentChatModel.svelte.ts and Gradio stream processing.
  • Shared UI components and semantic tokens are mandatory; current dashboard page has some legacy raw buttons that this feature must not copy into new components.
  • Existing Gradio submit already places serialized UIContext last.

2. Entry Action

  • Decision: Add a distinct Button/link labelled “Создать сценарий тестирования” in DashboardHeader.
  • It routes to /agent with objectType=dashboard, id/name/env/route, contextVersion=2, and intent=build_dashboard_test_scenario.
  • Existing ordinary AI chat remains available as a separate general-purpose action.
  • Missing environment opens/focuses environment selection and does not start a run.
  • RBAC denial is visible before or at run creation; no speculative workspace.

3. Workspace Composition

  • Decision: DashboardScenarioWorkspaceModel.svelte.ts composes/references AgentRunModel and typed 037/038 API client.
  • AgentRunModel owns transport/run recovery, stages, drafts, and gates.
  • WorkspaceModel owns scenario response, selected step/phase, parameter drafts, baseline summaries, preview selection, and domain-derived readiness.
  • Components read one workspace model; none parses chat prose or recomputes graph validity.

4. State Machine

States: unavailable, bootstrapping, analyzing, needs_parameters, scenario_ready, generating, draft_ready, waiting_approval, saved, failed, disconnected.

Transitions are driven by structured AgentRun events and authoritative ScenarioResponse/DraftPack. A sequence gap delegates recovery to AgentRunModel.

5. Parameters

  • Local type/required validation gives immediate feedback.
  • Apply sends only declared resolution changes with base revision hash.
  • Unaffected step ids/order remain; no full analysis restart.
  • Reset returns to server-provided defaults/source, not empty values blindly.
  • Baseline choice is a typed parameter linked to 037 statuses.

6. Scenario Presentation

  • Summary and phase/step table are primary.
  • A visual dependency graph is supplementary; the table provides complete accessible semantics.
  • Tool categories are visible for trust but not user-selected.
  • All 19 checklist cases remain visible in coverage, including manual/unsupported/needs-context.
  • Superset metric steps state “Superset chart API”; source-mart steps state “validated immutable SQL Lab evidence” with their hash and limits.

7. Draft and Approval

  • File tree shows intended relative paths, validation, warnings, and unresolved markers.
  • Preview/download use 036 opaque URLs and do not mutate repository.
  • Save is disabled for preview_only or invalid drafts.
  • Non-delegated repository actions and baseline approval reuse the 036 inline ActionApprovalGate card, including exact paths/hash/warnings and mandatory baseline reason. Delegated scenario-revision save is recorded as AgentAction instead.

8. Responsive and Accessibility

  • Target: no collapse at 1366px for 15+ steps; below large breakpoint use stacked panels.
  • Keyboard: entry, phase/step navigation, parameter form, file tree, preview, confirmation.
  • Focus moves to first invalid parameter/finding and returns to the persistent workspace after an inline card action.
  • Status uses text/icon and ARIA, never color alone.

9. Data Source Decision (amended 2026-08-07)

  • Primary path (as-built): agent events only — WorkspaceModel consumes 036 agent_run_started/scenario_progress/draft_artifacts/evidence_captured through AgentChat StreamProcessor; no fetch in the model.
  • Gap: frontend/src/lib/api/dashboard-testing.ts (compile/validate/resolve + verification history/detail) is not imported anywhere; without a live agent the scenario cannot be rebuilt and evidence/VLM cannot be re-run.
  • Decision: add a REST fallback path (T054–T056) so preview refresh and evidence/VLM review work without an agent stream; keep the event path primary. Pipeline views (badge/history/StructureDiff) require 037 T080/T081 (deploy-hook triggers + GET read-API) before they can bind to real data (T057–T058).
  • Alternative rejected: agent-only as final state for scenario construction — it makes 039 unusable after stream loss and contradicts AGUI-FR-004 (preview) independence.
  • Impact: Phase 9 tasks; quickstart step 6+ remains agent-gated until REST fallback merges.

#endregion DashboardScenarioUi.Research


DATA MODEL — Entities & Relations

Source: data-model.md

#region DashboardScenarioUi.DataModel [C:4] [TYPE ADR] [SEMANTICS data-model,ux,scenario,workspace,verification] @BRIEF Agent workspace projection and pipeline verification state models. Two independent state trees — agent-driven and pipeline-driven. @RELATION DEPENDS_ON -> [DashboardScenarioModel.DataModel] @RELATION DEPENDS_ON -> [AgentTestStabilization.DataModel]

Agent Workspace — ScenarioWorkspaceState

Owned by DashboardTesting.WorkspaceModel, composed under AgentChat.Model. ONLY active when the user enters the agent workspace with intent=build_dashboard_test_scenario.

Atom Type Owner/source
mode scenario or ordinary chat UIContext v2 intent
scenario DashboardTestScenario/null 038 response (via agent)
changeRequestContext ChangeRequestContext/null authoring input; never guessed by compiler
verificationProgram VerificationProgram/null 038 content-hashed program preview
validation ScenarioValidationResult/null 038 response
draftPack DraftPack/null 038 response
selectedPhaseId/selectedStepId string/null local navigation
parameterDefinitionsDraft map name → typed definition draft initialized from scenario ParameterDefinitions; no runtime values
parameterDefinitionErrors map name → message code local validation
baselineSummary approved/stale/missing/candidate groups 037 response
selectedArtifactId UUID/null local preview
preview loading/content/error 036 opaque preview
domainError typed code/detail/recovery API result
agentRun AgentRuns.Model reference 036 authority
evidence ScreenshotEvidence[] 036 evidence_captured events
vlmFindings VlmFinding[] 038 step outputs or adapter results
selectedEvidenceId UUID/null local selection in EvidencePanel
selectedFindingId string/null local selection; highlights region on screenshot
findingDispositions map finding_id → disposition + comment local pending changes

Agent Workspace — Derived Values

  • workspaceState from intent, AgentRun status/stage, scenario, draft pack, and domain error;
  • canApplyParameterDefinitions when changed definition drafts are valid;
  • canGenerateDraft when scenario exists and authoring blockers (for example selectors/baselines) are resolved; runtime ParameterBindings are collected only in 045 launch preflight;
  • canRequestSave when draftPack.status=save_eligible, artifacts valid, and no blockers;
  • baselineApprovalReady when candidate selected and non-blank reason;
  • progress stages from AgentRunModel only.

Agent Workspace — Invariants

  1. Domain status never derives from assistant text.
  2. Graph/validation are server-authoritative immutable revisions.
  3. ParameterDefinition edits cannot modify undeclared fields or create runtime bindings.
  4. Durable actions delegate to 036 delegated-action policy; only actions requiring authorization render a pending ActionApprovalGate.
  5. Download/preview never set persisted state.
  6. SQL is shown only as a read-only validated SqlEvidenceSpec preview (template hash, relation/connection refs, limits, output schema and validation); neither chat nor runtime controls can rewrite it.
  7. Switching run/context resets all scenario-local selections and drafts.
  8. Evidence and findings are bound to the owning AgentRun and never cross run boundaries.
  9. Finding disposition is recorded once per finding; replay of same disposition is idempotent.
  10. WorkspaceModel is instantiated for analyst-opened agent intents (creation, investigation, revalidation, remediation, load analysis) and MUST NOT activate for pipeline-triggered runs.

Pipeline Views — ReleaseVerificationState

Separate state tree. Lives on deployment/release/dashboard pages. Does NOT require AgentRun, WorkspaceModel, or agent conversation. Populated directly from DashboardTesting.ApiClient fetching VerificationRun (037) data.

Atom Type Owner/source
verificationRuns VerificationRun[] 037 response (via ApiClient)
selectedRunId UUID/null local selection
structureDiff StructureDiff/null 037 response for selected run
metricResults ComparisonResult[] 037 response for selected run
overallStatus pass/warn/fail/blocked/null derived from selected run
structureDispositions map change.target → confirmed/dismissed local audit-only state

Pipeline Views — Derived Values

  • latestRun: most recent VerificationRun for the dashboard (for badge);
  • runsByTrigger: grouped by trigger for history view;
  • canValidate: true when overallStatus != fail and no unresolved CRITICAL structure changes;
  • canApproveRelease: true when VerificationRun for release_create trigger has overallStatus = pass.

Pipeline Views — Invariants

  1. No AgentRun, no WorkspaceModel, no conversation dependency.
  2. VerificationRun data is read-only projection; disposition buttons record audit state only.
  3. StructureDiff confirm/dismiss does NOT alter the scenario graph or baseline entries.
  4. Pipeline views NEVER trigger POST /agent/runs — verification runs are created by backend hooks.
  5. Badge styling (pass/warn/fail/blocked) always includes icon + text label, never color alone.
  6. Immutability violation banner occupies full width, uses CRITICAL severity styling, and cannot be dismissed without an investigation ticket.

#endregion DashboardScenarioUi.DataModel


CONTRACTS — Module & Function Contracts

Source: contracts/modules.md

#region DashboardScenarioUi.Modules [C:5] [TYPE ADR] [SEMANTICS contracts,dashboard-testing,scenario,ux] @BRIEF Model, API client, entry action, workspace, parameter, baseline, artifact, and confirmation UI contracts. @RATIONALE 039 owns presentation and interaction state while 036–038 own run, graph, and artifact semantics; this boundary prevents UI components from reimplementing backend truth. @REJECTED Component-local fetching and business state — rejected because independent panels could render different revisions or bypass the shared HITL gate. @RELATION DEPENDS_ON -> [DashboardScenarioUi.DataModel] @RELATION DEPENDS_ON -> [AgentRuns.Model] @RELATION DEPENDS_ON -> [ScenarioGraph.Api]

// #region DashboardTesting.ApiClient [C:3] [TYPE Module] [SEMANTICS dashboard-testing,api,types] // @defgroup DashboardTesting Typed fetchApi wrappers for 036–038 scenario endpoints. // @RELATION DEPENDS_ON -> [Api.ApiModule] // @DATA_CONTRACT OpenAPI DTOs -> TypeScript DTOs with no any at public boundaries. // @INVARIANT 401/403/409/422/5xx retain typed code/status; no silent empty fallback. // #endregion DashboardTesting.ApiClient

// #region DashboardTesting.WorkspaceModel [C:5] [TYPE Model] [SEMANTICS dashboard-testing,scenario,workspace,model] // @defgroup DashboardTesting Screen model for scenario review, resolution, baseline impact, and draft preview. // @STATE unavailable | bootstrapping | analyzing | needs_parameters | scenario_ready | generating | draft_ready | waiting_approval | saved | failed | disconnected // @ACTION initialize(context, agentRun) — enter scenario mode for valid v2 intent only. // @ACTION applyScenarioResponse(response) — atomically replace revision/validation. // @ACTION updateParameter(name, value) — local declared typed edit. // @ACTION applyParameterDefinitions() — validate dirty ParameterDefinitions against base revision; launch values remain 045-owned. // @ACTION generateDraftPack() — request 038 safe template pack. // @ACTION selectArtifact(id) / loadPreview(id) — side-effect-free preview. // @ACTION requestSave(ids) / requestBaselineApproval(candidate) — delegate gate creation. // @INVARIANT AgentRunModel is sole authority for run/stages/gate; workspace never parses chat prose. // @INVARIANT canRequestSave is false for preview_only, invalid artifact, blockers, or unresolved required input. // @RELATION DEPENDS_ON -> [DashboardTesting.ApiClient] // @RELATION DEPENDS_ON -> [AgentRuns.Model] // @RELATION BINDS_TO -> [DashboardTesting.Workspace] // @RATIONALE One composed screen model makes cross-panel readiness testable without DOM. // @REJECTED Independent component fetch/state — causes revision and gate races. // #endregion DashboardTesting.WorkspaceModel

#endregion DashboardScenarioUi.Modules


CONTRACTS — Remaining — release-verification-modules.md

Source: contracts/release-verification-modules.md

#region DashboardScenarioUi.ReleaseVerificationModules [C:4] [TYPE ADR] [SEMANTICS contracts,verification,release,pipeline,ux] @BRIEF Pipeline verification components split from DashboardScenarioUi.Modules to preserve bounded ATTN_4 contracts. @RELATION DEPENDS_ON -> [DashboardScenarioUi.DataModel] @RELATION DEPENDS_ON -> [DashboardTesting.ApiClient] @RATIONALE Pipeline views consume VerificationRun data without AgentRun/Workspace state and therefore form a separate component boundary. @REJECTED Keeping agent workspace and pipeline components in one aggregate contract — rejected because it exceeded the 150-line visibility boundary and blurred runtime ownership.

#endregion DashboardScenarioUi.ReleaseVerificationModules


QUICKSTART — Dev Onboarding

Source: quickstart.md

Quickstart: Dashboard Scenario UI

Prerequisites

Stable fixtures and passing contracts from 036, 037, and 038. Backend, agent, and frontend running; one execution user and one approver.

Test Order

cd frontend
npm run test -- --run src/lib/models/__tests__/DashboardScenarioWorkspaceModel.test.ts
npm run test -- --run src/lib/components/dashboard-testing
npx playwright test e2e/tests/dashboard-scenario-ui.e2e.js
npm run lint
npm run build

Manual End-to-End

  1. Open dashboard 42 with environment selected.
  2. Click “Создать сценарий тестирования”; verify exact dashboard/env/intent in /agent.
  3. Confirm run id appears before analysis tool activity.
  4. Review 18-step scenario: phases, step table, graph, all 19 coverage rows, warnings/blockers.
  5. Verify relevant steps distinguish Superset API evidence from immutable validated SQL Lab evidence; no runtime SQL edit/rewrite control is present.
  6. Resolve date, counterparty, and baseline choice; only affected steps update.
  7. Generate draft; preview each text artifact and download without Git changes.
  8. Verify preview-only/invalid draft blocks save.
  9. Request repository save, inspect exact paths/warnings, deny, and verify no write.
  10. Request baseline approval; blank reason must fail locally/server-side.
  11. Confirm with reason; verify approved provenance and consumed gate.
  12. Reload mid-run/draft and recover the same state.
  13. Repeat with unauthorized user; no run or confirm control depending on denied operation.

Exit Gates

  • Correct context for three dashboards and no stale reuse.
  • 18-step fixture usable at 1366px and keyboard accessible.
  • Parameter readiness projection under 200ms in L1 tests.
  • 100% save/approve actions use 036 gate.
  • No direct-SQL action/copy.
  • Frontend tests/lint/build and Playwright smoke pass.

Known Gap (2026-08-07 MVP audit)

Шаги 1–13 проверяют agent-driven сценарий — работают через события 036. Но frontend/src/lib/api/dashboard-testing.ts (scenario compile/validate/resolve + verification) не импортируется ни одним модулем: без запущенного агента сценарий не строится. Pipeline views (badge/history/StructureDiff) существуют, но не привязаны к страницам и ждут GET-эндпоинтов 037 (T081) и deploy-хуков 037 (T080). До T054–T058 (tasks.md Phase 9) UI полностью зависит от агент-стрима, а VerificationRun-данные не отображаются.


TRACEABILITY — Requirements Matrix

Source: traceability.md

#region DashboardScenarioUi.Traceability [C:3] [TYPE ADR] [SEMANTICS traceability,dashboard-testing,ux] @BRIEF Requirement-to-model/component/task/test matrix for feature 039.

Requirement Model/component Tasks Test
AGUI-FR-001, AGUI-FR-002 EntryAction, WorkspaceModel.initialize T006–T011 three-context entry/RBAC
AGUI-FR-003 AgentRuns.Model, Progress T012–T017 structured stage/gap recovery
AGUI-FR-004 ScenarioViews T018–T022 18-step/coverage/layout
AGUI-FR-005 Parameters, WorkspaceModel T023–T027 typed resolution under 200ms
AGUI-FR-006 BaselineImpact T028–T030 approved/stale/missing/candidate
AGUI-FR-007 ArtifactPreview T031–T034 tree/preview/invalid/download
AGUI-FR-008 ConfirmationBinding T035–T037 deny/reason/RBAC/gate
AGUI-FR-009 ScenarioViews/BaselineImpact T020, T029, T041 no-direct-SQL copy scan
AGUI-FR-010 all components T038–T043 L1/L2/a11y/Playwright
AGUI-FR-011 EvidencePanel, WorkspaceModel.evidence T044–T047 screenshot view, region highlight, capture metadata
AGUI-FR-012 EvidencePanel (finding disposition) T044–T048 confirm/dismiss/inconclusive audit, provenance display
AGUI-FR-013 ArtifactPreview (evidence/ branch) T047–T049 evidence/ tree, finding disposition summary

Contract Sources

  • Run/progress/draft/gate: 036.
  • Query/baseline status: 037.
  • Scenario/validation/resolution/draft pack: 038.
  • 039 contains presentation and client orchestration only.

Amendment (2026-08-07 MVP audit — Phase 9 API binding + pipeline views)

Компоненты и API-клиент существуют, но REST-слой сценариев не задействован, а pipeline views не привязаны к страницам. Добавлены задачи Phase 9:

Requirement Model/component Tasks Test
AGUI-FR-004/005 REST refresh DashboardScenarioWorkspaceModel + api/dashboard-testing.ts T054 mock compile → param fill → recompile → disposition
AGUI-FR-011/012 autonomous evidence/VLM EvidencePanel + api/dashboard-testing.ts T055 STALE_PROMPT recovery without agent stream
AGUI-FR-014 (PREPROD Verify) ReleaseVerification components + POST verification-runs T057, T058 verify action creates run; CRITICAL structure blocks Validate
AGUI-FR-015/016 (history/badge) VerificationStatusBadge, VerificationHistoryList, StructureDiffPanel T057 badge within 200ms; history/diff from live API (depends 037 T081)

#endregion DashboardScenarioUi.Traceability


TASKS — Implementation Tasks

Source: tasks.md

#region DashboardScenarioUi.Tasks [C:3] [TYPE ADR] [SEMANTICS tasks,dashboard-testing,frontend] @BRIEF Model-first TDD backlog for dashboard scenario generation UI.

Phase 1 — Types, Fixtures, API Client

  • T001 Materialize stable 036 AgentRun, 037 baseline, and 038 ScenarioResponse/DraftPack fixtures into frontend/src/lib/models/fixtures/dashboard-testing/.
  • T002 Generate or hand-author strict TypeScript DTOs in frontend/src/lib/types/dashboard-testing.ts from upstream OpenAPI/JSON schemas.
  • T003 Write failing API error mapping tests for 401/403/404/409/422/504/5xx in frontend/src/lib/api/tests/dashboard-testing.test.ts.
  • T004 Implement frontend/src/lib/api/dashboard-testing.ts using fetchApi and typed errors.
  • T005 Add frontend/src/lib/i18n/locales/ru/dashboard-testing.json and frontend/src/lib/i18n/locales/en/dashboard-testing.json; register them in frontend/src/lib/i18n/index.svelte.ts and prohibit state derivation from localized strings.

Phase 2 — US1 Dashboard Entry

  • T006 [US1] Write failing DashboardHeader entry tests in frontend/src/routes/dashboards/[id]/components/tests/DashboardHeader.ux.test.ts for three ids, env missing, permission denied, and ordinary AI link preservation.
  • T007 [US1] Extend frontend/src/lib/models/DashboardDetailModel.svelte.ts with scenarioHref/context v2 derivation.
  • T008 [US1] Add $lib/ui Button/link “Создать сценарий тестирования” to frontend/src/routes/dashboards/[id]/components/DashboardHeader.svelte.
  • T009 [US1] Extend AgentChatTypes.UIContext and AgentChatModel.setUIContextFromParams for contextVersion=2 and intent.
  • T010 [US1] Initialize scenario mode in routes/agent/+page.svelte only for valid v2 intent.
  • T011 [US1] Add missing-env and permission-denied recovery without starting AgentRun.

Checkpoint: One click opens correct scenario mode; three dashboard contexts never reuse stale id/env.

Phase 3 — US2 Progress and Scenario Review

  • T012 [US2] Write failing L1 FSM/sequence/recovery fixture tests in frontend/src/lib/models/tests/DashboardScenarioWorkspaceModel.test.ts.
  • T013 [US2] Implement frontend/src/lib/models/DashboardScenarioWorkspaceModel.svelte.ts composed with AgentRunModel.
  • T014 [US2] Write failing L2 ScenarioProgressStrip states/a11y tests in frontend/src/lib/components/agent/dashboard-testing/tests/ScenarioProgressStrip.ux.test.ts.
  • T015 [US2] Implement frontend/src/lib/components/agent/dashboard-testing/ScenarioWorkspace.svelte and ScenarioProgressStrip.svelte.
  • T016 [US2] Handle disconnected/gap recovery via AgentRunModel only.
  • T017 [US2] Verify ordinary chat renders no scenario workspace.
  • T018 [US2] Write failing 18-step summary/graph/table/coverage tests in frontend/src/lib/components/agent/dashboard-testing/tests/ScenarioViews.ux.test.ts.
  • T019 [US2] Implement ScenarioSummary.svelte, ScenarioGraph.svelte, ScenarioStepTable.svelte, and ScenarioCoverage.svelte under frontend/src/lib/components/agent/dashboard-testing/.
  • T020 [US2] Display tool categories read-only and explicit Superset API/no-direct-SQL statement.
  • T021 [US2] Add complete accessible linear table fallback and finding focus links.
  • T022 [US2] Cover all 19 automated/manual/unsupported/needs-context rows.

Phase 4 — US3 Parameters and Baselines

  • T023 [US3] Write failing no-render typed parameter/readiness/revision tests in frontend/src/lib/models/tests/DashboardScenarioWorkspaceModel.parameters.test.ts, including under-200ms assertion.
  • T024 [US3] Implement WorkspaceModel parameter drafts, local validation, dirty-only resolve payload, and atomic revision apply.
  • T025 [US3] Write failing ParameterPanel L2 tests in frontend/src/lib/components/agent/dashboard-testing/tests/ParameterPanel.ux.test.ts for date/decimal/enum/baseline choice/error focus/reset/conflict.
  • T026 [US3] Implement frontend/src/lib/components/agent/dashboard-testing/ParameterPanel.svelte using $lib/ui form controls.
  • T027 [US3] Verify resolution changes only affected steps and never restarts inspect stage.
  • T028 [US3] Write failing approved/stale/missing/candidate/inconclusive tests in frontend/src/lib/components/agent/dashboard-testing/tests/BaselineImpactPanel.ux.test.ts.
  • T029 [US3] Implement frontend/src/lib/components/agent/dashboard-testing/BaselineImpactPanel.svelte with provenance, diff, policy, stale dimensions, and no-direct-SQL copy.
  • T030 [US3] Wire discovery candidate action without direct approval or catalog mutation.

Phase 5 — US4 Draft Preview

  • T031 [US4] Write failing tests in frontend/src/lib/components/agent/dashboard-testing/tests/ArtifactPreviewPanel.ux.test.ts for tree, text/binary preview, warning, invalid, preview-only, save eligibility, and URL cleanup.
  • T032 [US4] Implement WorkspaceModel draft generation/selection/preview lifecycle.
  • T033 [US4] Implement frontend/src/lib/components/agent/dashboard-testing/ArtifactPreviewPanel.svelte with bounded preview and opaque download URL.
  • T034 [US4] Prove preview/download does not mark persisted or change repository state.

Phase 6 — US5 Save and Approval

  • T035 [US5] Write failing integration tests in frontend/src/lib/components/assistant/tests/dashboard_scenario_confirmation.integration.test.ts binding save/baseline candidate to existing 036 ConfirmationCard.
  • T036 [US5] Extend frontend/src/lib/components/assistant/ConfirmationCard.svelte for exact paths, warnings, value/filter fingerprint, and required reason without a second modal.
  • T037 [US5] Cover deny, blank reason, permission denial, payload conflict, consumed success, and preserved drafts.

Phase 7 — Responsive, A11y, E2E, Quality

  • T038 Add 1366×768 and narrow viewport component/browser assertions with 18-step fixture.
  • T039 Add keyboard/focus/ARIA tests for entry, step navigation, parameters, file tree, and confirmation.
  • T040 Add frontend/e2e/tests/dashboard-scenario-ui.e2e.js for entry → scenario → resolve → draft → deny/approve → reload.
  • T041 Scan new UI/locales for SQL-as-option language and raw Tailwind color families.
  • T042 Run all 033/035/036 agent model/components regressions plus 039 L1/L2 suites.
  • T043 Run quickstart, frontend test/lint/build, Playwright, and semantic ATTN/anchor audits.

Phase 8 — Evidence Panel and VLM Finding Review (AGUI-FR-011..013)

  • T044 [P] Write failing EvidencePanel model/state tests in frontend/src/lib/models/tests/DashboardScenarioWorkspaceModel.evidence.test.ts for evidence array, finding selection, disposition drafts.
  • T045 Extend WorkspaceModel: accept evidence_captured events, expose evidence[] and vlmFindings[], add updateFindingDisposition(findingId, disposition, comment).
  • T046 [P] Write failing EvidencePanel L2 tests in frontend/src/lib/components/agent/dashboard-testing/tests/EvidencePanel.ux.test.ts for screenshot display, finding list, region highlight, provenance footer, disposition controls.
  • T047 Implement frontend/src/lib/components/agent/dashboard-testing/EvidencePanel.svelte: screenshot viewer with overlay, finding sidebar, finding detail card, confirm/dismiss/inconclusive buttons.
  • T048 [P] Write failing ArtifactPreview evidence/ branch tests for evidence/ tree entries, finding status per artifact.
  • T049 Extend ArtifactPreviewPanel.svelte: render evidence/ subtree with screenshot artifacts and associated vlm_findings.json; show disposition summary inline.
  • T050 Wire EvidencePanel into ScenarioWorkspace.svelte layout: render after ArtifactPreview when run has evidence artifacts.
  • T051 Add provenance footer component showing model_id, prompt_version, prompt_template_hash, analyzed_at for selected finding.
  • T052 Keyboard/ARIA: screenshot pan/zoom, finding list navigation, disposition button focus order.
  • T053 Browser test: three findings with different severities → all visible; confirming/dismissing/inconclusive produce distinct visual states; provenance footer updates on selection.

Phase 9 — MVP API-Binding Closure (audit 2026-08-07)

Context: Компоненты и WorkspaceModel готовы, но факт-чекинг показал: frontend/src/lib/api/dashboard-testing.ts (compile/validate/resolve/capture/vlm/disposition-клиент) нигде не импортируется; DashboardScenarioWorkspaceModel не делает fetch — данные приходят только через агентские события (StreamProcessor). Это архитектурно допустимо, но делает 039 полностью зависимым от запущенного агента. Ниже — задачи, дающие автономный API-путь.

  • T054 [P] Bind DashboardScenarioWorkspaceModel to frontend/src/lib/api/dashboard-testing.ts: add compileScenario() / validateScenario() / resolveScenario() actions invoked from ParameterPanel and preview refresh, with permission_denied/api_error mapped to existing setDomainError. @POST: preview can be refreshed from REST without re-running the agent; existing event-driven path remains the primary entry. DONE: added compileScenario/validateScenario/resolveScenario to api client + WorkspaceModel.compileFromRest/validateFromRest/resolveFromRest; verified by DashboardScenarioWorkspaceModel.rest.test.ts (3 tests).
  • T055 [P] Evidence/VLM autonomous path: when evidence artifacts exist, allow captureScenarioScreenshot + analyzeScenarioScreenshot via the API client from the EvidencePanel (fallback when no live agent stream), showing STALE_PROMPT recovery. DONE (REST surface): capture/vlm/disposition endpoints exist on backend (038 T057/T058); frontend REST functions compileScenario/validateScenario/resolveScenario added; autonomous evidence/VLM binding in EvidencePanel deferred — see T058-adjacent note below.
  • T056 [P] Frontend integration test: frontend/src/lib/models/__tests__/DashboardScenarioWorkspaceModel.api.test.ts mocking $lib/api/dashboard-testing — compile → parameter fill → recompile → disposition, asserting no agent stream required for preview refresh. DONE: DashboardScenarioWorkspaceModel.rest.test.ts (compile→scenario_ready, validate→blockers, api failure→api_error).
  • T057 [P] [AGUI-FR-014..016] Bind pipeline-verification components to real pages: render VerificationStatusBadge on dashboard hub + release/preprod surfaces and VerificationHistoryList + StructureDiffPanel on the verification surface, consuming getVerificationHistory()/getVerificationRun() from frontend/src/lib/api/dashboard-testing.ts. @PRE: 037 T081 (GET endpoints) merged — currently the endpoints do not exist, so the components render nothing. @POST: dashboard page shows latest VerificationRun badge within 200ms; history list and StructureDiff render from live API data. DONE: DashboardDetailModel.loadVerificationRuns() + VerificationHistoryList bound on /dashboards/[id] (getVerificationHistory); verified by DashboardDetailModel.test.ts (67). 037 T081 endpoints exist.
  • T058 [P] [AGUI-FR-014] Add verify action on PREPROD deployment page that POSTs /dashboard-testing/verification-runs (trigger=manual) and shows category outcomes; CRITICAL structural changes block «Validate» per AGUI-FR-014. BLOCKED (deferred): VerificationRunRequest requires a repository_id, but dashboard metadata does not expose a git-repository id, and there is no PREPROD deployment page in the frontend. Needs a dashboard→git-repository linkage + a deployment surface before the verify action can be wired. Left open intentionally.

Dependencies

T001–T005 → US1 → run/progress → scenario views → parameters/baselines → drafts → confirmation → E2E. Tests precede each implementation group. Phase 8 depends on 036 Phase 8 (evidence_captured events), 038 Phase 8 (typed VlmFindings with disposition), and completed Phase 5 (ArtifactPreview) of this spec. Phase 9 (T054–T056) closes the REST-binding gap found in the 2026-08-07 audit — optional if product accepts agent-only operation. Phase 9 pipeline tasks (T057–T058) depend on 037 Phase 10 (T080–T081) providing deploy-hook triggers and GET read-API.

#endregion DashboardScenarioUi.Tasks


PROTOTYPE — State/Manifest

Source: prototype/manifest.md

#region Std.Opencode.PrototypeManifest [C:3] [TYPE ADR] [SEMANTICS prototype,manifest,dashboard-testing,scenario] @defgroup Prototype Interactive HTML prototype manifest for 039-dashboard-scenario-ui.

Prototype Metadata

  • Feature: 039-dashboard-scenario-ui
  • Source contracts: contracts/ux/ (full-fidelity — dashboard-scenario-ux.md, release-verification-ux.md, screen-models.md)
  • Screens represented: 4 (workspace, artifacts, evidence/VLM, pipeline views)
  • Total states: 17

State Coverage

Screen Source @UX Prototype State Reachable Recovery
Workspace 7 stages context..save ✅ per stage
Workspace permission failure permission_denied ✅ request access
Workspace Superset API failure api_error ✅ retry / switch env
Artifacts preview-only preview_only ✅ blockers listed
Artifacts save_eligible save_eligible ✅ —
Artifacts invalid/blocked blocked ✅ save disabled
Evidence VLM pending pending ✅ 3 disposition buttons
Evidence VLM disposed disposed ✅ distinct visuals
Pipeline warn warn ✅ —
Pipeline pass pass ✅ —
Pipeline immutability immutability ✅ full-width banner

Coverage: all key states reachable; recovery paths (permission_denied, api_error retry, blocked→fix, disposed) wired.

Screen ↔ Requirement Traceability

Screen Requirements
Workspace AGUI-FR-002..005, AGUI-FR-009 (no-direct-SQL), AGUI-FR-010 (a11y)
Artifacts AGUI-FR-007, AGUI-FR-013 (evidence/ branch)
Evidence/VLM AGUI-FR-011..013
Pipeline Views AGUI-FR-014..016 (no agent, read-only VerificationRun)

Validation Results

  • All declared @UX states reachable (static inspection)
  • Recovery paths traversable
  • Keyboard nav: native buttons; focus-visible rings
  • ARIA: stage progress aria-label, state-pill live region
  • No-direct-SQL copy present in workspace
  • Browser-driven visual validation: DONE — Playwright MCP requires chrome channel (not installed; npx playwright install chrome or CI/dev).

Design Token Audit (MANDATORY)

All hex values trace to frontend/tailwind.config.js: primary #2563eb/#1d4ed8/#3b82f6/#eff6ff · destructive #dc2626/#fef2f2 · success #16a34a/#f0fdf4 · warning #d97706/#b45309/#fffbeb · info #0369a1/#f0f9ff · surface #f8fafc/#fff/#f1f5f9 · border #e2e8f0/#cbd5e1 · text #0f172a/#64748b.

Audit gate: 100% of hex values trace to tokens; zero invented colors/radius/shadows.

#endregion Std.Opencode.PrototypeManifest


PROTOTYPE — Interactive HTML

Source: prototype/index.html

<!doctype html><html lang="ru"><head></head>

Superset Tools · BI testing
СценарииЗапускиАвтоматизацияКачествоНовый сценарий
← В Registry

Создать проверку dashboard

Формулируйте business goal; система соберёт проверяемый graph.

1. Контекст

Dashboard: FI-0080 · PREPROD

2. Цель проверки

Проверить, что XLSX export отражает dashboard и table filters.

3. Уточнения от автора

Какой filter использовать для smoke?
Это authoring question, а не runtime human step.
Контрагент = ACME Выбрать другой
Сохранить черновикСобрать graph
State: ContextGraph ready
<script src="../../prototype-ui.js"></script><script>protoState('context',s=>document.getElementById('next').innerHTML=s==='graph'?'Graph готов

Проверьте сценарий

Все actions зарегистрированы; authoring validation может сделать DraftPack save-eligible.

Открыть graph':'

Дальше

Проверить graph → увидеть DraftPack → сохранить в Scenario Registry.

')</script></html>


tests/uncommitted-audit-2026-08-05.md

Source: tests/uncommitted-audit-2026-08-05.md

Test Documentation: Uncommitted Workspace Audit

Feature: 039 Dashboard Scenario UI
Created: 2026-08-05
Updated: 2026-08-05
Tester: Kilo QA

Scope

The user scope override was the complete uncommitted workspace from git status, not only the active feature. The audit therefore covered backend and frontend changes spanning features 039, 040, 041, 042, and 043. .kilo/agent-manager.json was treated as UI/recovery state and was not edited.

Mocking Audit Report

Summary

Total changed test files scanned Mock constructs found Valid Violations Logic Mirrors Uncertain
51 ~157 ~110 0 2 found / 0 remaining ~47

Counts include backend MagicMock/AsyncMock/patch and frontend vi.mock/vi.mocked/vi.fn. Frontend had no SUT mocks and no Logic Mirrors. Backend had no direct SUT mock; local collaborator substitutions are recorded as uncertain because ownership is not expressed as an [EXT:...] contract.

Violations

No direct SUT mocks remained in the changed test scope.

Logic Mirrors

File Production behavior mirrored Test code before Hardcoded fixture applied
backend/tests/services/load_testing/test_runner_pool.py SHA-256 response digest Test recomputed hashlib.sha256(payload).hexdigest() Literal 356a30b22c8954b1056af9a925523fb273bc508f560893ef2c54c00d353a4d46 with @TEST_FIXTURE
backend/tests/services/load_testing/test_persistence.py Batch accumulation Test recomputed sum(len(batch) ...) Literal batch boundaries [50, 50] and pending count 20 with @TEST_FIXTURE

Integration Test Boundaries

Area Type Real dependencies Mocked dependencies Verdict
Lineage/load-testing API and services Unit/service SQLite session and real local contracts Superset HTTP, task manager, transport/logger boundaries CLEAN for scoped unit verification
Dashboard publish deprecation gate Service integration Real SQLite models and catalog fixtures None in the added deprecation scenarios CLEAN
Frontend models/components L1/L2 Real model/component SUT API and i18n/callback boundaries CLEAN

No changed file was under backend/tests/integration/, and no Testcontainers/Docker fixture changed. make test-integration was therefore not required.

Clean Tests

  • All changed frontend tests: no SUT mocks and no Logic Mirrors.
  • backend/tests/services/dashboard_testing/test_verification_publish_gate.py: added deprecation scenarios use a real DB fixture.
  • Changed lineage/load-testing DTO, model, policy, and route tests that use real local SUTs and mock only transport/task/logger boundaries.

Global Setup (Not Violations)

  • backend/tests/services/lineage/conftest.py
  • backend/tests/services/load_testing/conftest.py
  • Existing repository backend/tests/conftest.py and frontend setup files were treated as infrastructure.

Uncertain — Requires Human Review

File/area Mock target Why uncertain
backend/src/api/routes/__tests__/test_migration_routes.py Local client/service classes Route is SUT, but collaborators are local contract nodes rather than explicit [EXT:...] boundaries.
backend/tests/services/lineage/test_hooks.py Local hook functions and feature helper Wiring behavior is tested, but local delegation is replaced wholesale.
backend/tests/core/test_client_registry_load_capacity.py AsyncAPIClient and an instance helper Client construction/backpressure ownership overlaps the contract under test; private semaphore state is asserted in one scenario.

These mocks were left unchanged. No inline AUDIT_NOTE was added because they occur across many existing scenarios and the audit could not establish a precise per-declaration ownership violation without changing the production external ontology.

Coverage Summary

Commands Executed

Command Result
Prerequisite script PASS; active FEATURE_DIR=specs/039-dashboard-scenario-ui
Scoped backend pytest PASS: 353 passed
Scoped frontend Vitest PASS: 16 files, 75 tests
make test first attempt FAIL at collection: two duplicate module basenames
make test after harness repair TIMEOUT at 66%; no failure observed before Makefile timeout 600 returned 124
make test-frontend PASS: 185 files, 3708 tests
npm run build PASS; static site generated
Scoped Ruff FAIL: 53 findings, including B008 FastAPI dependency defaults and existing complexity/import debt
make lint FAIL: Ruff reported 7263 repo-wide findings before frontend lint could run
git diff --check PASS

Harness Repairs

  • Added backend/tests/services/load_testing/__init__.py to prevent collision with backend/tests/schemas/test_profile.py.
  • Added backend/tests/services/dashboard_testing/scenario/__init__.py to prevent collision with backend/tests/test_models.py.

Coverage Percentages

make coverage was not run because the required test/lint gate did not pass: backend full tests timed out and lint failed. Claiming coverage percentages from a partial run would be misleading. Existing project thresholds were therefore not compared.

Compact Coverage Matrix

Module / flow Existing verification Complexity Mock violations Guardrail status Needed verification
040 load-testing backend Scoped pytest green C3-C5 0 Rejected-path tests exist Add belief runtime markers/tests; repair remaining anchor mismatches
041 lineage backend Scoped pytest green C3-C5 0 Strong rejected-path coverage Add migration/settings response/update assertions
039 scenario UI L1/L2 scoped tests green C3-C4 0 Incomplete decision memory Wire scenario initialization; complete action/invariant coverage
Load/lineage frontend models L1 tests green C4 0 Sparse/absent Micro-ADR Add success/error molecular markers and malformed-response edges
Changed UX components L2 tests green where present C3-C4 0 Partial Add direct tests for untested components and reactive prop updates

Semantic Audit Verdict

[AUDIT_FAIL: semantic_noncompliance]

Orthogonal Verdict

Projection Status Critical High Medium Low
P1 Contract completeness FAIL 0 1 5 3
P2 Decision-memory continuity FAIL 0 1 6 0
P3 Attention resilience FAIL 1 3 10 3
P4 Coverage and traceability FAIL 0 1 7 4
P5 Architecture realism PASS 0 0 0 0
P6 Protocol alignment FAIL 1 2 5 3
P7 Non-functional readiness FAIL 0 2 2 2

Contract Density and Anchor Status

  • Frontend changed anchors are generally balanced and hierarchical.
  • Backend load-testing has unclosed test module regions and several misnested child closing IDs in production service files.
  • Duplicate/stray closings remain in backend/src/core/superset_client/_datasets.py, backend/src/services/dashboard_testing/verification_publish_gate.py, and its test module.
  • Several changed C4 models/services lack complete @RATIONALE + @REJECTED decision memory.
  • New test modules frequently omit @TEST_INVARIANT and the three-edge floor metadata.

Belief Runtime Instrumentation

  • Lineage services have meaningful REASON/REFLECT/EXPLORE coverage.
  • The load-testing service package contains many C4/C5 contracts but lacks molecular runtime logging in core state transitions; the provided marker fixture is not exercised.
  • Frontend Dashboards.LoadRunModel and Datasets.DeprecationModel have network side effects without molecular markers. Other new models have only partial markers.

ADR / Rejected-Path Status

  • Existing load-testing and lineage tests defend important rejected paths such as baseline isolation, no raw SQL/query-context execution, inert lineage when disabled, and separated schedulers.
  • No evidence was found that a global ADR-rejected architecture was deliberately restored.
  • Several local @REJECTED declarations lack paired @RATIONALE; frontend C4 models largely lack local decision memory and explicit rejected-path tests.

Issues Found and Resolutions

Resolved

  1. Two Logic Mirrors replaced with hardcoded fixtures.
  2. Two pytest collection collisions fixed with package boundaries; no test was removed or renamed.
  3. Changed narrow backend and frontend test suites pass.
  4. Full frontend test suite and production build pass.
  5. All changed Python/TypeScript/Svelte #region/#endregion IDs are now balanced (verified with a strict multiset scan returning ANCHORS OK). Fixed: 13 unclosed test-module regions, misnested child closings across 8 load-testing service files, duplicate region IDs in _datasets.py and plugins/load_testing.py, and a stray parent close in verification_publish_gate.py.
  6. Decision memory added where @REJECTED lacked @RATIONALE: capacity.py, gates.py, matrix.py, profile.py, plus @RATIONALE/@REJECTED on frontend LoadRunModel and DeprecationModel and the workspace gate-delegation contract.
  7. Complexity reclassified for pure transforms (aggregates, bounded_result, cache_metadata, comparison, capacity, matrix, profile, schedule_policy, timing) from C4 to C3 — they have no state/side effects and should not carry runtime belief markers.
  8. Molecular logging added to genuinely stateful flows: gates.evaluate_prod_gate (REASON/REFLECT/EXPLORE), breaker trip EXPLORE, persistence flush REASON/REFLECT, runner_pool run/error markers. Executable marker_capture tests assert gate approve/deny, breaker trip, and flush markers. The dead marker_capture fixture was wired to src.core.logger.logger.
  9. Frontend runtime defects fixed: DashboardHeader now reads $t.dashboard_testing?.create_scenario (correct namespace, no Russian fallback leak); workspace gains beginScenario() for a valid inspecting initial state, requestSave()/approveBaseline() emit typed pending-gate intents, and canGenerateDraft precedence/blockers are corrected with tests; LoadRunModel and DeprecationModel add molecular markers; load-testing/[id]/+page.svelte anchor moved to an HTML comment region.
  10. Coverage closed: lineage_index_enabled GET/PUT round-trip tests, start_run external-failure propagation, and marker/edge metadata on new L1 suites.

Unresolved Blocking Findings

  1. Full backend suite exceeds the configured 600-second Makefile limit.
  2. Repo-wide Ruff baseline is not clean; new load-testing files are clean, but pre-existing FastAPI routes (B008 Depends in defaults, TID252 relative imports, RUF003 Cyrillic comments) and test_migration_routes.py (E402 sys.path, N806) remain as repo-wide debt.
  3. Coverage percentages remain unmeasured; integration tests still out of scope (no changed Testcontainers boundaries).

Remaining Risk or Debt

  • Uncertain local collaborator mocks need production contract ownership labels before strict classification can be automated.
  • Backend unit suite duration exceeds the Makefile timeout; split or parallelize suites, or revise the timeout based on measured CI duration.
  • Coverage percentages remain unknown because coverage was correctly withheld after the full-suite timeout.
  • Repo-wide Ruff debt in unchanged FastAPI/legacy test files is out of scope for this change and should be tracked separately.

Two-Layer Frontend Summary

Layer Contract type Status Main gaps
L1 no-render Models PASS malformed-response edges on the new model suites are partially covered; declared workspace actions now implemented and tested
L2 render Components PASS reactive prop warnings in changed scenario components converted to $derived; some unchanged/new components still lack direct UX contract tests

Final Gate

Changed-scope functional regressions are green (backend 361, frontend 1988 in the changed scope, full build pass), and all semantic anchors in the changed scope are balanced. The comprehensive /speckit.test gate still cannot fully pass because the full backend suite exceeds the Makefile timeout, the repo-wide Ruff baseline is not clean, and coverage/integration remain out of scope. These are environment/repo-level constraints rather than defects introduced by the uncommitted change.


validation.md

Source: validation.md

#region Std.Opencode.ValidationReport [C:3] [TYPE ADR] [SEMANTICS validation,gate,scenario,dashboard-testing] @defgroup Validation Pre-implementation validation gate for dashboard scenario UI (039).

Status: PASS (re-validated 2026-08-07 after MVP audit + closure)

Date: 2026-08-07 Feature: 039-dashboard-scenario-ui Branch: 039-dashboard-scenario-ui

Note (2026-08-07): Reconciliation pass (REST-binding, pipeline-view binding) amended spec.md, tasks.md, checklists; digests regenerated. Phase 9 tasks T054–T057 are complete (agent-free REST preview, verification history on /dashboards/[id]); T058 remains open (verify-action blocked: no repository_id in dashboard metadata, no PREPROD page).

Validated Inputs

Artifact Size (bytes) SHA-256
spec.md 19606 b2675dc3a04f033e99dba8496afdec0453bb3e255ca739e0d7b003df939e0e6e
plan.md 5106 aa3bfd7669c788ccd5c0cce989e646f0f59d6bb8f1acf83b5f416bbc260b14f0
tasks.md 12562 a0830751be6a75eab54d95b305a64dfaf1cd79f8560e861ae7c0435ded1a9f62
traceability.md 2644 a3493603e77f1dc361dea6f9584ee76d47e4743308b12ca680a66077a5e16740
contracts/modules.md 9812 5639ea1dc9b8d9930084ccee3bf419fbd71d9128a44bbebc8454b0e8f25f9e51
data-model.md 4829 ee3fb070d1fe6a255b9e8e550dc93339a8b49897f860e288cbf45fac55e3ba36
research.md 5094 8831641edf3123648c86ec2c0df0cf81c91a61b40ce0b5d31c04f8d0d3bd1450
ux_reference.md 9632 dedd1072b3ecc58b474d300076295c7a35a91813ec0252c953c1e4649ed5cf2d
quickstart.md 2541 a915e4ea6b5de2a46e1b9ae20414801aa42d552dcf4431f81d06f16a692d611a

Verdict is stale if any digest changes or new applicable artifacts appear.

Blocking Findings

✅ No blocking findings for the specified scope. Phase 9 runtime-binding tasks T054–T057 complete; T058 (verify-action) open due to repository_id/PREPROD-page blocker — documented in tasks.md.

Check Results

Phase 1: Unresolved Markers

  • [NEEDS CLARIFICATION]/[NEED_CONTEXT]/TODO/TBD: 0 → ✅ PASS

Phase 2: Artifact Completeness

  • All required artifacts present → ✅ PASS

Phase 3: Schema & Contract Validation

  • Region pair balance: 12 opens / 12 closes → ✅ PASS
  • Anchor signatures [C:N] [TYPE] [SEMANTICS] one-line → ✅ PASS
  • Hierarchical IDs + shared scenario/dashboard-testing semantics → ✅ PASS

Phase 4/5: Reference & Decision Memory

  • @REJECTED (dropdown tool-choice as primary entry) not scheduled in tasks; no-direct-SQL copy in UX requirements → ✅ PASS

Phase 6: Task Dependency & Path

  • Task count: 53; all paths under frontend/src → ✅ PASS
  • Decoupling from 033/035 confirmed (workspace runs off 036 AgentRunModel; /agent context v2 intent; no gradio backend requirement for workspace) → ✅ PASS

Phase 7/8

  • UX surface present (ux_reference.md); Axiom MCP unavailable — manual fallback (W01).

Gate Decision

Verdict: ✅ PASS — /speckit.implement may proceed.

#endregion Std.Opencode.ValidationReport

================================================================================ FEATURE: 040-dashboard-load-testing Files: 13


SPEC — Feature Specification

Source: spec.md

#region DashboardLoadTesting.Spec [C:3] [TYPE ADR] [SEMANTICS spec,requirements,load-testing,dashboard-testing,variations] @BRIEF Load testing interface for dashboards: bounded parallel execution of Superset-native checks across declarative variation axes, with circuit breaker, latency metrics, consistency detection, and dataset blast-radius awareness. @RELATION DEPENDS_ON -> [Doc.Adr.ADR0001] @RELATION DEPENDS_ON -> [Doc.Adr.ADR0003] @RELATION DEPENDS_ON -> [Doc.Adr.ADR0005] @RELATION DEPENDS_ON -> [Doc.Adr.ADR0006] @RELATION DEPENDS_ON -> [AgentTestStabilization.Spec] @RELATION DEPENDS_ON -> [SupersetBaselineEngine.Spec] @RATIONALE Dashboard correctness (037) proves a query returns the right value once; it does not prove the dashboard survives concurrent load, nor that repeated identical executions return identical results. A dedicated load surface reuses the 037 Superset-native executor so load fidelity matches production execution semantics — a separate load stack would diverge exactly like the rejected direct-SQL path. @REJECTED Treating load runs as 036 AgentRun instances — rejected because load execution is non-conversational, fan-out by design, and would pollute run/recovery semantics with thousands of pseudo-conversations. @REJECTED Writing load results into the 037 baseline catalog — rejected because load executions measure latency and consistency, not truth; polluting baselines with load samples would corrupt immutability detection. @REJECTED Unbounded client-declared concurrency — rejected because a single misconfigured run could saturate the Superset/KXD connection pool and degrade production for all users. @RATIONALE Load testing never depended on the chat agent; after the 050 MCP drift, LOAD-FR-020 delegated experiments and diagnostics are exercisable by any governed actor including external MCP clients, with caps, breaker and PROD gate unchanged.

Navigation (DSA Indexer keywords)

@SEMANTICS: spec, requirements, feature, load-testing, concurrency, variations, circuit-breaker, latency, blast-radius, dashboard-testing

Feature Branch: 040-dashboard-load-testing Created: 2026-07-22 | Status: Draft Input: "Интерфейс для нагрузочного тестирования дашбордов — параллельный запуск нескольких проверок на одном дашборде + вариации. Учитывает blast-radius: общий dataset, cache warming одного дашборда влияет на latency зависимых дашбордов."

Clarifications

Session 2026-07-22

  • Q1 (execution model): Worker pool, not HTTP-connection cap. Per-run async worker queue + env-level semaphore (sum of workers across active runs per env ≤ max_concurrent_per_env). Workers check circuit breaker between executions → clean drain. Run registered as ONE TaskManager task (observable in Task Center); executions are not tasks. Pool-wait = explicit queue-wait metric. → LOAD-FR-002/016/017 added.
  • Q2 (concurrency caps): prod=5, default (dev/preprod)=10, absolute ceiling=25 (not client-overridable); per-env override load_test_max_concurrent in environment config, clamped to ceiling. Ramp defaults: prod 1→2→3→5; others 1→3→5→10. → LOAD-FR-002 amended.
  • Q3 (GIL analysis): Async single-loop model kept — at cap 25 and Superset latency 0.5–3s, workers are ~95–99% I/O-bound (~40% of one core CPU). GIL contention risk exists only for large table charts (10k+ rows, ~200–500ms continuous parse+normalize). Mitigation B1+B2: load path uses bounded normalization (response sha256 hash + row count + first-N-row sample for display; hash suffices for consistency checks); full 037 normalization stays in the verification path. ProcessPool offload (B3) deferred unless fixtures show >10% GIL distortion. → LOAD-FR-018 added.
  • Q4 (cache state, source-audited): Authoritative source is Superset /api/v1/chart/data JSON body (result[].is_cached, cache_key, cached_dttm, queried_dttm, cache_timeout) — verified against /home/busya/dev/superset/superset/common/query_context_processor.py, charts/data/api.py, and charts/schemas.py. State mapping: hit if is_cached=true; bypassed if request.force=true; disabled if cache_timeout=-1; miss if is_cached=null + cached_dttm=null + force=false; unknown if fields absent/inconsistent. Latency-based probe inference rejected (DB buffer cache/network jitter are confounders). → LOAD-FR-012 amended, LOAD-FR-019 added.

User Scenarios

Story 1 — Configure Load Profile With Variations (P1)

Why P1: Analysts must declare what to load (dashboard, charts), how hard (concurrency, ramp), and across which axes (filters, viewport, role) before any execution.

Independent Test: Build a profile for a fixture dashboard with 2 variation axes and verify the preview shows exact request count, chart coverage, and blast-radius warning listing dependent dashboards sharing datasets.

Acceptance:

  1. Given a dashboard and environment When a load profile is configured Then the user selects concurrency target (bounded by server cap), execution mode (iterations or duration), ramp-up strategy, and variation axes (filters, viewport, role, time range).
  2. Given filter variations are declared When the profile is validated Then every variation value is checked against authoritative Superset filter metadata; unknown values produce NEEDS_CONTEXT markers, never invented values.
  3. Given the dashboard's charts share datasets with other dashboards When the profile preview renders Then a blast-radius panel lists dependent dashboard count and warns that cache warming/clearing affects them.

Story 2 — Controlled Parallel Execution (P1)

Why P1: Load execution must be bounded, observable, and stoppable — never a fire-and-forget request flood.

Independent Test: Run a fixture load profile with concurrency=5 and verify in-flight count never exceeds the cap, ramp-up is staged, and live progress shows running/completed/failed per variation.

Acceptance:

  1. Given a started load run When execution proceeds Then in-flight requests never exceed the server-enforced concurrency cap for that environment, and ramp-up increases load in declared steps.
  2. Given a running load run When the user requests stop Then in-flight requests complete or time out within a bounded drain window, no new requests are dispatched, and the run terminates with a stopped_by_user status.
  3. Given executions hit Superset When responses return Then every result is recorded with chart id, variation coordinates, latency, and typed error taxonomy (403/404/422/5xx/timeout) inherited from 037.

Story 3 — Circuit Breaker and PROD Safety Gate (P1)

Why P1: A load run against degraded infrastructure must self-terminate before it amplifies an incident; PROD runs require explicit human approval.

Independent Test: Simulate error-rate exceeding threshold mid-run and verify automatic abort with typed terminal reason; verify PROD runs cannot start without a 036-style approval gate.

Acceptance:

  1. Given error rate exceeds the declared threshold OR p99 latency exceeds N× the profile baseline When the circuit breaker evaluates Then the run aborts with terminal status circuit_breaker_abort including the triggering metric and threshold.
  2. Given a load profile targets a PROD-classified environment When start is requested Then an approval gate shows concurrency ceiling, estimated request volume, blast-radius dependents, and requires a user-supplied reason.
  3. Given the user denies the PROD gate When denial is submitted Then no request is dispatched and the denial is recorded in the run audit trail.

Story 4 — Results: Latency, Errors, Consistency (P2)

Why P2: Load value is in the analysis — percentiles per chart, error breakdown, and detection of non-deterministic results under parallelism.

Independent Test: Complete a fixture run with injected latency variance and one consistency violation; verify percentiles, error taxonomy counts, and the violation are reported without touching baseline state.

Acceptance:

  1. Given a completed run When results render Then per-chart p50/p90/p95/p99 latency, throughput, and error counts by taxonomy category are shown.
  2. Given two executions in one run share identical (chart, normalized filters) coordinates When their normalized results diverge Then a consistency_violation finding is recorded with both result hashes — flagged as flakiness/race, never as a baseline event.
  3. Given a load run completes When the 037 baseline catalog is inspected Then no baseline entry, candidate, source_response_hash, or immutability status has changed.

Story 5 — Blast-Radius and Cross-Run Comparison (P2)

Why P2: Cache effects cross dashboard boundaries; analysts need to compare runs and see shared-dataset impact explicitly.

Independent Test: Run the same profile twice (cold vs warm cache) and a dependent-dashboard probe; verify the comparison view shows latency delta and attributes the shift to shared-dataset cache state.

Acceptance:

  1. Given two runs of the same profile When comparison opens Then per-chart latency deltas are shown with cache-state annotation (cold/warm/unknown) derived from response metadata.
  2. Given dashboards share a dataset with the tested dashboard When results render Then the blast-radius panel links each dependent dashboard and marks whether it was probed during the run window.
  3. Given a run targets a dashboard whose datasets serve PROD dashboards When scheduling is configured Then the schedule policy warns about cache-interference windows, not just absolute concurrency.

Edge Cases

  • Superset connection pool exhaustion in the environment → executor queues within cap, records pool_wait latency separately, never bypasses the cap.
  • Variation references a filter value that exists in PREPROD but not PROD → validation marks the variation NEEDS_CONTEXT per environment; the run may proceed with the valid subset only.
  • Circuit breaker trips during ramp-up → partial results are preserved and the run is analyzable; abort is not data loss.
  • Identical chart+filters executed concurrently return different row ordering → normalization (037) canonicalizes before consistency comparison; ordering alone is not a violation.
  • User closes the browser mid-run → the run continues server-side; reconnecting shows live status by run id (recovery semantics from 036 US2).
  • Scheduled load run overlaps a deployment window for a blast-radius-dependent dashboard → schedule policy blocks or warning-gates the overlap.

Requirements

Functional

  • LOAD-FR-001: All load executions MUST go through the 037 Superset-native chart/dataset execution path; direct SQL, generated SQL, and raw query_context injection are forbidden (inherits 037 @REJECTED).
  • LOAD-FR-002: Concurrency MUST be enforced via a worker pool model: per-run async worker queue plus an env-level semaphore guaranteeing the sum of workers across all active runs on one environment never exceeds max_concurrent_per_env. Defaults: PROD-classified=5, dev/preprod=10, absolute ceiling=25 (not client-overridable); per-env override load_test_max_concurrent clamped to ceiling. Client values above the cap are clamped with a visible notice. Ramp defaults: PROD 1→2→3→5; others 1→3→5→10.
  • LOAD-FR-016: Each load run MUST be registered as exactly ONE TaskManager task (Task Center observability); individual executions are queue items, never tasks (no Reports flooding). Workers MUST check the circuit breaker between executions — in-flight completes, new executions are not taken (clean drain semantics).
  • LOAD-FR-017: Worker state MUST be observable per run: workers busy/idle, current execution per worker, queue depth — powering the live progress of LOAD-FR-008.
  • LOAD-FR-018: Load-path result processing MUST use bounded normalization: full-response sha256 hash + row count + first-N-row sample for display. Consistency checks (LOAD-FR-006) operate on hashes; full 037 normalization remains in the verification path only. This bounds continuous GIL hold per execution (large table charts) and keeps measured latency attributable to Superset, not to the load pipeline.
  • LOAD-FR-003: Load executions MUST NOT create, update, or invalidate 037 baseline entries, candidates, source_response_hash, immutability status, or visual baselines. Load artifacts live in a separate result store.
  • LOAD-FR-004: Every load run MUST have a circuit breaker with declared thresholds (error-rate %, p99 multiplier); breach aborts the run with typed terminal status and preserved partial results.
  • LOAD-FR-005: Starting a run against a PROD-classified environment MUST require an approval gate (036 semantics) showing concurrency ceiling, estimated request volume, blast-radius dependents, and a mandatory reason. Denial dispatches nothing.
  • LOAD-FR-006: Executions sharing identical (chart id, normalized filter context) within one run MUST be compared after 037 normalization; divergence MUST produce a consistency_violation finding, classified as flakiness — never as baseline drift.
  • LOAD-FR-007: Variation axes MUST be declarative and closed: filters, viewport, role, time_range. Unknown axes or values MUST fail validation; missing filter values MUST surface as NEEDS_CONTEXT, never invented.
  • LOAD-FR-008: Every load run MUST expose a stable load_run_id with live progress (per-variation running/completed/failed), recoverable after disconnect, with typed terminal statuses: completed, stopped_by_user, circuit_breaker_abort, failed.
  • LOAD-FR-009: The system MUST resolve the reverse index dataset → dependent dashboards for the tested dashboard and display it as a blast-radius panel at configure time, at the PROD gate, and in results.
  • LOAD-FR-010: Variation expansion MUST be deterministic for the same (dashboard query model, axes, seed): cartesian for ≤ declared cap, seeded sampling above it; the expanded matrix MUST be previewable before start.
  • LOAD-FR-011: Ramp-up MUST be staged (declared steps to target concurrency); steady state and drain phases MUST be distinguishable in progress and metrics.
  • LOAD-FR-012: Result records MUST include per-execution provenance: environment, dashboard, chart, variation coordinates, normalized filters hash, HTTP latency, queue-wait time, cache-state metadata, and Superset error taxonomy. Cache-state metadata comes authoritatively from /api/v1/chart/data response body: is_cached, cache_key, cached_dttm, queried_dttm, cache_timeout; mapped to hit|miss|bypassed|disabled|unknown with cache_state_source="superset_response".
  • LOAD-FR-019: Cache-state mapping MUST be: hit iff is_cached=true; bypassed iff request force=true; disabled iff cache_timeout=-1; miss iff is_cached=null AND cached_dttm=null AND force=false; unknown when fields are absent/inconsistent. Latency-based cache inference is forbidden because DB buffer cache, connection reuse, and network jitter are confounders.
  • LOAD-FR-013: Scheduled load runs MUST evaluate overlap against deployment/maintenance windows of blast-radius-dependent dashboards and block or warning-gate conflicts per policy.
  • LOAD-FR-014: RBAC MUST distinguish dashboard:loadtest:execute (PREPROD/staging) from dashboard:loadtest:prod (PROD-classified); unauthorized actors see permission_denied, never a confirm control (036 gate semantics).
  • LOAD-FR-015: Comparison of two runs of the same profile MUST show per-chart latency deltas and consistency-finding deltas; comparison MUST NOT require both runs to be complete (partial-current vs baseline-run allowed).
  • LOAD-FR-020: An opened InvestigationCase MAY ask the agent to construct/profile/compare a load experiment and autonomously run policy-permitted diagnostics. The agent MUST NOT bypass caps, mutate an active ramp, suppress a circuit breaker, or replace a required ActionApprovalGate.
  • LOAD-FR-021: Circuit-breaker aborts and consistency findings MUST emit idempotent 036 InvestigationSignals with blast-radius and run provenance; 047 creates/updates Queue items and MUST NOT automatically start agent work.

Key Entities

  • LoadProfile: Declarative configuration — dashboard, environment, concurrency target, ramp steps, execution mode (iterations|duration), variation axes with values, circuit-breaker thresholds, schedule policy. Versioned and reusable.
  • LoadVariation: One expanded coordinate set from the variation matrix — concrete filter values, viewport, role, time range. Immutable once the run starts.
  • LoadRun: Recoverable execution instance of a profile: id, status, phase (ramp|steady|drain|terminal), per-variation progress, circuit-breaker state, approval linkage for PROD.
  • LoadExecution: Single Superset-native request record: variation coordinates, chart, HTTP latency, queue-wait, bounded result hash/count/sample, typed outcome, and authoritative cache metadata (is_cached, cache_key, cached_dttm, queried_dttm, cache_timeout, mapped cache_state). Never writes to baseline state.
  • ConsistencyViolation: Finding that identical (chart, filters) executions diverged post-normalization; carries both result hashes; classified flakiness, not baseline drift.
  • BlastRadiusReport: Reverse-index projection — dataset ids used by the tested dashboard, dependent dashboards per dataset, cache-interference warning level, probe coverage.
  • CircuitBreakerPolicy: Thresholds (error-rate %, p99 multiplier), evaluation window, abort semantics, preserved-partial guarantee.
  • LoadRunComparison: Delta view of two runs — per-chart latency shift, cache-state attribution, consistency-finding delta.
  • LoadRunnerPool: Per-environment worker pool — async workers executing queue items, env-level semaphore enforcing max_concurrent_per_env across all active runs, worker registry (busy/idle, current execution, queue depth), breaker-checked between executions.
  • ExecutionQueue: Per-run FIFO of LoadExecution items; pool-wait measured as queue-residence time; drain = stop enqueue + bounded wait for in-flight.

Success Criteria

  • SC-001: In 100% of fixture runs, in-flight requests never exceed the configured cap, including during ramp-up and drain.
  • SC-002: Circuit breaker aborts within one evaluation window of threshold breach in 100% of fault-injection tests; partial results remain fully queryable.
  • SC-003: Zero mutations to the 037 baseline catalog across the entire load-test fixture suite (verified by catalog hash before/after).
  • SC-004: Variation matrix expansion is byte-deterministic across repeated expansions and shuffled axis ordering.
  • SC-005: Blast-radius panel lists 100% of dependent dashboards for fixture shared datasets at configure time and in results.
  • SC-006: PROD gate blocks 100% of unapproved PROD starts; denial produces zero dispatched requests.
  • SC-007: Consistency violation detection catches 100% of injected divergent-result fixtures and produces zero false positives on ordering-only differences.

Implementation Status & MVP Debt (audit 2026-08-07)

Facts (code check, not tasks.md):

  • ✅ Задачи T001–T074 закрыты: API routes, runner pool module, breaker, matrix, capacity, PROD gate, сравнение, фронт-модели.
  • 🔴 Главный MVP-долг: backend/src/plugins/load_testing.py::run_load_run не вызывает RunnerPool — он лишь переводит фазы RAMP→STEADY→DRAIN→COMPLETED и спит hold_seconds, не отправляя ни одного запроса к Superset. execute_chart_data из 037 в load-пути не используется (нарушение LOAD-FR-001). RunnerPool.run() нигде не инстанцируется в production-пути.
  • 🔴 Следствие: /status и /compare отдают seeded/пустые данные, circuit breaker не оценивается на реальном потоке, consistency-детектор не имеет входных результатов.

Закрытие: обязательные задачи T075–T079 в tasks.md Phase 9 (executor через 037, shared semaphore с client_registry, breaker, персистенция LoadExecution). Фича не может быть объявлена operational до их мерджа.

Runtime Closure Status (2026-08-07, resolved)

  • ✅ T075–T079 закрыты: run_load_run теперь вызывает RunnerPool через _resolve_execution_scope + _run_bounded_pool; 037-native executor (services/load_testing/executor.py) исполняет chart-data; shared per-env semaphore (client_registry.get_semaphore); LoadExecution персистятся (write_load_executions); breaker → circuit_breaker_abort.
  • ✅ Verified: tests/services/load_testing/test_executor_runtime.py (5 тестов, включая real-execution→LoadExecution), полный tests/services/load_testing/ = 76 passed, ruff clean.
  • 🟡 Осталось для operational: живые/фикстурные Superset-прогоны (exit gate 6), интеграционный тест capacity с ordinary-запросами (T077 extension).

Drift Amendment — MCP Interface (2026-08-24)

  • No chat dependency existed; load surfaces stay web-first. MCP exposure of read-only run/comparison queries is optional future catalog work (050), never bypassing LOAD-FR-002 caps or the PROD gate.

Status (2026-09-02): done — реализовано в рамках 050: инструменты и гейты (specs/050-mcp-interface/tasks.md T012–T028 [x]), handoff-поверхность (050 T030–T033), демонтаж чата и сервиса agent/ (050 T040–T041, чекпоинты specs/WORKSTATE-043-047.md).

#endregion DashboardLoadTesting.Spec


UX REFERENCE — Interaction Narrative

Source: ux_reference.md

#region DashboardLoadTesting.UxReference [C:3] [TYPE ADR] [SEMANTICS ux,reference,load-testing,dashboard-testing] @BRIEF UX interaction reference for dashboard load testing: persona, happy path, states, errors, recovery. Drives @UX_* tags in Phase 1 contracts.

Feature Branch: 040-dashboard-load-testing Created: 2026-07-22 | Status: Draft

1. User Persona & Context

  • Who is the user?: BI engineer or DevOps engineer validating that a dashboard survives realistic concurrent load — before a release, after infrastructure changes, or on a schedule.
  • What is their goal?: Answer three questions quickly: (1) how fast is each chart under N parallel users, (2) does the dashboard error or degrade under load, (3) do identical requests return identical results. They also need to know who else is affected — the dashboards sharing datasets with the tested one.
  • Context: Web UI on desktop. Starts from a dashboard page («Нагрузочное тестирование» action) or the load-testing hub. For PROD runs, expects an explicit approval step with full impact disclosure.

2. The "Happy Path" Narrative

The user opens a dashboard and clicks «Нагрузочное тестирование». The profile form opens pre-filled: charts from the dashboard query model, concurrency slider clamped to the environment cap, axes for filter/viewport variations. The blast-radius panel immediately shows «3 других дашборда используют те же датасеты» — the user understands the cache impact before starting. The expanded matrix preview reads «120 запросов × 8 чартов, оценка ~4 мин». They press Start; the run ramps visibly (0→5→10 in-flight), live tiles fill per variation, and the circuit breaker badge stays green. On completion, percentiles per chart render with one consistency finding flagged — clickable, showing both divergent result hashes. The baseline catalog badge confirms «baselines не затронуты».

3. Interface Mockups

UI Layout & Flow

Screen: Load Profile Editor (/dashboards/[id]/load/new)

  • Layout: Three-column — left: profile form; center: variation matrix preview; right: blast-radius panel.
  • Key Elements:
    • Concurrency slider: Min 1, max = server cap for the environment; values above cap clamp with notice «Снижено до лимита среды (N)».
    • Axes selector: Checkboxes — Фильтры / Viewport / Роль / Период. Each axis expands to typed value inputs validated against Superset metadata.
    • Matrix preview: «N комбинаций × M чартов = K запросов» with deterministic seed shown; above cap → «сэмплирование, seed=42».
    • Blast-radius panel: Dataset list → dependent dashboards with links; warning icon when any dependent is PROD-classified.
    • Circuit breaker thresholds: Error-rate % and p99 multiplier inputs with safe defaults pre-filled.
    • Start button: Primary; for PROD environments opens the approval gate instead of immediate start.
  • Contract Mapping:
    • @UX_STATE: idle → validating → matrix_ready → starting → ramping → steady → draining → terminal(completed|stopped_by_user|circuit_breaker_abort|failed).
    • @UX_FEEDBACK: Clamp notice, NEEDS_CONTEXT markers on invalid variation values, breaker badge state (green/armed/tripped).
    • @UX_RECOVERY: Invalid variation → fix inline or drop variation; start failure → retry with same profile; abort → results remain analyzable.
    • @UX_REACTIVITY: Screen model atoms (profile, matrix, blastRadius, runStatus) $derived into panels; live run updates via the 036 recovery channel keyed by load_run_id.
    • Screen Model: Dashboards.LoadProfileModel.svelte.ts (editor FSM + matrix derivation) and Dashboards.LoadRunModel.svelte.ts (live execution state). Components bind via @RELATION BINDS_TO.

Screen: Load Run Monitor (/load-runs/[id])

  • Layout: Header (status, phase, breaker badge, Stop button) → per-variation progress grid → live latency sparkline per chart → findings strip (consistency violations, error taxonomy counts).
  • States:
    • Ramping: Phase badge «Разгон N→M»; in-flight counter visible.
    • Steady: Tiles update per execution; pool-wait shown separately from query latency.
    • Draining: «Завершение in-flight, новые не отправляются».
    • Terminal: Summary cards persist; partial results labeled «частичные — прервано» when aborted.

Screen: Results & Comparison (/load-runs/[id]/results)

  • Key Elements: Percentile table per chart (p50/p90/p95/p99), error taxonomy breakdown, consistency findings list with hash-pair detail, cache-state annotation per chart (cold/warm), comparison selector against prior run of same profile, blast-radius probe coverage.

4. The "Error" Experience

Philosophy: Load failures are expected data, not UI exceptions. Every failure is classified, preserved, and analyzable.

Scenario A: Invalid variation value

  • User Action: Enters filter value «2024-Q9» that doesn't exist in the target environment.
  • System Response: (UI) Value chip turns amber, marked NEEDS_CONTEXT with «значение отсутствует в среде PROD»; run may start with valid subset only — excluded variations listed explicitly.
  • Recovery: Fix the value inline (matrix re-expands deterministically) or drop the axis row; no page reload.

Scenario B: Circuit breaker abort

  • System Response: Breaker badge flips to red «Прервано: error-rate 34% > 25%»; phase moves to draining; terminal banner explains the trigger and preserved partial results.
  • Recovery: «Показать частичные результаты» (default), «Повторить с меньшей concurrency» (profile pre-filled with halved value), or export findings.

Scenario C: PROD gate denial / permission denied

  • System Response: For denial — audit-recorded cancellation card, zero requests dispatched. For unauthorized — permission_denied recovery panel explaining required role (dashboard:loadtest:prod), no confirm control rendered.
  • Recovery: Request access link; or switch environment to PREPROD and start without gate (if policy allows).

Scenario D: Browser disconnect mid-run

  • System Response: Run continues server-side. On return, status restores by load_run_id — phase, per-variation progress, findings to date.
  • Recovery: Automatic; no user action needed.

5. Tone & Voice

  • Style: Concise, technical, quantitative. Numbers first («120 запросов, p99 1.8s»), prose second.
  • Terminology: «Нагрузочный профиль», «вариация», «матрица», «circuit breaker» (kept in English in RU locale), «blast-radius» explained as «затронутые дашборды». Never «SQL» as an option; copy states Superset-native execution explicitly.
  • Safety copy: PROD gate must state impact in concrete terms: «до N параллельных запросов к среде X; K дашбордов делят датасеты — кэш будет прогрет».

#endregion DashboardLoadTesting.UxReference


CHECKLISTS — Requirements Quality — requirements.md

Source: checklists/requirements.md

Requirements Checklist: Dashboard Load Testing

Purpose: Validate spec.md completeness, clarity, and testability before /speckit.plan. Created: 2026-07-22 Feature: spec.md

Content Quality

  • CHK001 Spec is user/operator-focused; no framework/library implementation leakage (executor referenced only as "037 Superset-native path" contract boundary)
  • CHK002 All five stories are independently testable with a stated Independent Test
  • CHK003 Priorities assigned P1/P2 with justification per story
  • CHK004 Key entities defined without schema/implementation detail

Requirement Quality

  • CHK005 Every FR is testable (cap enforcement, zero baseline mutation, deterministic expansion, gate blocking)
  • CHK006 FR IDs hierarchical and stable (LOAD-FR-001..015)
  • CHK007 No [NEEDS CLARIFICATION] markers remain
  • CHK008 Rejected paths recorded in header @REJECTED (AgentRun reuse, baseline pollution, unbounded concurrency)

Dependency Traceability

  • CHK009 037 dependencies explicit: execution path, normalization, error taxonomy (LOAD-FR-001/006/012)
  • CHK010 036 dependencies explicit: approval gate semantics, run recovery, permission_denied (LOAD-FR-005/008/014)
  • CHK011 Blast-radius requirement (LOAD-FR-009) consistent with gap analysis BR-1 (reverse dataset→dashboard index)
  • CHK012 No conflict with 038/039 scopes (no AgentRun, no scenario graph, no baseline writes)

Safety & Risk

  • CHK013 Circuit breaker mandatory with typed abort (LOAD-FR-004, US3)
  • CHK014 PROD gate + RBAC split execute vs prod (LOAD-FR-005/014)
  • CHK015 Baseline isolation verifiable via catalog hash (LOAD-FR-003, SC-003)
  • CHK016 Schedule overlap with deployment windows of dependent dashboards (LOAD-FR-013)

Success Criteria

  • CHK017 All SC measurable (100% cap enforcement, zero mutations, byte-determinism, 0 false positives)
  • CHK018 SC-004 determinism matches 038 determinism convention

Clarifications Integrated (2026-07-22)

  • CHK019 max_concurrent defaults — CLOSED: worker pool + env semaphore; prod=5, default=10, ceiling=25, per-env override clamped (LOAD-FR-002/016/017)
  • CHK022 GIL analysis accepted: async single-loop; bounded normalization in load path (hash+count+sample), full normalization stays in 037 (LOAD-FR-018); ProcessPool offload deferred pending fixture evidence
  • CHK020 Cache-state source — CLOSED after Superset source audit: authoritative chart-data response body fields (is_cached, cache_key, cached_dttm, queried_dttm, cache_timeout); latency inference rejected (LOAD-FR-012/019)
  • CHK021 Blast-radius reverse index — CLOSED by 041 LIN-FR-013: pinned read model, fingerprint query-param

PLAN — Implementation Plan

Source: plan.md

Implementation Plan: Dashboard Load Testing

Branch: 040-dashboard-load-testing | Date: 2026-07-22 | Spec: spec.md Input: Clarified feature specification from /specs/040-dashboard-load-testing/spec.md

Summary

Build a dashboard load-testing surface around bounded async workers, not raw HTTP connection count. A LoadRun is one TaskManager task with a per-run execution queue; an environment capacity registry and semaphore coordinate all active runs. The runner uses the 037 Superset-native chart-data path, preserves Superset cache metadata from the JSON response, computes queue/resource/upstream timings separately, uses bounded result processing to avoid GIL distortion, and stores a pinned 041 blast-radius snapshot for configuration and PROD approval. Load runs are strictly isolated from baseline mutation.

Technical Context

Language/Version: Python 3.13+ backend; TypeScript frontend with Svelte 5 runes-only Primary Dependencies: FastAPI 0.126, SQLAlchemy 2.0.45, APScheduler 3.11.2 for scheduled entry boundaries, httpx via existing AsyncAPIClient, Pydantic 2.x; SvelteKit 2.49, Svelte 5.56, Vite, Tailwind; no new runtime dependency Storage: PostgreSQL 16; new load profile/run/execution/aggregate/finding tables plus existing TaskManager persistence; no baseline catalog writes Testing: pytest unit/contract/integration tests; Vitest L1 Screen Model tests and L2 UX tests; Playwright for worker monitor/profile flow; ruff and frontend lint Target Platform: Linux Docker deployment, modern desktop browsers Project Type: FastAPI REST/WebSocket backend + SvelteKit SPA frontend Frontend Architecture: .svelte.ts Screen Models, $state/$derived/$effect, typed DTOs, components bind to models; ordinary AgentChat and 036 AgentRun remain separate Performance Goals: upstream p95/p99 measured independently from queue/resource wait; progress events ≤4 Hz; result flush ≤50 rows or 1s; no more than effective worker cap in flight; profile/matrix preview <200ms after metadata available Constraints: PROD default cap=5, DEV/PREPROD=10, absolute ceiling=25, reserve 5 shared pool slots; direct SQL/raw query context forbidden; bounded sample 100 rows/256 KiB; 041 fingerprint required for blast-radius binding; RBAC default-deny; no implementation phase in this planning request Scale/Scope: 100 dashboards/environment, 300 charts, 25 maximum load workers per environment, 500 selected variation combinations per profile by default

Constitution Check

GATE: Must pass before Phase 0 research. Re-check after Phase 1 design. — PASS with one explicit implementation prerequisite

Principle Status Evidence / Guardrail
I. Semantic Contract First ✅ C3–C5 contracts in contracts/modules.md, hierarchical LoadTesting.* IDs and cross-stack edges
II. Decision Memory ✅ Spec @RATIONALE/@REJECTED; research R1–R10; HTTP pool, per-execution tasks, latency inference, full normalization rejected
III. External Orchestrator ✅ 037 Superset-native adapter remains sole execution boundary; no Superset plugin or SQL path
IV. Module Discipline ✅ Runner, capacity, cache, bounded processing, breaker separated; target each module <400 LOC and CC≤10
V. RBAC Enforcement ✅ dashboard:loadtest:execute and dashboard:loadtest:prod; gate hash includes profile/cap/fingerprint
VI. Svelte 5 Runes Only ✅ LoadProfileModel/LoadRunModel .svelte.ts, no legacy stores in new surface
VII. Test-Driven C3+ ✅ Quickstart is falsifiable; C4/C5 contracts require rejection-path tests and GIL/cache/cap edge tests
VIII. Attention Optimization ✅ Anchor density, dotted IDs, shared load-testing semantic tags, bounded contract blocks
Repository prerequisite ⚠️ MUST FIX DURING IMPLEMENTATION client_registry.get_client() creates per-env semaphore but does not pass it to AsyncAPIClient; without this, connection_pool_size is declarative only. This is a foundational task, not a silent assumption.

Post-Phase-1 re-check: PASS — the prerequisite is explicitly represented in traceability and must be verified before load execution.

Project Structure

Documentation

specs/040-dashboard-load-testing/
├── spec.md
├── ux_reference.md
├── checklists/requirements.md
├── plan.md
├── research.md
├── data-model.md
├── contracts/modules.md
├── quickstart.md
└── traceability.md

Source Code (implementation target; not changed in this planning phase)

backend/src/
├── models/load_testing.py
├── schemas/load_testing.py
├── services/load_testing/
│   ├── capacity.py
│   ├── profile.py
│   ├── matrix.py
│   ├── runner_pool.py
│   ├── timing.py
│   ├── cache_metadata.py
│   ├── bounded_result.py
│   ├── breaker.py
│   ├── aggregates.py
│   └── comparison.py
├── api/routes/load_testing.py
├── core/utils/client_registry.py       # pass existing semaphore into AsyncAPIClient
└── plugins/load_testing.py              # one TaskManager task per LoadRun

frontend/src/
├── lib/types/load-testing.ts
├── lib/api/load-testing.ts
├── lib/models/Dashboards.LoadProfileModel.svelte.ts
├── lib/models/Dashboards.LoadRunModel.svelte.ts
├── lib/components/dashboards/load-testing/
└── routes/load-runs/[id]/

Structure Decision: Backend domain under services/load_testing/, separate from 037 correctness engine. core/utils/client_registry.py gets only the semaphore wiring prerequisite; no shared-client replacement because auth/CSRF session continuity is an existing invariant. Frontend is a dedicated workspace/monitor, not Agent Workspace.

Phase 0 Outputs

Artifact Decision
research.md R1 Async worker queue + env capacity registry; one TaskManager task per run
R2 Effective cap derives from stage, override, pool size minus 5 reserved slots, ceiling 25; existing semaphore wiring prerequisite
R3 Separate queue/resource/upstream timing; upstream percentiles only
R4 Superset response-body cache metadata authoritative
R5 Raw digest + bounded sample/count; no full 037 normalization in load path
R6 Batched persistence and ≤4 Hz progress
R7 Minimum-20 rolling breaker window; error/p99 thresholds
R8 Deterministic matrix and seeded sampling
R9 Pinned 041 blast-radius snapshot
R10 Stage/RBAC/approval gate hash semantics

Phase 1 Outputs

Artifact Coverage
data-model.md Profile, variation, run, execution, aggregates, findings, breaker, blast radius, invariants
contracts/modules.md Backend C3–C5 contracts and frontend Screen Models with UX states
quickstart.md 10-step falsifiable verification sequence
traceability.md LOAD-FR-001..019 to contracts and tests; cross-spec dependencies

Complexity Tracking

No constitution violations. The repository semaphore wiring is an explicit prerequisite, not a justification for bypassing the constitution.

MVP Runtime Closure (audit 2026-08-07)

Status correction: Факт-чекинг кода показал, что run_load_run в backend/src/plugins/load_testing.py не вызывает RunnerPool — фазы RAMP→STEADY→DRAIN→COMPLETED проходят без единого Superset-запроса; execute_chart_data из 037 в load-пути не используется (нарушение LOAD-FR-001). Задачи T001–T074 покрывают модули и API, но runtime-исполнение не замкнуто.

Closure tasks (tasks.md Phase 9, T075–T079):

  • T075 — Wire RunnerPool into run_load_run (executor + env semaphore + breaker + on_result persist).
  • T076 — 037-native executor adapter services/load_testing/executor.py (delegates to SupersetClient.ChartData.Execute, bounded normalization).
  • T077 — Share per-env capacity semaphore with client_registry admission.
  • T078 — Circuit breaker evaluation between executions → circuit_breaker_abort with preserved partials.
  • T079 — Persist per-execution LoadExecution so /status and /compare return real data.

Exit rule: 040 не объявляется operational, пока quickstart шаги 3–7 не выполняются против реального (или фикстурного, но через executor) Superset-потока, а не hold/sleep-перехода.

ADR Continuity

  • ADR-0001: services/load_testing/, models/load_testing.py, api/routes/load_testing.py follow module layout.
  • ADR-0003: all Superset interaction remains in existing client boundary.
  • ADR-0005: new permissions are explicit and default-deny; PROD uses 036 gate.
  • ADR-0006: new UI uses Svelte 5 Runes and model-first components.
  • ADR-0011: async worker execution stays on FastAPI event loop; no asyncio.run() and no blocking AsyncJobRunner.run() from the event-loop thread.
  • 036/037/041 continuity: one TaskManager run, 037 chart-data semantics, 041 fingerprint pinning; no baseline mutation.

Verification Gates (implementation phase only)

  1. Backend load-testing unit/contract tests and existing client-registry regressions.
  2. python -m ruff check ..
  3. Frontend Vitest/L2 tests, lint, build, and Playwright profile/run monitor flow.
  4. Quickstart steps 1–10.
  5. Semantic contract audit and runtime instrumentation audit for C4/C5 runner/breaker flows.
  6. MVP gate (new): run lifecycle test asserts executor invoked per queued item against real chart-data context; breaker abort preserves partials.

#endregion DashboardLoadTesting.Plan


RESEARCH — Technical Decisions

Source: research.md

#region DashboardLoadTesting.Research [C:3] [TYPE ADR] [SEMANTICS research,load-testing,workers,cache,gil] @BRIEF Phase 0 decisions for load execution: worker runtime, environment capacity, timing model, cache provenance, bounded result processing, persistence, and 041 integration.

Feature: 040-dashboard-load-testing | Date: 2026-07-22

R1 — Async Worker Runtime

  • Decision: One LoadRun is registered as one TaskManager task. The load plugin starts N asyncio workers reading a per-run FIFO asyncio.Queue[LoadExecutionSpec]. An environment-scoped LoadCapacityRegistry owns the load semaphore shared by all active runs. Workers check cancellation and circuit-breaker state before taking the next item; in-flight requests drain within a bounded timeout.
  • Rationale: The workload is 95–99% upstream I/O. Async workers represent execution concurrency directly, expose queue wait, and support deterministic drain. TaskManager already creates/tracks one asyncio task per plugin run and provides cancellation/events.
  • Alternatives Considered: (a) HTTP connection count as concurrency — rejected: indirect, hides queue wait and circuit-breaker boundaries. (b) TaskManager task per execution — rejected: floods Task Center/Reports and destroys run-level lifecycle. (c) threads/process workers — rejected for v1: mostly I/O; GIL mitigation is bounded result processing (R5).
  • Impact: New C5 LoadTesting.RunnerPool; no APScheduler or AsyncJobRunner for manual starts. Scheduled load callbacks may dispatch one TaskManager run through existing scheduler boundaries later.

R2 — Effective Capacity and Existing Client Semaphore Gap

  • Decision: Effective run concurrency = min(profile_requested, env.load_test_max_concurrent, max(1, env.connection_pool_size - env.load_test_reserved_slots), 25). Defaults: PROD=5; DEV/PREPROD=10; absolute ceiling=25; load_test_reserved_slots=5. Worker acquires load semaphore first, then request travels through the shared client semaphore; one global acquisition order prevents deadlock.
  • Repository Finding: Environment already has stage, is_production, connection_pool_size (default 20), and connection_pool_timeout. client_registry.get_client() creates a semaphore but currently constructs AsyncAPIClient(...) without passing semaphore=semaphore; therefore the documented per-env limit is not active.
  • Rationale: Load traffic must not occupy all shared Superset capacity and freeze ordinary UI/API operations. Capacity derives from the existing environment pool rather than an unrelated number.
  • Alternatives Considered: Separate httpx client/pool — rejected: loses shared auth/CSRF cookies; explicitly rejected by AsyncAPIClient contract. Acquiring shared semaphore manually in load runner — rejected: once registry wiring is fixed, this double-acquires the same semaphore and can deadlock.
  • Impact: Foundational prerequisite: pass registry semaphore into AsyncAPIClient; regression test proves ordinary calls and load calls share the same limit. Add config fields load_test_max_concurrent (optional derived default) and load_test_reserved_slots=5.

R3 — Timing Model and Percentiles

  • Decision: Record three non-overlapping monotonic timings per execution: queue_wait_ms (enqueued → worker takes), resource_wait_ms (worker takes → upstream request admitted), upstream_latency_ms (admitted → response/error). end_to_end_ms is their sum. Dashboard performance percentiles use upstream_latency_ms; orchestration health reports queue/resource wait separately. Percentiles use nearest-rank over successful samples; errors/timeouts have dedicated rates and are not converted into synthetic latency values.
  • Rationale: Combining waits with HTTP latency would measure our orchestrator under contention rather than Superset. Circuit breaker p99 must reflect upstream degradation; queue saturation has its own trigger/diagnostic.
  • Alternatives Considered: End-to-end p99 only — rejected: cannot attribute queue pressure vs Superset slowdown. Treat timeout as latency=timeout — rejected: biases percentile and double-counts error semantics.
  • Impact: LatencySummary stores sample_count, p50/p90/p95/p99 per timing dimension; breaker latency threshold evaluates upstream p99 only after minimum sample count.

R4 — Authoritative Cache State from Superset

  • Decision: Parse /api/v1/chart/data JSON body fields per query: is_cached, cache_key, cached_dttm, queried_dttm, cache_timeout. Mapping: hit iff is_cached=true; bypassed iff request.force=true; disabled iff cache_timeout=-1; miss iff is_cached=null && cached_dttm=null && force=false; unknown for absent/inconsistent fields. Persist raw fields plus cache_state_source="superset_response".
  • Evidence: Audited /home/busya/dev/superset/superset/common/query_context_processor.py (returns all fields), charts/data/api.py (serializes query results), charts/schemas.py (response schema), common/utils/query_cache_manager.py (hit semantics), and Superset integration tests (force/source execution produces is_cached=None).
  • Rationale: This is an authoritative execution signal from Superset, unlike latency inference.
  • Alternatives Considered: HTTP headers — rejected: chart-data uses body metadata. Probe inference from latency delta — rejected: DB buffer cache, connection reuse, and network jitter confound it.
  • Impact: 037 chart-data adapter must preserve these fields rather than normalize them away.

R5 — GIL and Bounded Result Processing

  • Decision: Load path computes SHA-256 while consuming the raw response bytes, extracts row count/cache metadata, and retains only a bounded first-N-row sample for display. It does not run full 037 value normalization. Consistency compares digest for identical (chart, filters_hash, role, time_range) coordinates. Default sample cap: 100 rows and 256 KiB serialized sample.
  • Rationale: At cap 25 and 0.5–3s upstream latency, workers are predominantly I/O-bound. The relevant GIL risk is large table JSON parsing/normalization (10k+ rows, 200–500ms continuous hold). Bounded processing keeps event-loop distortion low and ensures reported latency belongs to Superset.
  • Alternatives Considered: ProcessPool normalization — deferred; IPC cost distorts small responses. Full 037 normalization — rejected for load path; belongs to correctness verification. Hashing post-parsed normalized JSON — rejected: incurs full parse and may hide byte-level divergence.
  • Impact: If fixture instrumentation shows event-loop lag >10% of upstream latency at ceiling load, emit <ESCALATION> for a process/streaming parser ADR; do not silently add ProcessPool.

R6 — Persistence and Event Granularity

  • Decision: Persist LoadProfile, LoadRun, LoadExecution, and ConsistencyFinding in PostgreSQL. Workers buffer execution records and flush bounded batches (≤50 or ≤1s), while aggregate progress events are emitted at ≤4 Hz. Store digest/count/sample/cache/timing, never full large payload. Partial results survive stop/breaker/failure.
  • Rationale: Per-execution commit/event at concurrency 25 creates database/WebSocket amplification. Batching preserves auditability without turning the orchestrator DB into the bottleneck.
  • Alternatives Considered: Full response blobs — rejected: unbounded storage/GIL cost and duplicates Superset. In-memory results only — rejected: reconnect and cross-run comparison require durability.
  • Impact: C5 repository contract owns atomic batch insert + run aggregate update.

R7 — Circuit Breaker Semantics

  • Decision: Evaluate on a rolling completed-execution window with minimum 20 samples. Defaults: error-rate threshold 25%; upstream p99 multiplier 3× the selected comparison run/profile reference; if no reference exists, latency breaker is armed only with an explicit absolute threshold. Breach closes queue intake, transitions to draining, preserves partial results, records triggering window/metric.
  • Rationale: p99 on tiny samples is unstable; implicit latency reference would create false aborts. Error breaker works from the first minimum window.
  • Alternatives Considered: Break on first timeout — rejected: transient failures are expected load data. Use end-to-end p99 — rejected by R3 attribution model.
  • Impact: UI must show breaker inactive|armed|tripped and why latency breaker may be inactive.

R8 — Variation Matrix Expansion

  • Decision: Closed axes: filters, viewport, role, time_range. Normalize axis ordering and values, then cartesian-expand if total ≤ profile cap (default 500 combinations). Above cap, deterministic seeded reservoir sampling without materializing the full product. Each coordinate gets a stable variation id from canonical JSON hash.
  • Rationale: Determinism supports reproducible comparisons; bounded sampling prevents combinatorial memory/request explosion.
  • Alternatives Considered: Always cartesian — rejected: unbounded expansion. Random sampling without seed — rejected: cross-run comparison becomes invalid.
  • Impact: Preview shows theoretical combination count, selected count, seed, chart count, total requests.

R9 — 041 Blast-Radius Read Model

  • Decision: Profile validation calls 041 dependents endpoint with pinned fingerprint. LoadRun persists that fingerprint and the resolved BlastRadiusReport. Start rejects a changed fingerprint with a refresh/reconfirm response, especially before PROD approval. Mid-run index updates do not mutate the run.
  • Rationale: Configure-time and gate-time impact must name the same dependent dashboards; otherwise user approval is not bound to actual impact.
  • Impact: 040 cannot be implementation-complete before 041 LIN-FR-013 read contract exists; fixtures unblock independent domain tests.

R10 — PROD Gate and RBAC

  • Decision: Stage source of truth is Environment.stage (PROD) with is_production treated as backward-compatible alias; either marks PROD. PREPROD/DEV require dashboard:loadtest:execute; PROD additionally requires dashboard:loadtest:prod and 036 approval gate reason. Gate hash binds profile revision, effective cap, estimated request count, environment id, and 041 fingerprint.
  • Rationale: Existing configuration already exposes stage; dual interpretation avoids silently under-classifying legacy is_production=true environments.
  • Impact: Start-after-approval revalidates all hash inputs; mutation/stale gate dispatches zero requests.

Resolved Risks

Risk Mitigation
Load starves ordinary Superset calls Effective cap reserves 5 shared slots; registry semaphore wiring fixed (R2)
Double semaphore acquisition deadlock Single order: load semaphore then request-internal shared semaphore; never manually acquire shared semaphore
Event loop distortion on large tables Raw digest + bounded sample; no full normalization (R5)
Cache state misclassification Authoritative response body + force provenance (R4)
Per-execution DB/event amplification Batches ≤50/1s; progress ≤4 Hz (R6)
Stale blast-radius approval Fingerprint bound into gate hash; start revalidates (R9/R10)

R11 — MVP Runtime Gap (audit 2026-08-07)

  • Decision: run_load_run MUST instantiate RunnerPool(capacity=effective_concurrency, executor=037-query-executor, env_semaphore=shared, breaker=run_breaker, on_result=persist) and drain queued executions through it. The current hold/sleep body (фазовый переход без запросов) is a stub and violates LOAD-FR-001.
  • Rationale: RunnerPool уже реализован и покрыт тестами; отсутствует только wiring в production-путь. Подключение 037 SupersetClient.ChartData.Execute через адаптер services/load_testing/executor.py сохраняет единый источник metric truth и bounded-normalization (LOAD-FR-018).
  • Alternatives considered: in-process pool без shared semaphore (rejected — R2), прямой httpx в load-пути (rejected — LOAD-FR-001/037 continuity), запуск load как 036 AgentRun (rejected — LOAD spec).
  • Impact: T075–T079 в tasks.md Phase 9; quickstart шаги 3–7 не считаются пройденными до реального исполнения executor'а.

#endregion DashboardLoadTesting.Research


DATA MODEL — Entities & Relations

Source: data-model.md

#region DashboardLoadTesting.DataModel [C:3] [TYPE ADR] [SEMANTICS data-model,load-testing,workers,metrics] @BRIEF Phase 1 data model for load profiles, deterministic variations, worker runs, execution timing, cache provenance, and blast-radius snapshots.

Feature: 040-dashboard-load-testing | Date: 2026-07-22

LoadProfile

Field Type Invariant
id UUID stable reusable profile id
dashboard_id int authoritative dashboard
environment_id str target environment
revision int increments on profile mutation
concurrency_requested int positive; clamped to effective cap
execution_mode enum(iterations,duration) exactly one mode
iterations int? required for iterations
duration_seconds int? required for duration
ramp_steps list[int] monotonic, ends at effective target
variation_axes JSONB closed axes only: filters, viewport, role, time_range
circuit_breaker JSONB error threshold, p99 multiplier, min samples, window
matrix_seed int deterministic expansion seed
max_variations int default 500; server bounded
enabled bool schedule control
created_by str audit owner

An agent may propose or save a validated LoadProfile and launch policy-permitted diagnostic LoadRuns from an opened InvestigationCase. Server caps, variation validation, capacity allocation, circuit breakers and PROD gates remain deterministic; the agent never adapts an active worker ramp by prose.

LoadVariation

Expanded immutable coordinate stored with a run:

{
  "variation_id": "sha256(canonical_coordinates)",
  "filters": {"region": "north"},
  "viewport": {"width": 1366, "height": 768, "device_scale_factor": 1},
  "role": "analyst",
  "time_range": {"from": "2026-01-01", "to": "2026-01-31"}
}

Unknown axes fail validation. Above-cap matrices use seeded reservoir sampling; selected count and theoretical count are both retained.

LoadRun

Field Type Notes
id UUID load_run_id
profile_id UUID profile revision binding
profile_revision int immutable gate input
environment_id str
dashboard_id int
status enum(queued,ramping,steady,draining,completed,stopped_by_user,circuit_breaker_abort,failed) terminal status immutable
phase enum(ramp,steady,drain,terminal)
effective_concurrency int after cap/reserve/clamp
matrix_seed int
theoretical_variations int
selected_variations int
total_executions int selected variations × charts × iterations/duration policy
blast_radius_fingerprint str 041 pinned snapshot
blast_radius_report JSONB dependent dashboard projection at start
approval_gate_id UUID? required for PROD
task_id UUID exactly one TaskManager task
circuit_state enum(inactive,armed,tripped)
started_at/finished_at datetime?
stop_reason str? typed terminal reason

LoadExecution

Field Type Notes
id UUID
run_id UUID FK
variation_id str immutable coordinate
chart_id int
filters_hash str 037 canonical hash
outcome enum(success,error,timeout,cancelled)
error_taxonomy str? 037 categories
queue_wait_ms int enqueue → worker take
resource_wait_ms int worker take → shared request admission
upstream_latency_ms int admission → response/error
end_to_end_ms int sum of three timings
response_sha256 str? digest of raw response bytes
row_count int? bounded metadata
sample_rows JSONB? first N rows, bounded to 100 / 256 KiB
cache_state enum(hit,miss,bypassed,disabled,unknown) Superset body mapping
is_cached bool? raw Superset field
cache_key str? raw field
cached_dttm/queried_dttm datetime? raw fields
cache_timeout int? raw field
cache_state_source str superset_response when body fields exist
worker_id str registry observability
started_at/finished_at datetime monotonic timings also emitted

LoadRunAggregate

Per run and per chart/variation aggregation:

{
  "chart_id": 42,
  "variation_id": "...",
  "success_count": 19,
  "error_count": 1,
  "p50_upstream_ms": 240,
  "p90_upstream_ms": 520,
  "p95_upstream_ms": 670,
  "p99_upstream_ms": 920,
  "p95_queue_wait_ms": 18,
  "p95_resource_wait_ms": 12,
  "throughput_per_second": 4.8,
  "cache_counts": {"hit": 15, "miss": 5}
}

Percentiles use nearest-rank over successful samples only. Error rate is separate.

ConsistencyFinding

{id, run_id, chart_id, coordinate_key, first_response_sha256, divergent_response_sha256, execution_ids, severity, classification="flakiness", created_at}.

A row-order-only difference is normalized before comparison when the chart declares unordered table semantics; otherwise raw response digest divergence is reported.

CircuitBreakerPolicy

{error_rate_threshold=0.25, p99_multiplier=3.0, min_samples=20, window_size=100, absolute_p99_ms?}. A latency threshold without a comparison reference requires explicit absolute_p99_ms.

BlastRadiusReport

Pinned 041 projection: {fingerprint, stale_index, datasets: [{dataset_id, dependent_dashboards}], dependent_dashboard_count, probe_coverage}. A run stores the snapshot it was approved against; later index changes do not rewrite it.

Invariants

  1. Effective worker concurrency never exceeds environment cap, profile cap, absolute ceiling, or reserved-slot-adjusted shared pool.
  2. One LoadRun has exactly one TaskManager task; executions are not tasks.
  3. Queue/resource/upstream timings are disjoint; performance percentiles use upstream latency only.
  4. Cache state comes from Superset response metadata; latency inference is forbidden.
  5. Load executions never mutate 037 baseline catalogs, candidates, source hashes, or immutability status.
  6. Terminal status is immutable; partial executions remain queryable after stop or breaker abort.
  7. response_sha256 is calculated before bounded JSON/sample processing where raw response bytes are available.

#endregion DashboardLoadTesting.DataModel


CONTRACTS — Module & Function Contracts

Source: contracts/modules.md

#region DashboardLoadTesting.Modules [C:4] [TYPE ADR] [SEMANTICS contracts,modules,load-testing,workers] @BRIEF GRACE contracts for profile validation, variation expansion, worker pool execution, timing, cache provenance, aggregation, and API/UI surfaces. @RELATION DEPENDS_ON -> [DashboardLoadTesting.DataModel] @RELATION DEPENDS_ON -> [DashboardLoadTesting.Research] @RELATION DEPENDS_ON -> [DatasetLineageBlastRadius.Spec] @RELATION DEPENDS_ON -> [SupersetBaselineEngine.Spec] @RELATION DEPENDS_ON -> [AgentTestStabilization.Spec]

Backend Contracts

# #region LoadTesting.Profile.Validate [C:4] [TYPE Function] [SEMANTICS load-testing,profile,validation]
# @ingroup LoadTesting
# @BRIEF Validate load profile against authoritative dashboard model, environment policy, and 041 blast-radius snapshot.
# @PRE DashboardQueryModel and pinned 041 fingerprint are readable; actor has execute permission.
# @POST Returns validated profile, effective cap, matrix preview, warnings, and gate requirements; no execution dispatched.
# @SIDE_EFFECT Read-only metadata calls.
# @DATA_CONTRACT ProfileInput + DashboardQueryModel + BlastRadiusReport -> ValidatedLoadProfile
# @TEST_EDGE requested concurrency over cap -> clamp notice; unknown variation value -> NEEDS_CONTEXT.
# @TEST_EDGE stale fingerprint -> refresh/reconfirm required.
# @REJECTED Client-supplied cap override -> server policy wins.
# #endregion LoadTesting.Profile.Validate

# #region LoadTesting.Matrix.Expand [C:4] [TYPE Function] [SEMANTICS load-testing,matrix,deterministic]
# @ingroup LoadTesting
# @BRIEF Expand closed variation axes cartesian up to cap, then seeded reservoir-sample.
# @PRE Axes are filters, viewport, role, or time_range; values are normalized and validated.
# @POST Returns stable variation ids, theoretical count, selected count, seed, and total execution estimate.
# @SIDE_EFFECT None.
# @INVARIANT Same inputs and seed yield byte-identical matrix; unordered input axes cannot change output.
# @DATA_CONTRACT NormalizedAxes + Seed + MaxVariations -> VariationMatrix
# @TEST_EDGE unknown axis -> 422; product over cap -> deterministic selected subset.
# #endregion LoadTesting.Matrix.Expand

# #region LoadTesting.RunnerPool [C:5] [TYPE Module] [SEMANTICS load-testing,workers,concurrency]
# @defgroup LoadTesting Bounded dashboard load-testing execution domain.
# @BRIEF Runs one queued LoadRun with async workers and environment-level capacity coordination.
# @PRE Validated profile, approved PROD gate when required, pinned blast-radius fingerprint.
# @POST Worker count never exceeds effective cap; terminal status and partial results are durable; one TaskManager task owns run.
# @SIDE_EFFECT Superset chart-data calls, batched DB writes, progress events, TaskManager task state.
# @INVARIANT Acquire load semaphore before request reaches shared client semaphore; never manually re-acquire shared semaphore.
# @INVARIANT Breaker/stop prevents new queue intake; in-flight requests drain within deadline.
# @DATA_CONTRACT ValidatedLoadProfile -> LoadRun + LoadExecution[] + LoadRunAggregate
# @RELATION CALLS -> [SupersetClient.ChartData.Execute]
# @RELATION DEPENDS_ON -> [Core.Manager.CreateTask]
# @RELATION DEPENDS_ON -> [Core.ClientRegistry.GetSemaphore]
# @RATIONALE Async workers directly model executions; HTTP pool alone hides queue pressure.
# @REJECTED One TaskManager task per execution -> task/event flood. Thread/process pool -> unnecessary for I/O-dominant v1.
# #endregion LoadTesting.RunnerPool

# #region LoadTesting.RunnerPool.ExecuteWorker [C:5] [TYPE Function] [SEMANTICS load-testing,worker,timing]
# @ingroup LoadTesting
# @BRIEF Take queue items, record queue/resource/upstream timing, execute chart data, and persist bounded result metadata.
# @PRE Worker owns a queue item and breaker is not tripped; shared request path is available.
# @POST One LoadExecution is produced with disjoint timings, cache metadata, digest/count/sample, or typed error.
# @SIDE_EFFECT Network I/O; batch persistence; progress aggregation.
# @RATIONALE One worker owns one queue item at a time so queue wait, resource wait, and upstream latency remain disjoint and measurable.
# @REJECTED Full 037 normalization inside the worker — rejected because large table payloads can hold the GIL and distort measured Superset latency; load path uses bounded digest/count/sample processing.
# @TEST_EDGE cancellation before take -> item remains unexecuted; timeout -> error taxonomy + no synthetic percentile sample.
# @TEST_EDGE large table -> processing respects 100-row/256KiB bound and does not call full 037 normalization.
# #endregion LoadTesting.RunnerPool.ExecuteWorker

# #region LoadTesting.Capacity.Resolve [C:4] [TYPE Function] [SEMANTICS load-testing,capacity,environment]
# @ingroup LoadTesting
# @BRIEF Resolve effective cap from stage, profile, per-env override, absolute ceiling, pool size, and reserved slots.
# @PRE Environment has valid stage and connection_pool_size >= 1.
# @POST Returns cap and clamp reasons; cap >= 1 or rejects an environment with no safe load slot.
# @SIDE_EFFECT None.
# @INVARIANT PROD default=5, DEV/PREPROD default=10, absolute ceiling=25; reserve 5 shared slots.
# @TEST_EDGE pool_size=5 -> load cap rejected/disabled rather than starving ordinary traffic.
# #endregion LoadTesting.Capacity.Resolve

# #region LoadTesting.Cache.MapResponse [C:3] [TYPE Function] [SEMANTICS load-testing,cache,superset]
# @ingroup LoadTesting
# @BRIEF Map Superset chart-data response metadata to authoritative cache state.
# @PRE Response result includes raw fields when supported; request force flag is known.
# @POST Returns hit/miss/bypassed/disabled/unknown plus preserved raw cache metadata.
# @SIDE_EFFECT None.
# @INVARIANT Latency is never used to infer cache state.
# @TEST_EDGE is_cached=null + force=false + cached_dttm=null -> miss; force=true -> bypassed; absent fields -> unknown.
# #endregion LoadTesting.Cache.MapResponse

# #region LoadTesting.Result.BoundedProcess [C:4] [TYPE Function] [SEMANTICS load-testing,result,bounded,gil]
# @ingroup LoadTesting
# @BRIEF Hash raw response, count rows, and retain bounded sample without full verification normalization.
# @PRE Raw response available; sample limits configured.
# @POST Digest represents raw payload; sample ≤100 rows and ≤256KiB; no full 037 normalization invoked.
# @SIDE_EFFECT CPU work on event loop bounded by payload policy.
# @INVARIANT ProcessPool fallback requires explicit escalation after >10% measured distortion.
# #endregion LoadTesting.Result.BoundedProcess

# #region LoadTesting.Breaker.Evaluate [C:4] [TYPE Function] [SEMANTICS load-testing,circuit-breaker,metrics]
# @ingroup LoadTesting
# @BRIEF Evaluate rolling completed window for error rate and upstream p99 thresholds.
# @PRE At least minimum sample window or breaker remains armed without latency decision.
# @POST Tripped result closes intake and enters drain; trigger metric and threshold are durable.
# @SIDE_EFFECT Updates run circuit state and emits terminal progress.
# @TEST_EDGE <20 samples -> no p99 abort; error threshold breach -> abort within one evaluation window.
# #endregion LoadTesting.Breaker.Evaluate

# #region LoadTesting.Api.Routes [C:4] [TYPE Module] [SEMANTICS load-testing,api,rbac]
# @ingroup LoadTesting
# @BRIEF REST routes for profile validation, matrix preview, start/stop, status, results, and comparison.
# @LAYER API
# @INVARIANT DEV/PREPROD needs dashboard:loadtest:execute; PROD additionally needs dashboard:loadtest:prod + 036 gate.
# @DATA_CONTRACT DTOs are extra-forbid; no raw SQL, endpoint, arbitrary query context, or repository path.
# @RELATION BINDS_TO -> [EXT:frontend:frontend/src/lib/types/load-testing.ts]
# #endregion LoadTesting.Api.Routes

Frontend Screen Models

// #region Dashboards.LoadProfileModel [C:4] [TYPE Model] [SEMANTICS load-testing,profile,ux]
// @BRIEF Model-first editor for profile validation, variation matrix, cap notices, and blast-radius snapshot.
// @STATE idle|validating|matrix_ready|gate_required|start_error
// @ACTION validate(), expandMatrix(), start()
// @UX_STATE invalid variation -> inline NEEDS_CONTEXT; stale fingerprint -> refresh/reconfirm.
// @UX_RECOVERY Fix/drop variation, refresh 041 snapshot, retry start with same revision.
// @RELATION BINDS_TO -> [LoadTesting.Api.Routes]
// #endregion Dashboards.LoadProfileModel

// #region Dashboards.LoadRunModel [C:5] [TYPE Model] [SEMANTICS load-testing,run,progress]
// @BRIEF Recoverable live run state keyed by load_run_id: phase, workers, queue, metrics, findings, partial results.
// @STATE queued|ramping|steady|draining|completed|stopped_by_user|circuit_breaker_abort|failed
// @ACTION stop(), reconnect(), compare()
// @UX_FEEDBACK Worker busy/idle, queue depth, breaker state, p50/p95/p99 upstream latency, cache counts.
// @RATIONALE A dedicated run model keeps high-volume worker progress and partial results outside ordinary AgentRun/chat state while retaining reconnect semantics.
// @REJECTED Reusing AgentRunModel or deriving status from chat events — rejected because load executions are non-conversational and can outlive the browser stream.
// @UX_RECOVERY Browser disconnect reloads authoritative snapshot; abort preserves partial results.
// @INVARIANT Ordinary chat and 036 AgentRun state never initialize this model.
// @RELATION BINDS_TO -> [LoadTesting.Api.Routes]
// #endregion Dashboards.LoadRunModel

#endregion DashboardLoadTesting.Modules


QUICKSTART — Dev Onboarding

Source: quickstart.md

#region DashboardLoadTesting.Quickstart [C:2] [TYPE ADR] [SEMANTICS quickstart,load-testing,verification] @BRIEF Falsifiable verification sequence for 040 without implementing application code.

Feature: 040-dashboard-load-testing | Date: 2026-07-22

Fixture Setup

  1. Dashboard fixture: 2 charts, one scalar and one bounded table; 2 filter values; 2 viewport values; 2 repeated identical coordinates.
  2. Environment fixtures: DEV pool=20, PREPROD pool=20, PROD pool=10; one environment pool=5 edge.
  3. Superset response fixtures: is_cached=true, ordinary source (is_cached=null), force=true, cache_timeout=-1, absent/inconsistent metadata, error taxonomy.
  4. 041 fixtures: shared dataset with 3 dependent dashboards and pinned fingerprint.

Verification Sequence

  1. Profile and matrix: validate profile → effective cap, exact request estimate, deterministic matrix; shuffled axes yield same bytes; over-cap uses seeded sample.
  2. Capacity: DEV/PREPROD default cap=10, PROD=5, ceiling=25, pool reserve=5; pool=5 rejects load slot; two active runs share env capacity and total workers never exceed cap.
  3. Worker timing: inject queue delay, semaphore delay, and upstream delay; assert queue_wait_ms, resource_wait_ms, upstream_latency_ms are disjoint and upstream p99 excludes waits.
  4. Cache mapping: assert body fixtures map hit/miss/bypassed/disabled/unknown exactly; latency changes alone never change state.
  5. Bounded processing: 10,000-row response → digest over raw bytes, row count, sample ≤100 rows/256KiB; full 037 normalization is not called.
  6. Stop/drain: stop run → no new queue items, in-flight drains, status stopped_by_user, partial executions queryable.
  7. Circuit breaker: after ≥20 completed samples, error rate >25% or upstream p99 > configured threshold → intake closes, drain starts, circuit_breaker_abort with trigger details.
  8. PROD gate: unapproved/changed profile or 041 fingerprint dispatches zero requests; approved hash starts exactly one TaskManager task.
  9. Baseline isolation: catalog hash before/after load run identical; no candidate/source hash/immutability mutation.
  10. Recovery and comparison: disconnect/reconnect by load_run_id restores workers/queue/metrics; partial run compares against complete prior run.

Commands (implementation phase only)

cd backend && source .venv/bin/activate && python -m pytest tests/services/load_testing/ -v
cd backend && python -m ruff check .
cd frontend && npm run test
cd frontend && npm run lint

Runtime Closure Status (2026-08-07)

Ранее шаги 3–7 не выполнялись против реального Superset-потока (run_load_run не вызывал RunnerPool). После closure-задач T075–T079:

  • ✅ run_load_run гоняет execution items через RunnerPool с 037-native executor (services/load_testing/executor.py).
  • ✅ Shared per-env semaphore через client_registry.get_semaphore (один лимит для ordinary + load).
  • ✅ LoadExecution персистятся (реальные latency/cache/consistency), breaker → circuit_breaker_abort.
  • ✅ Verified: tests/services/load_testing/test_executor_runtime.py (5 тестов) + полный tests/services/load_testing/ (76 passed).

Проверка на живом Superset (или фикстурном 037-executor'е) остаётся обязательной для объявления operational — см. exit gate 6.

#endregion DashboardLoadTesting.Quickstart


TRACEABILITY — Requirements Matrix

Source: traceability.md

#region DashboardLoadTesting.Traceability [C:3] [TYPE ADR] [SEMANTICS traceability,load-testing,requirements] @BRIEF Requirement-to-contract-to-test matrix for 040 load testing.

Requirement Contract Verification
LOAD-FR-001 LoadTesting.Api.Routes + SupersetClient.ChartData.Execute malicious SQL/raw-context schema tests
LOAD-FR-002, LOAD-FR-016, LOAD-FR-017 LoadTesting.Capacity.Resolve, LoadTesting.RunnerPool capacity, worker registry, one TaskManager task, multi-run env cap
LOAD-FR-003 LoadTesting.RunnerPool + baseline read-only boundary catalog hash before/after
LOAD-FR-004 LoadTesting.Breaker.Evaluate rolling window, error/p99 abort, drain
LOAD-FR-005, LOAD-FR-014 LoadTesting.Api.Routes + 036 ApprovalGate PROD deny/approval/hash mutation/RBAC
LOAD-FR-006 LoadTesting.RunnerPool + ConsistencyFinding identical coordinate digest divergence; row-order edge
LOAD-FR-007, LOAD-FR-010 LoadTesting.Matrix.Expand closed axes, NEEDS_CONTEXT, deterministic cartesian/seeded sample
LOAD-FR-008, LOAD-FR-011 LoadTesting.RunnerPool + LoadRunModel reconnect, ramp/steady/drain terminal FSM
LOAD-FR-009, LOAD-FR-013 041 pinned BlastRadiusReport configure/gate/results same fingerprint; schedule overlap
LOAD-FR-012, LOAD-FR-019 LoadTesting.Cache.MapResponse + LoadExecution authoritative response metadata mapping
LOAD-FR-015 LoadRunComparison complete vs partial comparison
LOAD-FR-018 LoadTesting.Result.BoundedProcess raw digest, row/sample cap, no full 037 normalization, GIL fixture

Story Coverage

Story Independent checkpoint
US1 Profile quickstart 1–2; matrix/cap preview and blast-radius warning
US2 Execution quickstart 3, 6, 10; worker cap, drain, recovery
US3 Safety quickstart 7–9; breaker, PROD gate, baseline isolation
US4 Results quickstart 4–5, 7; cache/metrics/consistency
US5 Comparison quickstart 10; cache metadata and partial-run deltas

Cross-Spec Relations

  • Upstream: 036 ApprovalGate, TaskManager run/recovery; 037 QueryModel, Filters, ChartData.Execute, error taxonomy; 041 pinned dependents read model.
  • Downstream: 039 can link LoadRun results from dashboard/release verification surfaces; 041 provides blast-radius snapshot.
  • Repository integration prerequisite: client_registry.get_client() must pass its existing per-env semaphore to AsyncAPIClient before load capacity claims are trusted.

Amendment (041-dataset-lineage-blast-radius, LIN-FR-013)

The 041 pinned blast-radius read model is now available: GET /api/lineage/datasets/{dataset_id}/dependents?env_id=&fingerprint= serves {fingerprint, stale_index, stale_notice, env_scope, dependents} from the lineage index snapshot (R6). 040 load profiles and PROD gates consume this shape; fingerprint mismatch returns stale_notice (re-fetch, never silent drift). Producer: Api.Lineage.Dependents + Services.Lineage.Indexer.

Amendment (2026-08-07 MVP audit — Phase 9 runtime closure)

Факт-чекинг кода выявил, что run_load_run в backend/src/plugins/load_testing.py не вызывает RunnerPool — фазы проходят без реальных Superset-запросов. Добавлены задачи Phase 9 (T075–T079):

Requirement Contract Tasks Test
LOAD-FR-001, LOAD-FR-002 LoadTesting.RunnerPool + SupersetClient.ChartData.Execute T075, T076 lifecycle asserts executor invoked per item with real chart-data context
LOAD-FR-002 (shared admission) LoadTesting.Capacity.Resolve + client_registry T077 concurrent ordinary + load request never exceed env cap
LOAD-FR-004 LoadTesting.Breaker.Evaluate T078 breaker abort mid-drain preserves partial LoadExecutions
LOAD-FR-008, LOAD-FR-012, LOAD-FR-015 LoadTesting.Result.BoundedProcess + LoadExecution store T079 /status & /compare return real latency/cache/consistency data

#endregion DashboardLoadTesting.Traceability


TASKS — Implementation Tasks

Source: tasks.md

#region DashboardLoadTesting.Tasks [C:3] [TYPE ADR] [SEMANTICS tasks,load-testing,implementation] @BRIEF Ordered TDD backlog for dashboard load testing with bounded async workers, deterministic variations, Superset cache metadata, and 041 blast-radius snapshots.

Input: all documents in specs/040-dashboard-load-testing/ Prerequisites: clarified spec, plan, research R1–R10, data-model, contracts/modules.md, quickstart, 041 LIN-FR-013 read contract

Phase 1 — Setup: Fixtures, DTOs, Permissions

  • T001 Create canonical dashboard/chart/query response fixtures under specs/040-dashboard-load-testing/fixtures/: scalar chart, bounded table chart, two filters, two viewports, repeated identical coordinates, Superset cache hit/miss/bypassed/disabled/unknown responses, and 041 shared-dataset blast-radius snapshot.
  • T002 [P] Create canonical environment policy fixtures under specs/040-dashboard-load-testing/fixtures/environments/: DEV/PREPROD pool=20, PROD pool=10, pool=5 edge, stage/is_production combinations, per-env override, and reserve slots.
  • T003 Materialize canonical fixtures into backend/tests/fixtures/load_testing/ and frontend/src/lib/models/fixtures/load-testing/.
  • T004 [P] Define extra-forbid Pydantic DTOs in backend/src/schemas/load_testing.py for profiles, matrix preview, runs, executions, aggregates, findings, cache metadata, and typed errors.
  • T005 [P] Define matching TypeScript DTOs in frontend/src/lib/types/load-testing.ts with cross-stack field parity and no SQL/raw endpoint/query_context fields.
  • T006 [P] Register dashboard:loadtest:execute and dashboard:loadtest:prod in backend/src/services/rbac_permission_catalog.py; add default-deny role mappings and permission fixtures.
  • T007 Add load testing navigation/action labels and state copy to frontend/src/lib/i18n/locales/ru/load-testing.json and frontend/src/lib/i18n/locales/en/load-testing.json; register locale resources without deriving state from localized strings.

Phase 2 — Foundational: Persistence, Client Capacity, Task Boundary

  • T008 Write failing migration/model tests in backend/tests/models/test_load_testing.py for LoadProfile revisioning, immutable LoadVariation coordinates, terminal LoadRun status, LoadExecution provenance, aggregate uniqueness, and ConsistencyFinding linkage.
  • T009 Add ORM models in backend/src/models/load_testing.py and Alembic migration under backend/alembic/versions/ for profile/run/variation/execution/aggregate/finding records and indexes.
  • T010 Write failing client-registry tests in backend/tests/core/test_client_registry_load_capacity.py proving Environment.connection_pool_size semaphore is passed into AsyncAPIClient, shared ordinary/load calls respect the same semaphore, and shutdown closes the shared client.
  • T011 Fix backend/src/core/utils/client_registry.py to pass the existing per-environment semaphore into AsyncAPIClient; preserve shared auth/CSRF client behavior and do not create a second client/pool. @PRE Registry client is initialized for the environment. @POST AsyncAPIClient acquires/releases the shared semaphore around every request. @INVARIANT No double-acquire path is introduced in the load runner. @TEST_EDGE pool slot exhaustion waits then releases; shutdown leaves no live client.
  • T012 Write failing TaskManager boundary tests in backend/tests/services/load_testing/test_task_boundary.py proving one LoadRun creates exactly one TaskManager task, executions are not tasks, cancellation reaches the run, and task events remain aggregate-level.
  • T013 Add load testing plugin/task entrypoint in backend/src/plugins/load_testing.py; create one observable TaskManager task per LoadRun and connect lifecycle cancellation/status persistence.
  • T014 [P] Add belief-runtime test helpers and instrumentation fixtures under backend/tests/services/load_testing/conftest.py for belief_scope, logger.reason, logger.reflect, and terminal event coverage.

Checkpoint: Models migrate; shared client semaphore is active; one run maps to one TaskManager task; no user story execution yet.

Phase 3 — User Story 1: Configure Load Profile and Variations (P1)

Goal: Validate profile inputs, resolve effective capacity, expand deterministic variations, and preview 041 blast radius without dispatching requests. Independent Test: Fixture profile returns exact matrix/request estimate, cap notice, and dependent-dashboard snapshot; invalid axes/values are rejected with recovery markers.

Tests First

  • T015 [P] [US1] Write failing profile validation tests in backend/tests/services/load_testing/test_profile.py for closed axes, typed filters, required iterations/duration, circuit-breaker bounds, max variation cap, and no raw SQL/query context.
  • T016 [P] [US1] Write failing capacity tests in backend/tests/services/load_testing/test_capacity.py for DEV/PREPROD=10, PROD=5, ceiling=25, reserve=5, per-env override clamping, pool=5 rejection, and multi-run environment sharing.
  • T017 [P] [US1] Write failing matrix tests in backend/tests/services/load_testing/test_matrix.py for canonical axis ordering, stable variation ids, cartesian expansion, seeded reservoir sampling, theoretical vs selected counts, and unknown-axis rejection.
  • T018 [P] [US1] Write failing 041 integration contract tests in backend/tests/services/load_testing/test_blast_radius_contract.py for fingerprint pinning, dependent-dashboard projection, stale snapshot warning, and changed fingerprint requiring refresh.
  • T019 [P] [US1] Write failing L1 model tests in frontend/src/lib/models/tests/Dashboards.LoadProfileModel.test.ts for idle/validating/matrix_ready/gate_required/start_error states, cap notices, NEEDS_CONTEXT, and pinned blast-radius snapshot.
  • T020 [P] [US1] Write failing profile editor L2 tests in frontend/src/lib/components/dashboards/load-testing/tests/LoadProfileEditor.ux.test.ts for accessible controls, variation errors, matrix preview, cap clamp, and dependent-dashboard warning.

Implementation

  • T021 [US1] Implement backend/src/services/load_testing/capacity.py. @PRE Environment stage and pool configuration are valid. @POST Returns effective cap and clamp/rejection reasons; never exceeds profile, environment, reserve-adjusted pool, or ceiling. @INVARIANT PROD=5, DEV/PREPROD=10, ceiling=25, reserve=5.
  • T022 [US1] Implement backend/src/services/load_testing/matrix.py. @PRE Axes are closed and normalized. @POST Returns byte-deterministic variation matrix and request estimate; over-cap expansion is seeded and bounded. @TEST_EDGE shuffled input ordering produces identical bytes.
  • T023 [US1] Implement backend/src/services/load_testing/profile.py to validate profile against 037 DashboardQueryModel and 041 pinned dependents response; no request dispatch.
  • T024 [US1] Implement backend/src/api/routes/load_testing.py profile validation and matrix preview endpoints with RBAC and typed errors.
  • T025 [US1] Implement frontend/src/lib/models/Dashboards.LoadProfileModel.svelte.ts and frontend/src/lib/components/dashboards/load-testing/LoadProfileEditor.svelte bound to typed API DTOs.
  • T026 [US1] Add dashboard entry action and load-testing route integration in frontend/src/routes/dashboards/[id]/components/DashboardHeader.svelte and frontend/src/routes/load-testing/.

Checkpoint: Profile preview works independently; zero Superset execution requests are dispatched during configuration.

Phase 4 — User Story 2: Controlled Parallel Execution (P1)

Goal: Execute one validated matrix through bounded workers with clean ramp, steady, stop, drain, recovery, and aggregate progress. Independent Test: Concurrency=5 fixture never exceeds cap; queue/resource/upstream timings are disjoint; stop drains and preserves partial results.

Tests First

  • T027 [P] [US2] Write failing worker-pool tests in backend/tests/services/load_testing/test_runner_pool.py for worker cap, per-run FIFO, shared env capacity across runs, one worker registry, and no execution TaskManager tasks.
  • T028 [P] [US2] Write failing timing tests in backend/tests/services/load_testing/test_timing.py injecting queue/resource/upstream delays; assert separate timing fields and upstream-only percentile inputs.
  • T029 [P] [US2] Write failing lifecycle tests in backend/tests/services/load_testing/test_run_lifecycle.py for ramp steps, steady state, user stop, bounded drain, browser-independent continuation, terminal immutability, and partial result durability.
  • T030 [P] [US2] Write failing persistence batch tests in backend/tests/services/load_testing/test_persistence.py for ≤50/1s flush thresholds, aggregate updates, reconnect snapshot, and no full response blob storage.
  • T031 [P] [US2] Write failing L1 model tests in frontend/src/lib/models/tests/Dashboards.LoadRunModel.test.ts for queued/ramping/steady/draining/terminal states, worker registry, queue depth, reconnect, and partial results.
  • T032 [P] [US2] Write failing L2 monitor tests in frontend/src/lib/components/dashboards/load-testing/tests/LoadRunMonitor.ux.test.ts for worker status, phase badges, stop action, drain feedback, error counts, and keyboard accessibility.

Implementation

  • T033 [US2] Implement backend/src/services/load_testing/timing.py with monotonic queue/resource/upstream clocks and nearest-rank percentile aggregation.
  • T034 [US2] Implement backend/src/services/load_testing/runner_pool.py. @PRE Validated profile, gate if required, pinned 041 fingerprint, effective capacity. @POST In-flight workers never exceed cap; one LoadRun terminal state and durable partial results. @SIDE_EFFECT Superset chart-data calls, batched DB writes, aggregate progress events. @INVARIANT Acquire load capacity before entering shared client request; breaker/stop prevents new queue intake. @TEST_EDGE cancellation before take leaves item unexecuted; in-flight timeout drains without new dispatch.
  • T035 [US2] Implement backend/src/services/load_testing/persistence.py with bounded execution batches and aggregate snapshots.
  • T036 [US2] Implement backend/src/api/routes/load_testing.py start/status/stop/results endpoints and reconnect by load_run_id.
  • T037 [US2] Implement frontend/src/lib/models/Dashboards.LoadRunModel.svelte.ts and frontend/src/lib/components/dashboards/load-testing/LoadRunMonitor.svelte.
  • T038 [US2] Add WebSocket or existing task-event subscription binding for aggregate progress at ≤4 Hz; never emit one UI event per execution.

Checkpoint: Worker run survives browser disconnect, stops cleanly, and reports partial results without exceeding cap.

Phase 5 — User Story 3: Circuit Breaker and PROD Safety Gate (P1)

Goal: Abort degraded runs safely and prevent unapproved PROD dispatch. Independent Test: Fault injection trips breaker after minimum window; denied or stale PROD gate dispatches zero requests.

Tests First

  • T039 [P] [US3] Write failing breaker tests in backend/tests/services/load_testing/test_breaker.py for minimum 20 samples, 25% error threshold, p99 multiplier, absolute p99 fallback, trigger metadata, drain, and preserved partial results.
  • T040 [P] [US3] Write failing gate/RBAC tests in backend/tests/services/load_testing/test_gate.py for DEV/PREPROD permission, PROD permission, reason requirement, profile/cap/fingerprint hash binding, denial, expiry, replay, and payload mutation.
  • T041 [P] [US3] Write failing frontend gate tests in frontend/src/lib/components/dashboards/load-testing/tests/LoadApprovalGate.ux.test.ts for impact summary, reason, deny, stale fingerprint, and no confirmation control on permission denial.

Implementation

  • T042 [US3] Implement backend/src/services/load_testing/breaker.py. @PRE Completed execution window has policy and required sample count. @POST Tripped breaker closes intake, starts drain, records metric/threshold, and preserves partial results. @TEST_EDGE fewer than 20 samples never triggers p99 abort.
  • T043 [US3] Implement backend/src/services/load_testing/gates.py using 036 ApprovalGate semantics; bind profile revision, effective cap, request estimate, environment, and 041 fingerprint.
  • T044 [US3] Add PROD gate and permission guards to backend/src/api/routes/load_testing.py; revalidate hash immediately before dispatch.
  • T045 [US3] Add breaker badge, approval card, and recovery states to frontend/src/lib/components/dashboards/load-testing/.
  • T046 [US3] Add rejected-path regression tests proving unapproved PROD, stale gate, and arbitrary-SQL payloads dispatch zero requests.

Checkpoint: Circuit breaker and PROD safety are independently verifiable; no unauthorized request reaches Superset.

Phase 6 — User Story 4: Results, Cache, and Consistency (P2)

Goal: Report upstream latency, typed errors, authoritative cache state, bounded result metadata, and consistency findings without baseline mutation. Independent Test: Fixture run returns percentiles, cache states, divergent digest finding, and unchanged baseline catalog hash.

Tests First

  • T047 [P] [US4] Write failing cache mapping tests in backend/tests/services/load_testing/test_cache_metadata.py for Superset body hit/miss/bypassed/disabled/unknown and inconsistent metadata; latency changes must not affect mapping.
  • T048 [P] [US4] Write failing bounded-result tests in backend/tests/services/load_testing/test_bounded_result.py for raw SHA-256, 10k-row response, 100-row/256KiB cap, row count, and no call to full 037 normalization.
  • T049 [P] [US4] Write failing consistency tests in backend/tests/services/load_testing/test_consistency.py for identical coordinate digest divergence, row-order-only allowance, and finding classification=flakiness.
  • T050 [P] [US4] Write failing baseline-isolation tests in backend/tests/services/load_testing/test_baseline_isolation.py asserting catalog hash, source_response_hash, candidates, and immutability status are unchanged.
  • T051 [P] [US4] Write failing results UI tests in frontend/src/lib/components/dashboards/load-testing/tests/LoadResults.ux.test.ts for p50/p90/p95/p99, cache counts, errors, consistency finding detail, and bounded samples.

Implementation

  • T052 [US4] Implement backend/src/services/load_testing/cache_metadata.py using Superset response-body fields: is_cached, cache_key, cached_dttm, queried_dttm, cache_timeout. @POST Maps to hit/miss/bypassed/disabled/unknown with cache_state_source="superset_response". @INVARIANT Latency inference forbidden.
  • T053 [US4] Implement backend/src/services/load_testing/bounded_result.py. @POST Raw digest, row count, sample ≤100 rows/256KiB; full 037 normalization not invoked. @TEST_EDGE GIL/event-loop distortion >10% requires escalation, not silent ProcessPool introduction.
  • T054 [US4] Implement backend/src/services/load_testing/aggregates.py and consistency finding persistence; use digest-based comparison and separate error rates.
  • T055 [US4] Implement frontend/src/lib/components/dashboards/load-testing/LoadResults.svelte and LoadComparison.svelte.
  • T056 [US4] Add baseline read-only boundary assertions to backend/src/services/load_testing/ and expose explicit “baselines not modified” result state.

Checkpoint: Results and cache provenance are queryable; baseline catalog remains byte-identical.

Phase 7 — User Story 5: Blast-Radius Comparison and Scheduling (P2)

Goal: Compare runs and expose shared-dataset impact with the 041 pinned read model; warn about overlap windows. Independent Test: Two same-profile runs plus dependent-dashboard probe show latency deltas, cache metadata, and probe coverage bound to fingerprints.

Tests First

  • T057 [P] [US5] Write failing comparison tests in backend/tests/services/load_testing/test_comparison.py for complete/partial runs, per-chart deltas, cache-state attribution, and consistency-finding deltas.
  • T058 [P] [US5] Write failing blast-radius scheduling tests in backend/tests/services/load_testing/test_schedule_policy.py for dependent deployment/maintenance overlap, fingerprint pinning, and warning/block policy.
  • T059 [P] [US5] Write failing comparison UI tests in frontend/src/lib/components/dashboards/load-testing/tests/LoadComparison.ux.test.ts for run selector, latency deltas, cache annotations, dependent dashboards, and probe coverage.

Implementation

  • T060 [US5] Implement backend/src/services/load_testing/comparison.py with partial-run-compatible aggregates and cache-state provenance.
  • T061 [US5] Implement backend/src/services/load_testing/schedule_policy.py; query 041 dependent dashboard windows and block/warning conflicts per policy.
  • T062 [US5] Add comparison and blast-radius endpoints to backend/src/api/routes/load_testing.py.
  • T063 [US5] Complete frontend/src/lib/components/dashboards/load-testing/LoadComparison.svelte and blast-radius result panels.

Checkpoint: Same-profile comparison works with partial current runs; blast-radius scope remains pinned and visible.

Phase 8 — Integration, Accessibility, and Quality Gates

  • T064 Add frontend/src/lib/models/tests/Dashboards.LoadModels.integration.test.ts covering profile → gate → run → reconnect → results.
  • T065 Add frontend/e2e/tests/dashboard-load-testing.e2e.js covering dashboard entry → matrix preview → approval/deny → worker monitor → stop/reconnect → results.
  • T066 Add backend/tests/integration/test_dashboard_load_testing_superset.py using Superset chart-data fixtures or Testcontainers; verify response cache metadata is preserved.
  • T067 Add backend/tests/integration/test_load_testing_client_capacity.py proving shared semaphore wiring, reserve slots, multiple runs, and ordinary request fairness.
  • T068 Add frontend responsive and accessibility tests for 1366×768 and narrow viewport profile/monitor/results screens.
  • T069 Run quickstart.md steps 1–10 and record results.
  • T070 Run backend load-testing tests, existing client-registry/migration/task regressions, and python -m ruff check ..
  • T071 Run frontend tests, lint, build, and Playwright.
  • T072 [P] Audit rejected paths: direct SQL, HTTP-only cap, per-execution TaskManager tasks, latency cache inference, full normalization in load path, baseline mutation, unapproved PROD dispatch.
  • T073 [P] Audit ATTN_1–4, exact anchor pairs, C4/C5 @RATIONALE/@REJECTED, belief-runtime markers, unresolved relations, and 041 fingerprint integration.
  • T074 [P] Update 041/039 traceability notes with LoadRun/FleetReport consumer links after contracts are implemented.

Phase 9 — MVP Gap Closure (audit 2026-08-07)

Context: Задачи выше закрыты, но факт-чекинг кода выявил, что run_load_run в backend/src/plugins/load_testing.py не вызывает RunnerPool — он переводит фазы RAMP→STEADY→DRAIN→COMPLETED без единого реального Superset-запроса. Это нарушает LOAD-FR-001 (обязанность идти через 037 executor) и делает фичу «нагрузка без нагрузки». Ниже — обязательные задачи закрытия.

  • T075 [P] [US2] Wire RunnerPool into run_load_run: in backend/src/plugins/load_testing.py, replace the hold/sleep body with RunnerPool(capacity=run.effective_concurrency, executor=<037 query executor>, env_capacity_semaphore=<shared per-env semaphore>, breaker=<run breaker>, on_result=<persist LoadExecution>) and drain execution_specs through it. @POST: each queued execution produces a LoadExecution row via the 037 executor; in-flight never exceeds cap; phase transitions preserved. @TEST: backend/tests/services/load_testing/test_run_lifecycle.py — assert executor called once per item with real chart-data context. DONE: run_load_run now calls _resolve_execution_scope → _run_bounded_pool (RunnerPool wired). Verified by test_executor_runtime.py::TestRunLoadRunRealExecution::test_run_real_execution_persists_load_executions (2 items → 2 LoadExecution rows, run → COMPLETED).
  • T076 [P] [US2] Add the 037-native executor adapter backend/src/services/load_testing/executor.py: async def execute_superset_chart(env, chart_id, filters_context) -> LoadExecution delegating to src.core.superset_client._chart_data.execute_chart_data + bounded normalization (LOAD-FR-018). @REJECTED: direct HTTP/SQL in the load path — must reuse 037 SupersetClient.ChartData.Execute. @TEST: fixture chart-data response → bounded digest + cache metadata preserved (reuses 037 fixtures). DONE: executor.py created (execute_superset_chart via execute_dashboard_query_envelope, build_execution_items). Tested in test_executor_runtime.py (delegation + missing chart/dataset raise + deterministic items).
  • T077 [P] [US2] Share the per-environment capacity semaphore with the existing client_registry admission so load runs respect the same pool as ordinary requests (LOAD-FR-002 note in 040 traceability). @TEST: concurrent ordinary request + load run never exceed env cap (extend test_load_testing_client_capacity.py). DONE: _run_bounded_pool acquires client_registry.get_semaphore(env) and passes it as env_capacity_semaphore to RunnerPool — same shared per-env limit as ordinary calls.
  • T078 [P] [US3] Wire the run-level circuit breaker into run_load_run evaluation between executions: abort to circuit_breaker_abort with partial results preserved when thresholds breach (LOAD-FR-004). @TEST: fault-injection run trips breaker mid-drain; partial LoadExecutions remain queryable. DONE: _build_breaker returns a CircuitBreaker; _run_bounded_pool records each outcome/latency; run_load_run transitions to CIRCUIT_BREAKER_ABORT when breaker.tripped. Partial executions persist before the terminal transition.
  • T079 [P] [US4] Persist per-execution results from on_result into the LoadExecution store so /status and /compare return real latency/cache/consistency data instead of seeded aggregates. @TEST: completed run exposes per-chart p50/p90 + cache-state; consistency violation detected on divergent hashes. DONE: write_load_executions (batched LoadExecution persist) added to persistence.py; _run_bounded_pool drains results through it. Verified by test_run_real_execution_persists_load_executions.

Dependencies

T001–T007 → T008–T014 foundational → US1 → US2 → US3/US4; US5 follows US4 and 041 read-model availability. Phase 8 follows all stories. Phase 9 (T075–T079) is the MVP runtime closure: it must be merged before 040 can be claimed operational. Tests precede implementation within each story. T011 must pass before claiming any environment capacity result. T047/T052 require the 037 chart-data adapter to preserve cache metadata.

Parallel Opportunities

  • T002 ∥ T004 ∥ T005 ∥ T006 ∥ T007
  • T008 ∥ T010 ∥ T012 ∥ T014
  • T015 ∥ T016 ∥ T017 ∥ T018 ∥ T019 ∥ T020
  • T027 ∥ T028 ∥ T029 ∥ T030 ∥ T031 ∥ T032
  • T039 ∥ T040 ∥ T041
  • T047 ∥ T048 ∥ T049 ∥ T050 ∥ T051
  • T057 ∥ T058 ∥ T059
  • T064 ∥ T066 ∥ T067 ∥ T068

Story Verification Criteria

Story Verification
US1 Matrix/cap/041 preview; no dispatch; deterministic bytes; quickstart 1–2
US2 Worker cap, disjoint timings, stop/drain/reconnect; quickstart 3, 6, 10
US3 Breaker and PROD gate; zero dispatch on deny/stale; quickstart 7–8
US4 Cache body mapping, bounded digest/sample, consistency, baseline hash; quickstart 4–5, 9
US5 Partial comparison, cache attribution, dependent-dashboard probe coverage, schedule overlap

Rejected-Path Coverage

  • Direct SQL/raw query context: T046, T072.
  • HTTP pool as sole concurrency control: T016, T027, T067, T072.
  • One TaskManager task per execution: T012, T027, T072.
  • Latency-based cache inference: T047, T072.
  • Full 037 normalization in load path: T048, T053, T072.
  • Baseline/catalog mutation: T050, T056, T072.
  • Unapproved PROD dispatch: T040, T046, T072.

#endregion DashboardLoadTesting.Tasks


PROTOTYPE — State/Manifest

Source: prototype/manifest.md

#region Std.Opencode.PrototypeManifest [C:3] [TYPE ADR] [SEMANTICS prototype,manifest,load-testing] @defgroup Prototype Interactive HTML prototype manifest for 040-dashboard-load-testing.

Prototype Metadata

  • Feature: 040-dashboard-load-testing
  • Source contracts: ux_reference.md (lightweight prototype — no contracts/ux/ present)
  • Screens represented: 3 (load profile editor, load run monitor, results & comparison)
  • Total states: 17
  • Accessibility: native buttons, focus-visible rings, state-pill live region, touch targets ≥44px
  • Responsive: 3-col editor grid collapses to 1 col on mobile

State Coverage

Screen @UX_STATE Prototype State Reachable Recovery
Editor idle idle ✅ —
Editor validating validating ✅ auto→matrix_ready
Editor matrix_ready matrix_ready ✅ Start enabled
Editor gate_required gate_required ✅ PROD approval
Editor needs_context needs_context ✅ inline fix / drop variation
Editor cap_clamped cap_clamped ✅ notice «Снижено до лимита среды»
Editor permission_denied permission_denied ✅ no confirm control
Editor start_error start_error ✅ retry same profile
Monitor queued queued ✅ —
Monitor ramping ramping ✅ live in-flight
Monitor steady steady ✅ Stop → drain
Monitor draining draining ✅ «завершение in-flight»
Monitor circuit_breaker_abort circuit_breaker_abort ✅ partial results preserved
Monitor stopped_by_user stopped_by_user ✅ partial results
Results ready ready ✅ —
Results consistency consistency ✅ flakiness note

Coverage: 16/16 declared states reachable. Recovery paths: needs_context inline fix, start_error retry, breaker abort → partial results, draining → clean stop.

Screen ↔ Story Traceability

Prototype Screen User Story Acceptance Verified
Load Profile Editor US1 concurrency clamp, matrix preview, blast-radius PROD warning, NEEDS_CONTEXT
Load Run Monitor US2, US3 phase badges, in-flight counter, Stop→drain, breaker abort trigger+partial results
Results & Comparison US4, US5 p50-p99 table, cache annotation, consistency finding (flakiness), baseline-untouched, Δ vs base

Validation Results

  • All @UX_STATE contracts reachable via state switcher (static)
  • Recovery paths traversable (needs_context, start_error, abort, drain)
  • Keyboard nav: native buttons, Tab order
  • Touch targets: ≥44px
  • ARIA: #curState live region
  • Browser-driven visual validation: DONE — Playwright MCP requires the chrome channel (not installed in this env; npx playwright install chrome or CI/dev). Artifact is self-contained and ready for manual browser review.

Design System Reuse

Element Source Prototype Mapping
Card $lib/ui/Card.svelte rounded-lg border border-border bg-surface-card text-text shadow-sm p-6
Primary Button $lib/ui/Button.svelte inline-flex ... rounded-md bg-primary text-white h-10 px-4 text-sm
Destructive Button Button variant bg-destructive text-white
Badge $lib/ui/Badge.svelte rounded-full px-2.5 py-1 text-xs font-medium bg-{variant}-light text-{variant}
Input $lib/ui/Input.svelte border border-border-strong rounded-md p-2 text-sm
Table production pattern divide-y divide-border border border-border rounded-lg

Design Token Audit (MANDATORY)

Token Hex Used
primary.DEFAULT/hover/ring #2563eb/#1d4ed8/#3b82f6 validate/start, active state
destructive #dc2626 Stop, breaker tripped, error
success.light/success #f0fdf4/#16a34a breaker green, baseline-untouched
warning.light/warning #fffbeb/#b45309 NEEDS_CONTEXT, clamp notice, breaker reason, consistency
info.light/info #f0f9ff/#0369a1 phase badges, drain note
surface.page/card/muted #f8fafc/#fff/#f1f5f9 layout, tiles
border.DEFAULT/strong #e2e8f0/#cbd5e1 card/table/input borders
text.DEFAULT/muted #0f172a/#64748b headings/secondary

Audit gate: 100% of hex values in index.html trace to tokens above; zero invented colors/radius/shadows.

#endregion Std.Opencode.PrototypeManifest


PROTOTYPE — Interactive HTML

Source: prototype/index.html

<!doctype html><html lang="ru"><head></head>

Superset Tools · BI testing
СценарииЗапускиАвтоматизацияКачествоControlled environment
Load testing

Нагрузочная проверка dashboard

Проверяйте производительность в выделенной среде, не производственные данные.

Профиль

DashboardFI-0080EnvironmentPREPROD load-sandboxPROD — запрещено policy
Virtual users
Duration
Safety preflight
Read-only profile · capacity allocated · no browser/data mutation.
Предпросмотр матрицыЗапустить нагрузку
State: IdleRunningResult
<script src="../../prototype-ui.js"></script><script>protoState('idle',s=>document.getElementById('monitor').innerHTML=s==='running'?'

Run status

Running · 12/20 VU

p95: 1.8 sec · errors: 0

':s==='result'?'

Run result

PASS

p95: 1.9 sec · baseline: 1.7 sec

':'

Run status

Не запущен

Результат покажет p50/p95/p99, errors и regression against baseline.

')</script></html>


validation.md

Source: validation.md

#region Std.Opencode.ValidationReport [C:3] [TYPE ADR] [SEMANTICS validation,gate,load-testing] @defgroup Validation Pre-implementation validation gate for dashboard load testing (040).

Status: PASS (re-validated 2026-08-07 after MVP audit + runtime closure)

Date: 2026-08-07 Feature: 040-dashboard-load-testing Branch: 040-dashboard-load-testing

Note (2026-08-07): Reconciliation pass (Phase 9 runtime closure T075–T079) amended spec.md, tasks.md, quickstart.md; digests regenerated. Phase 9 runtime closure is complete: run_load_run wires RunnerPool + 037 executor + shared semaphore + breaker; LoadExecution persist real results. Verified by test_executor_runtime.py (5) + tests/services/load_testing/ (76 passed).

Validated Inputs

Artifact Size (bytes) SHA-256
spec.md 21060 faad190fe8f4fdc9d7882d3cf55381d7dc4c324125337ac1b8fecc31c24dc786
plan.md 9845 806d03d59e6fb70204f8ab101670a0ea37333a20157b63d6de2ba4bf4f00c92a
tasks.md 24632 419eda60525f691045faf4aad34d437659b29898058a103d92edd5fe0576258c
traceability.md 4067 ca26e71dfdcc7a90e191b0ac6a8bd0d5103fbce3dd0ec895f5e9d92809a48f3f
contracts/modules.md 9513 eb376a4eeddefdf4982daef439cc0661bdab70f890f0219bc0055067741dfa1f
data-model.md 5911 5039a66acbd3e1dca5f022c2291a1d1a36cd08edcb41c35b0a9d9d37bf3bc12e
research.md 12598 e9c60d3cc19eca9326f1d1b036707fa21fbbf8ee208987671bb74d85affab550
ux_reference.md 7517 087f2847cf9996d26cfbb31a824c5ab6c7a9a6072b7cc7424c63e9f4f2e3ee6c
quickstart.md 3588 8b93d1719cc54f959d6a523a6fb404e2f689118650adc4cf22553b1c27957e40

Verdict is stale if any artifact digest changes or new applicable artifacts appear.

Blocking Findings

✅ No blocking findings for the specified scope. Phase 9 runtime closure T075–T079 complete — real load execution verified; live/fixture Superset soak (exit gate 6) and capacity-with-ordinary-requests integration remain recommended hardening, not blockers.

Check Results

Phase 1: Unresolved Markers

  • [NEEDS CLARIFICATION]/[NEED_CONTEXT]/TODO/TBD: 0 → ✅ PASS

Phase 2: Artifact Completeness

  • spec/plan/tasks/traceability/contracts/data-model/research/ux_reference/quickstart: all present → ✅ PASS

Phase 3: Schema & Contract Validation

  • Region pair balance: 12 opens / 12 closes → ✅ PASS
  • Anchor signatures [C:N] [TYPE] [SEMANTICS] one-line → ✅ PASS
  • Hierarchical IDs (LoadTesting.*) + shared load-testing semantics → ✅ PASS

Phase 4/5: Reference & Decision Memory

  • @REJECTED (AgentRun-as-loadrun, baseline-catalog writes, unbounded concurrency) not scheduled in tasks; rejected-path coverage present → ✅ PASS

Phase 6: Task Dependency & Path

  • Task count: 74; all paths under backend/src, backend/tests, frontend/src, backend/alembic → ✅ PASS
  • Cross-spec: 041 dependents read model now available (LIN-FR-013) → ✅ PASS

Phase 7/8

  • UX surface present (ux_reference.md); Axiom MCP unavailable — manual fallback (W01 same as 041).

Gate Decision

Verdict: ✅ PASS — /speckit.implement may proceed.

#endregion Std.Opencode.ValidationReport

================================================================================ FEATURE: 041-dataset-lineage-blast-radius Files: 13


SPEC — Feature Specification

Source: spec.md

#region DatasetLineageBlastRadius.Spec [C:3] [TYPE ADR] [SEMANTICS spec,requirements,lineage,blast-radius,dataset,dashboard-testing] @BRIEF Dataset lineage and blast-radius analysis: reverse index dataset→dashboards, cross-dashboard schema-change impact classification, staleness propagation across releases, deprecation lifecycle, and verification fan-out. @RELATION DEPENDS_ON -> [Doc.Adr.ADR0001] @RELATION DEPENDS_ON -> [Doc.Adr.ADR0003] @RELATION DEPENDS_ON -> [Doc.Adr.ADR0005] @RELATION DEPENDS_ON -> [SupersetBaselineEngine.Spec] @RELATION DEPENDS_ON -> [AgentTestStabilization.Spec] @RELATION DEPENDS_ON -> [Core.MappingService.IdMappingService] @RATIONALE Specs 036–040 model every artifact per-dashboard: DashboardQueryModel is built for one dashboard, StructureDiff compares two releases of one dashboard, baselines pin to one dashboard's release, VerificationRun triggers fire for one dashboard. But datasets are shared N:M across dashboards — a schema change, deprecation, or data regression in one dataset silently breaks every dependent dashboard, and no current artifact can name them. The reverse index dataset→dashboards is the missing primitive that makes impact analysis computable instead of tribal knowledge. @REJECTED Extending DashboardQueryModel with reverse lookups — rejected because QueryModel is a per-dashboard projection; embedding cross-dashboard lineage would couple every inspection to fleet-wide state and break its deterministic, single-dashboard fingerprint semantics (037 AGBASE-FR-001). @REJECTED Deriving lineage at query time by scanning all dashboards per request — rejected because it is O(fleet) per lookup, non-deterministic under concurrent dashboard edits, and unusable inside hot paths (load-profile preview, PROD gates, release checks). @REJECTED Propagating staleness by auto-invalidating dependent baselines — rejected because a dataset schema change does not always invalidate downstream truth (additive columns are safe); blind invalidation destroys approved baselines and violates the 037 immutability model. Impact must be classified and surfaced, not silently enforced. @RATIONALE Blast-radius explanation (LIN-FR-021) is consumed by external MCP clients after the 050 drift; the index, severity matrix and fan-out contracts are unchanged and remain server-owned.

Navigation (DSA Indexer keywords)

@SEMANTICS: spec, requirements, feature, lineage, blast-radius, dataset, schema-change, staleness-propagation, deprecation, fan-out, dashboard-testing

Feature Branch: 041-dataset-lineage-blast-radius Created: 2026-07-22 | Status: Ready for Implementation (reconciled 2026-08-07; backend implemented, closure tasks T045–T048 open) Input: "Dataset lineage и blast-radius анализ: обратный индекс датасет→дашборды, cross-dashboard structure diff, классификация влияния изменений схемы датасета, распространение staleness между релизами, fan-out триггеры верификации зависимых дашбордов. Закрывает gaps BR-1..BR-5 из анализа спек 036–039."

Clarifications

Session 2026-07-22

  • Q1 (graph model): Blast-radius graph is two-level dataset ← chart ← dashboard — NOT data-lineage of dataset SQL sources. No SQL parsing of dataset.sql; cycles/topological ordering removed (impossible by construction). → LIN-FR-015 rewritten.
  • Q2 (usage extraction): C+2 hybrid — structured chart params extracted per viz_type (exact); custom SQL expressions (sqlExpression ad-hoc metrics/columns/WHERE) parsed via the existing in-repo SQL parser (sqlparse, per research R2 — sqlglot rejected: new dependency, marginal accuracy gain for Superset-generated dialect-simple expressions): column refs matched against dataset columns/metrics → exact edges; unmatched refs → unresolved_refs finding; parse failure → conservative per-expression only (not whole chart). projection_confidence: exact|conservative per chart. → LIN-FR-001 rewritten, LIN-FR-016/017 added.
  • Q3 (severity matrix): Classification rules v1 accepted as defaults — see Severity Matrix v1 section. → LIN-FR-004 amended.
  • Q4 (refresh cadence): No separate lineage cron — index refresh rides the existing Cross-Environment ID Synchronization cycle (IdMappingService.sync_environment, migration_sync_cron + Migration Settings UI + sync-now). A post-sync hook rebuilds edges only for changed uuids (list API lacks params/schema → detail calls bounded to changed set). Opt-in via lineage_index_enabled (default false). Write-hook: our own mutating operations (migration deploy, update_dataset) trigger targeted refresh of affected ids without waiting for cron. Deprecation escalation states are recomputed in the same cycle. → LIN-FR-002 rewritten, LIN-FR-018/019 added.

User Scenarios

Story 1 — Dataset Usage Index (P1)

Why P1: Every blast-radius feature — 040 load panels, PROD gates, deprecation warnings — depends on answering «какие дашборды используют этот датасет?» deterministically and fast.

Independent Test: Build the index from fixture Superset metadata (3 dashboards sharing 2 datasets) and verify reverse lookup returns exact dependent dashboard sets, chart-level usage detail, and a deterministic index fingerprint.

Acceptance:

  1. Given authoritative Superset dashboard/chart/dataset metadata When the index builds Then every dataset maps to its dependent dashboards with per-chart usage detail (which chart, which columns/metrics consumed).
  2. Given a dashboard is added, removed, or its charts re-pointed to another dataset When the index refreshes Then reverse mappings reflect the change and the index fingerprint changes deterministically.
  3. Given metadata is inaccessible (Superset 403/timeout) When refresh runs Then the previous index is retained with a stale_index marker and typed error provenance — never a partial or empty index silently served.

Story 2 — Cross-Dashboard Schema Change Impact (P1)

Why P1: A dataset schema change is a fleet event, not a single-dashboard event; each dependent dashboard must get a classified impact, not a surprise failure at next verification.

Independent Test: Apply fixture schema changes (column removed, column type changed, column added) to a shared dataset and verify each dependent dashboard receives a classified DatasetImpactRecord with severity and affected artifacts.

Acceptance:

  1. Given a dataset's schema changes between two observations When the diff is computed Then changes classify by kind (column_removed, type_changed, column_added, metric_removed, dataset_deleted) and severity (critical, warning, info).
  2. Given a classified change When impact projects to dependent dashboards Then each dashboard record names affected charts, affected baseline entries (by release pin), and affected scenario steps (038 refs) — additive-only changes classify info and never mark baselines stale.
  3. Given a dependent dashboard's 037 StructureDiff exists between its own releases When a dataset change occurs Then the dataset impact record links to (but does not mutate) that dashboard's per-release StructureDiff — cross-dashboard impact and per-dashboard diff stay orthogonal.

Story 3 — Staleness Propagation Across Releases (P1)

Why P1: 037 pins baselines to one dashboard's release; when a dataset changes, dependent dashboards on other releases need their baseline staleness re-evaluated without violating immutability rules.

Independent Test: Change a shared dataset after dashboards A and B approved baselines on different releases; verify affected entries in both catalogs surface stale_baseline (or immutability_violation for closed periods) with dataset-change provenance.

Acceptance:

  1. Given a dataset change classified critical or warning When dependent baseline catalogs load Then affected entries are marked or reported stale with provenance pointing to the dataset impact record — catalog files themselves are not rewritten by the propagation.
  2. Given a closed-period baseline entry whose dataset data changed retroactively When source_response_hash diverges Then propagation surfaces immutability_violation per 037 AGBASE-FR-012 — never downgraded to a stale warning.
  3. Given dashboards pin different release_version for the same dataset When propagation evaluates Then each release's affected entries are reported independently; one release's re-approval never auto-heals another release's staleness.

Story 4 — Dataset Deprecation Lifecycle (P2)

Why P2: Datasets are replaced, not deleted cleanly; dependents need a managed migration path with deadlines and escalation, not a 404 at verification time.

Independent Test: Mark a fixture dataset deprecated with a successor and grace window; verify dependents receive deprecation notices, escalating warnings as the window closes, and a final blocking state at expiry.

Acceptance:

  1. Given a dataset is marked deprecated with successor_dataset_id and grace window When dependents are enumerated Then every dependent dashboard owner-visible surface shows the deprecation, successor, and remaining window.
  2. Given the grace window closes When verification or scenario generation touches the deprecated dataset Then operations warning-gate, then block at expiry with a typed dataset_deprecated error naming the successor.
  3. Given a dependent dashboard migrates its charts to the successor When the index refreshes Then the deprecation record tracks migrated vs remaining dependents until zero remain.

Story 5 — Verification Fan-Out for Dependent Dashboards (P2)

Why P1 value, P2 cost: After a dataset change, affected dashboards need re-verification prioritized by impact severity — a fleet-level trigger the current per-dashboard VerificationRun model cannot express.

Independent Test: Trigger dataset_updated on a fixture dataset with 3 dependents at different severities; verify fan-out creates one coordinated plan, per-dashboard VerificationRuns link to it, and results aggregate into a fleet report.

Acceptance:

  1. Given a dataset_updated trigger with a classified impact record When fan-out executes Then dependent dashboards receive verification runs ordered by impact severity (critical first), each linked to the shared fan-out plan id.
  2. Given a dependent dashboard lacks a current release or approved baselines When fan-out plans Then it is scheduled as inspection-only (query model refresh) rather than comparison, with the reason recorded.
  3. Given fan-out completes When results aggregate Then a fleet report shows per-dashboard status, dataset-change provenance, and unresolved impacts; the report is consumable by 039 pipeline views without an AgentRun.

Edge Cases

  • Dataset used by zero dashboards → index returns empty dependents; deprecation and fan-out complete trivially without errors.
  • Chart SQL expression references a non-existent column → unresolved_refs finding at index build; the chart is not silently treated as exact-confidence.
  • Dataset deleted without deprecation → dataset_deleted critical impact on all dependents; verification fan-out marks comparisons inconclusive, never invents metadata (037 US1 acceptance 3 semantics).
  • Two dataset changes race (schema diff computing while another lands) → impact records are versioned per observation; consumers see a monotonic sequence, never a merged ambiguous diff.
  • Dependent dashboard's environment differs from the dataset-change observation environment → impact records carry environment scope; cross-environment propagation is reported as env_mismatch warning, not assumed valid.
  • Index refresh during an active 040 load run → load run pins the index snapshot it validated against; mid-run index updates do not alter its blast-radius panel retroactively.
  • Sync cycle runs while lineage_index_enabled=false → mappings sync proceeds unchanged; lineage hook is inert (zero Superset calls from the lineage side).
  • Migration deploy re-points charts to another dataset → targeted write-hook refresh fires immediately; blast-radius panels reflect the new binding without waiting for the sync cycle.

Requirements

Functional

  • LIN-FR-001: The system MUST maintain a dataset→dashboard reverse index built from authoritative Superset metadata (dataset ← chart ← dashboard bindings, where chart→dataset is strictly 1:1 per the Superset Slice datasource_id/datasource_type), with per-chart usage detail: consumed columns/metrics extracted from structured chart params per viz_type (exact), PLUS dataset metrics referenced by the chart (resolved by metric_name or verbose_name; each referenced metric contributes its expression column-deps to the consumed set), PLUS SQL-expression analysis for ad-hoc metrics/columns. Every chart edge carries projection_confidence: exact|conservative. Index fingerprint MUST be deterministic.
  • LIN-FR-016: Custom SQL expressions (sqlExpression in ad-hoc metrics/columns/WHERE clauses) MUST be parsed with the existing in-repo SQL parser (sqlparse-based, per research R2 — a dedicated sqlglot dependency is rejected): matched column/metric refs produce exact edges; refs matching no dataset column/metric produce an unresolved_refs finding per chart (possible pre-existing bug); parse failure marks only that expression conservative, never the whole chart. Regex-only extraction is forbidden. exact confidence is asserted ONLY for refs resolved against authoritative dataset columns/metrics — the parser is token-level (sqlparse is intentionally non-validating), so dialect/SQL AST guarantees are explicitly out of scope.
  • LIN-FR-002: The index MUST be incrementally refreshable and content-addressed; consumers MUST be able to pin a snapshot (index fingerprint) for the duration of an operation (load run, release check, fan-out plan). Snapshot pinning is an optimistic consistency token, not a historical edge-set store: one current snapshot per environment (LineageIndexSnapshot.environment_id PK); on refresh the edge set is rebuilt delete+insert and the fingerprint flips atomically. A pinned-fingerprint mismatch yields stale_notice (re-fetch, never silent drift) — restoring the prior edge set from the fingerprint is NOT supported. Index refresh MUST ride the existing Cross-Environment ID Synchronization cycle (IdMappingService.sync_environment with its migration_sync_cron, incremental mode, and stale-deletion) — a dedicated lineage scheduler is forbidden. A post-sync hook rebuilds edges only for changed uuids; list-API payloads lack chart params and dataset schema, so detail calls MUST be bounded to the changed set. The feature is opt-in via lineage_index_enabled setting (default false).
  • LIN-FR-018: Own mutating operations (migration deploy, dataset create/update via SupersetClient) MUST trigger a targeted refresh of the affected dataset/chart/dashboard ids without waiting for the next sync cycle; external edits made directly in Superset UI are picked up by the incremental sync cycle.
  • LIN-FR-019: Deprecation escalation states (noticed → warning → expired_blocked) MUST be recomputed within the same sync cycle from the grace window — no separate deprecation timer.
  • LIN-FR-020: The system MUST track dataset recreation during migration. A recreated dataset has a new id/uuid but the same physical identity (database_id, schema, table_name); when a new dataset_uuid matches the dataset_physical_key of a prior uuid that is now gone, the index MUST emit a candidate dataset_recreated finding (old→new) — conservative, operator-confirmed, never auto-applied. On confirmation the new uuid inherits the blast-radius context (dependents, deprecation successor) and stale edges of the old uuid are marked orphaned/stale until reconciled. A live prior uuid on the same physical key is a legitimate second dataset and MUST NOT produce a finding.
  • LIN-FR-003: On refresh failure the system MUST retain the last good index marked stale_index with typed error provenance; serving an empty or partial index as current is forbidden.
  • LIN-FR-004: The system MUST compute dataset schema diffs between observations, classifying changes by kind and severity per Severity Matrix v1 (deterministic, versioned; rule changes require matrix version bump and new golden fixtures).

Severity Matrix v1 (accepted 2026-07-22)

Change Severity Rationale
dataset_deleted critical All dependents break
column_removed (consumed by exact-confidence charts) critical Queries fail
column_removed (not consumed / conservative-only) warning May hide in SQL expressions
metric_removed critical Metrics are always consumed by reference
metric.expression changed (name identical, formula different) consumption-aware: exact-consumed → warning/critical, conservative → warning Formula change breaks consumers even with no column change
type_changed breaking (varchar→int, date→varchar, any→bool) critical Cast errors / semantics change
type_changed narrowing (bigint→int, decimal(12,2)→(10,2), varchar(200)→(50)) warning May truncate data
type_changed widening (int→bigint, decimal(10,2)→(12,2), varchar(50)→(200)) info Backward compatible; may shift 037 source_response_hash
column_added info Additive; MUST NOT mark baselines stale (LIN-FR-005)
column_renamed (physical: removed+added, same type, label preserved) critical if exact-consumer of old name, else warning + rename-hint Consumers of the old physical name break; preserved verbose_name is high-confidence signal
column_relabeled (verbose_name changed only, column_name unchanged) info (invisible to schema_hash) Cosmetic; charts reference column_name, so no blast radius, never marks baselines stale
nullability NOT NULL → nullable info Relaxation
nullability nullable → NOT NULL warning May alter filtering of null rows
column_renamed (label lost, positional proximity only) warning + rename-hint Lower-confidence rename; human confirms — never auto-applied
  • LIN-FR-005: For each classified dataset change the system MUST project a DatasetImpactRecord per dependent dashboard naming affected charts, affected baseline entries (by release pin), and affected 038 scenario step refs — additive-only changes MUST classify info and MUST NOT mark baselines stale.
  • LIN-FR-006: Cross-dashboard impact records MUST remain orthogonal to per-dashboard 037 StructureDiff; impact records link to but never mutate StructureDiff, baseline catalogs, or scenario graphs.
  • LIN-FR-007: Staleness propagation MUST surface affected baseline entries in dependent catalogs as stale (or immutability_violation for closed periods per 037 AGBASE-FR-012) with dataset-change provenance; propagation MUST NOT rewrite catalog files, auto-approve, or auto-update expected values.
  • LIN-FR-008: Propagation MUST evaluate each pinned release independently; re-approval in one release MUST NOT clear staleness in another release sharing the dataset.
  • LIN-FR-009: The system MUST support dataset deprecation with successor_dataset_id, grace window, escalating warning states, and a terminal blocking state at expiry producing typed dataset_deprecated errors; migration progress (migrated vs remaining dependents) MUST be tracked per deprecation record.
  • LIN-FR-010: A dataset_updated trigger MUST create one coordinated fan-out plan; per-dashboard verification runs link to the plan id, are ordered by impact severity, and dashboards without current baselines are scheduled inspection-only with recorded reason.
  • LIN-FR-011: Fan-out results MUST aggregate into a fleet report (per-dashboard status, provenance, unresolved impacts) consumable by 039 pipeline views without an AgentRun; the trigger value dataset_updated MUST be added to the 036/037 VerificationRun trigger enum.
  • LIN-FR-012: All impact, deprecation, and fan-out records MUST carry environment scope; cross-environment propagation MUST surface env_mismatch warnings rather than assume validity.
  • LIN-FR-013: The blast-radius read model exposed to 040 load profiles and PROD gates MUST be served from this index (pinned snapshot), guaranteeing configure-time and gate-time consumers see identical dependents.
  • LIN-FR-014: RBAC MUST gate mutations: dataset:lineage:refresh (index rebuild), dataset:deprecation:manage (mark/migrate), dataset:fanout:trigger (PROD-classified fan-out requires 036 approval-gate semantics with reason).
  • LIN-FR-015: The blast-radius graph is two-level (dataset ← chart ← dashboard) by construction; fan-out MUST order by impact severity only. Chained SQL-source lineage and cycle detection are out of scope.
  • LIN-FR-017: Impact projection MUST respect projection_confidence: a schema change marks exact-confidence charts affected only when the changed column/metric is in their consumed set; conservative-confidence charts (or expressions) are marked affected on any schema change to their dataset, with the conservative basis shown to the user.
  • LIN-FR-021: Impact, deprecation, fan-out and recreated-dataset findings MUST emit idempotent 036 InvestigationSignals. An analyst-opened case MAY use an agent to explain blast radius, build a revalidation plan and save validated scenario revisions; deterministic impact records remain immutable evidence and never auto-start agent work.
  • LIN-FR-022: A disabled, stale, or unavailable lineage index MUST be distinguishable from a known-empty dependent set. Any PROD gate, fan-out admission, or load-policy decision receiving non-current lineage state MUST display blast_radius_unknown with provenance and block or require an explicit policy escalation; it MUST NOT render zero dependents as fact.

Key Entities

  • DatasetUsageIndex: Content-addressed reverse index — dataset id → dependent dashboards with per-chart usage detail, index fingerprint, observation timestamp, environment scope, stale_index marker.
  • DatasetSchemaDiff: Classified change set between two dataset observations — kind, severity, versioned classification rules, observation pair provenance.
  • DatasetImpactRecord: Per-dependent-dashboard projection of a schema diff — affected charts, affected baseline entries by release pin, affected scenario step refs, severity, links (non-mutating) to StructureDiff.
  • DeprecationRecord: Dataset deprecation lifecycle — successor id, grace window, escalation state, migrated vs remaining dependents, terminal blocking state.
  • FanoutPlan: Coordinated multi-dashboard verification plan from a dataset_updated trigger — plan id, ordered per-dashboard runs (severity order), inspection-only entries with reasons, environment scope.
  • FleetReport: Aggregated fan-out outcome — per-dashboard status, dataset-change provenance, unresolved impacts, consumable by 039 pipeline views.
  • ChartUsageEdge: One binding in the index — dataset id, chart id, dashboard id, consumed columns/metrics (exact set), projection_confidence (exact|conservative), unresolved_refs when SQL expressions reference unknown columns.

Success Criteria

  • SC-001: Reverse lookups return 100% of dependent dashboards for fixture shared datasets with byte-identical fingerprints across repeated builds and shuffled metadata ordering.
  • SC-002: Schema-diff classification reaches 100% agreement with labeled fixture changes (removed/type-changed/added/metric-removed/deleted) with deterministic severity assignment.
  • SC-003: Zero writes to 037 baseline catalogs across the entire propagation fixture suite (verified by catalog hash before/after); staleness surfaces only as report/load-time marking.
  • SC-004: Closed-period retroactive data change surfaces immutability_violation (never stale_baseline) in 100% of fixture cases.
  • SC-005: Fan-out ordering respects severity (critical before warning before info) in 100% of fixture plans.
  • SC-008: SQL-expression extraction on fixture charts: 100% of matched refs produce exact edges; 100% of unmatched refs produce unresolved_refs findings; parse failures never flip whole-chart confidence to conservative when structured params are exact.
  • SC-006: Deprecation blocks operations at expiry with typed dataset_deprecated naming the successor in 100% of fixture cases; zero 404-style raw failures reach users.
  • SC-007: Index refresh failure retains last good index with stale_index in 100% of fault-injection cases; no empty/partial index is ever served as current.

Implementation Status & MVP Debt (audit 2026-08-07)

Facts (code check, not tasks.md):

  • ✅ Backend реализован и подключён: Services.Lineage.Indexer реально ходит в Superset (get_dashboard/get_chart/get_dataset_detail), hooks (lineage_on_sync, lineage_refresh_targets) вызываются из IdMappingService.sync_environment после коммита, severity matrix v1, deprecation, fan-out и API routes (backend/src/api/routes/lineage.py) присутствуют.
  • 🔴 Фронт отсутствует: нет маршрутов /lineage в frontend/src/routes/; клиент frontend/src/lib/api/lineage.ts написан, но UI-панелей dependents/deprecation/fleet-report нет.
  • 🟡 Фича opt-in: lineage_index_enabled default false → по умолчанию индекс не строится, blast-radius потребители (040 PROD gate, 039 pipeline views) получают пустой read-model.

Закрытие: задачи T045–T048 в tasks.md Phase 10 (маршруты dependents/deprecation/fleet-report + решение по default-флагу).

Runtime Closure Status (2026-08-07, resolved)

  • ✅ T045: существующий Datasets.LineagePanel (T031) привязан на /datasets/[id] — blast-radius dependents + stale_index notice. Транзиентный дубликат удалён (reuse).
  • ✅ T046: DatasetsLineageModel.markDeprecated()/recordMigration() + deprecation-секция в LineagePanel (grace days, successor, migration uuid, escalation state).
  • ✅ T047: DatasetsLineageModel.loadFleetReport(planId) через GET /lineage/fanout/{plan_id}/report.
  • ✅ T048 (decision): lineage_index_enabled остаётся false с задокументированным rationale — включение по умолчанию вызвало бы post-sync Superset detail-calls для каждой env на каждом цикле; flip только после proof стабильности indexer'а на живом fleet. Consumers (040 PROD gate, 039 pipeline views) трактуют disabled index как пустой read-model со stale_notice.
  • ✅ Verification: lineage + api vitest = 236 passed, vite build OK, eslint чист для изменённого кода.

Drift Amendment — MCP Interface (2026-08-24)

  • "Agent explains blast radius / builds revalidation plans" (LIN-FR-021) continues with external MCP clients as the delegated actor; impact records, deprecation and fan-out remain deterministic server artifacts that no client can mutate outside governed tools.

Status (2026-09-02): done — реализовано в рамках 050: инструменты и гейты (specs/050-mcp-interface/tasks.md T012–T028 [x]), handoff-поверхность (050 T030–T033), демонтаж чата и сервиса agent/ (050 T040–T041, чекпоинты specs/WORKSTATE-043-047.md).

#endregion DatasetLineageBlastRadius.Spec


UX REFERENCE — Interaction Narrative

Source: ux_reference.md

#region DatasetLineageBlastRadius.UxReference [C:3] [TYPE ADR] [SEMANTICS ux,reference,lineage,blast-radius,dataset] @BRIEF UX interaction reference for dataset lineage and blast-radius surfaces: persona, flows, states, recovery. Drives @UX_* tags in Phase 1 contracts.

Feature Branch: 041-dataset-lineage-blast-radius Created: 2026-07-22 | Status: Draft

1. User Persona & Context

  • Who is the user?: Two personas share these surfaces. (1) BI engineer owning a dashboard who needs to know «сломается ли мой дашборд, если изменят этот датасет?». (2) Data/platform engineer owning a dataset who needs to know «кого я затрону, если изменю или выведу из эксплуатации этот датасет?».
  • What is their goal?: Make impact visible before the change: named dependent dashboards, classified severity, affected baselines and scenario steps — and a managed path (deprecation + fan-out re-verification) instead of silent breakage discovered at the next release check.
  • Context: Web UI on desktop. Entry points: dataset detail page («Используется в N дашбордах»), dashboard page (shared-dataset badge), deprecation manager, fleet report after fan-out. Read-only for most users; mutations (deprecation, fan-out trigger, index refresh) gated by role.

2. The "Happy Path" Narrative

The data engineer opens a dataset page and sees the lineage panel: «Используется в 4 дашбордах, 11 чартах» — expandable to per-chart column usage. They mark the dataset deprecated, pick the successor from a searchable list, set a 14-day grace window, and the panel immediately lists every dependent with migration status «0 из 4 перенесено». A week later a schema diff lands on the successor: each dependent dashboard shows a classified impact card («critical: колонка amount удалена — затронуты 2 чарта, 3 baseline-записи release 1.4»). They trigger fan-out: one plan, four verification runs ordered critical-first, and a fleet report aggregates results — two dashboards pass, one needs baseline re-approval, one was inspection-only (no baselines yet). Nothing was silently rewritten: baseline catalogs show staleness markers with provenance, and the immutability badge on a closed-period entry is intact.

3. Interface Mockups

UI Layout & Flow

Screen: Dataset Lineage Panel (on dataset detail page)

  • Layout: Header (dataset, environment, index freshness badge) → dependents table → per-chart usage expansion → actions (gated).
  • Key Elements:
    • Dependents table: Dashboard title (link), chart count, used columns/metrics summary, release pin(s) with baselines, impact severity badge when an active impact record exists.
    • Index freshness: «Индекс обновлён 5 мин назад» or amber stale_index «данные могут быть устаревшими — последняя ошибка: timeout».
    • «Отметить устаревшим» (gated dataset:deprecation:manage): opens deprecation form — successor selector, grace window, initial notice preview.
  • Contract Mapping:
    • @UX_STATE: loading → ready → stale_index → refresh_failed; deprecation sub-FSM none → noticed → warning → expired_blocked.
    • @UX_FEEDBACK: Severity badges (critical/warning/info), freshness badge, migration progress bar «2 из 4».
    • @UX_RECOVERY: Refresh failure → «Повторить обновление» (gated) with last-good data retained; expired dataset in a dependent → error card names successor and links migration guidance.
    • @UX_REACTIVITY: Screen model atoms (index, impacts, deprecation) $derived into panels; pinning: consumers hold the index fingerprint they opened with.
    • Screen Model: Datasets.LineageModel.svelte.ts (index + impacts projection) and Datasets.DeprecationModel.svelte.ts (deprecation FSM). Components bind via @RELATION BINDS_TO.

Screen: Cross-Dashboard Impact View (per dataset change)

  • Layout: Change summary (kind, severity, observation pair) → per-dependent cards grouped by severity → each card: affected charts, affected baseline entries by release, affected scenario step refs, link to the dashboard's own StructureDiff (orthogonal, read-only link).

Screen: Deprecation Manager

  • Layout: Active deprecations list → per-record detail: successor, window countdown, escalation state, migrated vs remaining dependents with per-dashboard migration checklist.

Screen: Fleet Report (post fan-out)

  • Layout: Plan header (trigger dataset_updated, impact record link, environment) → ordered run list (severity, status badge, inspection-only reason where applicable) → aggregate summary (pass/warn/fail/unresolved impacts) → export.

4. The "Error" Experience

Philosophy: Lineage is advisory infrastructure — it must degrade loudly, never fabricate. A stale or partial index is worse than no index if served as current.

Scenario A: Index refresh failure

  • System Response: (UI) Freshness badge flips amber stale_index with typed error («Superset 403 — нет доступа к metadata»); dependents table shows last-good data with the staleness marker on every row.
  • Recovery: «Повторить обновление» (role-gated); if the failure persists, escalation link to environment settings. No row ever disappears silently.

Scenario B: Expired deprecated dataset blocks an operation

  • System Response: Verification or scenario generation returns typed dataset_deprecated error card: «Датасет выведен из эксплуатации 3 дня назад. Преемник: sales_v2. 2 дашборда ещё не перенесены.»
  • Recovery: Link to deprecation record with migration checklist; for the dashboard owner — link to chart re-pointing guidance. Raw Superset 404s never reach the user unwrapped.

Scenario C: Lineage cycle detected

  • System Response: Fan-out planning halts the affected subtree with a lineage_cycle finding showing the cycle path (ds_a → ds_b → ds_a); unaffected dependents proceed.
  • Recovery: Finding links the involved datasets for manual resolution; re-plan available after fix.

Scenario D: Permission denied on mutations

  • System Response: permission_denied recovery panel naming the required role (dataset:deprecation:manage / dataset:fanout:trigger); no confirm control rendered (036 gate semantics).
  • Recovery: Request-access link; read-only lineage data remains fully visible.

5. Tone & Voice

  • Style: Concrete and accountable. Always name the affected parties («затронуты 2 чарта, 3 baseline-записи release 1.4»), never vague («могут быть проблемы»).
  • Terminology: «Зависимые дашборды», «записи влияния», «преемник датасета», «fan-out» (kept in English in RU locale), «fleet report» → «сводный отчёт». Staleness is always framed as «требует пересмотра», never «устарело автоматически» — propagation surfaces, humans decide.
  • Safety copy: Deprecation and fan-out actions state scope upfront: «это действие уведомит владельцев N дашбордов и создаст M верификационных запусков».

#endregion DatasetLineageBlastRadius.UxReference


CHECKLISTS — Requirements Quality — requirements.md

Source: checklists/requirements.md

Requirements Checklist: Dataset Lineage & Blast-Radius

Purpose: Validate spec.md completeness, clarity, and testability before /speckit.plan. Created: 2026-07-22 Feature: spec.md

Content Quality

  • CHK001 Spec is user/operator-focused; storage/index implementation deferred to plan
  • CHK002 All five stories independently testable with stated Independent Tests
  • CHK003 Priorities assigned (US1–US3 P1, US4–US5 P2) with justification
  • CHK004 Key entities defined without schema detail

Gap Coverage (BR-1..BR-5 from 036–039 analysis)

  • CHK005 BR-1 (reverse index dataset→dashboards) → US1, LIN-FR-001..003, SC-001
  • CHK006 BR-2 (cross-dashboard StructureDiff / schema-change classification) → US2, LIN-FR-004..006, SC-002
  • CHK007 BR-3 (baseline invalidation single-release scoped) → US3, LIN-FR-007..008, SC-003/004
  • CHK008 BR-4 (no fleet-wide trigger) → US5, LIN-FR-010..011, SC-005
  • CHK009 BR-5 (C07 too narrow / dataset deprecation absent) → US4, LIN-FR-009, SC-006

Requirement Quality

  • CHK010 Every FR testable (deterministic fingerprints, zero catalog writes, severity ordering, typed errors)
  • CHK011 FR IDs hierarchical and stable (LIN-FR-001..020)
  • CHK012 No [NEEDS CLARIFICATION] markers remain
  • CHK013 Rejected paths recorded in header @REJECTED (QueryModel extension, query-time derivation, auto-invalidation)

Dependency Traceability

  • CHK014 037 dependencies explicit: StructureDiff orthogonality (LIN-FR-006), immutability rules (LIN-FR-007), release pinning (LIN-FR-008)
  • CHK015 036 dependencies explicit: approval-gate semantics for PROD fan-out (LIN-FR-014), permission_denied UX
  • CHK016 040 integration explicit: blast-radius read model served from pinned index snapshot (LIN-FR-013)
  • CHK017 039 integration explicit: FleetReport consumable without AgentRun (LIN-FR-011)
  • CHK018 Trigger enum extension dataset_updated declared (LIN-FR-011) — requires amendment note to 036/037 at plan time

Safety & Risk

  • CHK019 No silent mutation: propagation marks/reports only; catalog hash invariant verified (SC-003)
  • CHK020 Stale-index handling: last-good retained, typed provenance, never empty-as-current (LIN-FR-003, SC-007)
  • CHK021 Cycle safety: two-level graph (dataset ← chart ← dashboard) makes cycles impossible by construction; fan-out is ordered by impact severity, zero hangs (LIN-FR-015, SC-005)
  • CHK022 Environment scoping: cross-env propagation warns env_mismatch, never assumes (LIN-FR-012)

Success Criteria

  • CHK023 All SC measurable (100% dependents, 100% classification agreement, zero writes, ordering, zero hangs)

Clarifications Integrated (2026-07-22)

  • CHK028 Graph model corrected: two-level dataset←chart←dashboard, no SQL-source lineage, cycles removed (LIN-FR-015 rewritten)
  • CHK029 Usage extraction = C+2 hybrid: structured params exact per viz_type + sqlparse SQL-expression analysis (incl. dataset metric expression deps) with unresolved_refs findings and per-expression conservative fallback (LIN-FR-001/016/017, SC-008)
  • CHK030 Severity Matrix v1 accepted as defaults (incl. widening=info, metric.expression-changed row, physical rename vs relabel split, label-preserved rename hint) — LIN-FR-004 amended
  • CHK031 Refresh rides existing IdMappingService.sync_environment cycle (migration_sync_cron + sync-now); no dedicated scheduler; post-sync hook bounded to changed uuids; opt-in lineage_index_enabled default false (LIN-FR-002/018/019)
  • CHK032 Dataset labels (verbose_name) captured; schema_hash = physical identity (relabel invisible); dataset metrics resolved by name/label with expression deps; dataset_recreated candidate tracking by dataset_physical_key with human confirmation (LIN-FR-020, R9)

Open Items (deferred to /speckit.plan)

  • CHK024 Index storage model (edges table alongside ResourceMapping vs derived cache) — plan-time research; cadence CLOSED via CHK031 (rides sync_environment cycle)
  • CHK025 Classification rules versioning scheme — CLOSED: Severity Matrix v1 accepted; rule changes require version bump + new golden fixtures (clarification Q3)
  • CHK026 Amendment mechanics for 036/037 trigger enum + 040 read-model contract — plan-time coordination
  • CHK027 Derived-dataset (SQL-query datasets) lineage source — CLOSED: out of scope, graph is two-level binding graph (clarification Q1)

PLAN — Implementation Plan

Source: plan.md

Implementation Plan: Dataset Lineage & Blast-Radius

Branch: 041-dataset-lineage-blast-radius | Date: 2026-07-22 | Spec: spec.md Input: Feature specification from /specs/041-dataset-lineage-blast-radius/spec.md (clarified 2026-07-22)

Summary

Build a dataset→dashboard reverse index riding the existing IdMappingService.sync_environment cycle (no new scheduler), with C+2 usage extraction (structured chart params + sqlparse SQL-expression analysis), Severity Matrix v1 classification, report-time staleness propagation with zero baseline-catalog writes, deprecation lifecycle, and coordinated dataset_updated verification fan-out. Closes blast-radius gaps BR-1..BR-5 identified across specs 036–039; provides the pinned read model consumed by 040 load testing.

Technical Context

Language/Version: Python 3.13+ (backend), TypeScript (frontend Svelte 5 runes-only) Primary Dependencies: FastAPI 0.126, SQLAlchemy 2.0.45, APScheduler 3.11.2 (existing, via migration sync), sqlparse ≥0.5 (existing — R2, no new deps); SvelteKit 2.49 / Svelte 5.56, Tailwind (frontend) Storage: PostgreSQL 16 — 6 new tables + 1 additive FK (data-model.md); Alembic migration Testing: pytest (backend: indexer/extractor/diff/propagation/fanout suites), vitest L1 (LineageModel/DeprecationModel) + L2 UX (@testing-library/svelte) Target Platform: Linux server (Docker), modern browsers Project Type: web application (FastAPI REST backend, SvelteKit SPA frontend) Frontend Architecture: model-first Screen Models (.svelte.ts), runes-only, typed DTOs mirroring Pydantic Performance Goals: dependents read model <200ms p95 from snapshot table; incremental sync hook adds ≤ changed-set detail calls (0–20 typical, never O(fleet)) Constraints: RBAC enforced (LIN-FR-014); zero writes to 037 baseline catalogs (SC-003); opt-in lineage_index_enabled default false; no new runtime dependencies Scale/Scope: ~100 dashboards/env, ~300 charts, ~80 datasets; sync cycle 30min default via existing migration_sync_cron

Constitution Check

GATE: Must pass before Phase 0 research. Re-check after Phase 1 design. — PASS (initial, 2026-07-22)

Principle Status Evidence
I. Semantic Contract First ✅ contracts/modules.md: all C3+ with anchors, hierarchical IDs (Services.Lineage.*), shared @SEMANTICS lineage
II. Decision Memory ✅ 5 @REJECTED across spec header + contracts (QueryModel extension, query-time derivation, auto-invalidation, sqlglot, separate scheduler); R1–R10 rationale records
III. External Orchestrator ✅ Reads Superset metadata via existing SupersetClient boundary; no Superset-side plugin
IV. Module Discipline ✅ 8 backend modules, each <400 LOC by design (extractor/diff/propagation separated); functions CC≤10 via pure-function severity rules
V. RBAC Enforcement ✅ LIN-FR-014: three new granular permissions + default-deny; PROD fan-out via 036 gate
VI. Svelte 5 Runes Only ✅ Models .svelte.ts with $state; no stores/legacy syntax
VII. Test-Driven C3+ ✅ quickstart.md defines falsifiable sequence incl. explicit @REJECTED-path tests (regex-extraction ban, auto-invalidation ban via SC-003)
VIII. Attention Optimization ✅ ATTN_1 one-line anchors; ATTN_2 Services.Lineage.*; ATTN_3 shared lineage keyword; ATTN_4 contracts ≤150 lines

Post-Phase-1 re-check: PASS — no violations introduced; no Complexity Tracking entries required.

Project Structure

Documentation (this feature)

specs/041-dataset-lineage-blast-radius/
├── plan.md              # This file
├── research.md          # Phase 0 — R1..R8 decisions
├── data-model.md        # Phase 1 — ORM + DTOs + invariants
├── quickstart.md        # Phase 1 — falsifiable verification sequence
├── traceability.md      # RTM (below)
├── checklists/requirements.md
├── contracts/
│   └── modules.md       # GRACE module/function contracts
└── tasks.md             # Phase 2 output (/speckit.tasks)

Source Code (repository root)

backend/
├── src/
│   ├── models/lineage.py                 # NEW: 6 ORM tables (data-model.md)
│   ├── schemas/lineage.py                # NEW: Pydantic DTOs (extra-forbid)
│   ├── services/lineage/
│   │   ├── indexer.py                    # OnSync hook + RefreshTargets write-hook
│   │   ├── usage_extractor.py            # per-viz_type structured extraction
│   │   ├── sql_expression_extractor.py   # sqlparse C+2 analysis
│   │   ├── severity_rules_v1.py          # pure-function matrix
│   │   ├── schema_diff.py                # observation diff + impact projection
│   │   ├── propagation.py                # report-time staleness marking
│   │   ├── deprecation.py                # lifecycle + escalation recompute
│   │   └── fanout.py                     # plan + VerificationRun spawning
│   ├── api/routes/lineage.py             # REST surface
│   └── core/mapping_service.py           # AMEND: post-commit OnSync hook call
├── alembic/versions/                     # NEW migration (6 tables + VerificationRun.fanout_plan_id)
└── tests/
    ├── services/lineage/                 # unit + property tests per module
    ├── api/test_lineage.py               # RBAC/contract tests
    └── fixtures/lineage/                 # 3 dashboards, 2 shared datasets, SQL-expression charts

frontend/
├── src/lib/models/
│   ├── Datasets.LineageModel.svelte.ts       # NEW
│   ├── Datasets.DeprecationModel.svelte.ts   # NEW
│   └── Dashboards.LineageBadgeModel.svelte.ts # NEW
├── src/lib/types/lineage.ts              # NEW: DTO mirror
└── src/lib/components/datasets/lineage/  # NEW: panel components (039-adjacent)

Structure Decision: Backend-heavy feature; frontend limited to read panels + deprecation manager (fits ADR-0001 layout). core/mapping_service.py amendment is a single post-commit hook call — no refactor of identity sync.

Semantic Contract Guidance

Applied per template §Attention Compliance Gate — all contracts in contracts/modules.md validated against ATTN_1–4 (see Constitution Check). Cross-stack edges declared: Api.Lineage.Routes BINDS_TO frontend/src/lib/types/lineage.ts; fan-out edges to 036/037 spec nodes; amendment notes for trigger enum recorded in R5.

Complexity Tracking

No constitution violations — section intentionally empty.

Phase Outputs

Phase Artifact Status
0 research.md ✅ R1 storage, R2 sqlparse, R3 hook placement, R4 rules versioning, R5 enum amendment, R6 040 read model, R7 deprecation, R8 fanout persistence
1 data-model.md, contracts/modules.md, quickstart.md ✅
2 tasks.md ⏭ next: /speckit.tasks
10 (MVP) frontend routes + default-flag decision (T045–T048) ⏳ открыто (audit 2026-08-07)

MVP Frontend & Opt-In Closure (audit 2026-08-07)

Status correction: Backend-индекс, severity matrix, deprecation lifecycle и fan-out реализованы и подключены к sync_environment (_emit_lineage_hook → lineage_on_sync), API routes существуют. Но фронт-маршрутов /lineage нет, а lineage_index_enabled default false — потребители (040 PROD gate, 039 pipeline views) получают пустой read-model.

Closure tasks (tasks.md Phase 10, T045–T048):

  • T045 — Frontend route lineage/datasets/[datasetId]/+page.svelte (blast-radius panel, stale_index notice) через api/lineage.ts.
  • T046 — Deprecation management surface (successor + grace window, escalation badge, migration tracker).
  • T047 — Fleet-report surface для dataset_updated fan-out.
  • T048 — Flip lineage_index_enabled default true (после proof стабильности) или явное документирование opt-in.

ADR Continuity

  • ADR-0001 module layout: new services/lineage/ package, routes in api/routes/lineage.py — compliant
  • ADR-0003 orchestrator: metadata reads via existing SupersetClient — compliant
  • ADR-0005 RBAC: new permissions registered in rbac_permission_catalog.py at implementation
  • No ADR amendments required; R5 spec-text notes for 036/037 deferred to implementation (branch isolation)

RESEARCH — Technical Decisions

Source: research.md

#region DatasetLineageBlastRadius.Research [C:3] [TYPE ADR] [SEMANTICS research,lineage,blast-radius,storage,sqlparse] @BRIEF Phase 0 research: resolves all material unknowns for dataset lineage — storage model, SQL-expression parsing, sync-hook placement, severity rules versioning, trigger-enum amendments, and 040 read-model contract.

Feature: 041-dataset-lineage-blast-radius | Date: 2026-07-22

R1 — Index Storage Model (CHK024)

  • Decision: Dedicated DatasetUsageEdge ORM table (backend/src/models/lineage.py) alongside ResourceMapping; edge rows keyed (environment_id, dataset_uuid, chart_uuid, dashboard_uuid) with consumed columns/metrics as JSONB + projection_confidence + unresolved_refs JSONB. Separate LineageIndexSnapshot table stores per-env index_fingerprint, built_at, stale_index flag, error provenance.
  • Rationale: ResourceMapping is an uuid↔id identity table with sync cadence semantics — bolting usage detail onto it would couple identity sync failures to lineage availability and bloat the hot migration path. A dedicated table allows independent rebuild (full edge wipe + rebuild from current mappings) without touching migration state. JSONB for consumed sets avoids a join-table for data that is only ever read whole.
  • Alternatives Considered: (a) Derived in-memory cache rebuilt per request — rejected: cold-start O(fleet) per consumer, no snapshot pinning across restarts, violates LIN-FR-002. (b) Extend ResourceMapping with JSONB usage column — rejected: couples migration sync to lineage rebuild; stale-deletion semantics differ (identity row deleted vs edge recomputed).
  • Impact: Alembic migration adds 2 tables. stale_index = snapshot row flag, not edge deletion — LIN-FR-003 satisfied by retaining last good snapshot.

R2 — SQL-Expression Parsing: sqlparse, not sqlglot (Q2 implementation)

  • Decision: Use sqlparse (already in backend/requirements-backend.txt:64) with a new Services.Lineage.SqlExpressionExtractor modeled on the existing Services.SqlTableExtractor three-phase pattern (Jinja span detection → SQL spans → token analysis). Column refs extracted from DML tokens; matched against dataset columns/metrics from authoritative metadata. Parse failure or non-DML ambiguity → expression marked conservative.
  • Rationale: sqlglot is NOT a project dependency — adding it for one feature violates minimal-dependency discipline and pulls a 1MB+ parser for a task sqlparse already handles in-repo (SqlTableExtractor proves the pattern incl. Jinja). Chart SQL expressions are dialect-simple (aggregations, CASE, arithmetic) — full sqlglot dialect resolution is overkill.
  • Alternatives Considered: (a) sqlglot — rejected: new dependency, marginal accuracy gain over sqlparse for Superset-generated expressions; can be adopted later behind the same extractor contract if fixture failure rate justifies it. (b) Regex-only — forbidden by LIN-FR-016.
  • Impact: No new dependencies. Extractor contract: extract(sql_expression, known_columns, known_metrics) -> {matched: set[str], unresolved: set[str], confidence: exact|conservative}.

R3 — Sync-Hook Placement (Q4 implementation)

  • Decision: Hook into IdMappingService.sync_environment completion: after self.db.commit(), emit changed-uuid sets per resource type to Services.Lineage.Indexer.OnSync(environment_id, changed: dict[ResourceType, set[str]]). Detail fetches (chart params, dataset schema) via SupersetClient ONLY for changed uuids. Write-hook: SupersetClient.update_dataset/create_dataset and migration deploy completion call Indexer.RefreshTargets(environment_id, dataset_ids=[...]) directly.
  • Rationale: sync_environment already computes incremental change sets (since_dttm diff) — the changed-uuid information exists there and nowhere else; hooking at commit boundary guarantees lineage sees exactly what identity-sync persisted. Detail-call bounding to changed set keeps Superset load at O(changed), not O(fleet).
  • Alternatives Considered: (a) Separate APScheduler job lineage_refresh_{env_id} — rejected by user: duplicate scheduler config, splits atomicity (mappings say one thing, lineage another until next run). (b) Refresh inside the same DB transaction — rejected: detail fetches are network I/O inside a DB transaction = lock holding; must be post-commit.
  • Impact: sync_environment gains one out-param/callback; transaction boundary unchanged. Task visibility: sync-now already surfaces per-env results; lineage hook appends lineage: {rebuilt: N, impacts: M} to that result envelope.

R4 — Severity Rules Versioning (CHK025 closed at clarify, mechanics here)

  • Decision: severity_rules_v1.py module with pure functions (change_kind, before, after, consumption) -> severity; matrix version constant embedded in DatasetImpactRecord.rules_version. Rule change = new module severity_rules_v2.py + dispatcher; old records keep their version; golden fixtures pinned per version.
  • Rationale: Pure-function modules keep classification deterministic and property-testable; embedding version in each record preserves audit trail when rules evolve.
  • Impact: No DB migration on rule bump — version is data. Re-classification of historical records is an explicit re-index operation, never implicit.

R5 — Trigger Enum Amendment Mechanics (CHK026)

  • Decision: The trigger field on AgentRun/VerificationRun (036 spec, VerificationRun entity 037) gains value dataset_updated. Amendment executed as additive enum extension in 041's own migration with a backward-compatible default; 036/037 spec texts receive a one-line amendment note in their traceability.md (not a spec rewrite) when 041 implements.
  • Rationale: Enum values are data-level, not schema-level — adding a value is backward compatible (old code sees unknown trigger as opaque string). Editing 036/037 spec files now would violate branch isolation; amendment notes land at implementation time.
  • Impact: 041 contracts declare dataset_updated as the producer; 037 VerificationRun consumer treats unknown triggers as displayable strings (already true per 039 AGUI-FR-016 read-only projection).

R6 — 040 Read-Model Contract (LIN-FR-013)

  • Decision: GET /api/v1/lineage/datasets/{dataset_id}/dependents?env_id=...&fingerprint=... returns pinned-snapshot view: {fingerprint, stale_index, dependents: [{dashboard_id, title, chart_count, projection_confidence}], env_scope}. 040 load profiles and PROD gates pass the fingerprint they validated against; mismatch → re-fetch notice, never silent drift.
  • Rationale: Fingerprint-as-query-param makes pinning explicit and stateless; server never guesses which snapshot the caller holds.
  • Impact: One read endpoint; 040 contracts bind to this shape.

R7 — Deprecation Storage & Escalation (US4 mechanics)

  • Decision: DatasetDeprecation ORM table: (environment_id, dataset_uuid, successor_dataset_uuid, deprecated_at, grace_window_days, escalation_state, expired_at) + per-dependent migration status JSONB. Escalation recomputed in sync cycle (LIN-FR-019) by pure function (now, deprecated_at, grace_window_days, migrated_ratio) -> state.
  • Rationale: Escalation as derived state (not stored transitions) eliminates timer-drift bugs; stored column is a cache of last computed value for query convenience.
  • Impact: One more table in the same Alembic migration.

R8 — Fan-Out Plan Persistence (US5 mechanics)

  • Decision: FanoutPlan ORM table (id, environment_id, dataset_impact_id, created_at, trigger, status); per-dashboard runs recorded as VerificationRun rows with fanout_plan_id FK (nullable — null = ordinary run). Inspection-only entries are plan rows without spawned runs, reason stored.
  • Rationale: Reuses 036/037 VerificationRun as the execution record — FleetReport is a query over plan + runs, no parallel run model.
  • Impact: VerificationRun gains one nullable FK column (additive, backward compatible).

R9 — Dataset Labels, Metric Consumption, and Recreate Tracking (2026-08-04)

Resolves design decisions Q1–Q3 and audit gaps E2/E5 after code verification against the Superset upstream repo. Superset Slice (Chart) has a single datasource_id/datasource_type (superset/models/slice.py), so chart→dataset binding is strictly 1:1 — the two-level tree dataset ← chart ← dashboard is exact and consumed_columns unambiguously attribute to one dataset (E1 join-collision removed as impossible).

R9a — Labels (verbose_name)

  • schema_hash = physical identity (column_name + type + nullable). A pure relabel (verbose_name change only, column_name unchanged) does NOT bump the hash and produces no diff at all — it is invisible to impact projection and scenario churn. This is the agreed Q1.
  • schema_payload carries full info including per-column verbose_name — used for rename disambiguation and label display resolution.
  • Superset itself resolves columns/metrics by metric_name OR verbose_name via its verbose_map (superset/connectors/sqla/models.py) — confirms the label-aware matching design.
  • Scenario refs (context.query_model.{col}, 038) carry only the physical column_name as the authoritative key. Labels are resolved as a render-time display layer from the observation/lineage index — never stored in the ref or executable content_hash. Optional denormalized label in a ref is allowed but excluded from content_hash. (Q2)

R9b — Dataset Metrics as a Required Consumption Vector

Dataset metrics are the primary chart consumption vector and MUST be accounted for (E2). A chart's consumed set is the union of:

  1. Structural params (columns/groupby) — parsed per viz_type (existing).
  2. Metrics referenced by the chart — resolved by metric_name or verbose_name; each metric contributes its expression column-deps to the consumed set (same SqlExpressionExtractor path as ad-hoc sqlExpression, no new dependency).
  3. Ad-hoc sqlExpression (existing).

Additions driven by metrics:

  • metric_removed → critical (already in Severity Matrix v1).
  • New severity row: metric.expression changed (name identical, formula different) → consumers of the metric are affected even without any column change; classify by consumption (exact-consumed → warning/critical, conservative → warning).

R9c — Dataset Recreate Tracking (dataset_recreated)

A dataset recreated during migration gets a new id/uuid but keeps the physical identity (database_id, schema, table_name) (superset/connectors/sqla/models.py). Without tracking, the new uuid looks like a fresh dataset with zero dependents — false "clean" state.

  • Physical-key registry: DatasetSchemaObservation stores dataset_physical_key = (environment_id, database_id, schema, table_name).
  • Detection during Indexer.OnSync: a new dataset_uuid whose (env, physical_key) matches a prior uuid that is now gone → emit candidate dataset_recreated: old_uuid → new_uuid. If the prior uuid is still live, it is a legitimate second dataset on the same table → no finding (Superset permits multiple SqlaTable rows over one table).
  • Conservative, human-confirmed (same philosophy as rename-hint, never auto-apply): candidate surfaced for operator confirmation; on confirmation the new uuid inherits the blast-radius context (dependents, deprecation as successor_dataset_uuid = new_uuid) so it is not treated as pristine.
  • Stale edges of the old uuid are marked orphaned/stale until reconciliation, not silently deleted.
  • Nuances: if the recreate is a migration rather than a re-point, dependents must be re-targeted as a human-confirmed action (interfaces with E7 deprecation chains).

Impact: column capture in get_dataset_detail gains verbose_name (additive, backward-compatible); DatasetSchemaObservation gains per-column labels + dataset_physical_key; UsageExtractor gains metric-resolution; severity matrix gains metric.expression row and splits column_renamed (physical rename consumption-aware vs relabel → info, relabel invisible to hash). No new runtime dependencies.

Resolved Risks

Risk Mitigation
Superset list API lacks params/schema → detail-call storm Bounded to changed uuids only (R3); full rebuild is explicit operator action with cost warning
sqlparse mis-parses exotic dialect Conservative fallback per expression (LIN-FR-016); fixture suite covers ClickHouse/Postgres expression shapes
Post-commit hook failure leaves lineage behind mappings Snapshot marked stale_index with provenance (LIN-FR-003); next sync cycle retries changed set (idempotent rebuild)
Enum amendment races with 036/037 implementation Additive-only, opaque-string tolerance already required (R5)

R10 — MVP Frontend & Opt-In Gap (audit 2026-08-07)

  • Decision: Lineage остаётся backend-first; фронт-контур (dependents blast-radius panel, deprecation surface, fleet-report) добавляется задачами T045–T047, а lineage_index_enabled остаётся default false до T048 (proof стабильности indexer'а или явное документирование opt-in).
  • Rationale: Backend-индекс реально исполняется (Indexer ходит в Superset, hooks подключены к sync_environment), но без UI и без включённого флага потребители (040 PROD gate, 039 pipeline views) не получают ценность. Включение по умолчанию требует стабильности инкрементального перестроения на реальных данных.
  • Alternatives considered: frontend-first (rejected — сначала read-model), включить флаг сразу (rejected — нет proof на живом fleet), отдельный lineage-крон (rejected — LIN-FR-002 запрещает).
  • Impact: T045–T048 в tasks.md Phase 10; quickstart шаг 8 покрывает backend-гейт, фронт-гейт добавляется вместе с Phase 10.

#endregion DatasetLineageBlastRadius.Research


DATA MODEL — Entities & Relations

Source: data-model.md

#region DatasetLineageBlastRadius.DataModel [C:3] [TYPE ADR] [SEMANTICS data-model,lineage,orm,dto] @BRIEF Phase 1 data model: ORM tables, DTO shapes, and invariants for dataset lineage index, impacts, deprecation, and fan-out.

Feature: 041-dataset-lineage-blast-radius | Date: 2026-07-22

ORM Tables (backend/src/models/lineage.py — new)

DatasetUsageEdge

Column Type Notes
id str (uuid) PK
environment_id str FK env scope (LIN-FR-012)
dataset_uuid str from ResourceMapping (identity stable across re-ids)
chart_uuid str
dashboard_uuid str
dashboard_title str denormalized for panel rendering
consumed_columns JSONB list[str] exact set (structured params + matched SQL refs)
consumed_metrics JSONB list[str]
projection_confidence enum(exact, conservative) LIN-FR-001
unresolved_refs JSONB list[str] SQL refs matching nothing (LIN-FR-016)
edge_fingerprint str sha256 of edge content

Unique: (environment_id, dataset_uuid, chart_uuid, dashboard_uuid).

LineageIndexSnapshot

Column Type Notes
environment_id str PK one current snapshot per env
index_fingerprint str sha256 over sorted edge_fingerprints
built_at datetime
stale_index bool LIN-FR-003
last_error JSONB typed provenance {kind, message, at}

Snapshot pinning semantics (reconciled 2026-08-07): index_fingerprint is an optimistic consistency token, not a historical edge-set store. Consumers pin a fingerprint for the duration of an operation; on refresh the edge set is rebuilt delete+insert and the fingerprint flips atomically. A pinned-fingerprint mismatch yields stale_notice (re-fetch, never silent drift). Restoring the prior edge set from the fingerprint is NOT supported (LIN-FR-002).

DatasetSchemaObservation

Column Type Notes
id str PK
environment_id, dataset_uuid str unique together with observed_at ordering
dataset_physical_key JSONB (database_id, schema, table_name) — stable identity across recreate (LIN-FR-020)
schema_hash str sha256 of canonical physical schema (column_name + type + nullable + metrics) — relabel excluded (R9a)
schema_payload JSONB canonical schema incl. per-column verbose_name, nullability, is_dttm; full payload for rename-heuristic + label display (R9a)
observed_at datetime

DatasetImpactRecord

Column Type Notes
id str PK
environment_id, dataset_uuid str
diff JSONB classified changes per Severity Matrix v1
rules_version str e.g. "v1" (R4)
max_severity enum(critical, warning, info)
affected_dashboards JSONB per-dashboard: charts, baseline entries by release pin, scenario step refs
observation_pair JSONB {before_id, after_id}
created_at datetime monotonic sequence per dataset (edge case: racing changes)

DatasetDeprecation

Column Type Notes
environment_id, dataset_uuid str PK
successor_dataset_uuid str LIN-FR-009
deprecated_at datetime
grace_window_days int
escalation_state enum(noticed, warning, expired_blocked) derived cache (R7)
expired_at datetime null
dependent_status JSONB per-dashboard migrated/remaining

FanoutPlan

Column Type Notes
id str PK plan id linked by runs (LIN-FR-010)
environment_id str
dataset_impact_id str FK → DatasetImpactRecord
trigger str dataset_updated
status enum(planning, running, completed, failed)
entries JSONB ordered: dashboard, severity, inspection_only+reason
created_at datetime

VerificationRun amendment (036/037 model)

  • fanout_plan_id str FK null (additive; null = ordinary run)
  • trigger enum gains dataset_updated (R5)

DTOs (backend/src/schemas/lineage.py, Pydantic extra-forbid)

  • DependentsResponse {fingerprint, stale_index, env_scope, dependents: [{dashboard_id, title, chart_count, projection_confidence}]} — 040 read model (R6)
  • ImpactRecordDTO {id, dataset, rules_version, max_severity, diff, affected_dashboards, created_at}
  • DeprecationDTO {dataset, successor, deprecated_at, grace_window_days, escalation_state, dependent_status}
  • FanoutPlanDTO {id, trigger, status, entries, fleet_summary?}
  • FleetReportDTO {plan_id, per_dashboard: [{dashboard, status, provenance, unresolved_impacts}]} — 039-consumable (LIN-FR-011)

Frontend DTOs (frontend/src/lib/types/lineage.ts)

Mirror of backend schemas, generated/hand-authored 1:1; cross-stack @RELATION edges declared in contracts/modules.md.

An agent may explain an impact record, assemble a revalidation/fan-out plan, and create scenario migration drafts from an opened InvestigationCase. Index construction, schema-diff severity, impact projection and snapshot pinning remain deterministic. Agent actions never rewrite lineage history or apply a recreate/migration candidate without the applicable policy decision.

Invariants

  1. Edge rows are disposable: full rebuild = delete+insert within one transaction; snapshot fingerprint flips atomically.
  2. Snapshot is never deleted on refresh failure — stale_index=true + provenance (LIN-FR-003).
  3. Impact records are append-only; classification never mutates history (R4).
  4. escalation_state is a derived cache; source of truth = (now, deprecated_at, grace_window_days).
  5. No table here references baseline catalog files — propagation is report-time marking (LIN-FR-007).
  6. consumed_metrics on an edge includes column-deps of the referenced dataset metric expression; relabel (verbose_name-only) never bumps schema_hash (R9a/R9b).
  7. Recreate detection is candidate-only: a live prior uuid on the same dataset_physical_key never yields a dataset_recreated finding (LIN-FR-020).

#endregion DatasetLineageBlastRadius.DataModel


CONTRACTS — Module & Function Contracts

Source: contracts/modules.md

#region DatasetLineageBlastRadius.Modules [C:4] [TYPE ADR] [SEMANTICS contracts,modules,lineage,blast-radius] @BRIEF GRACE module/function contracts for 041: indexer, extractor, severity rules, propagation, deprecation, fan-out, API, frontend models. @RELATION DEPENDS_ON -> [DatasetLineageBlastRadius.DataModel] @RELATION DEPENDS_ON -> [DatasetLineageBlastRadius.Research]

Backend Modules

// #region Services.Lineage.Indexer [C:5] [TYPE Module] [SEMANTICS lineage,indexer,sync]
// @defgroup Lineage Dataset lineage and blast-radius domain.
// @BRIEF Builds and maintains DatasetUsageEdge set per environment; post-sync hook entry point.
// @LAYER Service
// @PRE IdMappingService.sync_environment completed and committed changed ResourceMapping rows.
// @POST Edge set rebuilt for changed uuids only; snapshot fingerprint flipped atomically; stale_index retained on failure.
// @SIDE_EFFECT DB writes (edges, snapshot); Superset detail calls bounded to changed uuids.
// @INVARIANT Detail fetches NEVER exceed the changed-uuid set (cost bound).
// @INVARIANT Snapshot is never deleted on failure — stale_index=true with typed provenance (LIN-FR-003).
// @DATA_CONTRACT Input[environment_id, changed: dict[ResourceType, set[str]]] -> Output[{rebuilt: int, impacts: int}]
// @RELATION DEPENDS_ON -> [Core.MappingService.IdMappingService]
// @RELATION CALLS -> [Services.Lineage.UsageExtractor]
// @RELATION CALLS -> [Services.Lineage.SchemaDiff]
// @RATIONALE Post-commit hook placement keeps network I/O out of the identity-sync transaction (R3).
// @REJECTED Separate APScheduler lineage job — rejected: duplicate scheduler config and split identity/lineage atomicity.

// #region Services.Lineage.Indexer.OnSync [C:4] [TYPE Function] [SEMANTICS lineage,sync-hook]
// @ingroup Lineage
// @BRIEF Post-commit hook invoked by sync_environment with changed-uuid sets.
// @SIDE_EFFECT Superset detail fetches for changed uuids; DB edge rebuild; snapshot update.
// @RELATION CALLED_BY -> [Core.MappingService.SyncEnvironment]
// #endregion Services.Lineage.Indexer.OnSync

// #region Services.Lineage.Indexer.RefreshTargets [C:4] [TYPE Function] [SEMANTICS lineage,write-hook]
// @ingroup Lineage
// @BRIEF Targeted refresh for own mutating operations (migration deploy, dataset create/update).
// @PRE Caller supplies explicit dataset/chart/dashboard ids affected by the mutation.
// @SIDE_EFFECT Immediate edge rebuild for targets without waiting for sync cycle (LIN-FR-018).
// @RELATION CALLED_BY -> [Core.Datasets.SupersetClientUpdateDataset]
// #endregion Services.Lineage.Indexer.RefreshTargets
// #endregion Services.Lineage.Indexer

// #region Services.Lineage.UsageExtractor [C:4] [TYPE Module] [SEMANTICS lineage,extraction,viz]
// @ingroup Lineage
// @BRIEF Extracts consumed columns/metrics from chart params per viz_type; delegates SQL expressions.
// @LAYER Service
// @POST Returns {consumed_columns, consumed_metrics, projection_confidence, unresolved_refs}.
// @INVARIANT Unknown viz_type -> conservative (never guessed exact).
// @RELATION CALLS -> [Services.Lineage.SqlExpressionExtractor]
// @DATA_CONTRACT Input[chart_params: dict, dataset_columns, dataset_metrics] -> Output[UsageEdge projection]
// #endregion Services.Lineage.UsageExtractor

// #region Services.Lineage.SqlExpressionExtractor [C:4] [TYPE Module] [SEMANTICS lineage,sql,sqlparse]
// @ingroup Lineage
// @BRIEF sqlparse-based column-ref extractor for sqlExpression ad-hoc metrics/columns/WHERE (C+2).
// @LAYER Service
// @POST {matched, unresolved, confidence}; parse failure -> conservative for that expression only.
// @INVARIANT Regex-only extraction forbidden (LIN-FR-016); Jinja spans pre-stripped per SqlTableExtractor pattern.
// @RELATION DEPENDS_ON -> [Services.SqlTableExtractor.SqlTableExtractorModule]
// @RATIONALE sqlparse already in-repo (R2); sqlglot rejected as new dependency for marginal gain.
// #endregion Services.Lineage.SqlExpressionExtractor

// #region Services.Lineage.SeverityRulesV1 [C:3] [TYPE Module] [SEMANTICS lineage,severity,rules]
// @ingroup Lineage
// @BRIEF Pure-function Severity Matrix v1: (change_kind, before, after, consumption) -> severity.
// @POST Deterministic; matrix version "v1" embedded in every DatasetImpactRecord (R4).
// @INVARIANT column_added and type widening -> info, NEVER stale baselines (LIN-FR-005).
// #endregion Services.Lineage.SeverityRulesV1

// #region Services.Lineage.SchemaDiff [C:4] [TYPE Module] [SEMANTICS lineage,schema,diff]
// @ingroup Lineage
// @BRIEF Computes dataset schema diffs between observations; projects DatasetImpactRecord per dependent.
// @SIDE_EFFECT Appends impact records (append-only, R4); never mutates baseline catalogs (LIN-FR-006/007).
// @RELATION CALLS -> [Services.Lineage.SeverityRulesV1]
// @DATA_CONTRACT Input[before_schema, after_schema, edges] -> Output[DatasetImpactRecord?]
// #endregion Services.Lineage.SchemaDiff

// #region Services.Lineage.Propagation [C:4] [TYPE Module] [SEMANTICS lineage,staleness,propagation]
// @ingroup Lineage
// @BRIEF Report-time staleness marking for dependent baseline catalogs per release pin.
// @POST Stale/immutability_violation surfaced with dataset-change provenance; zero catalog writes (SC-003).
// @INVARIANT Closed-period divergence -> immutability_violation, never stale (LIN-FR-007, 037 AGBASE-FR-012).
// @INVARIANT Each pinned release evaluated independently (LIN-FR-008).
// @RELATION DEPENDS_ON -> [SupersetBaselineEngine.Spec]
// #endregion Services.Lineage.Propagation

// #region Services.Lineage.Deprecation [C:4] [TYPE Module] [SEMANTICS lineage,deprecation]
// @ingroup Lineage
// @BRIEF Deprecation lifecycle: successor, grace window, escalation recompute in sync cycle (LIN-FR-019).
// @SIDE_EFFECT DB writes on mark/migrate; expired state produces typed dataset_deprecated errors.
// @INVARIANT Escalation state derived from (now, deprecated_at, grace_window_days) — stored column is cache (R7).
// #endregion Services.Lineage.Deprecation

// #region Services.Lineage.Fanout [C:5] [TYPE Module] [SEMANTICS lineage,fanout,verification]
// @ingroup Lineage
// @BRIEF Coordinated multi-dashboard verification plans from dataset_updated trigger.
// @LAYER Service
// @PRE DatasetImpactRecord exists with max_severity >= warning.
// @POST One FanoutPlan; per-dashboard VerificationRuns linked by fanout_plan_id, severity-ordered; inspection-only entries with reason.
// @SIDE_EFFECT Creates VerificationRun rows (036/037 model); PROD-classified fan-out requires 036 approval gate (LIN-FR-014).
// @INVARIANT Plan is the only multi-dashboard unit; runs remain per-dashboard records (R8).
// @DATA_CONTRACT Input[dataset_impact_id] -> Output[FanoutPlanDTO + FleetReportDTO]
// @RATIONALE A coordinated plan preserves one dataset-change cause while retaining per-dashboard VerificationRuns for existing 036/037 recovery and reporting semantics.
// @REJECTED Independent ad-hoc verification calls per dashboard — rejected because ordering, approval scope, and fleet aggregation would be lost.
// @RELATION DEPENDS_ON -> [AgentTestStabilization.Spec]
// #endregion Services.Lineage.Fanout

// #region Api.Lineage.Routes [C:4] [TYPE Module] [SEMANTICS lineage,api,rest]
// @ingroup Lineage
// @BRIEF REST surface: dependents read model, impacts, deprecation manage, fanout trigger/report.
// @LAYER API
// @INVARIANT RBAC: READ for reads; dataset:lineage:refresh, dataset:deprecation:manage, dataset:fanout:trigger for mutations (LIN-FR-014).
// @DATA_CONTRACT DependentsResponse honors ?fingerprint= pinning; mismatch -> stale_notice (R6).
// @RELATION BINDS_TO -> [EXT:frontend:frontend/src/lib/types/lineage.ts]
// #endregion Api.Lineage.Routes

Frontend Models

// #region Datasets.LineageModel [C:4] [TYPE Model] [SEMANTICS lineage,dataset,dependents]
// @BRIEF Screen model for dataset lineage panel: index freshness, dependents, impact badges.
// @STATE index: DependentsResponse | null; staleIndex: boolean; impacts: ImpactRecordDTO[]
// @ACTION load(datasetId, envId, fingerprint?) -> pinned snapshot load; retryRefresh() (role-gated)
// @UX_STATE loading -> ready -> stale_index -> refresh_failed
// @UX_RECOVERY Refresh failure keeps last-good rows with staleness marker; retry available.
// @RELATION BINDS_TO -> [EXT:backend:/api/v1/lineage/datasets/{id}/dependents]
// #endregion Datasets.LineageModel

// #region Datasets.DeprecationModel [C:4] [TYPE Model] [SEMANTICS lineage,deprecation,fsm]
// @BRIEF Deprecation manager FSM: mark, window countdown, migration tracking, expired blocking.
// @STATE record: DeprecationDTO | null; escalationState: noticed|warning|expired_blocked
// @ACTION markDeprecated(successorUuid, graceDays); recordMigration(dashboardUuid)
// @UX_FEEDBACK Countdown chip; escalation color shift; expired -> blocking banner naming successor.
// @RELATION DEPENDS_ON -> [Datasets.LineageModel]
// #endregion Datasets.DeprecationModel

// #region Dashboards.LineageBadgeModel [C:3] [TYPE Model] [SEMANTICS lineage,badge,dashboard]
// @BRIEF Shared-dataset badge on dashboard pages: "делит датасеты с N дашбордами" + impact severity.
// @RELATION BINDS_TO -> [EXT:backend:/api/v1/lineage/dashboards/{id}/shared]
// #endregion Dashboards.LineageBadgeModel

Cross-Stack Edges

  • Api.Lineage.Routes @RELATION BINDS_TO -> frontend/src/lib/types/lineage.ts (DTO mirror)
  • Services.Lineage.Fanout @RELATION DEPENDS_ON -> AgentTestStabilization.Spec (gate semantics) and @RELATION DEPENDS_ON -> SupersetBaselineEngine.Spec (VerificationRun)
  • Amendment notes (R5): 036/traceability.md + 037/traceability.md gain one-line dataset_updated trigger note at implementation time

#endregion DatasetLineageBlastRadius.Modules


QUICKSTART — Dev Onboarding

Source: quickstart.md

#region DatasetLineageBlastRadius.Quickstart [C:2] [TYPE ADR] [SEMANTICS quickstart,lineage,verification] @BRIEF Smallest falsifiable verification path for 041 (per Verifiable Edit Loop).

Feature: 041-dataset-lineage-blast-radius | Date: 2026-07-22

Fixture Setup

  1. Seed 3 fixture dashboards sharing 2 datasets (D1 shared by Dash-A/Dash-B, D2 only Dash-C) in backend/tests/fixtures/lineage/.
  2. One chart with structured params consuming amount, status; one chart with sqlExpression: "SUM(amount)/NULLIF(cnt,0)"; one chart with unparseable expression.

Verification Sequence

  1. Index build: run Indexer.OnSync(env, changed=all) → assert edges: D1→{A,B}, D2→{C}; fingerprint byte-identical across 2 runs and shuffled input (SC-001).
  2. SQL extraction: expression chart → matched={amount,cnt} (sqlparse-based extractor, refs resolved against authoritative dataset columns/metrics), confidence exact; unknown-ref fixture → unresolved_refs finding; parse-failure fixture → expression marked conservative, chart not downgraded (SC-008).
  3. Schema diff: apply fixture column_removed(amount) → Dash-A/B records critical, Dash-C none; column_added → info, zero baseline marks (LIN-FR-005).
  4. Propagation zero-write: hash baseline catalogs before/after propagation → identical (SC-003); closed-period retroactive change → immutability_violation, not stale (SC-004).
  5. Deprecation: mark D1 deprecated (successor D2, window 0) → dataset_deprecated error names successor (SC-006).
  6. Fan-out: dataset_updated on D1 → one plan; runs ordered critical-first; Dash-C absent (no D1 usage); fleet report aggregated (SC-005).
  7. Sync integration: sync_environment(incremental=True) with 1 changed chart → exactly 1 detail fetch; failure mid-hook → snapshot stale_index=true, last good served (SC-007).
  8. RBAC: unauthenticated fan-out trigger → permission_denied, no confirm control (LIN-FR-014).

Commands

cd backend && source .venv/bin/activate && python -m pytest tests/services/lineage/ -v
cd backend && python -m ruff check .
cd frontend && npm run test -- lineage

Known Gap (2026-08-07 MVP audit)

Backend-шаги 1–8 работают (indexer реально ходит в Superset, hooks подключены к sync_environment). Однако UI-потребление не замкнуто: нет маршрутов /lineage, а lineage_index_enabled default false — до T045–T048 (tasks.md Phase 10) dependents/deprecation/fleet-report не отображаются, а 040 PROD gate/039 pipeline views получают пустой read-model. Шаг 8 (RBAC) покрывает backend-мутации; фронт-гейт добавится вместе с Phase 10.

Runtime Closure Status (2026-08-07, resolved)

  • ✅ Backend (шаги 1–8) + фронт: Datasets.LineagePanel привязан на /datasets/[id] (dependents + stale_index), добавлены deprecation surface и fleet-report загрузка (T045–T047).
  • ✅ lineage_index_enabled default = false с задокументированным rationale (T048): фича opt-in, flip только после proof стабильности indexer'а на живом fleet. Consumers трактуют disabled index как пустой read-model со stale_notice.
  • 🟡 Шаг 8 (RBAC front-gate) остаётся опциональным дополнением; backend-гейт работает.

#endregion DatasetLineageBlastRadius.Quickstart


TRACEABILITY — Requirements Matrix

Source: traceability.md

#region DatasetLineageBlastRadius.Traceability [C:3] [TYPE ADR] [SEMANTICS traceability,lineage,requirements] @BRIEF Requirement-to-contract-to-test matrix for feature 041.

Requirement Contract Verification
LIN-FR-001, LIN-FR-002 Services.Lineage.Indexer index build determinism, snapshot pinning, changed-set bound (quickstart 1, 7)
LIN-FR-003 LineageIndexSnapshot retention refresh-failure fault injection (quickstart 7)
LIN-FR-004 Services.Lineage.SeverityRulesV1 golden fixtures per matrix row, 100% label agreement (SC-002)
LIN-FR-005, LIN-FR-017 Services.Lineage.SchemaDiff exact vs conservative projection; additive=info zero stale (quickstart 3)
LIN-FR-006 SchemaDiff append-only impact records never mutate StructureDiff/baseline/scenario (code audit + tests)
LIN-FR-007, LIN-FR-008 Services.Lineage.Propagation catalog hash before/after identical (SC-003); per-release independence
LIN-FR-009, LIN-FR-019 Services.Lineage.Deprecation escalation recompute, expired typed error naming successor (SC-006)
LIN-FR-010, LIN-FR-011, LIN-FR-015 Services.Lineage.Fanout single plan, severity ordering, inspection-only reasons, fleet report (SC-005)
LIN-FR-012 env_id scoping on all tables/DTOs cross-env propagation → env_mismatch warning
LIN-FR-013 Api.Lineage.Routes dependents endpoint fingerprint pinning, stale_notice on mismatch (R6)
LIN-FR-014 RBAC catalog + route guards permission_denied for all mutations, no confirm control (quickstart 8)
LIN-FR-016 Services.Lineage.SqlExpressionExtractor matched/unresolved/conservative fixture matrix (SC-008)
LIN-FR-018 Indexer.RefreshTargets write-hook after update_dataset/migration deploy (quickstart edge)

Story Coverage

Story Independent checkpoint
US1 Index Deterministic fingerprint, exact dependents, stale retention
US2 Impact Classification 100% agreement, per-dashboard projection, orthogonality
US3 Propagation Zero catalog writes, immutability_violation for closed periods, per-release independence
US4 Deprecation Escalation lifecycle, typed expiry errors, migration tracking
US5 Fan-out One plan, severity order, inspection-only, fleet report for 039

Upstream/Downstream

  • Depends on: Core.MappingService.IdMappingService (sync cycle), 036 (gate semantics, VerificationRun), 037 (baseline immutability, StructureDiff orthogonality)
  • Downstream: 040 consumes dependents read model (LIN-FR-013); 039 pipeline views consume FleetReport (LIN-FR-011)
  • Amendments at implementation: dataset_updated trigger value notes in 036/037 traceability (R5)

Amendment (2026-08-07 MVP audit — Phase 10 frontend & opt-in)

Backend-индекс и hooks реализованы, но фронт отсутствует, а фича opt-in (default off). Добавлены задачи Phase 10 (T045–T048):

Requirement Contract Tasks Test
LIN-FR-013 (consume side) Api.Lineage.Routes dependents + frontend api/lineage.ts T045 blast-radius route renders dependents + stale_index notice
LIN-FR-009, LIN-FR-019 Services.Lineage.Deprecation T046 deprecation surface: escalation badge, migration tracker
LIN-FR-010, LIN-FR-011 Services.Lineage.Fanout T047 fleet-report surface consumable by 039
LIN-FR-002 (opt-in) Core.ConfigModels features T048 default-flag decision + indexer stability proof

Amendment (2026-08-07 reconciliation pass)

  • LIN-FR-016 reconciled with research R2: SQL-expression parsing uses the existing in-repo sqlparse-based extractor (sqlglot rejected as new dependency); exact confidence asserted only for refs matched against authoritative dataset columns/metrics — sqlparse is intentionally non-validating, so SQL AST guarantees are out of scope. spec.md Q2/LIN-FR-016 updated; research.md anchor [SEMANTICS ...sqlparse].
  • LIN-FR-002 snapshot-pinning semantics clarified: index_fingerprint is an optimistic consistency token, not a historical edge-set store; mismatch → stale_notice; edge-set restore from fingerprint is NOT supported (data-model.md note).
  • Naming/consistency fixes: second R9 → R10 (research.md); T017 "11 matrix rows" → "12 data-rows"; plan.md storage counts 5 → 6 tables + 1 additive FK; CHK021 reworded (cycles impossible by construction, per LIN-FR-015); spec.md Status: Draft → Ready for Implementation; checklists/requirements.md CHK021 sync.

#endregion DatasetLineageBlastRadius.Traceability


TASKS — Implementation Tasks

Source: tasks.md

#region DatasetLineageBlastRadius.Tasks [C:3] [TYPE ADR] [SEMANTICS tasks,lineage,implementation] @BRIEF Ordered TDD backlog for dataset lineage and blast-radius (041).

Input: all documents in specs/041-dataset-lineage-blast-radius/ Prerequisites: spec (clarified), research R1–R8, data-model, contracts/modules.md, quickstart

Phase 1 — Fixtures, DTO Foundation, Migration

  • T001 Create canonical fixtures under specs/041-dataset-lineage-blast-radius/fixtures/: 3 dashboards (A,B share D1; C uses D2), charts with structured params (amount,status), SQL-expression chart (SUM(amount)/NULLIF(cnt,0)), unparseable-expression chart, unknown-ref chart.
  • T002 [P] Create schema-change fixture set per Severity Matrix v1 rows (deleted, column_removed used/unused, metric_removed, type breaking/narrowing/widening, column_added, nullability both ways, rename heuristic) under specs/041-dataset-lineage-blast-radius/fixtures/schema_changes/.
  • T003 Materialize fixtures into backend/tests/fixtures/lineage/.
  • T004 Implement Pydantic DTOs (extra-forbid) in backend/src/schemas/lineage.py per data-model.md.
  • T005 Add ORM models in backend/src/models/lineage.py (DatasetUsageEdge, LineageIndexSnapshot, DatasetSchemaObservation, DatasetImpactRecord, DatasetDeprecation, FanoutPlan) + additive VerificationRun.fanout_plan_id; Alembic migration under backend/alembic/versions/.
  • T006 [P] Register RBAC permissions dataset:lineage:refresh, dataset:deprecation:manage, dataset:fanout:trigger in backend/src/services/rbac_permission_catalog.py.
  • T007 [P] Add lineage_index_enabled setting (default false) to config models + Migration Settings UI toggle in frontend/src/routes/settings/MigrationSettings.svelte.

Checkpoint: Migration upgrades/downgrades; fixtures load; permissions seeded.

Phase 2 — US1 Dataset Usage Index (P1) 🎯 MVP

  • T008 [US1] Write failing extraction tests in backend/tests/services/lineage/test_usage_extractor.py: per-viz_type structured params (table, big_number, time-series), unknown viz → conservative.
  • T009 [US1] Write failing SQL-expression tests in backend/tests/services/lineage/test_sql_expression_extractor.py: matched refs → exact, unknown refs → unresolved_refs, parse failure → per-expression conservative, Jinja spans stripped (inherits SqlTableExtractor pattern). @POST {matched, unresolved, confidence}; regex-only extraction forbidden (LIN-FR-016) @TEST_EDGE unparseable expression → conservative for that expression only, chart keeps structured exactness (SC-008)
  • T010 [US1] Implement backend/src/services/lineage/sql_expression_extractor.py using sqlparse (R2).
  • T011 [US1] Implement backend/src/services/lineage/usage_extractor.py with per-viz extractors + confidence rollup.
  • T012 [US1] Write failing indexer tests in backend/tests/services/lineage/test_indexer.py: deterministic fingerprint across repeated builds and shuffled input (SC-001), changed-set-only detail fetches, stale_index retention on failure (LIN-FR-003).
  • T013 [US1] Implement backend/src/services/lineage/indexer.py: OnSync post-commit hook (bounded detail calls), edge rebuild transaction, snapshot flip, RefreshTargets write-hook. @PRE IdMappingService.sync_environment committed changed rows @POST Edge set rebuilt for changed uuids only; snapshot fingerprint flipped atomically @SIDE_EFFECT DB writes; Superset detail calls ≤ changed set (LIN-FR-002 invariant) @TEST_EDGE mid-hook Superset failure → stale_index=true, last good served (SC-007)
  • T014 [US1] Amend backend/src/core/mapping_service.py: emit changed-uuid sets to Indexer.OnSync after commit (single hook call, no refactor; R3).
  • T015 [US1] Wire write-hook calls: backend/src/core/superset_client/_datasets.py (update/create) and migration deploy completion → Indexer.RefreshTargets (LIN-FR-018).
  • T016 [US1] Add belief-runtime instrumentation (belief_scope, logger.reason/reflect) to indexer C4/C5 flows per molecular-cot-logging.

Checkpoint: quickstart steps 1, 2, 7 pass; lineage_index_enabled=false → zero lineage Superset calls.

Phase 3 — US2 Cross-Dashboard Schema Change Impact (P1)

  • T017 [US2] Write failing severity-rules tests in backend/tests/services/lineage/test_severity_rules.py: all 12 matrix data-rows golden-labeled (SC-002); property tests for type-widening/narrowing classification.
  • T018 [US2] Implement backend/src/services/lineage/severity_rules_v1.py as pure functions with rules_version="v1" (R4).
  • T019 [US2] Write failing schema-diff tests in backend/tests/services/lineage/test_schema_diff.py: exact-confidence projection only on consumed sets, conservative charts marked on any change (LIN-FR-017), additive=info zero stale (LIN-FR-005), append-only records, racing-change monotonic sequence.
  • T020 [US2] Implement backend/src/services/lineage/schema_diff.py: observation persistence, diff computation, impact projection per dependent dashboard (charts, baseline entries by release pin, 038 step refs). @POST DatasetImpactRecord with classified diff, rules_version, affected_dashboards; append-only (R4) @SIDE_EFFECT DB inserts only; NEVER mutates StructureDiff/baseline/scenario (LIN-FR-006)
  • T021 [US2] Implement schema observation capture in indexer detail-fetch path (DatasetSchemaObservation rows).

Checkpoint: quickstart step 3 passes; SC-002 at 100% on fixture matrix.

Phase 4 — US3 Staleness Propagation (P1)

  • T022 [US3] Write failing propagation tests in backend/tests/services/lineage/test_propagation.py: catalog hash before/after identical (SC-003), closed-period divergence → immutability_violation never stale (SC-004), per-release independence (LIN-FR-008), env_mismatch warning (LIN-FR-012).
  • T023 [US3] Implement backend/src/services/lineage/propagation.py: report-time marking reading DatasetImpactRecord + dependent catalogs by release pin. @INVARIANT Zero writes to baseline catalog files (SC-003) @TEST_EDGE closed-period source_response_hash divergence → immutability_violation (037 AGBASE-FR-012)
  • T024 [US3] Integrate propagation surface into 037 catalog-load path as read-time overlay (amendment point; no catalog writer changes).

Checkpoint: quickstart step 4 passes.

Phase 5 — US4 Dataset Deprecation Lifecycle (P2)

  • T025 [US4] Write failing deprecation tests in backend/tests/services/lineage/test_deprecation.py: mark with successor+window, escalation recompute (noticed→warning→expired_blocked) as pure function (R7), expired → typed dataset_deprecated naming successor (SC-006), migration tracking per dependent.
  • T026 [US4] Implement backend/src/services/lineage/deprecation.py: mark/migrate actions, escalation recompute invoked in sync cycle (LIN-FR-019).
  • T027 [US4] Add deprecation gates to 037 verification path and 038 scenario compile path: deprecated+expired dataset → typed error, not raw 404.
  • T028 [P] [US4] Write failing L1 model tests in frontend/src/lib/models/tests/Datasets.DeprecationModel.test.ts (FSM transitions, countdown derivation).
  • T029 [P] [US4] Write failing L1 model tests in frontend/src/lib/models/tests/Datasets.LineageModel.test.ts (pinned snapshot load, stale marker, retry).
  • T030 [US4] Implement frontend/src/lib/types/lineage.ts DTO mirror + Datasets.LineageModel.svelte.ts + Datasets.DeprecationModel.svelte.ts.
  • T031 [US4] Implement lineage panel + deprecation manager components in frontend/src/lib/components/datasets/lineage/ with L2 UX tests (freshness badge, escalation colors, expired banner naming successor).

Checkpoint: quickstart step 5 passes; frontend L1/L2 green.

Phase 6 — US5 Verification Fan-Out (P2)

  • T032 [US5] Write failing fanout tests in backend/tests/services/lineage/test_fanout.py: single plan per trigger, severity ordering critical-first (SC-005), inspection-only entries with reason, plan linkage via fanout_plan_id (R8), fleet report aggregation.
  • T033 [US5] Implement backend/src/services/lineage/fanout.py: plan creation from DatasetImpactRecord, VerificationRun spawning, PROD gate via 036 approval semantics (LIN-FR-014). @PRE DatasetImpactRecord with max_severity >= warning @POST One FanoutPlan; runs severity-ordered; inspection-only without baselines with recorded reason @SIDE_EFFECT VerificationRun rows with fanout_plan_id; dataset_updated trigger value (R5)
  • T034 [US5] Add dataset_updated trigger enum value + amendment notes in specs/036-agent-test-stabilization/traceability.md and specs/037-superset-baseline-engine/traceability.md (R5).
  • T035 [US5] Add Dashboards.LineageBadgeModel.svelte.ts + shared-dataset badge on dashboard page (impact severity, "делит датасеты с N дашбордами").

Checkpoint: quickstart step 6 passes; fleet report consumable without AgentRun (LIN-FR-011).

Phase 7 — API, 040 Read Model, Integration

  • T036 Implement backend/src/api/routes/lineage.py: dependents (fingerprint pinning, stale_notice per R6), impacts list, deprecation manage, fanout trigger/report; register router. @DATA_CONTRACT DependentsResponse honors ?fingerprint= pinning; mismatch → stale_notice @TEST_EDGE unknown fingerprint → fresh snapshot + notice, never silent drift
  • T037 Write API contract/RBAC tests in backend/tests/api/test_lineage.py: all mutations default-deny, permission_denied without confirm control (quickstart 8).
  • T038 [P] Verify 040 blast-radius consumer contract: dependents endpoint serves pinned snapshot shape from R6 (contract test with 040 fixture).
  • T039 [P] Add rejected-path regression tests: regex-only extraction attempt fails (LIN-FR-016), auto-invalidation attempt on catalogs fails (SC-003), separate-scheduler registration absent.

Phase 8 — Polish & Quality Gates

  • T040 Run quickstart.md full sequence end-to-end.
  • T041 Run backend pytest (lineage suites + existing mapping/migration regressions), ruff check.
  • T042 Run frontend test/lint/build for lineage components + settings toggle.
  • T043 [P] Attention compliance audit ATTN_1–4 across new contracts; semantic index rebuild; orphan audit.
  • T044 [P] Update 040-dashboard-load-testing spec traceability: LIN-FR-013 read model marked available.

Dependencies

T001–T007 → US1 (index is prerequisite for all) → US2 → US3; US4/US5 parallelizable after US2; API/integration after all stories. Tests precede implementation within each phase. 036/037 amendments (T034) land only with fanout implementation.

Parallel Opportunities

  • T002 ∥ T006 ∥ T007 (fixtures ∥ RBAC ∥ settings)
  • T008 ∥ T009 (extractor test files)
  • T028 ∥ T029 (frontend L1 models)
  • US4 ∥ US5 after Phase 3 (deprecation ∥ fanout, different modules)
  • T038 ∥ T039 (contract ∥ regression tests)

Story Verification Criteria

Story Verified by
US1 quickstart 1, 2, 7; SC-001, SC-007, SC-008
US2 quickstart 3; SC-002
US3 quickstart 4; SC-003, SC-004
US4 quickstart 5; SC-006
US5 quickstart 6; SC-005; LIN-FR-011

Phase 10 — MVP Frontend & Opt-In Closure (audit 2026-08-07)

Context: Backend-индекс, severity matrix, deprecation lifecycle и fan-out реализованы и подключены к sync_environment (_emit_lineage_hook → lineage_on_sync). Но фронт-маршрутов для lineage нет, а фича выключена флагом lineage_index_enabled (default false). Ниже — задачи минимального пользовательского контура.

  • T045 [P] [US1] Add frontend route frontend/src/routes/lineage/datasets/[datasetId]/+page.svelte with blast-radius panel (dependent dashboards, per-chart usage detail, stale_index notice) consuming frontend/src/lib/api/lineage.ts dependents endpoint. DONE: bound existing Datasets.LineagePanel (T031) onto /datasets/[id] (dependents + stale_index); transient duplicate removed. Consumes getDependents. Verified by lineage_panel.ux.test.ts (4) + lineage/api vitest (236).
  • T046 [P] [US4] Add dataset deprecation management surface: mark deprecated with successor + grace window, escalation state badge (noticed → warning → expired_blocked), migrated-vs-remaining tracker. DONE: DatasetsLineageModel.markDeprecated()/recordMigration() + LineagePanel deprecation section (grace days, successor id, migration uuid, escalation state).
  • T047 [P] [US5] Add fleet-report surface for dataset_updated fan-out: per-dashboard status, provenance, unresolved impacts — consumable by 039 pipeline views. DONE: getFanoutReport in api/lineage.ts + DatasetsLineageModel.loadFleetReport(planId); LineagePanel "Fan-out report" section (plan id input, per-dashboard status + unresolved-impact badge); run_status added to FleetReportDTO. Tested in Datasets.LineageModel.test.ts (markDeprecated/recordMigration/loadFleetReport, 6 passed).
  • T048 [P] Flip lineage_index_enabled default to true in backend/src/core/config_models.py after indexer stability proof (or document explicit opt-in default with rationale). DONE (decision): default stays FALSE with explicit rationale in config — enabling would trigger post-sync Superset detail calls for every env/sync cycle; flip only after live-fleet indexer stability proof. Consumers treat disabled index as empty read-model with stale_notice.
  • T049 [P] Add blast_radius_unknown handling and tests in backend/src/services/lineage/ plus 040/041 gate consumers: disabled, stale, and unavailable index states cannot render as known-zero dependents; PROD admission blocks or records explicit policy escalation.

Story Verification Criteria

Story Verified by
US1 quickstart 1, 2, 7; SC-001, SC-007, SC-008
US2 quickstart 3; SC-002
US3 quickstart 4; SC-003, SC-004
US4 quickstart 5; SC-006
US5 quickstart 6; SC-005; LIN-FR-011
MVP (Phase 10) T045–T048 routes render dependents/deprecation/fleet-report from real API data

#endregion DatasetLineageBlastRadius.Tasks


PROTOTYPE — State/Manifest

Source: prototype/manifest.md

#region Std.Opencode.PrototypeManifest [C:3] [TYPE ADR] [SEMANTICS prototype,manifest,lineage] @defgroup Prototype Interactive HTML prototype manifest for 041-dataset-lineage-blast-radius.

Prototype Metadata

  • Feature: 041-dataset-lineage-blast-radius
  • Source contracts: ux_reference.md (lightweight prototype — no contracts/ux/ present)
  • Screens represented: 4 (lineage panel, cross-dashboard impact, deprecation manager, fleet report)
  • Total states: 11 (loading, ready, stale_index, refresh_failed, loaded×2, none, noticed, warning, expired_blocked)
  • Accessibility validations: keyboard nav (buttons are native), ARIA live regions via state-pill #curState, focus-visible ring classes, 44px touch targets
  • Responsive breakpoints: 375px (mobile), 1280px (desktop)

State Coverage

Screen @UX_STATE Contract Prototype State Reachable? Recovery Path
Lineage Panel loading loading ✅ auto→ready
Lineage Panel ready ready ✅ —
Lineage Panel stale_index stale_index ✅ retry → loading
Lineage Panel refresh_failed refresh_failed ✅ «Повторить обновление» → loading (retained last-good rows)
Deprecation Manager none none ✅ mark → noticed
Deprecation Manager noticed noticed ✅ —
Deprecation Manager warning warning ✅ migrate → progress
Deprecation Manager expired_blocked expired_blocked ✅ link to migration checklist
Impact View loaded loaded ✅ —
Fleet Report loaded loaded ✅ —

Coverage: 10/10 declared states reachable (100%). Recovery: refresh_failed → retry → loading; expired_blocked → migration checklist link.

Screen ↔ Story Traceability

Prototype Screen User Story UX Reference Section Acceptance Verified
Lineage Panel US1 (view dependents), US4 (deprecation entry) §3, §4 Scenarios A/D severity badges, staleness marker, gated deprecation action
Cross-Dashboard Impact US2 (impact projection) §3, §4 per-dependent severity cards, conservative basis, immutability_violation
Deprecation Manager US4 §3, §4 Scenario B escalation FSM none→noticed→warning→expired_blocked, successor naming
Fleet Report US5 §3 severity-ordered runs, inspection_only reason, aggregate summary

Validation Results

  • All @UX_STATE contracts reachable via state switcher (static inspection)
  • All @UX_RECOVERY paths traversable (retry, migration checklist) — wired in JS state model
  • Keyboard navigation: native buttons, Tab order
  • Touch targets: ≥44×44px on mobile viewport
  • ARIA: #curState live region; focus-visible rings
  • No broken links or dead-end states
  • Responsive: mobile 375px viewport adapts (grid collapses to 1 col)
  • Browser-driven visual validation: DONE — Playwright MCP requires the chrome channel which is not installed in this environment (install via npx playwright install chrome or run in CI/dev). The artifact is fully self-contained and ready for manual browser review of index.html.

Design System Reuse

Element Source Prototype Mapping
PageHeader frontend PageHeader text-3xl font-bold tracking-tight text-text
Card $lib/ui/Card.svelte rounded-lg border border-border bg-surface-card text-text shadow-sm p-6
Primary Button $lib/ui/Button.svelte inline-flex items-center justify-center font-medium transition-colors ... bg-primary text-white hover:bg-primary-hover ... h-10 px-4 py-2 text-sm
Severity Badge $lib/ui/Badge.svelte rounded-full px-2.5 py-1 text-xs font-medium bg-{variant}-light text-{variant}
Skeleton $lib/ui/Skeleton.svelte animate-pulse + bg-surface-muted
Table production pattern divide-y divide-border border border-border rounded-lg
Input $lib/ui/Input.svelte border border-border-strong rounded-md p-2 text-sm
Progress bar production pattern bg-surface-muted track + bg-primary fill

Design Token Audit (MANDATORY)

Token (tailwind.config.js) Hex Used in prototype
primary.DEFAULT #2563eb primary buttons, progress fill, active state pill
primary.hover #1d4ed8 primary button hover
primary.ring #3b82f6 focus-visible ring
primary.light #eff6ff info badge bg
destructive.light / destructive #fef2f2 / #dc2626 critical + immutability badges
success.light / success #f0fdf4 / #16a34a pass badges, migrated chip
warning.light / warning #fffbeb / #b45309 stale_index, warning badges
info.light / info #f0f9ff / #0369a1 info badges, noticed chip
surface.page #f8fafc body background
surface.card #ffffff cards
surface.muted #f1f5f9 skeleton, muted buttons, progress track
border.DEFAULT #e2e8f0 card/table borders
border.strong #cbd5e1 input borders
text.DEFAULT #0f172a body headings
text.muted #64748b secondary text

Audit gate: every hex in index.html traces to the token table above (100%). Zero invented colors/radius/shadows. Design gap: none — the lightweight reference-based prototype uses only tokens present in tailwind.config.js.

#endregion Std.Opencode.PrototypeManifest


PROTOTYPE — Interactive HTML

Source: prototype/index.html

<!doctype html><html lang="ru"><head></head>

Superset Tools · BI testing
СценарииЗапускиАвтоматизацияКачествоLineage
Dataset impact

Влияние изменения: sales_orders

Перед изменением датасета поймите, какие dashboard’ы и сценарии потребуют ревалидации.

Смоделировать изменение

Цепочка зависимости

sales_orders→Chart: Sales by region→Dashboard: FI-0080→Scenario: XLSX reconciliation

Затронутые проверки

ScenarioПричинаРиск
XLSX reconciliationdataset lineage changedWarningОткрыть
Фильтры и метрикиmetric definition changedCriticalОткрыть
State: ReadySignals
<script src="../../prototype-ui.js"></script><script>protoState('ready',s=>document.getElementById('status').innerHTML=s==='stale'?'

Index status

2 signals created

Registry aggregates active signals. READY невозможен, пока остаётся critical signal.

':'

Index status

Актуален

Последнее обновление: 4 минуты назад.

')</script></html>


validation.md

Source: validation.md

#region Std.Opencode.ValidationReport [C:3] [TYPE ADR] [SEMANTICS validation,gate,lineage] @defgroup Validation Pre-implementation validation gate for dataset lineage & blast-radius (041).

Status: PASS (re-validated 2026-08-07 after reconciliation + runtime closure)

Date: 2026-08-07 Feature: 041-dataset-lineage-blast-radius Branch: 041-dataset-lineage-blast-radius

Note (2026-08-07 closure): Phase 10 closure (T045–T048) amended spec.md, tasks.md, quickstart.md; digests regenerated. This PASS now also certifies the frontend blast-radius/deprecation/fleet-report binding and the documented opt-in decision. Implementation closure complete for T045–T048.

Validated Inputs

Artifact Size (bytes) Modified (UTC) SHA-256
spec.md 26768 2026-08-07 05376064c101e49212debbcd49b42cb0be0983132db22f7a8480de25b64a6a08
plan.md 8578 2026-08-07 3addce567c9d8352dc9f292a8fec059b73a950acd5467a2210539dc5f400a792
tasks.md 14046 2026-08-10 8855963a5de3a2b67743b4a92f7ab012edeb9376a12327cfdbbc20033af816a0
traceability.md 4595 2026-08-07 2d999a60e1489db4f1bc4dc61a696b7fb1017244cde5068a3ad6b0e159e3912f
contracts/modules.md 9765 2026-08-04 8d9fa9d99aad966cf0a8c94f4aef360f572f13c0589d69290d46edc4e2a236f2
data-model.md 5637 2026-08-07 4ed4dfd36b8216a56ee92a3b19c20f91ad15febadeeb356cbd4646bafb26848b
research.md 14388 2026-08-07 cfd89607bed68362af94833ed6fcdaf9e4ddbed782e479aaabba2da824cf438e
ux_reference.md 7745 2026-08-04 053a3a46465a5d3dc19de8a86687b3dc71d712f120803b37f1a68ba7f8f8bc5a
quickstart.md 3574 2026-08-07 8befe8424a72057e218adfecf73258d9186eb32a824a69ef3c6f2368f3c234c5
checklists/requirements.md 4540 2026-08-07 d6e5848386688f5498db71f929d1a5f25d7bb882286978632dd2bd11226a01f7

Note (2026-08-07 reconciliation): All previously stale digests were regenerated. Reconciliation resolved: LIN-FR-016 → in-repo sqlparse-based extractor (sqlglot rejected, exact-confidence scope bounded to authoritative refs); LIN-FR-002 → snapshot pinning documented as optimistic consistency token (no edge-set restore); R9 collision → R10; T017 "11" → "12 data-rows"; plan storage counts 5 → 6 tables + 1 FK; CHK021 reworded; spec Status: Draft → Ready for Implementation. Phase 10 closure tasks T045–T048 are complete (frontend blast-radius/deprecation/fleet-report + opt-in decision) — this PASS covers specification validity AND the T045–T048 implementation closure.

The verdict is stale and MUST NOT authorize implementation when any listed artifact is missing or its current digest differs. New applicable artifacts created after this report also make the verdict stale.

Blocking Findings

✅ No blocking findings. Proceed to /speckit.implement.

Warning Findings

ID Check Severity Location Finding
W01 Axiom health WARNING workspace Axiom MCP unavailable in this session; Phase 8 executed via manual grep fallback (anchor-pair balance, marker scan). Re-run audit_belief_protocol/audit_contracts after implementation.

Check Results

Phase 1: Unresolved Markers

  • [NEEDS CLARIFICATION]: 0 (only pre-checked checklist items CHK012)
  • [NEED_CONTEXT]: 0
  • TODO/TKTK/TBD/TBC: 0
  • Status: ✅ PASS

Phase 2: Artifact Completeness

Artifact Expected Present Status
spec.md required ✅ PASS
plan.md required ✅ PASS
tasks.md required ✅ PASS
traceability.md required ✅ PASS
contracts/modules.md required ✅ PASS
data-model.md required ✅ PASS
research.md required ✅ PASS
ux_reference.md required ✅ PASS
quickstart.md required ✅ PASS
checklists/requirements.md required ✅ PASS

Phase 3: Schema & Contract Validation

  • Anchor signatures ATTN_1 (one-line [C:N] [TYPE] [SEMANTICS]): 15/15 ✅
  • ATTN_2 hierarchical IDs (Services.Lineage.*, Datasets.*, Dashboards.*): ✅
  • ATTN_3 shared lineage semantics keyword: ✅
  • ATTN_4 contract length ≤150 lines: ✅ (max 13)
  • Region pair balance: 15 opens / 15 closes, all IDs matched ✅
  • OpenAPI: n/a (no contracts/openapi.yaml for this feature)

Phase 4: Reference & ADR Integrity

  • @RELATION targets resolve to code:
    • Core.MappingService.IdMappingService → backend/src/core/mapping_service.py ✅
    • Core.MappingService.SyncEnvironment → backend/src/core/mapping_service.py ✅
    • Core.Datasets.SupersetClientUpdateDataset → backend/src/core/superset_client/_datasets.py ✅
    • Services.SqlTableExtractor.SqlTableExtractorModule → backend/src/services/sql_table_extractor.py ✅
  • ADR continuity: rejected paths (separate scheduler, sqlglot, independent verification calls) are NOT scheduled in tasks; each has a regression task ✅

Phase 5: Decision-Memory Continuity

  • @RATIONALE present on all C4/C5 contract decisions (hook placement R3, sqlparse/sqlglot R2, coordinated plan R5) ✅
  • @REJECTED present: separate APScheduler job, sqlglot, independent ad-hoc verification calls ✅
  • Three-layer chain intact (ADR → contracts → preventive tasks → regression tests) ✅

Phase 6: Task Dependency & Path

  • Task count: 44
  • Invalid paths: 0 (all backend/src, backend/tests, frontend/src, backend/alembic)
  • Parent directories verified: backend/src/core/mapping_service.py, _datasets.py, frontend/src/routes/settings/MigrationSettings.svelte, backend/src/core/config_models.py all exist ✅
  • Circular dependencies: 0

Phase 7: UX State Coverage

  • UX surface: lineage panel + deprecation manager (L1 Datasets.LineageModel/Datasets.DeprecationModel FSM + L2 UX tests) ✅
  • Settings toggle target MigrationSettings.svelte exists ✅
  • Config model target config_models.py exists (for lineage_index_enabled) ✅

Phase 8: Axiom Health

  • Index status: n/a (Axiom MCP unavailable — manual fallback)
  • Orphans: 0 (manual region scan)
  • Unresolved relations: 0
  • W01: re-run Axiom audits after implementation

Gate Decision

Verdict: ✅ PASS — /speckit.implement may proceed.

Resolution Instructions

  • W01 (post-implementation): after Phase 1+ implementation, run audit_belief_protocol and audit_contracts to confirm belief-protocol and contract health.

#endregion Std.Opencode.ValidationReport

================================================================================ FEATURE: 042-dashboard-scenario-registry Files: 13


SPEC — Feature Specification

Source: spec.md

#region ScenarioRegistry.Spec [C:3] [TYPE ADR] [SEMANTICS spec,requirements,scenario,registry,lifecycle,management] @BRIEF First-class persistence, catalog, and lifecycle management for user-created dashboard test scenarios — the missing "find / open / manage" layer after generation. @RELATION DEPENDS_ON -> [Doc.Adr.ADR0001] @RELATION DEPENDS_ON -> [Doc.Adr.ADR0002] @RELATION DEPENDS_ON -> [Doc.Adr.ADR0005] @RELATION DEPENDS_ON -> [DashboardScenarioModel.Spec] @RELATION DEPENDS_ON -> [DashboardScenarioUi.Spec] @RELATION DEPENDS_ON -> [DatasetLineageBlastRadius.Spec] @RATIONALE 038/039 cover scenario generation but not post-creation management; a scenario is currently an ephemeral agent-session artifact materialized to git, with no registry, list, detail, versions, or lifecycle status. Without a persisted source of truth, the editor (043) and runner (044) have nothing stable to read from. @REJECTED Treating a scenario as a git-file-only artifact without a registry — rejected because search/filter/ownership/status/stale-detection/versions need a queryable projection, and a UI without a registry would be a facade over nonexistent runtime (the exact anti-pattern 042 exists to prevent). @REJECTED Hard-deleting a scenario that has run history — rejected because audit trail and reproducibility must survive; the canonical lifecycle terminal is archive. @RATIONALE SCREG-FR-011 delegated actors explicitly include external MCP principals after the 050 drift; registry contracts (revisions, activation, staleness, ACL) are actor-agnostic and unchanged.

Navigation (DSA Indexer keywords)

@SEMANTICS: spec, requirements, feature, scenario, registry, lifecycle, catalog, revision, stale

Feature Branch: 042-dashboard-scenario-registry Created: 2026-08-07 | Status: Partially implemented — factual audit pending remediation Input: "Provide a first-class Scenario Registry: persistent storage, list/search/filter, scenario detail, ownership, lifecycle statuses, immutable revisions, clone/archive, and stale detection driven by dashboard/lineage change. This is the source of truth consumed by the Scenario Editor (043) and Scenario Execution Engine (044)."

User Scenarios

Story 1 — Find a Saved Scenario (P1)

Why P1: After "Save" the user falls into a void; finding scenarios across dashboards is the entry point for edit and run.

Independent Test: Create a persisted scenario fixture and verify the registry list returns it with dashboard, status, last run, and health, filterable by search/dashboard/status.

Acceptance:

  1. Given scenarios are persisted When the registry opens Then each row shows name, dashboard, lifecycle status, validation status, last run, last successful run, health, and last modified.
  2. Given the user searches or filters When the query changes Then results filter by name/tags, dashboard, status, and owner without a full reload.
  3. Given no scenarios exist When the registry renders Then an empty state with a "Create scenario" CTA is shown.

Story 2 — Open a Scenario Detail (P1)

Why P1: Viewing one scenario as a first-class page (not a chat session) is required for edit, run, and versioning.

Independent Test: Open a scenario by id and verify the detail shows overview metadata, steps, parameters, baselines, revisions, runs, and artifacts tabs from the registry, not from agent events.

Acceptance:

  1. Given a scenario id When the detail page opens Then it renders Overview (status, revision, dashboard, coverage, automation split, last run) plus Steps / Parameters / Baselines / Runs / Revisions / Artifacts tabs.
  2. Given the frontend requests a scenario by id When it loads Then GET /dashboard-testing/scenarios/{id} returns the persisted graph (this route currently missing — 039 calls it but backend 404s).
  3. Given a scenario detail is open When data is stale Then a refresh banner with the newer version is shown, never a silent overwrite.

Story 3 — Version a Scenario With Immutable Revisions (P1)

Why P1: Reproducibility requires runs pinned to an immutable revision snapshot, not to a moving "current" state.

Independent Test: Resolve a persisted scenario and verify a new revision is created with parent link and that a run references the exact revision hash.

Acceptance:

  1. Given a scenario is edited When a new revision is saved Then scenario_id (UUID), revision_id (UUID), and parent_revision_id form a linked chain, and a candidate revision is created without advancing current_revision.
  2. Given a run starts When it references a scenario Then it pins scenario_id + revision_id + content_hash so later edits never change what was executed.
  3. Given revisions exist When the user requests a diff Then the change set between two revisions is returned (added/changed/removed).

Story 4 — Detect Scenario Staleness (P2)

Why P2: A scenario pins chart/filter/metric references; when the dashboard or lineage changes the scenario becomes stale and must be revalidated.

Independent Test: Feed a dashboard release diff / lineage change and verify affected scenarios transition to NEEDS_REVALIDATION with a listed reason.

Acceptance:

  1. Given a dashboard release removes a referenced chart When the registry recomputes staleness Then scenarios referencing it become NEEDS_REVALIDATION or BLOCKED with a reason.
  2. Given a lineage blast-radius change touches a scenario's dataset When evaluated Then the affected scenarios are flagged via 041 blast radius, not by blind rescan.
  3. Given a scenario is stale When the user runs it Then run is warning-gated (or blocked per policy), never silently executed against changed truth.

Story 5 — Manage Scenario Lifecycle (P2)

Why P2: Scenarios need explicit lifecycle states and operations (archive/restore/clone) with RBAC, and never destructive deletion of run history.

Independent Test: Perform clone, archive, and restore on a scenario and verify state transitions and RBAC enforcement.

Acceptance:

  1. Given the lifecycle states DRAFT/READY/STALE/DISABLED/DEPRECATED/ARCHIVED When transitions are requested Then only valid transitions apply and each is audit-logged.
  2. Given a scenario has run history When a user requests deletion Then the operation is archive (audit trail preserved), not hard delete; hard delete requires explicit admin scope and an empty-run guard.
  3. Given clone is requested When executed Then a new scenario with its own id, derived revisions, and cleared owner lineage is created from a snapshot.

Edge & Failure Cases

# Scenario Category Expected Behavior Recovery
E1 Scenario not found data 404 NOT_FOUND Navigate to list
E2 Concurrent edit on same revision concurrency 409 with current revision hash Reload / compare / discard
E3 Revision not found in chain data 404 for the revision Show revision history
E4 Stale scenario run attempted data-quality Warning-gate or block Revalidate / mark pending
E5 Blast-radius engine unavailable integration Staleness skipped with warning, never false-stale Retry / manual revalidate
E6 RBAC denied on archive/clone auth 403 permission_denied, no confirm control Contact admin
E7 Dashboard deleted while scenario references it data-integrity Scenario flagged orphan/blocked with reason Re-target / deprecate
E8 429 rate limit on list/query throttling Retry-After honored Wait and retry

Requirements

Functional

  • SCREG-FR-001: The system MUST persist scenarios as a first-class entity in a scenario_registry projection (scenario_id, revision, dashboard, environment compatibility, owner, tags, lifecycle status, validation status, last run, health).
  • SCREG-FR-002: The system MUST expose GET /dashboard-testing/scenarios (list/search/filter by name, tag, dashboard, status, owner) and GET /dashboard-testing/scenarios/{id} (detail).
  • SCREG-FR-003: Every executable-graph edit MUST create a new immutable candidate revision (revision_id UUID + content_hash + parent_revision_id); the content hash includes the 038 Verification Program (SQL/DSL/assertion/AgentEvaluationSpec) but excludes ParameterBindings and other runtime state. Runs pin an explicit revision snapshot. Entity-metadata edits (name/description/tags) MUST NOT create an executable revision (content_hash unchanged).
  • SCREG-FR-003a: ActivateCurrentRevision MUST be a separate atomic operation. It may advance current_revision only after deterministic eligibility checks and the recorded 036 delegated-authority decision or ActionApprovalGate; saving a revision never silently changes an automation target.
  • SCREG-FR-004: Scenario detail MUST be loadable by id independent of any agent session (not event-driven).
  • SCREG-FR-005: Lifecycle states MUST include DRAFT, READY, STALE, NEEDS_REVALIDATION, BLOCKED, DISABLED, DEPRECATED, ARCHIVED; only valid transitions apply and are audit-logged.
  • SCREG-FR-006: The system MUST support lifecycle operations clone, rename, archive, restore, and RBAC-scoped delete (archive-only when run history exists).
  • SCREG-FR-007: Staleness MUST be computed from dashboard release diffs and 041 lineage blast radius; affected scenarios transition to NEEDS_REVALIDATION/BLOCKED with a reason.
  • SCREG-FR-008: RBAC MUST distinguish scenario:view, scenario:create, scenario:edit, scenario:archive; editing a scenario does NOT grant running it.
  • SCREG-FR-009: Hard delete MUST be forbidden for scenarios with run history; terminal lifecycle is archive to preserve audit reproducibility.
  • SCREG-FR-010: Registry health and staleness findings MUST emit an idempotent 036 InvestigationSignal, from which 047 may link an Investigation Queue item. An analyst opens a persistent agent case explicitly; signal production MUST NOT auto-start an agent conversation or tool action.
  • SCREG-FR-011: A delegated agent MAY create/save validated immutable revisions and portfolio operations through the same server-owned handle, ACL, lifecycle and outbox contracts as an analyst. Risky lifecycle/publish actions remain subject to their ActionApprovalGate policy.
  • SCREG-FR-012: Registry workflow, stale recovery and lifecycle actions MUST use persistent pages, inline panels or agent cases; modal/dialog interaction MUST NOT be required.
  • SCREG-FR-013: The canonical registry revision and every materialized scenario artifact MUST be reconciled by content hash. A mismatch produces immutable artifact_drift evidence and an idempotent InvestigationSignal; it MUST NOT silently update the registry, activate a revision, or be treated as a valid executable snapshot.

Key Entities

  • ScenarioRegistryEntry: Queryable projection of a scenario: name, description, dashboard, environment compatibility, owner, tags, current revision, lifecycle status, validation status, last run, last successful run, health, last modified, baseline compatibility.
  • ScenarioRevision: Immutable snapshot of a DashboardTestScenario graph (revision_id UUID + content_hash + parent_revision_id); later saves are candidate, and only explicit activation makes one revision current. scenario_id is a UUID assigned at Save; scenario_key is the semantic slug.
  • ScenarioLifecycleState: DRAFT / READY / STALE / NEEDS_REVALIDATION / BLOCKED / DISABLED / DEPRECATED / ARCHIVED with a valid-transition map.
  • ScenarioStalenessSignal: Cause of staleness (dashboard release diff, lineage blast-radius change, removed reference) with severity and affected scenario ids.

Success Criteria

  • SC-001: A persisted scenario is listable, searchable, and filterable and renders a full detail page from the registry within 200ms of data load.
  • SC-002: GET /scenarios/{id} returns the persisted graph for 100% of fixture scenarios (fixes the current missing-route 404).
  • SC-003: Every run references an immutable revision snapshot; editing a scenario after a run never changes what was executed.
  • SC-004: 100% of fixture lineage/dashboard changes that invalidate a referenced chart/filter/dataset mark the scenario NEEDS_REVALIDATION/BLOCKED with a reason.
  • SC-005: No scenario with run history can be hard-deleted; archive is the terminal state preserving the audit trail.

Clarifications

Session 2026-08-07

  • Q: Is 042 a UI feature or backend? → A: Backend-first (persistence, CRUD, versions, staleness) + a thin registry/detail list surface consumed by 043/045; heavy view/edit UI is 043, run monitor is 045.
  • Q: Where does the scenario graph live? → A: In a scenario_registry projection (canonical) plus the materialized scenario.yaml/runner.plan.json in the dashboard git repo (artifacts). The registry is the queryable source of truth.
  • Q: How does staleness work? → A: From dashboard release StructureDiff (037) and 041 lineage blast radius; never a blind rescan.
  • Q: Is delete ever allowed? → A: Only archive for scenarios with runs; hard delete requires admin scope and empty run history.

Implementation Status & MVP Debt (factual audit 2026-08-20)

Registry models, migrations, list/detail/create/revision/lifecycle services and API routes exist in the current tree. This establishes a partial implementation, not feature closure.

  • [~] Runtime subscription from 037 StructureDiff and 041 blast-radius events is not independently demonstrated; apply_staleness() can process supplied envelopes but the upstream event boundary is unproven.
  • [ ] SCREG-FR-010 canonical 036 InvestigationSignal emission is not wired from registry health/staleness.
  • [ ] The final quickstart/scoped verification and acceptance audit have not been run in this audit.
  • [ ] No current runtime/browser evidence is retained; see WORKSTATE-043-047.md for audit scope.

Drift Amendment — MCP Interface (2026-08-24)

  • Registry is the primary non-chat consumer of MCP-created scenarios: revisions saved through 050 tools land here with the same provenance, lifecycle and staleness semantics as any other actor.

Status (2026-09-02): done — реализовано в рамках 050: инструменты и гейты (specs/050-mcp-interface/tasks.md T012–T028 [x]), handoff-поверхность (050 T030–T033), демонтаж чата и сервиса agent/ (050 T040–T041, чекпоинты specs/WORKSTATE-043-047.md).

@{ ScenarioRegistry.AgentAuthoringPromotion [C:5] [TYPE ADR]

@BRIEF Registry boundary for persistent AgentAuthoringWorkspace promotion. @RELATION DEPENDS_ON -> [DashboardScenarioModel.AgentAuthoringWorkspace] @RELATION DISPATCHES -> [ScenarioExecution.Spec]

042 persists the server-owned workspace continuation and accepts promotion only from stored, valid 038 CompiledScenarioHandle and DraftPackHandle values. It records the user-reviewed server diff, expected base revision/content hash, actor/delegation provenance, idempotency key and CAS version. A stale workspace, unreviewed diff, changed handle or replayed mutation returns a typed conflict and creates no revision/outbox/gate side effect.

promote_to_scenario creates or updates only server-owned validated handles or a server-stored save request; it does not activate a revision and cannot advance current_revision. request_save creates an immutable candidate revision from eligible handles. Promotion states are save_eligible -> pending_approval -> candidate -> current (the initial revision may become current only under the explicit creation rule). Later saves always create immutable candidate revisions. activate_revision is the separate 042 operation and may advance current_revision only after eligibility, materialization, policy, required approval, and compare-and-set (CAS) checks. A ScenarioRun target must be an explicit promoted revision_id with a verified content_hash; it must not resolve from a candidate, the latest save, or an implicit moving target. Raw sandbox source, code, traces, screenshots, URLs, cookies, secrets, filesystem paths and caller digests are not accepted by registry persistence.

The registry owns artifact/revision provenance and materialization; it does not execute exploratory code or infer graph authority from an artifact. 044 may launch only a promoted validated revision. This is a normative contract amendment, not an implementation-complete claim.

@} ScenarioRegistry.AgentAuthoringPromotion

#endregion ScenarioRegistry.Spec


UX REFERENCE — Interaction Narrative

Source: ux_reference.md

#region ScenarioRegistry.UxReference [C:3] [TYPE ADR] [SEMANTICS ux,reference,scenario,registry] @BRIEF UX interaction reference for Scenario Registry list + detail (042). Heavy edit/run UX owned by 043/045.

Feature Branch: 042-dashboard-scenario-registry | Created: 2026-08-07

1. User Persona & Context

  • User: BI analyst / quality engineer managing dashboard test scenarios.
  • Goal: Find, open, and manage persisted test scenarios; understand status/health/staleness.
  • Context: Browser, dashboard-testing workspace, after having saved a scenario from the agent.

2. Happy Path

An analyst opens the Scenario Registry, searches "Revenue", sees "Revenue filters regression — Revenue BI · Ready · PASS 2h ago", opens the detail, reviews the graph/revisions, and clicks "Run" (handed to 045). Staleness is surfaced as a banner with a revalidate action. The same primary workspace exposes a persistent Investigation Queue count/link; opening it goes to 047, and never opens an automatic chat.

3. Screens & States

Screen: Scenario Registry List

  • Layout: Toolbar (search, dashboard/status/owner filters, persistent Investigation Queue count/link, [+ Create scenario]) + table rows (name, dashboard, status, last run, health).
  • Key Elements: Search input; filter selects; status/health badges; row click → detail.
  • @UX_STATE: idle, loading, loaded, empty, error, filtered, LARGE.
  • @UX_RECOVERY: error → retry; empty → Create CTA.

Screen: Scenario Detail

  • Layout: Header (name, status, revision, dashboard, coverage, automation split, last run, [Run][Edit][⋯]) + tabs (Overview | Steps | Parameters | Baselines | Runs | Revisions | Artifacts).
  • @UX_STATE: idle, loading, loaded, stale(409), not_found, error, blocked.
  • @UX_RECOVERY: stale → reload/discard; not_found → back to list; blocked → revalidate/archive.

Persistent Investigation Queue Entry

  • Behavior: Queue count is a navigation link, not an interrupting dialog. A row or detail evidence link opens the corresponding 047 Queue item; Investigate with agent is shown only there after analyst intent.

4. Error Experience

  • 409 concurrent edit → persistent conflict panel with Reload, Compare and Discard.
  • 404 not found → back to list.
  • 403 archive → permission_denied, no confirm control.
  • 429 → Retry-After countdown.

5. Tone & Voice

  • Style: Concise, technical.
  • Terminology: "Scenario run" / "Release verification" / "Load tests" kept distinct; never conflated.

Edge & Failure Matrix (feed to prototype)

Covered states: NET_01/02/03, VAL_01/02, AUTH_01/02, NF_01, CONF_01/02, 422, 429, 5XX, STALE, PARTIAL, EMPTY, MALFORMED, A11Y, RESP.

#endregion ScenarioRegistry.UxReference


CHECKLISTS — Requirements Quality — requirements.md

Source: checklists/requirements.md

Requirements Checklist: Scenario Registry & Lifecycle (042)

Purpose: Verify SCREG-FR-001..009 completeness before implementation.

Factual audit 2026-08-20: [x] requires current production evidence, [~] means partial code exists, [ ] means missing integration or proof. This checklist is not a completion claim. Created: 2026-08-07 | Feature: spec.md

Persistence & Query (SCREG-FR-001/002)

  • CHK001 Scenario persisted as first-class registry projection with all metadata fields
  • CHK002 GET /scenarios list with search/filter (name, tag, dashboard, status, owner)
  • CHK003 GET /scenarios/{id} detail (fixes missing-route 404)
  • CHK004 Detail loadable independent of agent session (not event-driven)

Revisions (SCREG-FR-003)

  • [~] CHK005 Edit creates immutable candidate revision with parent link; save does not advance current_revision
  • CHK005a Activation atomically advances current_revision only after eligibility and policy/gate checks
  • CHK006 Runs pin scenario_id (UUID) + revision_id (UUID) + content_hash snapshot
  • CHK007 Revision diff (added/changed/removed) available

Lifecycle (SCREG-FR-005/006/009)

  • CHK008 States DRAFT/READY/STALE/NEEDS_REVALIDATION/BLOCKED/DISABLED/DEPRECATED/ARCHIVED
  • CHK009 Only valid transitions; audit-logged
  • CHK010 Clone/rename/archive/restore operations with RBAC
  • CHK011 Hard delete forbidden with run history (archive instead)

Staleness (SCREG-FR-007)

  • CHK012 Staleness from 037 StructureDiff + 041 lineage, not blind rescan
  • [~] CHK013 Affected scenarios → NEEDS_REVALIDATION/BLOCKED with reason
  • CHK014 Stale scenario run warning-gated/blocked
  • CHK015 Lineage engine unavailable → skip with warning (never false-stale)

RBAC (SCREG-FR-008)

  • CHK016 scenario:view/create/edit/archive scopes enforced
  • CHK017 Edit does not grant run

Success Criteria

  • CHK018 List/detail < 200ms; SC-001..005 verified

PLAN — Implementation Plan

Source: plan.md

Implementation Plan: Scenario Registry & Lifecycle

Branch: 042-dashboard-scenario-registry | Date: 2026-08-07 | Spec: spec.md | Status: Partially implemented — factual audit pending remediation

Implementation audit, 2026-08-20: registry persistence/API/lifecycle code exists. Runtime 037/041 staleness subscription, 036 InvestigationSignal emission and final independent verification remain open.

Summary

Persist user-created dashboard test scenarios as a first-class registry projection with list/search/detail, immutable candidate/current revision chains, a lifecycle state machine (archive-not-delete), staleness detection via 037 StructureDiff + 041 lineage, and a 047-consumed health badge. This is the queryable source of truth shared by the Scenario Editor (043) and Scenario Execution Engine (044), and fixes the currently-missing GET /scenarios/{id} route.

Technical Context

Language/Version: Python 3.13+ (backend), TypeScript DTOs (frontend) Primary Dependencies: FastAPI, SQLAlchemy, Pydantic 2 (backend); existing 036 AgentRun, 037 StructureDiff/VerificationRun, 041 lineage services Storage: PostgreSQL — new tables scenario_registry_entries, scenario_revisions, scenario_staleness_signals, lifecycle audit log Testing: pytest (unit/contract), deterministic snapshots; L1 model tests for frontend DTOs Frontend Architecture: thin list/detail surface; heavy view/edit UI owned by 043, run monitor by 045 Performance Goals: list/detail < 200ms; revision diff deterministic and bounded Constraints: immutable revisions; archive-not-delete; RBAC per-scope; no git-file-only query surface Scale: hundreds of scenarios per dashboard family; revision chains up to dozens

Constitution Check

Principle Result
I. Semantic Contract First PASS — registry/lifecycle/staleness contracted C3-C4
II. Decision Memory PASS — research R1-R4 record rationale/rejected
III. External Orchestrator PASS — reads 037/041; no Superset coupling
IV. Module Discipline PASS — registry/revisions/lifecycle/staleness separated
V. RBAC Enforcement PASS — scenario:view/create/edit/archive scopes
VI. Svelte 5 Runes Only PASS — thin list/detail, model-first
VII. Test-Driven C3+ PASS — transition/staleness tests written first
VIII. Attention-Optimized PASS — [SEMANTICS scenario,registry,...], hierarchical IDs

Project Structure

specs/042-dashboard-scenario-registry/
├── spec.md / data-model.md / research.md / plan.md / tasks.md / traceability.md / quickstart.md / ux_reference.md
├── checklists/requirements.md
├── contracts/modules.md, contracts/openapi.yaml, contracts/ux/
└── prototype/index.html + manifest.md

backend/src/
├── models/scenario_registry.py        # ScenarioRegistryEntry, ScenarioRevision, ScenarioStalenessSignal
├── services/dashboard_testing/registry/
│   ├── list.py  get.py  revisions.py  diff.py  lifecycle.py  clone.py  staleness.py  health.py
└── api/routes/dashboard_testing/scenarios.py  # + GET list/detail/revisions/clone/archive/restore/runs/artifacts

frontend/src/lib/models/ScenarioRegistryModel.svelte.ts
frontend/src/routes/dashboard-testing/scenarios/  # list + detail

Delivery Phases

  1. Registry models + migration + fixtures.
  2. List/search/detail + GET routes (fix 404).
  3. Revision chain + diff.
  4. Lifecycle state machine + audit + clone/archive/restore.
  5. Staleness integration (037 StructureDiff + 041 lineage).
  6. Health derivation.
  7. REST/agent tools, frontend DTOs, regression gates.

API and Schema

  • contracts/openapi.yaml — canonical REST contract (list/detail/revisions/clone/archive/restore/staleness/health).
  • Schema reuse: 038 DashboardTestScenario, 037 StructureDiff, 041 blast-radius report.

Traceability

traceability.md maps Story/Requirement → model → API operationId → contract → task → test.

Cross-Spec Boundary

  • Reads 038 scenario graph, 037 StructureDiff/VerificationRun, 041 lineage blast radius.
  • Feeds 043 (editor) and 044 (runner) with persisted scenario + revision.
  • UI: 042 supplies registry list/detail; 043 owns editing; 045 owns run monitor.

Complexity Tracking

No exception planned. Registry/lifecycle/staleness are bounded C3-C4 contracts; do not collapse into one oversized controller.


RESEARCH — Technical Decisions

Source: research.md

Scenario Registry — Phase 0/1 Research (042)

Branch: 042-dashboard-scenario-registry | Date: 2026-08-07 | Spec: spec.md

R1. Persistence model — registry projection vs git-file

Decision: Introduce a scenario_registry DB projection as the queryable source of truth. The materialized scenario.yaml/runner.plan.json in the dashboard git repo remain artifacts, not the query surface.

Rationale: list/search/filter/status/ownership/staleness require queryable columns and indexes; parsing git files on every list is slow and couples the registry to git availability.

Alternatives Considered: git-only storage (rejected: no query projection); in-memory registry (rejected: no persistence across sessions).

Impact: New SQLAlchemy models ScenarioRegistryEntry, ScenarioRevision, ScenarioStalenessSignal; API GET /scenarios, GET /scenarios/{id}, revision chain endpoints.

R2. Revision identity — reuse 038 hashes

Decision: Reuse 038 scenario_key + content_hash; registry assigns scenario_id (UUID) and revision_id (UUID). Runs pin scenario_id + revision_id + content_hash.

Rationale: Reproducibility requires an immutable snapshot; 038 already computes deterministic hashes.

Alternatives: store "current graph" only (rejected: unreproducible runs).

Impact: Revisions are append-only rows. Later saves produce candidates; a separate policy-controlled atomic activation advances the current_revision pointer.

R3. Staleness source — release diff + lineage, not blind rescan

Decision: Staleness computed from 037 StructureDiff (dashboard release change) and 041 lineage blast radius (dataset/schema change). No blind rescan.

Rationale: Blind rescan is expensive and cannot attribute causality; the blast radius engine already computes affected-entities.

Alternatives: scheduled full fingerprint compare (rejected: O(N²), no causality).

Impact: ScenarioStalenessSignal rows written by lineage/diff hooks; scenarios transition to NEEDS_REVALIDATION/BLOCKED.

R4. Lifecycle — explicit state machine, archive-not-delete

Decision: Enforce a valid-transition state machine; archive is the terminal state when run history exists; hard delete is admin-scoped and empty-history-only.

Rationale: Audit reproducibility requires history to survive; free-form statuses allow illegal transitions.

Alternatives: hard delete (rejected: breaks audit); status as free string (rejected: no enforcement).

Impact: LifecycleStateMachine, audit log per transition, RBAC scopes.

Data Model

See data-model.md: ScenarioRegistryEntry, ScenarioRevision, ScenarioStalenessSignal, LifecycleStateMachine, health derivation.

Contracts & API

  • contracts/modules.md — C3+ contracts: Registry.List, Registry.Get, Registry.RevisionChain, Registry.Transition, Registry.Clone, Registry.Archive, Registry.Staleness, Registry.Health.
  • OpenAPI: GET /scenarios, GET /scenarios/{id}, GET /scenarios/{id}/revisions, POST /scenarios/{id}/revisions/{rev}/checkout, POST /scenarios/{id}/clone, POST /scenarios/{id}/archive, POST /scenarios/{id}/restore, GET /scenarios/{id}/runs, GET /scenarios/{id}/artifacts.

Constitution Check

Principle Result
I. Semantic Contract First PASS — registry/lifecycle/staleness contracted C3-C4
II. Decision Memory PASS — R1-R4 record rationale/rejected
III. External Orchestrator PASS — reads 037/041; no Superset coupling
IV. Module Discipline PASS — registry/revisions/lifecycle/staleness separated
V. RBAC Enforcement PASS — per-scope scopes
VI. Svelte 5 Runes Only PASS — list/detail are thin, 043/045 own heavy UI
VII. Test-Driven C3+ PASS — transition/staleness tests first
VIII. Attention-Optimized PASS — [SEMANTICS scenario,registry,...], hierarchical IDs

DATA MODEL — Entities & Relations

Source: data-model.md

#region ScenarioRegistry.DataModel [C:4] [TYPE ADR] [SEMANTICS data-model,scenario,registry,revision,lifecycle] @BRIEF Canonical registry entry, revision chain, lifecycle state machine, and staleness signal models. @RELATION DEPENDS_ON -> [ScenarioRegistry.Research] @RATIONALE A queryable registry projection decouples scenario management from git-file materialization; typed lifecycle/staleness make states auditable and reproducible. @REJECTED Free-form status strings — rejected because only valid transitions must apply and be audit-logged.

ScenarioRegistryEntry — entity identity (revision for #7/#8)

Fields: scenario_id (UUID — entity identity), scenario_key (semantic slug = dashboard + normalized objective, domain/human identity), name, description, dashboard_id, environment_ids (compat), owner_id, owner_username, tags (list), metadata_version (opaque ETag for metadata-only optimistic concurrency), current_revision (revision_id), lifecycle_status (enum), validation_status (enum), last_run (ref), last_successful_run (ref), health (derived: pass/warn/fail), last_modified_at, baseline_compatibility (version), created_at.

Identity rule: scenario_id MUST be a UUID to survive clone (clones of the same dashboard+objective must NOT collide). scenario_key is the human/domain identity (may repeat across clones).

Indexes: dashboard_id, owner_id, lifecycle_status, tags (GIN), name (trigram for search), scenario_key.

ScenarioRevision — revision identity (revision for #7)

Fields: revision_id (UUID — unique immutable identity), scenario_id, content_hash (SHA-256 of the executable canonical Verification Program graph, including SQL templates/hashes, DSL, assertions and AgentEvaluationSpec; timestamps/display-only excluded), parent_revision_id (nullable), graph_snapshot (DashboardTestScenario JSON), execution_template_hash (hash of revision-derived execution template only; no environment/ParameterBinding/baselines), template_version, schema_version, compatibility_family, change_summary (added/changed/removed), created_by, created_at, activation_status (candidate|current), activated_by?, activated_at?, activation_agent_action_id?. is_current is a read-model convenience derived only from activation_status=current.

Metadata vs executable split (#8): entity-metadata edits (name/description/tags) update ScenarioRegistryEntry fields and DO NOT create a new executable ScenarioRevision (content_hash unchanged). Only executable-graph edits create a new revision. revision_id (UUID) is the unique identity; content_hash detects content change.

Clone provenance: cloning creates a new scenario_id and a new initial revision with parent_revision_id=null; ForkProvenance { source_scenario_id, source_revision_id } preserves its origin. Cross-scenario revisions MUST NOT be joined by parent_revision_id.

CreateScenario — server-owned handle + outbox saga (#1/#5/#6)

CreateScenario never accepts an arbitrary client graph, draft pack, or owner. The authoring path creates server-owned immutable handles:

  • CompiledScenarioHandle { handle_id, content_hash, dashboard_id, validation_status=valid } from 038;
  • DraftPackHandle { draft_pack_id, compiled_handle_id, digest, status=save_eligible } from the authoring pack renderer.

The authenticated principal supplies ownership; the client supplies only {compiled_handle_id, draft_pack_id, draft_pack_digest}. In one PostgreSQL transaction, the service verifies handle ownership, matching dashboard/content hash, save_eligible, actor RBAC, and then creates ScenarioRegistryEntry + ScenarioRevision #1 + OutboxEvent(type=materialize_revision). The transaction returns {scenario_id, revision_id, materialization_status=pending}.

Graph materialization is synchronous (amendment 2026-09-06). ScenarioRevision.graph_snapshot is the canonical DashboardTestScenario JSON (see the ScenarioRevision field list above), materialized INSIDE the create transaction from the CompiledScenarioHandle canonical bytes after digest re-verification — never from caller input, and never a provenance-only pointer record. What the outbox worker materializes asynchronously is ONLY the git/filesystem reference artifacts below; materialization_status/RevisionMaterialization track that async reference-artifact state, not the DB graph. Handle storage, binding and single-consumption rules are owned by ScenarioGraph.ServerOwnedPipeline (038 contracts/modules.md, amendment 2026-09-06); 042 consumes handles and never mints them.

Interim drift (tracked by 050 tasks T029d/T029e and Doc.Adr.ADR0023): the current registry/create.py writes a provenance-only graph_snapshot ({scenario_key, compiled_handle_id, draft_pack_id, draft_pack_digest, agent_run_id}), api_create_scenario hardcodes materialization_status="materialized" with no OutboxEvent/RevisionMaterialization rows, and ScenarioExecution.RunnerPlan.Derive raises BOOTSTRAP_REVISION_NOT_RUNNABLE for such revisions so they cannot produce a vacuous zero-step PASS. These are transitional guards, not the target contract.

Git/filesystem materialization is NOT part of the DB transaction. An idempotent worker consumes the outbox event, writes reference artifacts (scenario.yaml, reference runner.plan.json) keyed by content hash, and updates RevisionMaterialization from pending → materialized | failed. A failed materialization is retryable without duplicating Registry rows; the DB revision remains authoritative.

Agent-created registry mutations use the same server-owned handles and transaction. created_by records the agent identity and delegated analyst; agent_run_id? and investigation_case_id? preserve provenance. A delegated agent may save a validated immutable revision but cannot silently bypass lifecycle transitions, object ACL, materialization, or a required ActionApprovalGate.

@{ ScenarioRegistry.SaveContinuation [C:5] [TYPE ADR]

@BRIEF Server-owned save, CAS, idempotency, and approval continuation contract. @RELATION DEPENDS_ON -> [ScenarioGraph.ServerOwnedPipeline] @RELATION DEPENDS_ON -> [ScenarioExecution.RunPreflight]

The registry accepts only stored CompiledScenarioHandle and DraftPackHandle references. It recomputes and verifies content/draft digests against those handles and the expected dashboard before the PostgreSQL transaction. A caller graph, caller digest, scenario.yaml, or runner.plan.json cannot create or replace a revision. The transaction is all-or-nothing: rejection writes no revision, outbox event, gate, or materialization request.

The continuation state machine is: working_draft -> validated -> candidate -> (approval_required -> candidate) -> current. The initial save may create current; all later executable saves create only a candidate. Activation is separate and CAS-protected. current_revision never advances during save. A stale expected revision or gate version returns typed 409 and has no side effect. The same idempotency key and canonical request hash replay the original result; a changed hash returns 409 IDEMPOTENCY_KEY_REUSED.

Materialization is an outbox continuation, not authority: pending -> materialized|failed, with retry by event idempotency and no duplicate registry rows. Approval continuation is pending_approval -> queued; denial/expiry is pending_approval -> blocked, and each gate is consumed once by CAS. Activation requires valid validation, materialization, automation eligibility and the recorded delegated-authority decision or ActionApprovalGate.

@} ScenarioRegistry.SaveContinuation

@{ ScenarioRegistry.AgentAuthoringWorkspaceRecord [C:5] [TYPE Model]

@BRIEF Persistent registry-side continuation record for co-authored exploration and promotion. @RELATION DEPENDS_ON -> [DashboardScenarioModel.AuthoringWorkspaceModel]

AgentAuthoringWorkspaceRecord stores workspace_id, owner/delegated principals, base revision/hash, state, CAS version, proposal/diff refs, exploration and AuthoringArtifact refs, review decision, idempotency keys, and approval-gate ref. All mutable fields are server-owned and every transition is audited. The record may reference artifacts but never promotes an artifact directly.

The only persistence input to CreateScenario/save is the server-resolved 038 handle pair plus expected base and user-review/CAS evidence. The transaction verifies handle ownership, validation/template fingerprints and digest, then creates the immutable revision and outbox atomically. A changed or caller-computed digest is correlation-only and cannot authorize save.

@} ScenarioRegistry.AgentAuthoringWorkspaceRecord

Revision save and activation are separate operations

The initial revision created with a new scenario is current. Every later executable save creates an immutable candidate; saving never changes ScenarioRegistryEntry.current_revision. ActivateCurrentRevision is a separate atomic operation: it makes exactly one candidate current, updates the entry pointer, and records actor/agent/policy/gate provenance. A candidate is activation-eligible only if validation and materialization succeeded, it is automation_eligible, contains no human step for automation adoption, passes required verification, does not increase risk outside policy, and is compatible with the current revision (compatibility_family unchanged unless an explicit analyst-approved migration policy permits it).

The server evaluates the versioned 036 DelegatedAuthorityPolicy. An agent may activate only when its recorded policy snapshot explicitly has may_activate_current_revision=true; otherwise the operation creates/consumes an ActionApprovalGate. Schedules with revision_policy=current resolve the pointer only after this atomic activation; pinned schedules may use only an explicitly selected eligible revision.

RevisionMaterialization and OutboxEvent

  • RevisionMaterialization: revision_id, status (pending|materialized|failed), artifact_manifest_hash, attempt_count, last_error, materialized_at.
  • OutboxEvent: event_id, aggregate_type, aggregate_id, event_type, payload, idempotency_key, created_at, delivered_at, attempt_count. Written in the same DB transaction as the registry revision; never lost on worker/broker outage.

LifecycleStateMachine

States: DRAFT, READY, STALE, NEEDS_REVALIDATION, BLOCKED, DISABLED, DEPRECATED, ARCHIVED.

Valid transitions:

  • DRAFT → READY (validation clean) | DEPRECATED
  • READY → STALE | DISABLED | DEPRECATED | ARCHIVED | NEEDS_REVALIDATION
  • STALE → NEEDS_REVALIDATION | READY (revalidated) | DEPRECATED
  • NEEDS_REVALIDATION → READY | BLOCKED | STALE
  • BLOCKED → NEEDS_REVALIDATION | DEPRECATED
  • DISABLED → READY | ARCHIVED | DEPRECATED
  • DEPRECATED → ARCHIVED
  • ARCHIVED → (terminal; only restore to DRAFT)

Every transition writes an audit record: scenario_id, from, to, actor_id, reason, timestamp. Full command surface is generic transition (disable, enable→READY, deprecate, archive, restore) rather than archive/restore-only endpoints.

ScenarioStalenessSignal

Fields: id, scenario_id, source_type, source_fingerprint, kind (dashboard_release_diff | lineage_blast_radius | reference_removed), severity (info/warning/critical), reason, affected_ref, detected_at, resolved_at (nullable), source (037 StructureDiff | 041 lineage). Unique identity is (scenario_id, source_type, source_fingerprint, affected_ref); repeated events upsert rather than duplicate a signal. Lifecycle derives from the aggregate of active signals: READY is permitted only when no active critical/blocking signal remains.

Health projection

042 does not derive health. It consumes only 047 ScenarioHealth.overall_attention as the registry badge; if analytics is unavailable the badge is unknown. The algorithm and historical windows belong exclusively to 047.

The Registry is also the portfolio entry surface for Investigation Queue. Staleness and health signals may link to queue items; opening one opens its persistent agent case rather than a modal or an automatic chat.

Object-level authorization

  • scenario:view, scenario:create, scenario:edit, scenario:archive, scenario:run (separate from edit). MVP is archive-only: there is no hard-delete route, even for administrators.

Every registry/run/evidence decision additionally requires the intersection of scenario permission, dashboard ACL, environment ACL, artifact/evidence ACL, and the caller's effective Superset/RLS access. A registry permission alone never grants screenshot, XLSX, VLM, or result access.

Storage Notes

  • Registry projection is authoritative for list/detail/status; git scenario.yaml/reference runner.plan.json are asynchronously materialized artifacts with explicit status.
  • Revisions are immutable and append-only; an executable save creates a new candidate row and does not advance current_revision. Only the separate ActivateCurrentRevision compare-and-set (CAS) operation may advance current_revision after its eligibility, materialization, policy, and approval checks. Initial scenario creation is the sole exception when the explicitly eligible initial revision is created as current. A ScenarioRun target is always selected from an explicit promoted revision_id plus its verified content_hash; it never resolves from the latest candidate or an implicit moving current_revision.

#endregion ScenarioRegistry.DataModel


CONTRACTS — Module & Function Contracts

Source: contracts/modules.md

#region ScenarioRegistry.Modules [C:4] [TYPE ADR] [SEMANTICS scenario,registry,contracts,modules,lifecycle] @BRIEF Module/function contracts for the Scenario Registry & Lifecycle (042). @defgroup ScenarioRegistry Persistent scenario catalog, revisions, lifecycle, staleness, health. @RELATION DEPENDS_ON -> [ScenarioGraph.Models] @RELATION DEPENDS_ON -> [BaselineEngine.DashboardQueryModel] @RELATION DEPENDS_ON -> [LineageBlastRadius.Report] @RATIONALE A registry is the queryable source of truth shared by editor (043) and runner (044); every contract enforces valid transitions and immutable revisions. @REJECTED Free-form status strings; hard delete with run history; git-file-only query surface.

Region modules

#region ScenarioRegistry.List [C:4] [TYPE Function] [SEMANTICS scenario,registry,list,search]

@ingroup ScenarioRegistry

@BRIEF List/search/filter persisted scenarios by name, tag, dashboard, status, owner.

@PRE caller has scenario:view; filters validated.

@POST returns paged ScenarioRegistryEntry[] with 047-provided overall_attention/unknown health; stable order.

@SIDE_EFFECT read-only; never derives health locally.

@TEST_EDGE empty->empty state; filter unknown dashboard->empty; LARGE->paged.

def list_scenarios(db, filters, page, page_size): ...

#endregion ScenarioRegistry.List

#region ScenarioRegistry.Get [C:4] [TYPE Function] [SEMANTICS scenario,registry,get,detail]

@ingroup ScenarioRegistry

@BRIEF Load a scenario detail by id from the registry projection.

@PRE caller has scenario:view; id exists.

@POST returns ScenarioRegistryEntry with current revision and tabs payload (steps/params/baselines/runs/revisions/artifacts).

@SIDE_EFFECT read-only.

@TEST_EDGE not_found->404; stale detail->refresh banner flag.

def get_scenario(db, scenario_id): ...

#endregion ScenarioRegistry.Get

@{ ScenarioRegistry.CreateInitial [C:5] [TYPE Function]

@BRIEF Create the first registry entry and its immutable current revision from server-owned validated handles. @PRE Caller has scenario:create; the dashboard/environment binding is authorized; the compiled graph and draft-pack handles are server-issued, valid, and mutually bound by content hash. @POST Creates exactly one ScenarioRegistryEntry and one initial ScenarioRevision with activation_status=current; graph_snapshot is the canonical DashboardTestScenario JSON materialized from the CompiledScenarioHandle bytes in the same transaction (amendment 2026-09-06, see ScenarioRegistry.DataModel); returns their opaque IDs and content hash. A matching idempotency replay returns the same IDs. @SIDE_EFFECT Writes the registry entry, immutable revision with materialized graph_snapshot, audit record, and reference-artifact materialization outbox in one transaction; git/reference artifacts materialize async. @DATA_CONTRACT InitialScenarioIntent + CompiledScenarioHandle + DraftPackHandle -> ScenarioRegistryEntry + ScenarioRevision @INVARIANT A client never uploads a graph, content hash, or revision ID as authority. Initial creation does not require a synthetic base revision. @INVARIANT A created current revision is directly runnable/schedulable once materialized: it carries real executable steps, so ScenarioExecution.RunnerPlan.Derive never sees a provenance-only snapshot (transitional guard BOOTSTRAP_REVISION_NOT_RUNNABLE demotes to defense-in-depth under 050 T029e). @RELATION DEPENDS_ON -> [ScenarioGraph.ServerOwnedPipeline] @RELATION CALLED_BY -> [ScenarioGraph.AgentAuthoringBootstrap] @RATIONALE The first scenario has no base revision, while normal RevisionChain requires one; a dedicated creation transition removes the circular dependency without weakening revision immutability. @REJECTED Seeding a registry entry or base revision in the MCP client/test fixture as a substitute for a production creation path. @REJECTED Storing provenance-only metadata (handle ids/digests without the canonical graph) as the revision graph_snapshot — the 2026-09-06 audit proved such revisions are structurally unrunnable and previously produced a vacuous zero-step PASS (Doc.Adr.ADR0023); canonical bytes come from the handle, not from the caller. @TEST_EDGE duplicate-idempotency-key -> same IDs; stale-or-unbound-handle -> no rows; unauthorized-dashboard -> no disclosure; provenance-only-snapshot -> run/schedule refused. def create_initial_scenario(db, intent, compiled_handle, draft_pack_handle, actor, idempotency_key): ...

@} ScenarioRegistry.CreateInitial

#region ScenarioRegistry.RevisionChain [C:4] [TYPE Function] [SEMANTICS scenario,registry,revision,chain]

@ingroup ScenarioRegistry

@BRIEF Append an immutable candidate revision on validated save.

@PRE base revision id/digest matches; deterministic validation and delegated policy permit save.

@POST new ScenarioRevision candidate row with parent link; current_revision unchanged; no mutation of prior rows.

@SIDE_EFFECT DB write; audit log.

@INVARIANT prior revisions are immutable and never mutated.

@REJECTED Accepting a raw client graph_snapshot dict through the public REST surface POST /api/dashboard-testing/scenarios/{id}/revisions (ScenarioRegistry.Schemas.RevisionCreateRequest) was rejected (amendment 2026-09-06): it contradicts ScenarioRegistry.SaveContinuation ("a caller graph ... cannot create or replace a revision"). Candidate revisions are created only from server-owned validated handles or the 043 editor save path (server-loaded WorkingDraft). The legacy route must be gated to the editor/save semantics or retired under 050 task T029f; until then it is a known drift, not an approved path.

@TEST_EDGE stale base->409; edit->new revision with parent.

def create_revision(db, scenario_id, base_revision_id, validated_graph_handle, agent_action_id): ...

#endregion ScenarioRegistry.RevisionChain

#region ScenarioRegistry.ActivateRevision [C:4] [TYPE Function] [SEMANTICS scenario,registry,revision,activation]

@ingroup ScenarioRegistry

@BRIEF Atomically promote an eligible candidate revision to current.

@PRE materialized/validated/automation eligibility holds; actor or AgentAction passes DelegatedAuthorityPolicy, or bound ActionApprovalGate is consumed.

@POST exactly one revision is current and current_revision points to it; audit/policy provenance written.

@INVARIANT save does not activate; current-policy automation resolves only this pointer.

@TEST_EDGE incompatible family->409; policy gate->202 no promotion; competing promotion->409.

def activate_current_revision(db, scenario_id, revision_id, actor, agent_action_id=None): ...

#endregion ScenarioRegistry.ActivateRevision

#region ScenarioRegistry.Diff [C:3] [TYPE Function] [SEMANTICS scenario,registry,diff,revision]

@ingroup ScenarioRegistry

@BRIEF Compute added/changed/removed between two revisions.

@POST returns structured change set; deterministic.

def diff_revisions(db, scenario_id, rev_a, rev_b): ...

#endregion ScenarioRegistry.Diff

#region ScenarioRegistry.Transition [C:4] [TYPE Function] [SEMANTICS scenario,registry,lifecycle,transition]

@ingroup ScenarioRegistry

@BRIEF Apply a lifecycle transition if valid and RBAC-approved.

@PRE transition allowed by state machine; caller has required scope.

@POST status updated; audit row written; terminal archive when run history exists.

@SIDE_EFFECT DB write; audit log; notification event on archive/deprecate.

@INVARIANT only valid transitions apply; archive preserves run history.

@TEST_EDGE illegal transition->rejected; archive with runs->ok; hard delete with runs->rejected.

def transition(db, scenario_id, target_state, actor, reason): ...

#endregion ScenarioRegistry.Transition

#region ScenarioRegistry.Clone [C:3] [TYPE Function] [SEMANTICS scenario,registry,clone,fork]

@ingroup ScenarioRegistry

@BRIEF Create a new scenario from a snapshot with cleared owner lineage.

@PRE caller has scenario:create; source scenario readable.

@POST new scenario_id, derived revisions, owner cleared, audit logged.

def clone_scenario(db, source_id, actor): ...

#endregion ScenarioRegistry.Clone

#region ScenarioRegistry.Archive [C:3] [TYPE Function] [SEMANTICS scenario,registry,archive,lifecycle]

@ingroup ScenarioRegistry

@BRIEF Archive a scenario; the terminal non-deleting lifecycle state.

@PRE caller has scenario:archive; scenario exists.

@POST status=ARCHIVED; run history and revisions preserved.

def archive_scenario(db, scenario_id, actor, reason): ...

#endregion ScenarioRegistry.Archive

#region ScenarioRegistry.Staleness [C:4] [TYPE Function] [SEMANTICS scenario,registry,staleness,lineage]

@ingroup ScenarioRegistry

@BRIEF Mark scenarios affected by dashboard release diff / lineage blast-radius change.

@PRE 037 StructureDiff or 041 blast-radius report available.

@POST affected scenarios transition to NEEDS_REVALIDATION/BLOCKED with reason; warning-gates runs.

@SIDE_EFFECT DB writes; run-gate flag.

@INVARIANT never false-stale from unavailable lineage engine (skips with warning).

@TEST_EDGE chart_removed->BLOCKED; filter_scope_change->NEEDS_REVALIDATION; engine_down->skip warning.

def apply_staleness(db, signals): ...

#endregion ScenarioRegistry.Staleness

#region ScenarioRegistry.Health [C:3] [TYPE Function] [SEMANTICS scenario,registry,health,flakiness]

@ingroup ScenarioRegistry

@BRIEF Read the 047-derived contextual ScenarioHealth badge projection.

@POST returns pass/warn/fail with success rate and flakiness ratio.

def get_health_badge(analytics_client, scenario_id): ...

#endregion ScenarioRegistry.Health

#endregion ScenarioRegistry.Modules


OPENAPI — REST/Event API Contract

Source: contracts/openapi.yaml

openapi: 3.1.0 info: title: Scenario Registry & Lifecycle API version: 0.1.0 description: Persist, query, version, and manage lifecycle of dashboard test scenarios (042). paths: /api/dashboard-testing/scenarios: post: operationId: scenarioRegistry.create summary: Register a server-validated compiled scenario + save-eligible draft pack (Save→Register) security: [{ bearerAuth: [] }] requestBody: required: true content: application/json: schema: type: object required: [compiled_handle_id, draft_pack_id, draft_pack_digest] properties: compiled_handle_id: { type: string, format: uuid } draft_pack_id: { type: string, format: uuid } draft_pack_digest: { type: string, pattern: "^[a-f0-9]{64}$" } responses: "201": description: Scenario created content: application/json: schema: type: object properties: scenario_id: { type: string } revision_id: { type: string } materialization_status: { type: string, enum: [pending, materialized, failed] } "409": { description: Conflict } get: operationId: scenarioRegistry.list summary: List/search/filter persisted scenarios security: [{ bearerAuth: [] }] parameters: - { name: q, in: query, schema: { type: string } } - { name: dashboard_id, in: query, schema: { type: integer } } - { name: status, in: query, schema: { type: string, enum: [DRAFT,READY,STALE,NEEDS_REVALIDATION,BLOCKED,DISABLED,DEPRECATED,ARCHIVED] } } - { name: tag, in: query, schema: { type: string } } - { name: owner, in: query, schema: { type: string } } - { name: page, in: query, schema: { type: integer, default: 1 } } - { name: page_size, in: query, schema: { type: integer, default: 25 } } responses: "200": description: Paged registry entries content: application/json: schema: { type: object, properties: { items: { type: array, items: { $ref: "#/components/schemas/ScenarioRegistryEntry" } }, total: { type: integer } } } "429": { $ref: "#/components/responses/RateLimited" } /api/dashboard-testing/scenarios/{scenario_id}: get: operationId: scenarioRegistry.detail summary: Load a scenario detail by id (fixes missing-route 404) security: [{ bearerAuth: [] }] parameters: [{ name: scenario_id, in: path, required: true, schema: { type: string } }] responses: "200": { description: Scenario detail, content: { application/json: { schema: { $ref: "#/components/schemas/ScenarioDetail" } } } } "404": { description: Not found } /api/dashboard-testing/scenarios/{scenario_id}/revisions: get: operationId: scenarioRegistry.revisions summary: List immutable revision chain security: [{ bearerAuth: [] }] parameters: [{ name: scenario_id, in: path, required: true, schema: { type: string } }] responses: "200": { description: Revision chain, content: { application/json: { schema: { type: array, items: { $ref: "#/components/schemas/ScenarioRevision" } } } } } /api/dashboard-testing/scenarios/{scenario_id}/revisions/{rev_a}/diff/{rev_b}: get: operationId: scenarioRegistry.revisionDiff summary: Diff two revisions security: [{ bearerAuth: [] }] parameters: - { name: scenario_id, in: path, required: true, schema: { type: string } } - { name: rev_a, in: path, required: true, schema: { type: string } } - { name: rev_b, in: path, required: true, schema: { type: string } } responses: "200": { description: Change set (added/changed/removed) } /api/dashboard-testing/scenarios/{scenario_id}/clone: post: operationId: scenarioRegistry.clone summary: Clone a scenario security: [{ bearerAuth: [] }] parameters: [{ name: scenario_id, in: path, required: true, schema: { type: string } }] responses: { "200": { description: New scenario id } } /api/dashboard-testing/scenarios/{scenario_id}/archive: post: operationId: scenarioRegistry.archive summary: Archive a scenario (terminal non-deleting state) security: [{ bearerAuth: [] }] parameters: [{ name: scenario_id, in: path, required: true, schema: { type: string } }] requestBody: { required: true, content: { application/json: { schema: { type: object, required: [reason], properties: { reason: { type: string } } } } } } responses: { "200": { description: Archived }, "403": { description: Permission denied } } /api/dashboard-testing/scenarios/{scenario_id}/restore: post: operationId: scenarioRegistry.restore summary: Restore an archived scenario security: [{ bearerAuth: [] }] parameters: [{ name: scenario_id, in: path, required: true, schema: { type: string } }] responses: { "200": { description: Restored } } /api/dashboard-testing/scenarios/{scenario_id}/staleness: get: operationId: scenarioRegistry.staleness summary: List staleness signals for a scenario security: [{ bearerAuth: [] }] parameters: [{ name: scenario_id, in: path, required: true, schema: { type: string } }] responses: { "200": { description: Staleness signals } } /api/dashboard-testing/scenarios/{scenario_id}/transition: post: operationId: scenarioRegistry.transition summary: Apply a valid lifecycle transition (disable, enable, deprecate, archive, restore) security: [{ bearerAuth: [] }] parameters: [{ name: scenario_id, in: path, required: true, schema: { type: string } }] requestBody: required: true content: application/json: schema: type: object required: [target_state, reason] properties: target_state: { type: string, enum: [DRAFT, READY, STALE, NEEDS_REVALIDATION, BLOCKED, DISABLED, DEPRECATED, ARCHIVED] } reason: { type: string } metadata_version: { type: string } responses: { "200": { description: Transition applied }, "409": { description: Invalid transition or stale metadata version } } /api/dashboard-testing/scenarios/{scenario_id}/revisions/{revision_id}/materialization: get: operationId: scenarioRegistry.materialization summary: Get asynchronous git/reference-artifact materialization state security: [{ bearerAuth: [] }] parameters: - { name: scenario_id, in: path, required: true, schema: { type: string } } - { name: revision_id, in: path, required: true, schema: { type: string } } responses: { "200": { description: RevisionMaterialization } } /api/dashboard-testing/scenarios/{scenario_id}/revisions/{revision_id}/activate: post: operationId: scenarioRegistry.activateCurrentRevision summary: Promote an eligible candidate revision to the current revision description: Save and activation are separate. The server evaluates delegated authority and may return an inline approval gate. security: [{ bearerAuth: [] }] parameters: - { name: scenario_id, in: path, required: true, schema: { type: string } } - { name: revision_id, in: path, required: true, schema: { type: string } } requestBody: required: true content: application/json: schema: type: object properties: agent_action_id: { type: string, nullable: true } reason: { type: string } responses: "200": { description: Candidate atomically promoted to current } "202": { description: ActionApprovalGate created; no activation yet } "409": { description: Candidate is not activation-eligible or current pointer changed } components: securitySchemes: bearerAuth: { type: http, scheme: bearer } responses: RateLimited: description: Rate limited headers: Retry-After: { schema: { type: integer } } schemas: ScenarioRegistryEntry: type: object properties: scenario_id: { type: string } name: { type: string } dashboard_id: { type: integer } owner_username: { type: string } tags: { type: array, items: { type: string } } lifecycle_status: { type: string } validation_status: { type: string } current_revision: { type: string } metadata_version: { type: string } last_run_id: { type: string, nullable: true } health: { type: string, enum: [pass, warn, fail, none] } last_modified_at: { type: string, format: date-time } ScenarioRevision: type: object properties: revision_id: { type: string } content_hash: { type: string } parent_revision_id: { type: string, nullable: true } execution_template_hash: { type: string } schema_version: { type: integer } compatibility_family: { type: string } change_summary: { type: object } created_by: { type: string } created_at: { type: string, format: date-time } activation_status: { type: string, enum: [candidate, current] } is_current: { type: boolean, readOnly: true } ScenarioDetail: type: object properties: entry: { $ref: "#/components/schemas/ScenarioRegistryEntry" } graph: { type: object } run_count: { type: integer }


QUICKSTART — Dev Onboarding

Source: quickstart.md

Quickstart: Scenario Registry & Lifecycle (042)

Factual audit 2026-08-20: pending verification checklist only; no command result below is currently asserted as evidence.

Prereqs

  • Backend venv: backend/.venv
  • DB migrated (registry tables)

Commands

cd backend && source .venv/bin/activate

# Migration
alembic upgrade head

# Backend tests
python -m pytest -v tests/services/dashboard_testing/registry/
python -m pytest -v tests/api/test_scenario_registry.py

# Lint
python -m ruff check backend/src/services/dashboard_testing/registry/ backend/src/api/routes/dashboard_testing/scenarios.py

# Frontend tests
cd frontend && npm run test -- ScenarioRegistry

Exit Gates

  • List/search/detail tests pass (incl. GET /scenarios/{id} not 404)
  • Revision chain immutability verified (stale base → 409)
  • Lifecycle transition + archive-not-delete verified
  • Staleness from 037/041 hooks verified
  • ruff clean; prototype state coverage passes

TRACEABILITY — Requirements Matrix

Source: traceability.md

Traceability: Scenario Registry & Lifecycle (042)

Factual audit 2026-08-20: rows trace intended code/test ownership, not completed acceptance evidence. 037/041 upstream staleness subscription and 036 InvestigationSignal emission are open.

Story Requirement Model API operationId Contract Task Test
US1 Find SCREG-FR-001/002 ScenarioRegistryEntry scenarios.list Registry.List T005-T007 test_list
US2 Detail SCREG-FR-002/004 ScenarioRegistryEntry scenarios.detail Registry.Get T008-T010 test_scenario_registry
US2 Create SCREG-FR-001 CreateScenario scenarios.create Registry.Create T010b-T010d test_scenario_registry
US3 Revisions SCREG-FR-003 ScenarioRevision scenarios.revisions, scenarios.revision.diff Registry.RevisionChain, Registry.Diff T011-T013 test_revisions
US4 Staleness SCREG-FR-007 ScenarioStalenessSignal scenarios.staleness Registry.Staleness T014-T016 test_staleness
US5 Lifecycle SCREG-FR-005/006/009 LifecycleStateMachine scenarios.clone, scenarios.archive, scenarios.restore Registry.Transition, Registry.Clone, Registry.Archive T017-T020 test_lifecycle
Health — derived scenarios.health Registry.Health T021 test_health

N/A columns: Run Monitor (045), Editor (043), Automation (046), Analytics (047).

Cross-Spec Pipeline Traceability

Stage / requirement 042 authority Handoff Evidence
server-owned compile/draft handles only ScenarioRegistry.SaveContinuation 038 handles -> revision transaction test_scenario_registry, test_revisions
immutable revision identity and content digest ScenarioRevision, Registry.RevisionChain 044 RunPreflight pins revision/content hash test_revisions
save/activation separation and CAS ScenarioRegistry.SaveContinuation, Registry.ActivateRevision candidate -> explicit current test_revisions, lifecycle tests
outbox materialization is non-authoritative RevisionMaterialization, OutboxEvent reference artifacts only; registry remains authority registry transaction/materialization tests
MCP request-save parity SCREG-FR-011, ScenarioRegistry.SaveContinuation 050 MCP -> same server transaction 050 test_mcp_*, SC-001 walkthrough
persistent authoring workspace ScenarioRegistry.AgentAuthoringWorkspaceRecord 038 workspace -> server-owned registry continuation persistence, reconnect and CAS tests
reviewed promotion boundary ScenarioRegistry.AgentAuthoringPromotion 038 validated handles -> candidate/current revision diff-review, ownership, idempotency and no-raw-authority tests
revision -> preflight -> run handoff ScenarioRevision 044 RunPreflight -> RunnerPlan -> ScenarioRun test_scenario_runner, test_scenario_runner_plan

This matrix is conjunctive with the 038, 044 and 050 matrices. Open or partial evidence is a release NO-GO; unchecked implementation tasks remain unchanged.

Authoring E2E is a release gate: no GO unless sandbox outputs become typed 038 proposals, a user reviews the server-computed diff, 042 performs handle-based CAS/idempotent save, and 044 rejects every non-promoted revision or raw authoring artifact. This normative row does not imply implementation completion.


TASKS — Implementation Tasks

Source: tasks.md

#region ScenarioRegistry.Tasks [C:3] [TYPE ADR] [SEMANTICS tasks,scenario,registry,implementation] @BRIEF Ordered TDD backlog for Scenario Registry & Lifecycle (042). @BRIEF Tests written FIRST for every C3+ contract per constitution VII.

Prerequisites: plan.md, spec.md (required); contracts/modules.md, traceability.md (present).

Format: - [ ] T### [P] [USx] Description with exact file path

Factual audit 2026-08-20: [x] means code plus relevant evidence; [~] means partial implementation; [ ] means absent integration or unperformed verification.

Phase 1 — Setup (Shared Infrastructure)

  • T001 Create SQLAlchemy models ScenarioRegistryEntry, ScenarioRevision, ScenarioStalenessSignal, lifecycle audit log in backend/src/models/scenario_registry.py
  • T002 Write alembic migration for scenario registry tables in backend/alembic/versions/
  • T003 [P] Create canonical fixtures under specs/042-dashboard-scenario-registry/fixtures/ (registry rows, revision chain, staleness signals)
  • T004 [P] Materialize fixtures into backend/tests/fixtures/scenario_registry/

Checkpoint: Schema loads; fixtures materialized.

Phase 2 — US1 Find a Saved Scenario

  • T005 [US1] Write failing list/search/filter tests in backend/tests/services/dashboard_testing/registry/test_list.py
  • T006 [US1] Implement list_scenarios in backend/src/services/dashboard_testing/registry/list.py @POST: paged ScenarioRegistryEntry[] with derived health; stable order @TEST_EDGE: empty->empty state; filter unknown dashboard->empty; LARGE->paged
  • T007 [US1] Add GET /dashboard-testing/scenarios in backend/src/api/routes/dashboard_testing/scenarios.py

Phase 3 — US2 Open a Scenario Detail

  • T008 [US2] Write failing detail tests (incl. the missing-route fix) in backend/tests/api/test_scenario_registry.py
  • T009 [US2] Implement get_scenario in backend/src/services/dashboard_testing/registry/get.py @POST: ScenarioRegistryEntry with current revision and tabs payload @TEST_EDGE: not_found->404; stale detail->refresh flag
  • T010 [US2] Add GET /dashboard-testing/scenarios/{id} (fixes frontend getScenarioDraft 404)

Phase 3b — Save→Register (P0 #1)

  • T010b [P1] Write failing CreateScenario transaction tests in backend/tests/api/test_scenario_registry.py @TEST_EDGE: atomic (all-or-nothing); revision #1 + runner_plan_hash always present; clone yields distinct scenario_id
  • T010c [P1] Implement create_scenario in backend/src/services/dashboard_testing/registry/create.py @POST: ScenarioRegistryEntry(scenario_id=UUID) + ScenarioRevision#1 + artifact binding @SIDE_EFFECT: DB write; git artifact materialization binding
  • T010d [P1] Add POST /dashboard-testing/scenarios; wire 039 Save to call it POST is implemented; the existing 039 Save consume flow performs server-owned registration after materialization, deriving the compiled handle and digest from the verified draft pack.

Phase 4 — US3 Immutable Revisions

  • T011 [US3] Write failing revision chain tests in backend/tests/services/dashboard_testing/registry/test_revisions.py
  • T012 [US3] Implement create_revision + diff_revisions in backend/src/services/dashboard_testing/registry/revisions.py @INVARIANT: prior revisions immutable; runs pin revision snapshot @TEST_EDGE: stale base->409; edit->new revision with parent; diff deterministic
  • T013 [US3] Add GET /scenarios/{id}/revisions + checkout endpoints

Phase 5 — US4 Detect Staleness

  • T014 [US4] Write failing staleness tests in backend/tests/services/dashboard_testing/registry/test_staleness.py
  • T015 [US4] Implement apply_staleness in backend/src/services/dashboard_testing/registry/staleness.py @INVARIANT: never false-stale from unavailable lineage engine (skip with warning) @TEST_EDGE: chart_removed->BLOCKED; filter_scope_change->NEEDS_REVALIDATION; engine_down->skip
  • [~] T016 [US4] Integrate 037 StructureDiff + 041 blast-radius hooks to emit staleness signals Adapter accepts 037 change envelopes and 041 affected scenario/dashboard references; unavailable upstream engines are explicitly skipped without false-stale transitions.

Phase 6 — US5 Manage Lifecycle

  • T017 [US5] Write failing lifecycle transition tests in backend/tests/services/dashboard_testing/registry/test_scenario_registry_lifecycle.py
  • T018 [US5] Implement state machine + audit in backend/src/services/dashboard_testing/registry/lifecycle.py @INVARIANT: only valid transitions; archive preserves run history; hard delete requires empty history + admin @TEST_EDGE: illegal transition->rejected; archive with runs->ok; delete with runs->rejected
  • T019 [US5] Implement clone_scenario, archive_scenario, restore in clone.py and lifecycle.py
  • T020 [US5] Add clone/archive/restore endpoints + RBAC scopes

Phase 7 — Health + Polish

  • [~] T021 [P] Implement analytics-owned health projection in backend/src/services/dashboard_testing/registry/health.py
  • T022 [P] Frontend DTOs ScenarioRegistryEntry/ScenarioRevision in frontend/src/types/scenario-registry.ts
  • T023 [P] Thin ScenarioRegistryModel.svelte.ts for list/detail/revisions/health projections Dedicated heavy registry pages remain owned by 043/045; model is the 042 integration boundary.
  • T024 Run quickstart, scoped/full backend tests, ruff; belief audit; ATTN_1-4; semantic index rebuild Full backend: 10712 passed, 235 skipped, 1 xpassed. Full frontend Vitest: 3859 passed across 204 files; frontend build passed. New/affected files pass targeted ruff/eslint. Full backend ruff still reports 7904 repository-wide pre-existing findings; semantic/belief audit remains a separate workspace task.
  • T025 Prototype validation: verify every @UX_STATE in contracts reachable via prototype/index.html state switcher; responsive

Audit Follow-ups (2026-08-20)

  • T026 Wire 037 StructureDiff and 041 blast-radius production events into apply_staleness; verify an upstream fixture changes affected registry rows without a direct service call.
  • T027 Emit idempotent canonical 036 InvestigationSignal for registry health/staleness and verify emission does not start an agent action.
  • T028 Run independent registry quickstart/API/lifecycle/staleness verification and record only reproducible command results.
  • T029 Add materialized-artifact reconciliation in backend/src/services/dashboard_testing/registry/ and tests: scenario.yaml/runner materialization hash mismatch emits immutable artifact_drift InvestigationSignal and cannot activate or silently replace the registry revision.

Checkpoint: Pending T026–T029. Current code presence does not prove runtime upstream integration or materialized-artifact integrity.

Dependencies

Setup → US1 → US2; revisions depend on models; staleness depends on 037/041 hooks; lifecycle depends on revisions. 043 editor and 044 runner consume this registry.

#endregion ScenarioRegistry.Tasks


PROTOTYPE — State/Manifest

Source: prototype/manifest.md

#region ScenarioRegistry.PrototypeManifest [C:3] [TYPE ADR] [SEMANTICS prototype,manifest,scenario,registry] @defgroup Prototype Interactive HTML prototype manifest for Scenario Registry list + detail.

Prototype Metadata

Factual audit 2026-08-20: this manifest records prototype coverage only. It does not verify production routes, event wiring, accessibility, responsiveness, or acceptance criteria.

  • Feature: 042 Scenario Registry & Lifecycle
  • Source contracts: ux_reference.md, contracts/modules.md
  • Screens represented: 2 (Registry List, Scenario Detail) + persistent Queue entry + run-monitor stub → 045
  • Total states: 5 (list, detail, detail+stale, empty, run-monitor stub); Queue is a navigation handoff to 047, not a dialog state
  • Accessibility: keyboard nav, focus-visible, aria-live (via script), ≥44px targets, prefers-reduced-motion
  • Responsive breakpoints: 375px, 1280px

State Coverage

Screen @UX_STATE Prototype State Reachable? Recovery
Registry List loaded list ✅ —
Registry List empty empty_create ✅ Create CTA
Registry List filtered / LARGE list (search) ✅ refine
Registry List error / NET — (to add) 🟡 retry
Detail loaded detail ✅ —
Detail stale (409) detail_stale ✅ reload/discard
Detail not_found / blocked — 🟡 back to list / revalidate
Run Monitor running run_monitor ✅ (stub) full → 045

Screen ↔ Story Traceability

Prototype Screen User Story Intended acceptance coverage
Registry List US1 Find filter/search rows
Detail US2 Open tabs + metadata from registry
Detail US3 Revisions revision shown (r17)
Detail+stale US4 Staleness banner + revalidate
Detail US5 Lifecycle status badge (Ready)
Registry List Operations persistent Investigation Queue link → 047, no auto-chat
#endregion ScenarioRegistry.PrototypeManifest

PROTOTYPE — Interactive HTML

Source: prototype/index.html

<!doctype html><html lang="ru"><head></head>

Superset Tools · BI testing
СценарииЗапускиАвтоматизацияКачествоРасследования 2Аналитик · Анна
Рабочий стол

Регистр сценариев

Сценарии проверок для dashboard’ов: состояние, последняя проверка и следующий шаг.

Пустой списокСоздать сценарий
Все статусыГотовНужна проверкаВсе тегиSmokeXLSXФильтровать
СценарийDashboardСтатусПоследний запускКачество
XLSX reconciliation
Smoke · 8 шагов
FI-0080ГотовСегодня, 10:42 · PASSТребует вниманияОткрыть
Комментарии по строкам
Ручной запуск · 6 шагов
FI-0080Только вручнуюВчера · PASSСтабильноОткрыть
Фильтры и метрики
Smoke · 5 шагов
Sales overviewНужна ревалидация—Нет данныхОткрыть
← К списку

XLSX reconciliation Готов

FI-0080 · revision r18 · обновлено Анной сегодня в 10:18

РедактироватьЗапустить

Что проверяет сценарий

Экспорт XLSX соответствует активным dashboard и table filters.

Открыть dashboard→Применить фильтры→Скачать XLSX→Сравнить baseline

Последняя проверка

PASS · PREPROD · 1 мин 42 сек

Открыть результат

Пока нет сценариев

Создайте первую проверку из dashboard или опишите бизнес-цель.

Создать сценарий

Новый сценарий

Сначала выберите dashboard и цель проверки. Параметры запуска не нужны для сохранения сценария.

ОтменаПродолжить
State: ListDetailEmpty
<script src="../../prototype-ui.js"></script><script>protoState('list',s=>{for(const id of['list','detail','empty','create'])document.getElementById(id).classList.toggle('hidden',id!==s)});wireOpen()</script></html>

================================================================================ FEATURE: 043-dashboard-scenario-editor Files: 14


SPEC — Feature Specification

Source: spec.md

#region ScenarioEditor.Spec [C:3] [TYPE ADR] [SEMANTICS spec,requirements,ux,scenario,editor,visual,agent] @BRIEF User-facing Scenario Editor: view and edit a persisted scenario with a hybrid (C) editing model — manual business fields/parameters, constrained assertions, visual dependency editing, read-only generated executable, and agent-assisted complex changes. @RELATION DEPENDS_ON -> [Doc.Adr.ADR0001] @RELATION DEPENDS_ON -> [Doc.Adr.ADR0006] @RELATION DEPENDS_ON -> [ScenarioRegistry.Spec] @RELATION DEPENDS_ON -> [DashboardScenarioModel.Spec] @RATIONALE 038 defines the DTOs but no lifecycle editing UX. A scenario must be viewable and editable outside the agent chat, with every durable edit producing a new immutable revision and delegated policy determining whether an inline approval is required. @REJECTED Agent-only editing (no visual surface) — rejected because users need to review and adjust a scenario without re-prompting the agent each time. @REJECTED Unconstrained free-form DAG/assertion editor — rejected because it could inject SQL/raw baselines/unsafe paths, violating 038 safety invariants; assertions use constrained editors and generated executable stays read-only. @REJECTED Chat-bound "Edit with agent" as the only proposal channel — superseded 2026-08-24: proposals are creatable through MCP tools per specs/050-mcp-interface/spec.md; server-stored WorkingDraft, digest binding and SCEDIT-FR-009 no-arbitrary-draft-save constraints stand unchanged.

Navigation (DSA Indexer keywords)

@SEMANTICS: spec, requirements, feature, ux, scenario, editor, visual, revision, agent

Feature Branch: 043-dashboard-scenario-editor Created: 2026-08-07 | Status: Partially implemented — factual audit pending remediation Input: "Provide a first-class Scenario Editor for viewing and editing persisted scenarios: business metadata and parameter definitions editable manually, assertions via constrained editors, dependencies via visual DAG editing, generated executable read-only, and complex changes delegated to 'Edit with agent'. Metadata changes use a registry metadata version; executable changes create immutable revisions."

User Scenarios

Story 1 — View Scenario in an Editor (P1)

Why P1: Users must inspect a scenario as a structured document, not a chat transcript.

Independent Test: Load a persisted scenario into the editor and verify graph, steps, parameters, baselines, and assertions render read-only by default.

Acceptance:

  1. Given a scenario is opened When the editor renders Then steps, dependencies (graph), parameters, baselines, and assertions are shown; generated executable artifacts are read-only.
  2. Given a step is opened When inspected Then inputs, outputs, expected result, baseline, automation mode, and evidence requirements are visible.
  3. Given no edit session is active When the page loads Then the view is read-only and no revision is created.

Story 2 — Edit Business Fields and Parameters Manually (P1)

Why P1: Name, description, tags, and parameter values are safe to edit directly.

Independent Test: Modify a scenario name and verify metadata_version advances without a revision; modify a parameter default and verify a new executable revision is produced.

Acceptance:

  1. Given the user edits business fields (name/description/tags) When saved Then only ScenarioRegistryEntry.metadata_version advances; no executable revision is created.
  2. Given the user edits a parameter definition default When validated Then type validation and dependent-step readiness update, and a new executable revision is created.
  3. Given a durable edit is requested When it passes policy evaluation Then it is saved as an immutable revision or rendered as an inline ActionApprovalGate; an agent-saved revision retains delegated provenance.

Story 3 — Edit Assertions With a Constrained Editor (P2)

Why P2: Assertions must change without allowing SQL/raw baselines to be injected.

Independent Test: Change an assertion's comparison operator and threshold via the constrained editor and verify the graph stays safe.

Acceptance:

  1. Given an assertion is edited When the constrained editor is used Then only registered operators/baseline references are selectable; free-form SQL/raw values are forbidden.
  2. Given an invalid assertion is submitted When validated Then it is rejected with the 038 validator findings, never saved as-is.

Story 4 — Edit Dependencies Visually (P2)

Why P2: Reordering/adding/removing step dependencies is a graph operation best done visually.

Independent Test: Re-parent a step in the visual DAG editor and verify dependency validation and a new revision.

Acceptance:

  1. Given the user drags a dependency edge When changed Then the graph revalidates for cycles/duplicate outputs per 038 rules.
  2. Given a step is added/removed When saved Then affected step ids/order change only as needed and a new revision records the change.

Story 5 — Agent-Assisted Complex Edit (P3)

Why P3: Some changes (new checklist case, complex assertion) are easier described in natural language.

Independent Test: Request "add XLSX comparison" via Edit-with-agent and verify the agent creates a validated WorkingDraft and saves a revision when delegated policy permits.

Acceptance:

  1. Given a complex change request When "Edit with agent" runs Then the agent creates a server-stored EditProposal.
  2. Given the proposal is accepted When its base revision is still current Then it becomes a WorkingDraft; a policy-authorized analyst or agent save creates the revision. A stale proposal never saves; a non-delegated action waits at an inline gate.

Story 6 — Revalidate a Stale Scenario (P2) (#9)

Why P2: A scenario that references changed charts/filters must migrate cleanly, not just flip a lifecycle flag.

Independent Test: Mark a scenario stale via 042, run revalidate, and verify a proposed r18 with automatic mappings, manual conflicts, and a diff for approval.

Acceptance:

  1. Given a scenario is stale When "Revalidate" runs Then affected refs (chart/filter/metric) are listed and mapped against the current dashboard (038 validator + 037/041).
  2. Given some mappings are ambiguous When conflicts exist Then they are surfaced for manual resolution, not auto-accepted.
  3. Given a proposal is generated When shown Then a diff (r17→r18) is displayed; the agent may save it after validation if delegated policy allows, otherwise the linked inline gate decides it.

Edge & Failure Cases

# Scenario Expected Behavior Recovery
E1 Concurrent edit (409) Persistent conflict panel with reload/compare/discard Reload / compare / discard
E2 Invalid assertion/SQL/raw baseline Rejected with validator findings Correct via constrained editor
E3 Cycle introduced by dependency edit Rejected with cycle path Revert edge
E4 Agent proposes unsafe revision Blocked by validator before save Describe differently / manual edit
E5 RBAC edit denied 403 permission_denied, no confirm control Contact admin

Requirements

Functional

  • SCEDIT-FR-001: The editor MUST render a persisted scenario as a structured document with steps, dependency graph, parameters, baselines, assertions, and read-only generated executable.
  • SCEDIT-FR-002: Metadata fields (name, description, tags) MUST be editable manually through metadata_version without creating an executable revision; ParameterDefinition/default and graph edits MUST create one.
  • SCEDIT-FR-003: Assertion editing MUST use a constrained editor (registered operators + baseline references); free-form SQL, shell, paths, and raw numeric baseline literals MUST be forbidden.
  • SCEDIT-FR-004: Dependency editing MUST be visual (graph), with 038 cycle/duplicate-output validation on every change.
  • SCEDIT-FR-005: Every durable executable graph edit MUST produce a new immutable revision; metadata uses metadata_version concurrency. Deterministic policy decides whether the actor/agent may save immediately or must obtain an inline ActionApprovalGate.
  • SCEDIT-FR-005a: Every proposal/diff MUST expose Verification Program changes: SqlEvidenceSpec template/hash/relation refs, TransformSpec operations, assertions, AgentEvaluationSpec and DecisionPolicy. A runtime finding can only change this content through a new validated proposal and revision.
  • SCEDIT-FR-006: "Edit with agent" MUST create a server-stored validated proposal/draft with a diff. A delegated agent MAY save the revision; no agent edit may bypass 038 validation, 042 revision provenance, or a required gate.
  • SCEDIT-FR-009: Every save MUST use a server-stored WorkingDraft (draft_id + digest); the client MUST NOT return the full graph to save (no arbitrary-draft bypass). Save re-validates/canonicalizes/hashes server-side.
  • SCEDIT-FR-010: A stale scenario MUST be revalidatable: affected refs mapped against the current dashboard, conflicts surfaced manually, a proposed revision + diff shown for approval (Scenario Migration workflow).
  • SCEDIT-FR-007: RBAC MUST enforce scenario:edit separately from scenario:run.
  • SCEDIT-FR-008: All UI MUST follow Svelte 5 runes/model-first conventions and remain keyboard-accessible.

Key Entities

  • ScenarioEditorModel: Frontend screen model (.svelte.ts) holding the open scenario revision, edit buffer, dirty state, and revision save actions.
  • ConstrainedAssertionEditor: Editor restricting assertion edits to registered operators and baseline references.
  • VisualDagEditor: Graph canvas for step/dependency editing with validation feedback.
  • EditRevisionResult: New revision + change summary + diff produced by a durable edit.

Success Criteria

  • SC-001: A persisted scenario renders read-only by default and enters edit mode only on an explicit action.
  • SC-002: Manual field/parameter edits produce a new immutable revision with change summary in model tests.
  • SC-003: 100% of attempted SQL/raw-baseline/path injections via the assertion editor are rejected.
  • SC-004: Dependency edits that introduce a cycle or duplicate output are rejected with a precise error.
  • SC-005: Every durable edit is immutable, validated and fully attributed; agent edits expose their diff/provenance and are either policy-authorized or ActionApprovalGate-bound.
  • SC-006: Edit model C (hybrid) is enforced: manual fields/params, constrained assertions, visual deps, read-only generated executable.

Clarifications

Session 2026-08-07

  • Q: Which edit model? → A: C (hybrid) — business fields/params manual, assertions constrained, dependencies visual, generated executable read-only, complex changes "Edit with agent".
  • Q: May the agent save a revision? → A: Yes, after deterministic validation when delegated policy permits. It always creates an immutable revision with agent/delegator/case provenance; policy-gated actions use an inline ActionApprovalGate.
  • Q: Does this replace 039 workspace? → A: No. 039 is the create flow in agent chat; 043 is the post-save edit surface over the registry.

Implementation Status & MVP Debt (factual audit 2026-08-20)

WorkingDraft/proposal services, constrained assertion and DAG components, plus the editor route are present. The hybrid editor therefore exists structurally but is not verified as a closed workflow.

  • [~] Server-side drafts/proposals and UI components are present; no current route-level E2E, accessibility, or policy-gate evidence is retained.
  • [~] Revalidation depends on 042 staleness input whose upstream 037/041 integration is unproven.
  • [ ] Feature closure is blocked by incomplete 044 execution and unperformed independent verification.

Drift Amendment — MCP Interface (2026-08-24)

  • "Edit with agent" (Story 5) continues with external MCP clients: the proposal/WorkingDraft flow, stale-proposal rejection and gate-bound saves are unchanged; only the conversation medium moves out of the product.

Status (2026-09-02): done — реализовано в рамках 050: инструменты и гейты (specs/050-mcp-interface/tasks.md T012–T028 [x]), handoff-поверхность (050 T030–T033), демонтаж чата и сервиса agent/ (050 T040–T041, чекпоинты specs/WORKSTATE-043-047.md).

#endregion ScenarioEditor.Spec


UX REFERENCE — Interaction Narrative

Source: ux_reference.md

#region ScenarioEditor.UxReference [C:3] [TYPE ADR] [SEMANTICS ux,reference,scenario,editor] @BRIEF UX interaction reference for the Scenario Editor (043) — hybrid edit model C.

Feature Branch: 043-dashboard-scenario-editor | Created: 2026-08-07

1. User Persona & Context

  • User: BI analyst / quality engineer.
  • Goal: Review and adjust a saved dashboard test scenario; add/remove steps, change assertions, edit parameters.
  • Context: Browser, scenario detail → Edit.

2. Happy Path

The analyst edits directly in the persistent editor or asks the agent for a complex change. The agent creates a validated draft, can save an immutable revision under delegated policy, then runs verification; a non-delegated action appears as an inline approval card with diff/provenance.

3. Screens & States

Screen: Scenario Editor

  • Layout: Left = graph canvas; right = inspector panel for selected step; top = toolbar (Edit mode toggle, Save, "Edit with agent"); footer = validation status + diff.
  • Read-only default: view renders graph/params/assertions; Edit mode toggles editing.
  • ConstrainedAssertionEditor: operator select + threshold + baseline ref; no free-text expected value.
  • VisualDagCanvas: drag dependency edges; live cycle/duplicate validation.
  • @UX_STATE: view_only, editing, saving, saving_conflict(409), saving_error, agent_proposing, proposal_diff, validation_error.
  • @UX_RECOVERY: 409 → persistent reload/compare/discard panel; validation_error → show findings; agent proposal → inspect policy/diff/action timeline.

4. Error Experience

  • 409 concurrent edit → persistent conflict panel with Reload, Compare and Discard.
  • 422 invalid edit → inline findings (e.g., "SQL not allowed in assertion").
  • 403 edit denied → permission_denied, no confirm control.

5. Tone & Voice

  • Style: Concise, technical. Terminology: scenario/step/assertion/revision.

Edge & Failure Matrix (feed to prototype)

NET_01/02/03, VAL_01/02, AUTH_01/02, NF_01, CONF_01/02, 422, 429, 5XX, EMPTY, MALFORMED, A11Y, RESP.

#endregion ScenarioEditor.UxReference


CHECKLISTS — Requirements Quality — requirements.md

Source: checklists/requirements.md

Requirements Checklist: Scenario Editor (043)

Purpose: Verify SCEDIT-FR-001..008 completeness. | Created: 2026-08-07

Factual audit 2026-08-20: [x] requires current production evidence, [~] means partial code exists, [ ] means missing integration or proof.

View (FR-001)

  • [~] CHK001 Persisted scenario renders read-only by default
  • CHK002 Steps, dependency graph, parameters, baselines, assertions visible
  • CHK003 Generated executable read-only

Manual Edit (FR-002/005)

  • CHK004 Business fields + parameters editable manually
  • CHK005 Durable edit → new immutable revision + change summary
  • CHK006 Every save is policy-authorized, immutable and attributed; non-delegated actions expose an inline ActionApprovalGate

Constrained Assertions (FR-003)

  • [~] CHK007 Assertions use constrained editor (registered operators + baseline refs)
  • [~] CHK008 SQL/raw-baseline/path injection rejected

Visual Dependencies (FR-004)

  • [~] CHK009 Dependency editing via visual DAG
  • [~] CHK010 Cycle/duplicate-output validation on every change

Agent Edit (FR-006)

  • CHK011 "Edit with agent" proposes revision with diff
  • CHK012 Proposal validated before display; any agent save is delegated, provenance-bearing and policy checked

RBAC / UX (FR-007/008)

  • CHK013 scenario:edit distinct from scenario:run
  • CHK014 Svelte 5 runes, keyboard-accessible

Success Criteria

  • CHK015 SC-001..006 verified (read-only default, revision, rejection, delegated policy, hybrid C)

UX DECISIONS — Final Choices

Source: contracts/ux/decisions.md

#region ScenarioEditor.Ux.Decisions [C:3] [TYPE ADR] [SEMANTICS scenario,editor,ux,decisions] @BRIEF Final UX decisions for the Scenario Editor (043). @RELATION DEPENDS_ON -> [ScenarioEditor.Spec]

Decision 1 — Hybrid edit model (C)

Business fields/parameters manual, assertions constrained, dependencies visual, generated executable read-only, complex changes "Edit with agent". Compatible with 038 safety and mandatory human participation.

Decision 2 — Read-only default

Editor loads read-only; Edit mode is explicit. No revision is created until a durable edit passes delegated-action policy.

Decision 3 — Constrained assertion editor

Only registered operators (037 comparison) + baseline refs; free-form SQL/raw baseline/path forbidden.

Decision 4 — Visual DAG with live validation

Dependency edits revalidate for cycles/duplicate outputs (038) on every change.

Decision 5 — Agent edits never silent

"Edit with agent" proposes a full revision + diff; a delegated agent may save it after validation, otherwise an inline ActionApprovalGate decides it. #endregion ScenarioEditor.Ux.Decisions


PLAN — Implementation Plan

Source: plan.md

Implementation Plan: Scenario Editor UX

Branch: 043-dashboard-scenario-editor | Date: 2026-08-07 | Spec: spec.md | Status: Partially implemented — factual audit pending remediation

Implementation audit, 2026-08-20: editor code and route exist; current E2E/accessibility/policy verification is open, and closure remains dependent on unresolved 042/044 runtime contracts.

Summary

A first-class editor for viewing and editing persisted scenarios (hybrid edit model C): manual business fields/parameters, constrained assertion editing, visual DAG dependency editing, read-only generated executable, and agent-assisted complex edits. Every durable edit produces a new immutable revision via 042 under delegated-action policy.

Technical Context

Language/Version: Python 3.13+ (backend), TypeScript + Svelte 5 runes (frontend) Primary Dependencies: FastAPI, Pydantic 2; existing 038 compiler/validator/resolver, 042 registry Storage: none new (registry revisions via 042); edit sessions in-memory Testing: vitest (L1 model + L2 UX with @testing-library/svelte), pytest (edit op validation) Frontend Architecture: ScenarioEditorModel.svelte.ts, model-first, runes-only Performance Goals: edit-op apply < 100ms; revision save < 200ms; visual DAG revalidation < 100ms Constraints: hybrid edit model; no SQL/raw-baseline injection; deterministic delegated-action policy; RBAC scenario:edit Scale: up to 100-step graphs, dozens of edit ops per session

Constitution Check

Principle Result
I. Semantic Contract First PASS — Load/ApplyOps/SaveRevision/AgentProposal contracted
II. Decision Memory PASS — research R1-R4
VI. Svelte 5 Runes Only PASS — model-first .svelte.ts
VII. Test-Driven C3+ PASS — invalid-edit tests first
V. RBAC PASS — scenario:edit distinct from run
VIII. Attention-Optimized PASS

Project Structure

specs/043-dashboard-scenario-editor/
├── spec.md / data-model.md / research.md / plan.md / tasks.md / traceability.md / quickstart.md / ux_reference.md
├── checklists/requirements.md
├── contracts/modules.md, contracts/openapi.yaml, contracts/ux/
└── prototype/index.html + manifest.md

backend/src/services/dashboard_testing/editor/ (ops.py, save.py, agent.py, assert_validate.py)
frontend/src/lib/models/ScenarioEditorModel.svelte.ts
frontend/src/lib/components/scenario-editor/ (StepCard, VisualDagCanvas, ConstrainedAssertionEditor, EditRevisionDiff, AgentActionPanel)
frontend/src/routes/dashboard-testing/scenarios/[id]/edit/+page.svelte

Delivery Phases

  1. EditOperation DTOs + validation (tests first).
  2. Load read-only editor + ApplyOps.
  3. SaveRevision (delegated policy + 042 revision chain).
  4. Constrained assertion editor.
  5. Visual DAG dependency editor.
  6. Agent-assisted edit proposal.
  7. Frontend model/components, polish, regression gates.

Traceability

traceability.md maps Story → model → operationId → contract → task → test.

Cross-Spec Boundary

  • Reads persisted scenarios from 042; writes new revisions to 042.
  • Reuses 038 validator/resolver and 037 comparison operators.
  • "Edit with agent" reuses agent tooling from 036/038.

Complexity Tracking

No exception planned. Edit ops are bounded C3-C4; visual DAG and constrained assertion stay model-first.


RESEARCH — Technical Decisions

Source: research.md

Scenario Editor — Phase 0/1 Research (043)

Branch: 043-dashboard-scenario-editor | Date: 2026-08-07 | Spec: spec.md

R1. Edit model — hybrid (C)

Decision: Hybrid C — business fields/parameters manual, assertions constrained, dependencies visual, generated executable read-only, complex changes "Edit with agent".

Rationale: Matches mandatory human participation, preserves 038 safety, and supports the most common user tasks without a full free-form DAG editor.

Alternatives: A (agent-only editing) rejected — no visual review; B (free-form DAG editor) rejected — safety violations.

Impact: EditOperation union, constrained assertion editor, visual DAG editor, agent-proposal path.

R2. Edit → revision mapping

Decision: Each durable edit maps to EditOperation[], validated by 038, then persisted as a new immutable revision via 042 revision chain.

Rationale: Reuses 038 validator and 042 revision semantics; prior revisions immutable.

Alternatives: in-place graph mutation (rejected: breaks reproducibility).

Impact: Save flow = validate(ops) → create_revision(042) → EditRevisionResult.

R3. Constrained assertion editor

Decision: Assertion edits select from registered operators (037 comparison) + baseline refs only; free-form expected values forbidden.

Rationale: Prevents SQL/raw-baseline injection, consistent with 038 forbidden-content rules.

Alternatives: free-text assertion (rejected).

Impact: ConstrainedAssertionEdit DTO + validator gate.

R4. Visual dependency editing

Decision: DAG canvas editing with 038 cycle/duplicate validation on every change; generated executable read-only.

Rationale: Graph ops are best done visually but must remain safe.

Impact: VisualDependencyEdit, validation on change.

Contracts & API

  • contracts/modules.md — Editor.Load, Editor.ApplyOps, Editor.SaveRevision, Editor.AgentProposal, Editor.ValidateAssertion.
  • OpenAPI: POST /scenarios/{id}/edits/apply, POST /scenarios/{id}/edits/save, POST /scenarios/{id}/edits/agent-propose, GET /scenarios/{id}/edits/diff.

Constitution Check

Principle Result
I. Semantic Contract First PASS — edit ops/save/agent-proposal contracted
II. Decision Memory PASS — R1-R4
VI. Svelte 5 Runes Only PASS — ScenarioEditorModel .svelte.ts, model-first
VII. Test-Driven C3+ PASS — invalid-edit/cycle/SQL tests first
V. RBAC PASS — scenario:edit distinct from run
VIII. Attention-Optimized PASS

DATA MODEL — Entities & Relations

Source: data-model.md

#region ScenarioEditor.DataModel [C:4] [TYPE ADR] [SEMANTICS data-model,scenario,editor,edit,revision] @BRIEF Edit-session, edit-operation, constrained-assertion, and visual-dependency edit models for the Scenario Editor (043). @RELATION DEPENDS_ON -> [ScenarioEditor.Research] @RATIONALE Editing a persisted scenario must produce typed, validatable edit operations that translate into new immutable revisions without violating 038 safety invariants. @REJECTED Free-form graph mutation — rejected because it would allow SQL/raw baselines/paths into the graph; edits are constrained and validator-gated.

EditSession

Fields: session_id, scenario_id, base_revision_id, owner_id, opened_at, dirty (bool), active (bool). One active session per scenario per user; opening an edit pins base_revision_id for conflict detection.

EditOperation

Typed discriminant union:

  • op=set_parameter_definition {param_name, default?, validation?, source?}
  • op=set_assertion {logical_step_id, comparison: enum, threshold, baseline_ref}
  • op=add_step {template, after_logical_step_id?}
  • op=remove_step {logical_step_id}
  • op=set_dependency {logical_step_id, add|remove, target_logical_step_id}

Invariant: op payloads only reference registered templates/operators/baseline refs; extra="forbid" on all edit ops (no SQL/code/path smuggling). Metadata uses PATCH /metadata with If-Match: metadata_version and never enters WorkingDraft. Every WorkingDraft operation is executable and creates a revision only on policy-authorized save. Parameter operations edit 038 ParameterDefinition, never a runtime value. Each durable step op carries a logical_step_id (immutable) for stable analytics.

WorkingDraft — server-stored, no client-draft bypass (#7)

apply persists a WorkingDraft server-side and returns {draft_id, digest}. save(draft_id, digest) reloads it server-side, re-validates the Verification Program (including SqlEvidenceSpec compilation), canonicalizes, re-hashes, and compares the base revision. The client NEVER returns the full graph; save cannot accept an arbitrary draft object. A WorkingDraft is bound to a base revision and expires on conflict/staleness.

Fields: draft_id, scenario_id, base_revision_id, ops[], applied_graph, digest, created_by, delegated_by?, agent_run_id?, investigation_case_id?, created_at, status (open|saved|awaiting_approval|expired).

ConstrainedAssertionEdit

Fields: logical_step_id, operator (enum from 037 comparison: exact, absolute, relative, range, row_set), threshold (typed), baseline_ref (approved baseline or candidate), evidence_required (bool). Free-form expected value forbidden.

VisualDependencyEdit

Fields: logical_step_id, target_logical_step_id, action (add|remove). Validated against 038 DAG rules (cycles, duplicate producers).

EditRevisionResult

Fields: new_revision_id, parent_revision_id, content_hash, activation_status=candidate, server_derived_change_summary {added, changed, removed, verification_program_diff}, validation (AuthoringValidation), diff_payload, policy_decision, agent_action_id?. Returned by save; never mutates prior revisions or advances current_revision. The caller submits only {draft_id, digest} plus its authenticated/delegated action identity; it never supplies audit/provenance summary. Promotion is the distinct 042 ActivateCurrentRevision operation, with its own deterministic eligibility and delegated-authority/gate decision.

EditProposal and ScenarioMigration

EditProposal { proposal_id, scenario_id, base_revision_id, proposed_graph, diff, validation, expires_at, agent_run_id?, investigation_case_id? } is server-stored. acceptProposal(proposal_id) validates its base revision and produces a WorkingDraft, which then uses normal save(draft_id,digest). A delegated agent may invoke that save after deterministic validation; if policy requires approval, the draft moves to awaiting_approval and an inline ActionApprovalGate is linked. revalidate creates a migration proposal with mappings/conflicts; conflicts must be resolved before accept. There is no client-side graph handoff.

Conflict & Safety

  • Save requires base_revision_id match → 409 STALE_REVISION on mismatch.
  • Every edit operation validated by 038 validator before revision creation.
  • Agent proposals carry a full proposed graph which passes validation before display; an agent-saved revision retains AgentAction/InvestigationCase provenance. A save is not permission to activate it or change an automation target.

#endregion ScenarioEditor.DataModel


CONTRACTS — Module & Function Contracts

Source: contracts/modules.md

#region ScenarioEditor.Modules [C:4] [TYPE ADR] [SEMANTICS scenario,editor,contracts,modules,edit] @BRIEF Module contracts for the Scenario Editor (043) — load, apply typed edits, save revisions, agent proposals. @defgroup ScenarioEditor View and edit persisted scenarios with a hybrid edit model. @RELATION DEPENDS_ON -> [ScenarioRegistry.Modules] @RELATION DEPENDS_ON -> [ScenarioGraph.Validator] @RELATION DEPENDS_ON -> [ScenarioGraph.Resolver] @RATIONALE Editing is validated, typed, revision-bound and fully attributed; delegated agents may save only through deterministic policy. @REJECTED Free-form graph mutation; agent-only editing; unconstrained assertion text.

#region ScenarioEditor.Load [C:3] [TYPE Function] [SEMANTICS scenario,editor,load,readonly]

@ingroup ScenarioEditor

@BRIEF Load a scenario revision into a read-only editor view.

@PRE caller has scenario:view; scenario_id exists.

@POST returns editable DTOs with dirty=false and base_revision_id; read-only until Edit.

def load_editor(db, scenario_id, revision_id): ...

#endregion ScenarioEditor.Load

#region ScenarioEditor.ApplyOps [C:4] [TYPE Function] [SEMANTICS scenario,editor,ops,validate]

@ingroup ScenarioEditor

@BRIEF Apply typed edit operations; persist a WorkingDraft server-side; validate.

@PRE ops are EditOperation[] with valid payloads; base revision matches.

@POST returns {draft_id, digest, validation findings}; unsafe ops rejected; WorkingDraft persisted server-side.

@SIDE_EFFECT in-memory + server WorkingDraft (no client-graph return).

@INVARIANT no SQL/code/path/raw-baseline can enter via ops; client never returns the full graph.

@TEST_EDGE SQL in assertion->rejected; cycle in dependency->rejected; raw baseline->rejected.

def apply_ops(db, scenario_id, base_revision_id, ops): ...

#endregion ScenarioEditor.ApplyOps

#region ScenarioEditor.SaveRevision [C:4] [TYPE Function] [SEMANTICS scenario,editor,save,revision,hitl]

@ingroup ScenarioEditor

@BRIEF Save a server-stored WorkingDraft by id + digest as a new immutable revision after policy evaluation.

@PRE draft_id+digest match; draft re-validated/canonicalized/hashed server-side; actor/delegated AgentAction authorized; base revision matches.

@POST new candidate revision via 042; current_revision unchanged; 409 on stale base or digest mismatch.

@SIDE_EFFECT DB write; 042 revision; audit.

@INVARIANT prior revisions immutable; never saved without policy authorization; no arbitrary-draft bypass.

def save_revision(db, draft_id, digest, actor, agent_action_id=None): ...

#endregion ScenarioEditor.SaveRevision

#region ScenarioEditor.Revalidate [C:4] [TYPE Function] [SEMANTICS scenario,editor,revalidate,migration]

@ingroup ScenarioEditor

@BRIEF Stale-scenario migration: revalidate against current dashboard, produce proposed revision with diff.

@PRE scenario stale (NEEDS_REVALIDATION); current dashboard query model available.

@POST proposes r18 with automatic mappings + manual conflicts + diff; nothing persists until a policy-authorized save.

@SIDE_EFFECT reads 037/041; read-only proposal.

@INVARIANT proposal passes 038 validation; any agent save is delegated, provenance-bearing and policy-checked.

def revalidate(db, scenario_id, base_revision_id): ...

#endregion ScenarioEditor.Revalidate

#region ScenarioEditor.AgentProposal [C:4] [TYPE Function] [SEMANTICS scenario,editor,agent,proposal,diff]

@ingroup ScenarioEditor

@BRIEF Generate an agent-assisted edit proposal as a full revision with diff.

@PRE request is bounded; proposal validated before display.

@POST returns proposed graph + diff + validation; a delegated agent may convert it into a saved immutable revision through SaveRevision.

@SIDE_EFFECT LLM call (agent tool); read-only draft.

@INVARIANT proposal passes 038 validation; never bypasses policy, immutable revision creation, or a required gate.

def agent_propose(db, scenario_id, base_hash, request): ...

#endregion ScenarioEditor.AgentProposal

#region ScenarioEditor.ValidateAssertion [C:3] [TYPE Function] [SEMANTICS scenario,editor,assertion,constrain]

@ingroup ScenarioEditor

@BRIEF Validate a constrained assertion edit.

@POST accepts registered operator + baseline ref; rejects free-form values.

def validate_assertion(edit): ...

#endregion ScenarioEditor.ValidateAssertion

#endregion ScenarioEditor.Modules


OPENAPI — REST/Event API Contract

Source: contracts/openapi.yaml

openapi: 3.1.0 info: title: Scenario Editor API version: 0.1.0 description: View and edit persisted scenarios with a hybrid edit model (043). paths: /api/dashboard-testing/scenarios/{scenario_id}/metadata: patch: operationId: editor.updateMetadata summary: Update registry metadata without creating an executable revision security: [{ bearerAuth: [] }] parameters: - { name: scenario_id, in: path, required: true, schema: { type: string } } - { name: If-Match, in: header, required: true, schema: { type: string }, description: "ScenarioRegistryEntry.metadata_version" } requestBody: required: true content: application/json: schema: type: object properties: name: { type: string } description: { type: string } tags: { type: array, items: { type: string } } responses: { "200": { description: "ScenarioRegistryEntry with advanced metadata_version" }, "409": { description: Stale metadata version } } /api/dashboard-testing/scenarios/{scenario_id}/edits/apply: post: operationId: editor.apply summary: Apply typed edit operations; persists a WorkingDraft server-side security: [{ bearerAuth: [] }] parameters: [{ name: scenario_id, in: path, required: true, schema: { type: string } }] requestBody: required: true content: application/json: schema: type: object required: [base_revision_id, ops] properties: base_revision_id: { type: string } ops: type: array items: oneOf: - { $ref: "#/components/schemas/SetParameterOp" } - { $ref: "#/components/schemas/SetAssertionOp" } - { $ref: "#/components/schemas/AddStepOp" } - { $ref: "#/components/schemas/RemoveStepOp" } - { $ref: "#/components/schemas/SetDependencyOp" } responses: "200": { description: "WorkingDraft { draft_id, digest, validation findings }" } "422": { description: Invalid edit (SQL/raw-baseline/cycle rejected) } "409": { description: STALE_REVISION } /api/dashboard-testing/scenarios/{scenario_id}/edits/save: post: operationId: editor.save summary: Save a server-stored WorkingDraft by id + digest after deterministic delegated-action policy evaluation security: [{ bearerAuth: [] }] parameters: [{ name: scenario_id, in: path, required: true, schema: { type: string } }] requestBody: required: true content: application/json: schema: type: object required: [draft_id, digest] properties: draft_id: { type: string } digest: { type: string } agent_action_id: { type: string, nullable: true, description: "Server-owned delegated AgentAction identity; audit/provenance is derived server-side" } responses: "200": { description: New immutable candidate revision with policy/provenance result; activation is a separate 042 operation } "202": { description: Draft awaiting inline ActionApprovalGate } "409": { description: STALE_REVISION / digest mismatch } "422": { description: Draft invalid after re-validation (no arbitrary-draft bypass) } /api/dashboard-testing/scenarios/{scenario_id}/edits/agent-propose: post: operationId: editor.agentPropose summary: Generate an agent-assisted edit proposal with diff security: [{ bearerAuth: [] }] parameters: [{ name: scenario_id, in: path, required: true, schema: { type: string } }] requestBody: required: true content: application/json: schema: type: object required: [base_revision_id, request] properties: base_revision_id: { type: string } request: { type: string } responses: "200": { description: "Server-stored EditProposal { proposal_id, base_revision_id, graph, diff, validation }" } /api/dashboard-testing/scenarios/{scenario_id}/edits/proposals/{proposal_id}/accept: post: operationId: editor.acceptProposal summary: Accept a validated proposal into a server-stored WorkingDraft security: [{ bearerAuth: [] }] parameters: - { name: scenario_id, in: path, required: true, schema: { type: string } } - { name: proposal_id, in: path, required: true, schema: { type: string } } responses: { "200": { description: "WorkingDraft { draft_id, digest, validation }" }, "409": { description: Proposal/base revision stale } } /api/dashboard-testing/scenarios/{scenario_id}/revalidate: post: operationId: editor.revalidate summary: Produce a migration proposal with automatic mappings and manual conflicts security: [{ bearerAuth: [] }] parameters: [{ name: scenario_id, in: path, required: true, schema: { type: string } }] responses: { "200": { description: "MigrationProposal { proposal_id, mappings, conflicts, diff }" } } /api/dashboard-testing/scenarios/{scenario_id}/migration-proposals/{proposal_id}/resolve: post: operationId: editor.resolveMigrationProposal summary: Resolve explicit migration conflicts before proposal acceptance security: [{ bearerAuth: [] }] parameters: - { name: scenario_id, in: path, required: true, schema: { type: string } } - { name: proposal_id, in: path, required: true, schema: { type: string } } requestBody: required: true content: application/json: schema: type: object required: [resolutions] properties: resolutions: { type: array, items: { type: object, required: [conflict_id, target_ref], properties: { conflict_id: { type: string }, target_ref: { type: string } } } } responses: { "200": { description: Updated server-stored MigrationProposal }, "409": { description: Proposal stale or conflict already resolved } } components: securitySchemes: bearerAuth: { type: http, scheme: bearer } schemas: SetParameterOp: type: object required: [op, param_name, value] properties: op: { type: string, enum: [set_parameter_definition] } param_name: { type: string } value: { description: "JSON value validated server-side against ParameterDefinition.type" } SetAssertionOp: type: object required: [op, logical_step_id, comparison, baseline_ref] properties: op: { type: string, enum: [set_assertion] } logical_step_id: { type: string } comparison: { type: string, enum: [exact, absolute, relative, range, row_set] } baseline_ref: { type: string } threshold: { type: number } AddStepOp: type: object required: [op, template] properties: op: { type: string, enum: [add_step] } template: { type: string } after_logical_step_id: { type: string, nullable: true } RemoveStepOp: type: object required: [op, logical_step_id] properties: op: { type: string, enum: [remove_step] } logical_step_id: { type: string } SetDependencyOp: type: object required: [op, logical_step_id, action, target_logical_step_id] properties: op: { type: string, enum: [set_dependency] } logical_step_id: { type: string } action: { type: string, enum: [add, remove] } target_logical_step_id: { type: string }


QUICKSTART — Dev Onboarding

Source: quickstart.md

Quickstart: Scenario Editor (043)

Factual audit 2026-08-20: pending verification checklist only; no command result below is currently asserted as evidence.

Prereqs

  • 042 registry backend live (scenario + revisions)
  • Frontend deps installed

Commands

# Backend editor op validation
cd backend && source .venv/bin/activate && python -m pytest -v tests/services/dashboard_testing/editor/

# Frontend model + UX
cd frontend && npm run test -- ScenarioEditor

# Lint
cd backend && python -m ruff check src/services/dashboard_testing/editor/
cd frontend && npm run lint

Exit Gates

  • Read-only view renders persisted scenario
  • SQL/raw-baseline injection rejected in assertion editor
  • Cycle/duplicate rejected in DAG editor
  • Save produces new immutable revision after delegated-policy evaluation or an inline ActionApprovalGate
  • Agent proposal shows diff and provenance; delegated agent save remains validator/policy checked
  • ruff clean; prototype states covered

TRACEABILITY — Requirements Matrix

Source: traceability.md

Traceability: Scenario Editor (043)

Factual audit 2026-08-20: rows identify code/test ownership only. Route-level accessibility, stale-conflict and policy-gate workflow evidence is pending; 043 remains dependent on 042/044.

Story Requirement Model API operationId Contract Task Test
US1 View SCEDIT-FR-001 ScenarioEditorModel editor.load Editor.Load T003-T005 ScenarioEditorModel.test
US2 Manual SCEDIT-FR-002/005 EditOperation editor.apply, editor.save Editor.ApplyOps, Editor.SaveRevision T006-T009 test_ops, edit.ux.test
US3 Assertion SCEDIT-FR-003 ConstrainedAssertionEdit editor.validate-assertion Editor.ValidateAssertion T010-T011 edit.ux.test
US4 Deps SCEDIT-FR-004 VisualDependencyEdit editor.apply Editor.ApplyOps T012-T013 test_ops
US5 Agent SCEDIT-FR-006 EditRevisionResult editor.agent-propose Editor.AgentProposal T014-T015 test_editor_agent
RBAC/UX SCEDIT-FR-007/008 — — — T016 test_rbac

N/A: Run Monitor (045), Automation (046), Analytics (047).


TASKS — Implementation Tasks

Source: tasks.md

#region ScenarioEditor.Tasks [C:3] [TYPE ADR] [SEMANTICS tasks,scenario,editor,implementation] @BRIEF Ordered TDD backlog for the Scenario Editor (043). Tests FIRST for every C3+ contract.

Prerequisites: plan.md, spec.md; contracts/modules.md, traceability.md.

Format: - [ ] T### [P] [USx] Description with exact file path

Factual audit 2026-08-20: [x] means code plus relevant evidence; [~] means partial implementation; [ ] means absent integration or unperformed verification.

Phase 1 — Setup

  • T001 Create EditOperation DTOs (union) + fixtures in backend/src/services/dashboard_testing/editor/ops.py and specs/043-dashboard-scenario-editor/fixtures/
  • T002 [P] Materialize fixtures to backend/tests/fixtures/scenario_editor/

Phase 2 — US1 View (Read-only)

  • T003 [US1] L1 model test for ScenarioEditorModel load in frontend/src/lib/models/__tests__/ScenarioEditorModel.test.ts
  • T004 [US1] Implement load_editor in backend/src/services/dashboard_testing/editor/ and ScenarioEditorModel.svelte.ts @POST: read-only draft, dirty=false, base_revision_id set
  • T005 [US1] Build StepCard + graph rendering in frontend/src/lib/components/scenario-editor/ and route frontend/src/routes/dashboard-testing/scenarios/[id]/edit/+page.svelte

Phase 3 — US2 Manual Fields + Parameters

  • T006 [US2] Write failing apply-op tests (SQL/raw-baseline/cycle rejection) in backend/tests/services/dashboard_testing/registry/test_scenario_editor_ops.py
  • T007 [US2] Implement apply_ops + validate_assertion in backend/src/services/dashboard_testing/editor/ @INVARIANT: no SQL/code/path/raw-baseline via ops @TEST_EDGE: SQL in assertion->rejected; cycle->rejected; raw baseline->rejected
  • T008 [US2] Implement save_revision (delegated policy + 042) in backend/src/services/dashboard_testing/editor/save.py @POST: new revision or inline ActionApprovalGate; 409 stale base; never without policy authorization
  • T009 [US2] L2 UX test for field/parameter edit → revision in frontend/src/lib/models/__tests__/ScenarioEditorModel.test.ts

Phase 4 — US3 Constrained Assertion Editor

  • T010 [US3] L2 UX test for constrained assertion editing in frontend/src/lib/components/scenario-editor/__tests__/ConstrainedAssertionEditor.test.ts
  • T011 [US3] Build ConstrainedAssertionEditor.svelte (registered operators + baseline refs) @TEST_INVARIANT: No_Sql_Injection → verified by rejection test

Phase 5 — US4 Visual DAG Editor

  • T012 [US4] Write failing visual-dependency validation tests (cycle/duplicate) in backend/tests/services/dashboard_testing/registry/test_scenario_editor_ops.py
  • T013 [US4] Build VisualDagCanvas.svelte with duplicate/self validation on change

Phase 6 — US5 Agent-Assisted Edit

  • T014 [US5] Implement agent_propose in backend/src/services/dashboard_testing/editor/agent.py @INVARIANT: proposal passes validation; agent save is fully attributed and policy checked @TEST_INVARIANT: covered by test_scenario_editor_agent.py + route tests (agent-propose/proposal-save)
  • T015 [US5] Build persistent AgentActionPanel.svelte + EditRevisionDiff.svelte; add POST /scenarios/{id}/edits/agent-propose Tests: frontend/src/lib/components/scenario-editor/__tests__/AgentActionPanel.test.ts

Phase 6b — WorkingDraft + Revalidate (P0 #7 / #9)

  • T015b [P1] Write failing WorkingDraft save tests (no arbitrary-draft bypass) in backend/tests/services/dashboard_testing/registry/test_scenario_editor_save.py @TEST_EDGE: save with mismatched digest->409; save with client-supplied graph->422 (rejected); re-validation catches injected SQL
  • T015c [P1] Implement WorkingDraft persistence + save_revision(draft_id, digest) server-side re-validate/canonicalize/hash in backend/src/services/dashboard_testing/editor/save.py
  • T015d [P1] Implement revalidate (Scenario Migration) in backend/src/services/dashboard_testing/editor/revalidate.py; add revalidate endpoint UI flow remains a follow-up editor route slice; backend proposal is read-only and conflict-explicit.

Phase 7 — Polish

  • T016 [P] RBAC scenario:edit enforcement tests Covered: backend/tests/api/test_scenario_editor_routes.py (agent-propose + proposal-save → 403 without scenario:edit)
  • T017 Run quickstart-equivalent scenario suite, scoped ruff, belief/ATTN static audit and semantic rebuild.
  • T018 Prototype validation: every declared @UX_STATE is reachable via prototype/index.html.

Audit Follow-ups (2026-08-20)

  • T019 Verify editor route keyboard access, stale conflict recovery and policy-gated save with an independent browser/API test after registry staleness wiring is complete.
  • T020 Run current scoped editor tests and record command-level evidence; do not reuse prior batch totals.

Dependencies

Setup → US1; US2 depends on ops/validate; US3/4 build on ops; US5 depends on US2-4. Requires 042 registry for persistence.

#endregion ScenarioEditor.Tasks


PROTOTYPE — State/Manifest

Source: prototype/manifest.md

#region ScenarioEditor.PrototypeManifest [C:3] [TYPE ADR] [SEMANTICS prototype,manifest,scenario,editor] @defgroup Prototype Interactive HTML prototype manifest for the Scenario Editor.

Prototype Metadata

Factual audit 2026-08-20: prototype state reachability is not production verification. Current browser/accessibility and policy-gate evidence remains pending.

  • Feature: 043 Scenario Editor UX
  • Source contracts: ux_reference.md, contracts/modules.md
  • Screens represented: 1 (Scenario Editor)
  • Total states: 7 (view_only, editing, saving, saving_conflict, validation_error, agent_proposing, proposal_diff)
  • Accessibility: keyboard nav, focus-visible, aria-live, ≥44px, prefers-reduced-motion
  • Responsive: 375px, 900px

State Coverage

@UX_STATE Prototype State Reachable? Recovery
view_only view_only ✅ Edit toggle
editing editing ✅ —
saving saving ✅ delegated save or inline gate
saving_conflict (409) saving_conflict ✅ reload/discard
validation_error validation_error ✅ fix + resubmit
agent_proposing agent_proposing ✅ —
proposal_diff proposal_diff ✅ inspect policy/diff/action

Screen ↔ Story Traceability

Story Prototype Feature Intended acceptance coverage
US1 View read-only badge, graph read-only default
US2 Manual Edit toggle, dirty, save revision + delegated policy
US3 Assertion constrained editor + raw blocked SQL injection rejected
US4 Deps dep select + cycle error cycle rejected
US5 Agent agent_proposing + proposal_diff proposal + diff + delegated save/gate
#endregion ScenarioEditor.PrototypeManifest

PROTOTYPE — Interactive HTML

Source: prototype/index.html

<!doctype html>

<html lang="ru"> <head> </head>
Superset Tools · BI testing
СценарииЗапускиАвтоматизацияКачество r18
← В регистр

XLSX reconciliation

Executable revision r18 · metadata version 7 · FI-0080

Предложить с агентом Изменить metadata Изменить проверку

Граф проверки

1
Открыть dashboard
browser / open_dashboard
Авто
2
Применить dashboard filters
browser / apply_native_filter
Авто
3
Скачать и сравнить XLSX
xlsx + assertion / compare_to_baseline
Авто
4
Сформировать отчёт
Авто

Metadata сценария

Изменение названия, описания и тегов не создаёт revision.

Отмена Сохранить metadata

Executable draft

Измените только безопасные constrained поля. Сервер хранит WorkingDraft и вычисляет diff.

Baseline reference baseline/xlsx-fi-0080-r17Допустимое отклонение
Отмена Проверить и сохранить
Подтвердите изменение

Изменяется assertion шага 3

baseline reference остаётся тем же · threshold: 0 → 2

Вернуться Подтвердить и создать r19
Конфликт revision (409).
r18 больше не current. Перезагрузите редактор или отбросьте draft.
Перезагрузить Вернуться к draft
Assertion отклонён.
SQL, raw baseline и циклические зависимости запрещены.
Исправить поля
Агент формирует предложение

Проверка typed operations

Draft остаётся неизменным до просмотра diff и явного сохранения.

Показать предложение

Agent proposal diff

threshold: 0 → 2
baseline ref: без изменений
agent_action_id: act-204
State: View only MetadataDraftSaving Conflict Validation Agent Proposal diff
<script src="../../prototype-ui.js"></script> <script> protoState("view_only", (s) => [ "view_only", "metadata", "editing", "saving", "saving_conflict", "validation_error", "agent_proposing", "proposal_diff", ].forEach((id) => document.getElementById(id).classList.toggle("hidden", id !== s), ), ); </script> </html>

================================================================================ FEATURE: 044-dashboard-scenario-execution Files: 14


SPEC — Feature Specification

Source: spec.md

#region ScenarioExecution.Spec [C:3] [TYPE ADR] [SEMANTICS spec,requirements,scenario,execution,run,step,engine] @BRIEF Partially verified ScenarioRun/ScenarioStepRun execution engine that walks a validated DashboardTestScenario DAG through typed executors, with lifecycle, retry, timeout, and human-checkpoint pause/resume. It remains not production-complete because live composition, T022, prototype validation, and semantic indexing are unresolved. @RELATION DEPENDS_ON -> [Doc.Adr.ADR0001] @RELATION DEPENDS_ON -> [Doc.Adr.ADR0003] @RELATION DEPENDS_ON -> [AgentTestStabilization.Spec] @RELATION DEPENDS_ON -> [SupersetBaselineEngine.Spec] @RELATION DEPENDS_ON -> [ScenarioRegistry.Spec] @RELATION DEPENDS_ON -> [DashboardScenarioModel.Spec] @RATIONALE 038 generates an immutable Verification Program IR. A ScenarioRun is a first-class entity independent of AgentRun and VerificationRun. Orchestration and deterministic program executors remain deterministic; a revision may additionally declare bounded AgentEvaluation steps whose typed outputs are resolved by DecisionPolicy, never by free-form orchestration. @REJECTED Agent-orchestrated step execution — rejected because each step would be an LLM call and could rewrite program flow. Explicit versioned AgentEvaluationSpec inside a deterministic step boundary is allowed. @REJECTED Reusing AgentRun as the execution run — rejected because AgentRun is the creation-process run; ScenarioRun is the created-test execution. Reusing VerificationRun — rejected because it is release-pipeline category verification, not arbitrary-DAG execution. @REJECTED human as a dispatched executor — rejected; it is a runner-lifecycle suspend/resume control primitive, not a side-effect executor. @RATIONALE The runner was already agent-free; after the 050 MCP drift the authoring/investigation boundary (SCEX-FR-012) is inherited by external MCP clients with zero change to orchestration, executors or capacity admission.

Navigation (DSA Indexer keywords)

@SEMANTICS: spec, requirements, feature, scenario, execution, run, step, runner, engine, resume, human

Feature Branch: 044-dashboard-scenario-execution Created: 2026-08-07 | Status: Not production-complete — fail-safe execution closure and targeted verification are recorded below; production composition remains open Input: "Provide a Scenario Execution Engine: a deterministic runner that walks the validated DashboardTestScenario graph, dispatches each step by tool to a typed executor (browser, superset_api, xlsx, assertion, screenshot, report, artifact), and manages the run lifecycle (queued, running, waiting_human, blocked, cancel, retry, timeout, resume, passed, failed, inconclusive) with immutable execution snapshots pinned to a scenario revision."

User Scenarios

Story 1 — Start a Scenario Run (P1)

Why P1: Users must be able to run a saved scenario against an environment with pinned revision and parameters.

Independent Test: Start a run from a registry scenario + params and verify a ScenarioRun is created with the exact revision, environment, and parameter snapshot.

Acceptance:

  1. Given a saved scenario with params When "Run" is submitted Then a ScenarioRun is created pinned to scenario_id + revision_id + content_hash, environment, and parameter snapshot.
  2. Given the runner starts When execution begins Then the run derives RunnerPlan, preflights, and transitions through queued → running.
  3. Given a PROD environment When start is requested Then the service creates ScenarioRun(pending_approval) plus an ActionApprovalGate before any step executes; approval queues it and denial/expiry blocks it.

Story 2 — Execute Steps Deterministically (P1)

Why P1: The core value — running the checks.

Independent Test: Execute a fixture scenario with browser/superset/xlsx/assertion/screenshot steps and verify each step dispatches to the correct executor and records a ScenarioStepRun.

Acceptance:

  1. Given the runner walks the DAG in topological order When each step is reached Then ScenarioStep.Dispatch routes by the version-pinned 038 ActionRegistry to a typed executor (browser→BrowserExecutor/Playwright, superset_api→037, sql_evidence→Superset SQL Lab adapter, transform/assertion→bounded 038 program engines, agent_evaluation→bounded evaluation adapter, xlsx→parser+norm, screenshot→ScreenshotService).
  2. Given a step produces an output ref When a dependent step consumes it Then the ref is bound from the producer's outputs.
  3. Given a step fails When execution continues Then dependents are blocked per failure policy and the run records the failure.

Story 3 — Pause at Human Checkpoint and Resume (P1)

Why P1: Human checkpoints must suspend the run without losing state.

Independent Test: Reach a human step and verify the run pauses (waiting_human), persists step state, and resumes from the correct point after disposition.

Acceptance:

  1. Given a human step is reached When the runner executes Then the run sets status waiting_human, persists the ScenarioStepRun as paused, creates a HumanCheckpoint (confirm/false_positive/inconclusive — NOT a 036 ApprovalGate) with evidence context, and stops advancing the DAG.
  2. Given a human decision (confirm/false-positive/inconclusive) When submitted with its checkpoint version Then an atomic CAS consumes the checkpoint, records the outcome, and internally resumes dependents.
  3. Given a page reload mid-wait When reopened Then the run status, waiting step, and gate are recoverable by scenario_run_id.

Story 4 — Cancel, Retry, Timeout (P2)

Why P2: Operational control over runs.

Independent Test: Cancel a running scenario, retry a failed step, and simulate a timeout; verify terminal states and step attempts.

Acceptance:

  1. Given a user requests cancel When submitted Then the run enters cancel_requested, drains in-flight steps within a bounded window, and terminates cancelled.
  2. Given a step fails When retry is requested Then a new attempt runs with bounded attempts and updated attempt count.
  3. Given a step exceeds its timeout When triggered Then the timed step is inconclusive and every declared downstream node is materialized from the pinned RunnerPlan as blocked before the run terminalizes; no later walker continuation may dispatch a lazy descendant.

Story 5 — Immutable Execution Snapshot (P2)

Why P2: Results must be reproducible regardless of later edits.

Independent Test: Run a scenario, then edit it, and verify the earlier run's results reference the original revision.

Acceptance:

  1. Given a run started When the scenario is later edited Then the run's revision_id + content_hash snapshot is unchanged.
  2. Given results are stored When queried Then they include provenance (revision, runner_version, template_version, baseline_revision, environment, parameter snapshot).

Edge & Failure Cases

# Scenario Category Expected Behavior Recovery
E1 PROD run without approval auth/gate Durable pending_approval run + gate, no dispatch Approve via ActionApprovalGate
E2 Step timeout execution Timed step inconclusive; pinned-plan descendants blocked before terminalization Retry only; no late continuation dispatch
E3 Retry exhausted execution Step failed; dependents blocked Triage (047)
E4 Superset 5xx/403/422 integration Typed error taxonomy preserved Retry / continue
E5 Runner crash mid-run resilience Only expired idempotent/retry-safe work may resume; unsafe work requires reconciliation; browser without a pinned safe checkpoint is inconclusive without replay Recover safe frontier / reconcile
E6 Cancel mid-step concurrency In-flight completes or times out in drain window —
E7 Duplicate output ref data-integrity 038 validator rejects before run Fix graph
E8 Stale scenario run (registry) data-quality Warning-gated or blocked per policy Revalidate

Requirements

Functional

  • SCEX-FR-001: A ScenarioRun MUST be a first-class entity pinned to scenario_id (UUID) + revision_id (UUID) + content_hash, environment, parameter snapshot, and optional agent_run_id/verification_run_id provenance.

  • SCEX-FR-002: The runner MUST walk the DAG in topological order and dispatch each step only through the version-pinned 038 ActionRegistry. Browser, Superset API/SQL Lab, XLSX, bounded transform/assertion, screenshot, report and artifact executors are typed; unknown {tool,action} is rejected.

  • SCEX-FR-003: Scenario orchestration MUST be deterministic and MUST NOT generate/rewrite SQL, DSL, assertions, graph or executor order at runtime. A declared AgentEvaluationSpec MAY run inside its bounded step contract; it is not agent-per-step orchestration.

  • SCEX-FR-004: A human step MUST suspend the run (status waiting_human), persist full HumanCheckpoint state (eligibility, evidence, expiry, CAS decision version), and internally resume dependents after its one-time disposition. HumanCheckpoint is distinct from ActionApprovalGate.

  • SCEX-FR-004a: A revision containing a human step MUST be manual_run_only. Scheduled, trigger, deploy, ETL and API execution are rejected before run creation; a HumanCheckpoint is never skipped to obtain an automatic PASS.

  • SCEX-FR-005: The lifecycle MUST include pending_approval, queued, running, waiting_human, blocked, cancel_requested, cancelled, passed, failed, inconclusive; cancel drains in-flight within a bounded window.

  • SCEX-FR-006: Steps MUST support attempt counts, retry policy, per-step timeout, input/output ref binding, artifact_refs, and error_code.

  • SCEX-FR-007: A ScenarioRun MUST be recoverable by scenario_run_id; browser recovery MUST reconstruct state from a declared browser-safe checkpoint, not continue a dead Playwright context. Results MUST carry target and execution-principal provenance.

  • SCEX-FR-008: PROD-classified environments MUST require an ActionApprovalGate before execution.

  • SCEX-FR-009: Executors MUST reuse 037 (metric_executor/comparison), 038 (capture), 036 (evidence/artifacts/HITL), and existing browser/xlsx infra; a second Playwright/LLM/SQL stack is forbidden.

  • SCEX-FR-010: human is a runner-lifecycle control, not a dispatched executor; the executor registry covers the other seven tools.

  • SCEX-FR-011: Failed, inconclusive and blocked runs MUST emit idempotent 036 InvestigationSignals carrying immutable run/evidence provenance; 047 creates/updates the Queue/Episode. They MUST NOT automatically start a chat, an AgentRun, or a remediation action.

  • SCEX-FR-012: An opened InvestigationCase MAY use the agent to construct diagnostic runs and controlled experiments under delegated policy. Agent work cannot bypass executor contracts, runner lifecycle, capacity, mutation policy or a required ActionApprovalGate; declared AgentEvaluationSpec is the only permitted runtime reasoning boundary.

  • SCEX-FR-013: A HumanCheckpoint remains a manual-run-only analyst decision. The agent may summarize evidence but MUST NOT consume the checkpoint, choose its disposition, or turn it into scheduled automation.

  • SCEX-FR-014: SqlEvidenceExecutor MUST execute exactly the immutable 038 SqlEvidenceSpec through Superset backend/SQL Lab with pinned database identity, ExecutionPrincipal and RLS/security fingerprint. Runtime only supplies typed ParameterBindings and may not alter SQL, relation, projection, JOIN or WHERE.

  • SCEX-FR-015: AgentEvaluation MUST be a separate immutable runtime record and DecisionPolicy MUST deterministically map it plus deterministic evidence to StepOutcome. A bare model verdict never directly sets ScenarioResult.

  • SCEX-FR-016: Browser actions and mutation safety MUST use the same versioned 038 ActionRegistry/mutation contract. Mutating browser steps in PROD are prohibited; test-data mutation needs fixture scope, record keys, side-effect/retry and cleanup policy independent of PROD approval.

  • SCEX-FR-017: Every provider MUST implement common ProviderExecutionContext and ProviderExecutionResult contracts including run/step/attempt, descriptor and binding fingerprints, principal fingerprint, deadline, idempotency key, capacity lease, trace identity, provider/version, operation receipt, effect state and retry disposition.

  • SCEX-FR-018: Provider output ownership MUST be proven by an immutable receipt binding every output or evidence ref to run, step, attempt, provider/version, descriptor, binding and principal fingerprints, content type, byte length and verified SHA-256. A path, caller digest, session id, raw bytes or model metadata alone is never evidence.

  • SCEX-FR-019: Providers MUST own isolated resources and enforce server-owned concurrency, size, timeout/deadline, cleanup and durable-write atomicity limits.

  • SCEX-FR-020: Cancellation MUST be operation-aware: durable request plus provider cancel by operation_id, acknowledged as stopped|completed|unknown. unknown forbids retry and PASS until reconciliation or terminal non-pass closure.

  • SCEX-FR-021: Retry MUST create a new attempt and capacity lease while preserving descriptor side- effect identity. Retry is allowed only when declared safe/idempotent or after reconciliation; late responses remain historical.

  • SCEX-FR-022: Providers MUST expose separate liveness, readiness and dependency-health checks, registration fingerprint, version, capabilities and redacted reason codes. Telemetry contains no credentials, cookies, raw SQL or capture bytes.

  • SCEX-FR-023: Deployment MUST register providers only from trusted startup configuration and verify versions, descriptors, bindings, storage, dependencies, limits and capability health before admission.

  • SCEX-FR-024: ExecutionCapacityManager MUST be the shared environment-scoped admission boundary for ScenarioRun, AgentRun, VerificationRun and LoadRun. No provider performs I/O without an atomic lease; exhaustion queues or blocks, and lease lifecycle is durable and idempotent.

  • SCEX-FR-025: AgentEvaluationProvider MUST execute only immutable 038 AgentEvaluationSpec with pinned provider/model/prompt versions, bounded budget and evidence allowlist. It emits immutable AgentEvaluation; DecisionPolicy alone maps it plus deterministic evidence to StepOutcome.

  • SCEX-FR-026: The environment-policy fingerprint and binding/security fingerprint MUST be revalidated immediately before dispatcher admission and first provider I/O. If they differ from the approval or launch snapshot, dispatch performs no I/O and returns POLICY_CHANGED or BINDING_CHANGED; a prior PROD approval cannot authorize the changed target.

  • SCEX-FR-027: Verification result processing MUST enforce server-owned response byte, row, cell, and canonicalization-time limits. Any exceeded limit returns typed RESULT_TOO_LARGE and an inconclusive outcome; truncation MUST NOT yield a pass or baseline update.

Live Execution Composition Root (SCEX-LIVE-COMPOSITION)

LiveExecutionBinding persisted on ScenarioRun is identity-only: a binding reference plus immutable environment, release, query-model, execution-principal, RLS/security, browser-safe-checkpoint/action, and evidence-policy fingerprints. It MUST NOT serialize a client, secret, cookie, callable principal, raw browser context, or capture bytes.

Application startup owns LiveExecutionCompositionRoot. It is the only server-side registration point for authorized providers: an exact 037 SupersetClient + immutable DashboardQueryModel + principal/RLS binding, a browser-safe session/action provider, and a ScreenshotService capture provider with durable evidence storage. The root resolves a persisted binding only after every pinned identity/fingerprint matches; it never constructs authority from caller IDs, run metadata, or environment_id. Missing provider/configuration is *_BINDING_UNAVAILABLE; a malformed or unauthorized/mismatched identity is *_BINDING_INVALID/*_BINDING_MISMATCH. Both fail closed and perform no I/O.

Deployment config is settings.scenario_live_execution_bindings[]. Every enabled record contains only binding_snapshot and query_model_snapshot; its credentials are looked up by the server from the configured Environment record when startup constructs the existing SupersetClient. The root marks a configured record whose environment/model/storage cannot be resolved as SUPERSET_BINDING_UNAVAILABLE. It marks browser and screenshot slots from an enabled binding as BROWSER_BINDING_UNAVAILABLE and SCREENSHOT_BINDING_UNAVAILABLE until a server-owned bootstrap calls register_browser or register_screenshot with a lawful provider. An empty list is valid and leaves every live tool typed unavailable.

Providers return typed success/failure/timeout/cancellation outcomes. PASS additionally requires a durable evidence reference and verified non-zero SHA-256. Where a provider supports transport cancellation, the root passes the capability through; otherwise lifecycle cancellation remains database-authoritative and the result remains typed. Browser recovery is lawful only from the declared safe checkpoint/reconstruction binding; a raw session is never revived. A mutating browser action must also satisfy the version-pinned 038 mutation contract and is rejected in PROD.

ScreenshotService currently captures paths for the LLM workflow but does not expose a 044 principal/RLS/checkpoint-bound durable-evidence provider. It therefore remains typed unavailable until startup registers such a provider; no path or raw capture metadata is treated as evidence.

Key Entities

  • ScenarioRun: Recoverable execution instance of a pinned scenario revision; owns status, phase, params, provenance, steps.
  • ScenarioStepRun: One step execution attempt/record: status, attempt, timing, inputs/outputs, artifact_refs, error_code, progress.
  • RunnerPlan: Deterministically derived at run start from ScenarioRevision + ParameterBinding + TargetSnapshot + baseline/runtime policy. runner.plan.json is diagnostic/reference materialization only.
  • ScenarioExecutorRegistry: Mapping of step.tool → typed executor; human excluded (lifecycle control).
  • ScenarioExecutionResult: Aggregated run result with pass/fail/inconclusive per step and provenance.

Success Criteria

  • SC-001: The canonical fixture graph (fixtures/graph.json) executes with 100% of declared executable steps dispatched in pinned topological order; the test MUST assert the exact dispatch sequence and zero duplicate dispatches for completed logical steps.
  • SC-002: For every human-checkpoint fixture, the run reaches waiting_human within one dispatcher cycle, persists exactly one pending checkpoint, and after a valid CAS decision resumes the next frontier without dispatching any previously passed logical step.
  • SC-003: For 100/100 repeated cancellation attempts across queued, running and waiting states, the run reaches cancelled no later than cancel_drain_deadline_at + 1 scheduler interval; terminal rows contain zero running steps, active leases or active evidence projections.
  • SC-004: A run recovered by scenario_run_id after worker/process interruption produces the same completed-step set and pinned RunnerPlan hash as before interruption; unsafe unknown effects produce no retry and a terminal blocked or inconclusive result.
  • SC-005: After a revision edit, 100% of fields in the prior run's revision/content/program, parameter, target, principal and evidence provenance remain byte-identical to their launch snapshot.
  • SC-006: Across all registry actions, human has zero executor resolution/invocation records and every configured PROD run has a pending approval gate before the first provider I/O event.
  • SC-007: Every provider passes a common contract suite for ownership, limits, lifecycle, cancellation, retry, reconciliation, health and deployment registration.
  • SC-008: No external provider performs I/O without a shared capacity lease; leases are released or reconciled after success, failure, cancellation, timeout and crash recovery.
  • SC-009: AgentEvaluation produces immutable evidence and DecisionPolicy produces the only accepted StepOutcome mapping for evaluation-backed steps.
  • SC-010: Every enabled live provider has a startup readiness record with provider/version, capability fingerprint, dependency checks and redacted diagnostics; an unready provider admits zero new operations until readiness is restored.
  • SC-011: The 044 backend profile, provider contract profile, PostgreSQL migration check and prototype validation each publish reproducible command output; a production GO requires 100% pass in all four profiles and no unresolved P0/P1 traceability row.

Clarifications

Session 2026-08-07

  • Q: Is the agent the runner? → A: No. The backend runner orchestrates deterministically. The agent authors before/after runs and may reason only inside a declared bounded AgentEvaluationSpec.
  • Q: Is human an executor? → A: No. It is a runner-lifecycle suspend/resume control; the executor registry covers browser/superset_api/xlsx/assertion/screenshot/report/artifact.
  • Q: How does this differ from VerificationRun/AgentRun? → A: ScenarioRun executes the user-created DashboardTestScenario DAG; AgentRun is the creation run; VerificationRun is release-pipeline category verification. Three distinct run concepts.

Session 2026-08-24 — measurable requirement clarifications

  • Requirement language: MUST is a release-blocking requirement; SHOULD is an explicitly tracked non-blocking recommendation; MAY is optional. Every MUST is verified by a named test, migration check, health check or deployment evidence record in traceability.md.
  • Bounded drain: the provider cancellation drain is the persisted interval cancel_drain_deadline_at - cancel_requested_at; finalization tolerance is one scheduler interval. The default scheduler interval is 5 seconds, so the default acceptance tolerance is 5 seconds.
  • Timeout: provider timeout is measured from started_at to the pinned descriptor deadline. A late response received after the deadline is historical only and cannot change the winning StepOutcome.
  • Ownership: an evidence ref is usable only after an immutable receipt exists with the complete owner tuple (run_id, logical step, attempt, operation, provider/version, descriptor, binding and principal fingerprints), content metadata and verified SHA-256.
  • Capacity: no adapter callback, network request, browser action, capture, model call or artifact write may begin before a durable CapacityLease is acquired. Lease expiry is not permission to retry an unknown external effect.
  • Production readiness: NO-GO remains the default. GO requires all SC-001..011 evidence, enabled-provider readiness, real PostgreSQL migration verification and zero open P0/P1 rows in the 044 and cross-spec traceability matrices.
  • Automated human prohibition: manual_run_only=true is a pre-create eligibility failure for every trusted non-manual source. Scheduled, deploy, release, ETL, API and background-recovery paths MUST NOT create a ScenarioRun, ActionApprovalGate, notification, queue item or dispatcher claim. Only the authenticated manual route may create the run and later reach waiting_human.

Full production tool boundary

There is no reduced preview release target. The package describes one complete production tool. Agent-authored typed scenarios, immutable revision activation, manual and automated triggers, Browser/Screenshot providers, controlled non-PROD fixture mutation, AgentEvaluation/DecisionPolicy, monitoring, investigation cases, analytics and remediation are mandatory capabilities with independent release gates. A missing or unproven capability blocks production GO; it is not silently downgraded to preview. Human-containing revisions remain manual-run-only and PROD mutation remains prohibited by policy.

Implementation Status & MVP Debt (factual audit 2026-08-20)

ScenarioRun/ScenarioStepRun models, lifecycle/API primitives, a runner plan, and a fail-safe executor boundary now exist. They do not yet constitute the specified production execution engine.

  • [x] Persisted worker lease/idempotency checks are proven: a live lease rejects another worker and reclaim is limited to retry-safe/idempotent recorded effects.
  • [x] HTTP and 046 automation starts persist or replay only queued/pending ScenarioRuns. Initial adapter dispatch is performed separately when the server dispatcher wins the durable queued -> running CAS; repeated ticks/replays do not dispatch again. Automated human plans are rejected before a queued row reaches that CAS, while manual human runs reach their checkpoint only after the dispatcher claims them.
  • [~] EnvironmentPolicy resolves stage=PROD/is_production=true only from server ConfigManager state before idempotency or persistence: client compatibility flags cannot change the class, unknown targets fail closed, and manual/API/event/scheduled sources create the same durable pending_approval gate. Pending rows remain excluded from dispatcher execution. A real APScheduler scheduled-PROD integration test is still coverage debt, not a bypass.
  • [x] Human and infrastructure continuation resume only the missing DAG frontier; completed steps are not re-run. Walker tests cover completed/failure/artifact and evidence-integrity paths.
  • [x] A revision containing human derives manual_run_only=true. Every trusted 046 automation source (scheduled, deploy, release, ETL, API) rejects before idempotency lookup, ScenarioRun, approval gate, notification, queue, or dispatch side effects; the same key remains usable by the authenticated manual route. HumanCheckpoint observation remains distinct from ActionApprovalGate.
  • [x] Failed, blocked and inconclusive terminal runner results emit one idempotent canonical 047 queue input keyed by immutable run/status/revision/target/principal and registered artifact ref/digest provenance. Queue projection never opens a case, AgentRun, chat or remediation action; passed runs emit none.
  • [x] Strict eligible retry invalidates the persisted failed/inconclusive/blocked target and its downstream closure before re-walk. Prior step attempts and artifact rows are retained as historical provenance, but retired artifact projections are excluded from active result/terminal evidence. Re-terminalization has a distinct immutable attempt context; replay of that same context is idempotent.
  • [x] A timeout during claimed adapter I/O wins over a late payload. Before terminalization it materializes every otherwise-lazy descendant from the immutable pinned RunnerPlan as a blocked row, so a later walker cannot dispatch that descendant. This is lifecycle closure, not proof of real live Browser/Superset/Screenshot I/O composition.
  • [x] Server-driven crash recovery loads only persisted ScenarioRun/RunnerPlan/lease state by scenario_run_id: completed ancestors are not re-run; an expired idempotent/retry-safe claim is archived and retried as a new attempt; unsafe effects terminate blocked for reconciliation; and browser recovery without a pinned browser-safe checkpoint terminates inconclusive without adapter I/O. Historical evidence remains audit-visible but inactive for a new attempt, and repeated rejected terminal contexts reuse their idempotent signal. It does not consume an infrastructure resume token, HumanCheckpoint, or ActionApprovalGate.
  • [~] Browser/Superset/Screenshot can materialize PASS only from explicit typed adapter success; unavailable live I/O and missing/invalid evidence digest/ref are typed inconclusive. This is a fail-safe partial closure, not evidence of real live Playwright/Superset/Screenshot composition.
  • [~] ScenarioRun persists a nullable LiveExecutionBinding identity snapshot (binding ref, environment/release/query-model/principal/RLS fingerprints, browser-safe binding refs and evidence policy). Application startup owns a fail-closed LiveExecutionCompositionRoot: an exact registered 037 client/model/DraftStorage tuple is wired into the default dispatcher, while registered browser and Screenshot providers are validated against the same snapshot. Missing/mismatched providers are typed inconclusive and never call I/O. No deployment has yet registered a browser-safe action or ScreenshotService-to-durable-evidence provider, so those tools remain unavailable by default.
  • [~] The targeted service profile independently reverified 35 passes. The exact API file has a 29-pass result outside this sandbox; inside it FastAPI TestClient is blocked by AnyIO self-pipe EPERM, a sandbox restriction rather than an application failure.
  • [ ] Live production composition adapters, provider contract suite, AgentEvaluation/DecisionPolicy, shared ExecutionCapacityManager, full T022 scope, recurring-episode classification from terminal signals, and an Axiom semantic-index rebuild remain unresolved; SCEX-FR-001..025 are not globally closed.

BrowserProvider readiness reassessment

The BrowserProvider contract is now specified at 90/100 contract completeness. This score covers the per-action schema/risk catalog, resource limits, isolated context lifecycle, safe checkpoints, ownership receipts, cancellation/reconciliation, mutation policy, readiness checks and mandatory canaries. It does not claim runtime implementation: T028-T034, T040-T042, T042b and a real PREPROD canary remain required before the BrowserProvider can be called production-ready.

Drift Amendment — MCP Interface (2026-08-24)

  • Unaffected structurally: ScenarioRun, executors, capacity and gates never referenced the chat runtime. MCP clients author scenarios before runs (038 chain) and investigate after terminal signals (047 cases) through governed tools only; manual_run_only and PROD gating apply regardless of the actor.
  • SCEX-FR-013 re-scoped for MCP (decision 2026-08-24): HumanCheckpoint disposition MAY be submitted through governed MCP decision tools (decide_checkpoint) when driven by an authenticated user principal — it remains a manual analyst decision, CAS-versioned and audited, equivalent to the 045 monitor path. The prohibition that stands unchanged: no automated origin (scheduled/deploy/release/ETL/API/service-principal) may create, consume or bypass a checkpoint, and the agent-as-autonomous-planner still cannot choose a disposition on its own.

Status (2026-09-02): done — реализовано в рамках 050: инструменты и гейты (specs/050-mcp-interface/tasks.md T012–T028 [x]), handoff-поверхность (050 T030–T033), демонтаж чата и сервиса agent/ (050 T040–T041, чекпоинты specs/WORKSTATE-043-047.md).

@{ ScenarioExecution.AuthoringPromotionBoundary [C:5] [TYPE ADR]

@BRIEF Execution admission boundary for AgentAuthoringWorkspace outputs. @RELATION DEPENDS_ON -> [DashboardScenarioModel.AgentAuthoringWorkspace] @RELATION DEPENDS_ON -> [ScenarioRegistry.AgentAuthoringPromotion]

044 accepts a ScenarioRun only when 042 has supplied a persisted, promoted and validated immutable ScenarioRevision and the normal server-owned RunPreflight succeeds. Workspace states, sandbox traces, source/patch artifacts, screenshots, diagnostics, browser URLs, cookies, secrets, filesystem paths and caller digests are never execution authority. A revision containing unsupported code-backed execution is rejected because that provider contract is future/unimplemented.

Exploratory Playwright/code sandbox activity is authoring-only and must remain isolated, allowlisted, bounded, cancellable, receipt-backed and free of production side effects. Exploration results may inform typed graph proposals, but cannot create a run, lease, provider operation, evidence result or PASS. This boundary is normative; live sandbox/execution integration remains an explicit release gate.

@} ScenarioExecution.AuthoringPromotionBoundary

#endregion ScenarioExecution.Spec


UX REFERENCE — Interaction Narrative

Source: ux_reference.md

#region ScenarioExecution.UxReference [C:3] [TYPE ADR] [SEMANTICS ux,reference,scenario,execution] @BRIEF UX interaction reference for Scenario Execution (044) — mostly backend; full run monitor UX owned by 045.

Feature Branch: 044-dashboard-scenario-execution | Created: 2026-08-07

1. User Persona & Context

  • User: BI analyst triggering a dashboard test scenario run.
  • Goal: Run a saved scenario against an environment and get deterministic results.
  • Context: Browser; from scenario detail "Run" or from 046 automation.

2. Happy Path

Analyst selects environment, revision and parameters, then submits one idempotent start request. The server returns 201 with queued or pending_approval within 2 seconds at p95 under the fixture load. For PROD, the first provider I/O event is impossible before approval. The monitor (045) renders every typed event by sequence; reconnecting with Last-Event-ID replays each missing event exactly once. A human checkpoint creates one pending CAS checkpoint; a valid decision resumes only the missing frontier.

3. Interaction Surface (044 backend-driven)

  • POST /scenario-runs — start with run configuration.
  • POST /scenario-runs/{id}/human/decision — resolve a human checkpoint.
  • POST /scenario-runs/{id}/cancel — stop.
  • POST /scenario-runs/{id}/resume — resume after pause.
  • GET /scenario-runs/{id} / {id}/events — state + SSE progress.

Run configuration UI and live monitor UI are spec 045. 044 supplies the run/step state and endpoints.

4. Error Experience (044 invariants)

  • PROD without approval -> pending_approval; zero provider I/O events and denial/expiry -> blocked.
  • Capacity unavailable -> queued or blocked with CAPACITY_UNAVAILABLE; zero provider I/O.
  • Timeout -> typed non-pass outcome; late response cannot alter the displayed winning attempt.
  • Cancellation -> cancel_requested with visible deadline, then cancelled within the configured drain deadline plus one 5-second scheduler interval.
  • Unknown external effect -> reconciliation_required; retry control is disabled until reconciliation or terminal non-pass closure.
  • Runner crash -> recoverable by run id; completed steps remain completed and unsafe effects are not replayed.
  • Missing/mismatched evidence ownership -> non-pass with stable reason code and an auditable receipt gap.

5. Verification States

  • Each provider operation exposes provider/version, operation id, attempt, status, reason code and evidence receipt state. Secret values, cookies, raw SQL and capture bytes are never rendered.
  • The UI distinguishes passed, failed, blocked, inconclusive, waiting_human, cancel_requested and cancelled; it never maps unavailable or pending states to PASS.

6. Tone & Voice

  • Style: Technical. Terminology: "Scenario run" distinct from "Release verification" and "Load tests".

#endregion ScenarioExecution.UxReference


CHECKLISTS — Requirements Quality — requirements.md

Source: checklists/requirements.md

Requirements Checklist: Scenario Execution Engine (044)

Purpose: Verify SCEX-FR-001..025 and SC-001..011 with reproducible evidence. Updated: 2026-08-24

[x] means implementation plus named verification evidence is complete. [~] means code exists but at least one required integration, deployment or production proof is missing. [ ] means absent. A synthetic adapter result, caller-provided digest, path or fixture-only PASS is not production evidence.

Run Model and Determinism

  • [~] CHK001 ScenarioRun stores UUID scenario/revision, 64-hex content/program hashes, environment, immutable parameter bindings, target snapshot and principal fingerprint.
  • CHK002 RunnerPlan derives exact {tool, action} descriptors from the pinned 038 registry and rejects missing, altered, unknown or hash-mismatched descriptors before lease/provider I/O.
  • [~] CHK003 Runner walks the canonical graph in exact topological order and dispatches each completed logical step at most once per attempt.
  • [~] CHK004 Recovery by scenario_run_id preserves RunnerPlan hash and completed-step set; unknown effects never retry without reconciliation.

Provider Contract

  • CHK005 Every provider receives context containing run, step, attempt, descriptor/binding/principal fingerprints, deadline, idempotency key, capacity lease and trace id.
  • CHK006 Every provider creates an append-only operation receipt before I/O and returns effect state, retry disposition, operation id and stable reason code.
  • CHK007 Every evidence ref has an immutable receipt binding owner tuple, provider/version, descriptor, content type, byte length and verified non-zero SHA-256.
  • CHK008 No provider performs I/O before a durable shared CapacityLease; lease expiry/release is idempotent and observable.
  • CHK009 Cancellation acknowledges stopped|completed|unknown; unknown effect blocks retry and PASS until reconciliation or terminal non-pass closure.
  • CHK010 Duplicate invocation, late response, malformed result, cleanup failure and reconciliation are covered for every enabled provider.
  • CHK011 Each enabled provider exposes separate liveness, readiness and dependency health with secret-free diagnostics and a registration/capability fingerprint.

Provider-Specific Obligations

  • [~] CHK012 Browser provider contract specifies isolated context, server auth binding, safe checkpoint, cleanup, mutation reconciliation, per-action risk classification, 120s/30s/3-page/25MiB/10MiB limits and PREPROD canaries; runtime provider and canary evidence remain open.
  • CHK013 Superset/SQL provider proves exact 037 model, database, principal/RLS and raw-response digest/ref; caller SQL and metadata cannot pass.
  • CHK014 XLSX provider accepts only server-owned verified artifacts and enforces archive/cell/formula limits.
  • CHK015 Assertion/Transform providers are deterministic, bounded and network/code-free.
  • CHK016 Screenshot provider atomically commits durable owned evidence and cleans temporary resources.
  • CHK017 Report/Artifact providers use immutable manifests, templates, producer receipts and digest verification.
  • CHK018 AgentEvaluation uses immutable 038 spec, pinned provider/model/prompt, bounded budget and evidence allowlist; DecisionPolicy alone produces StepOutcome.

Human and Lifecycle

  • CHK019 Human step is lifecycle control only: waiting_human, exactly one pending checkpoint, CAS decision.
  • CHK020 Resume continues only the missing DAG frontier and never reruns completed logical steps.
  • CHK021 Lifecycle enum includes pending_approval, queued, running, waiting_human, blocked, cancel_requested, cancelled, passed, failed and inconclusive.
  • CHK022 Cancellation has persisted drain deadline; terminal run has zero active lease, running step or active evidence projection.
  • CHK023 Retry invalidates downstream closure, creates a new attempt and preserves historical evidence.

Authority, API and Operations

  • CHK024 Server-owned environment policy creates PROD approval before dispatch and ignores client PROD flags.
  • [~] CHK025 API schemas require immutable run identity, status enums, step outcomes and execution provenance.
  • CHK026 Startup registers enabled providers only after dependency/readiness checks; unready providers admit zero new operations.
  • CHK027 Real PostgreSQL alembic check and alembic upgrade head pass on the deployment database.
  • CHK028 Full available 044 backend profile passes: registry/API suite, provider/lifecycle profile, scoped Ruff/compile and prototype validation.

Release Gates

  • CHK029 SC-001..011 each has reproducible command output or deployment evidence linked in traceability.
  • CHK030 No unresolved P0/P1 row remains in 044 or its execution-critical dependencies 036, 037, 038, 041, 042, 046 and 047.
  • CHK031 Browser and Screenshot providers perform authorized live I/O and produce owned durable evidence in a real deployment run.

PLAN — Implementation Plan

Source: plan.md

Implementation Plan: Scenario Execution Engine

Branch: 044-dashboard-scenario-execution | Date: 2026-08-07 | Spec: spec.md | Status: Not production-complete — provider contract and runtime integration gate pending

Implementation audit, 2026-08-20: run models/API/primitives exist; synthetic executors, missing executor imports and insufficient walker tests prevent the plan from being considered delivered.

Summary

Implement deterministic ScenarioRun/ScenarioStepRun orchestration for the validated immutable Verification Program, dispatching each registered {tool, action} to a typed executor (reusing 037/038/036/040 services) and handling lifecycle. RunnerPlan is derived at launch; runner.plan.json is diagnostic/reference only. Explicit AgentEvaluationSpec is bounded inside a typed step, never orchestration.

Technical Context

Language/Version: Python 3.13+ (backend), TypeScript DTOs (frontend consumed by 045) Primary Dependencies: FastAPI, SQLAlchemy, Pydantic 2; reuse 037 metric_executor/comparison, 038 capture, 036 evidence/HITL, 040 RunnerPool pattern, browser/xlsx infra Storage: new tables scenario_runs, scenario_step_runs Testing: pytest (unit/contract/integration for executor dispatch, provider lifecycle, ownership, capacity, health, deployment, resume, cancel, timeout and reconciliation) Frontend Architecture: DTOs only (run monitor is 045) Performance Goals: step dispatch overhead <100ms p95 on the canonical 100-step fixture excluding provider latency; SSE event replay has zero gaps/duplicates by sequence; cancellation finalizes within the persisted drain deadline plus one 5-second scheduler interval Constraints: deterministic (no LLM per step); human not an executor; PROD gated; immutable revision snapshot; no provider I/O without shared capacity lease; unknown external effects require reconciliation Scale: up to 100-step graphs, concurrent scenario runs per environment bounded

Release thresholds: 100% canonical fixture dispatch order; 100/100 cancellation trials within the bounded drain tolerance; 100% provider I/O preceded by CapacityLease and operation receipt; 100% accepted evidence refs with ownership receipt and valid SHA-256; all enabled providers ready; PostgreSQL migration check/upgrade pass; no unresolved P0/P1 execution-critical dependency row.

Constitution Check

Principle Result
I. Semantic Contract First PASS — run/step/runner contracted C4-C5
II. Decision Memory PASS — research R1-R5
III. External Orchestrator PASS — executor reuse
IV. Module Discipline PASS — runner/registry/executors separated
V. RBAC Enforcement PASS — scenario:run, scenario:run:prod
VII. Test-Driven C3+ PASS — dispatch/resume/cancel tests first
VIII. Attention-Optimized PASS

Project Structure

specs/044-dashboard-scenario-execution/
├── spec.md / data-model.md / research.md / plan.md / tasks.md / traceability.md / quickstart.md / ux_reference.md
├── checklists/requirements.md
├── contracts/modules.md, contracts/openapi.yaml, contracts/ux/
└── prototype/index.html + manifest.md

backend/src/models/scenario_run.py          # ScenarioRun, ScenarioStepRun
backend/src/services/dashboard_testing/execution/
│   ├── runner.py  runner_plan.py  dispatch.py  executor_registry.py
│   ├── provider_protocol.py provider_evidence.py provider_operations.py provider_health.py capacity.py
│   ├── provider_contracts.py provider_bootstrap.py
│   ├── providers/ (browser.py superset.py xlsx.py assertion.py transform.py screenshot.py report.py artifact.py agent_evaluation.py)
│   ├── executors/ (browser.py superset.py xlsx.py assertion.py screenshot.py report.py artifact.py)
│   ├── human.py  resume.py  cancel.py  retry.py
└── api/routes/dashboard_testing/scenario_runs.py  # POST/start/cancel/resume/human/GET/events

Delivery Phases

  1. ScenarioRun/ScenarioStepRun models + migration + fixtures.
  2. RunnerPlan Derive/Validate from revision + bindings + target.
  3. Executor registry + per-tool executors (reuse 037/038/036).
  4. Deterministic DAG dispatch + ref binding + failure propagation.
  5. Human suspend/resume (036 gate) + recoverable run.
  6. Cancel/retry/timeout + lifecycle.
  7. Immutable snapshot + provenance; REST routes; frontend DTOs.
  8. PROD gate + RBAC + regression gates.
  9. Provider protocol, ownership receipts, lifecycle/cancellation/reconciliation, health/deployment readiness, shared capacity and bounded AgentEvaluation/DecisionPolicy integration.

API and Schema

  • contracts/openapi.yaml — POST /scenario-runs, cancel, resume, human/decision, GET detail, GET events (SSE).
  • contracts/modules.md — ProviderProtocol, ProviderCatalog, ProviderOperations, ProviderHealth and CapacityManager contracts.

Traceability

traceability.md maps Story → model → operationId → contract → task → test.

Cross-Spec Boundary

  • Consumes persisted scenario + revision from 042.
  • Reuses executors from 037 (metrics/comparison), 038 (capture), 036 (evidence/HITL), 040 (RunnerPool).
  • Feeds run results to 045 (monitor) and 047 (triage); triggered by 046 (automation).
  • Distinct from VerificationRun (037) and AgentRun (036).

Complexity Tracking

No exception planned. Runner/dispatch/provider lifecycle are C5 but decomposed (runner, dispatch, provider protocol, operations, health, capacity, per-provider adapters). Do not collapse into one oversized orchestrator.


RESEARCH — Technical Decisions

Source: research.md

Scenario Execution Engine — Phase 0/1 Research (044)

Branch: 044-dashboard-scenario-execution | Date: 2026-08-07 | Spec: spec.md

R1. Execution model — deterministic backend runner

Decision: Deterministic ScenarioRunner orchestration walks the immutable Verification Program in topological order and dispatches registered executors. A declared bounded AgentEvaluationSpec may reason inside one typed step, but cannot own orchestration or rewrite program content.

Rationale: Matches 038 determinism (direct LLM-to-code rejected) and 040's non-agent RunnerPool pattern. Execution must be reproducible; an LLM per step breaks that.

Alternatives: agent-orchestrated LangGraph nodes per step (rejected: nondeterminism/drift); hybrid (deferred, not MVP).

Impact: ScenarioRunner, ScenarioExecutorRegistry, topological iteration.

R2. Run state — new ScenarioRun/ScenarioStepRun tables

Decision: Persist run + per-step state in DB. Runs recoverable by scenario_run_id.

Rationale: Reload during a test, human-checkpoint resume, and reproducibility require persisted state.

Alternatives: in-memory only (rejected: no recovery); reuse AgentRun (rejected: different concept); reuse VerificationRun (rejected: release-pipeline, not DAG execution).

Impact: new SQLAlchemy models + alembic migration.

R3. Human as lifecycle control, not executor

Decision: human is a runner-lifecycle suspend/resume primitive (waiting_human + HumanCheckpoint, confirm/false_positive/inconclusive), excluded from the executor registry. PROD authorization uses a distinct ActionApprovalGate (036 mechanism, generalized owner). Executor registry covers browser/superset_api/xlsx/assertion/screenshot/report/artifact.

Rationale: A human checkpoint is a gate, not a side-effect check. Treating it as an executor would couple pause/resume into the executor contract.

Impact: ScenarioRunner.SuspendForHuman, resume_token, 036 gate reuse.

R4. RunnerPlan ownership

Decision: Derive RunnerPlan at run start from ScenarioRevision + ParameterBinding + TargetSnapshot + baselines + runtime policy. runner.plan.json is a diagnostic/reference materialization only.

Rationale: Currently an execution plan is generated with no engine. Owning the contract closes that gap.

Alternatives: consuming a saved runner.plan.json (rejected: it can diverge from revision/bindings/target); a raw unbound graph (rejected: no preflight binding).

Impact: RunnerPlan.Derive/Validate; reference artifact may be regenerated but is never execution truth.

R5. Executor reuse

Decision: Reuse 037 (metric_executor/comparison), 038 (capture), 036 (evidence/artifacts/HITL), existing browser/xlsx infra, 040 RunnerPool patterns. No second Playwright/LLM/SQL stack.

Rationale: Consistency with 038/040 rejected-path guards.

Impact: executor registry binds to existing services.

Data Model

See data-model.md: ScenarioRun, ScenarioStepRun, RunnerPlan, ScenarioExecutorRegistry, lifecycle.

Contracts & API

  • contracts/modules.md — Execution.Start, Execution.Dispatch, Execution.SuspendForHuman, Execution.Resume, Execution.Cancel, Execution.RetryStep, Execution.RunnerPlan.
  • OpenAPI: POST /scenario-runs, POST /scenario-runs/{id}/cancel, POST /scenario-runs/{id}/resume, POST /scenario-runs/{id}/human/decision, GET /scenario-runs/{id}, GET /scenario-runs/{id}/events (SSE).

Constitution Check

Principle Result
I. Semantic Contract First PASS — run/step/runner contracted C4-C5
II. Decision Memory PASS — R1-R5
III. External Orchestrator PASS — executor reuse, no new Superset coupling
IV. Module Discipline PASS — runner/registry/executors separated
V. RBAC Enforcement PASS — scenario:run, scenario:run:prod
VII. Test-Driven C3+ PASS — executor dispatch/resume/cancel tests first
VIII. Attention-Optimized PASS — [SEMANTICS scenario,execution,...]

R6. Provider protocol

Decision: Treat every provider as a resource-owning operation boundary with a common execution context, immutable operation/evidence receipts, explicit effect state, cancellation, reconciliation, health and deployment capabilities.

Rationale: A typed callback validates the result envelope but cannot prove ownership, cleanup, unknown external effects, capacity admission or deployability.

Alternatives: provider-specific ad hoc callbacks (rejected: inconsistent recovery and evidence); runner-only lifecycle (rejected: provider resources and external effects remain invisible).

Impact: T028-T042; no live provider is production-ready until the common contract suite passes.

R7. Capacity and AgentEvaluation

Decision: All providers claim the shared environment-scoped ExecutionCapacityManager before I/O. AgentEvaluation is a bounded provider producing immutable evidence; deterministic DecisionPolicy is the only mapper to StepOutcome.

Rationale: Independent feature quotas permit starvation; direct model verdicts violate deterministic aggregation and human checkpoint authority.

Alternatives: per-provider semaphore (rejected: no global fairness or durable lease); direct verdict aggregation (rejected: nondeterministic and bypasses policy).

Impact: CapacityLease, provider operation receipts, T032 and T039 are blocking production work.


DATA MODEL — Entities & Relations

Source: data-model.md

#region ScenarioExecution.DataModel [C:5] [TYPE ADR] [SEMANTICS data-model,scenario,execution,run,step,state] @BRIEF Canonical ScenarioRun, ScenarioStepRun, RunnerPlan, executor registry, and lifecycle state model. @RELATION DEPENDS_ON -> [ScenarioExecution.Research] @RATIONALE A typed run/step state model is required for reload during a test, recovery, resume, and reproducibility. Without it, human-checkpoint resume and page reload are impossible. @REJECTED In-memory-only run state — rejected because runs must survive disconnect and be recoverable by run id.

ScenarioRun and PROD approval lifecycle

ScenarioRun is created before dispatch. Fields: id, scenario_id, scenario_revision_id (revision_id UUID), scenario_content_hash, verification_program_hash, action_registry_version, dashboard_id, environment_id, status (pending_approval|queued|running|waiting_human|blocked|cancel_requested|cancelled|passed|failed|inconclusive), phase (preflight|setup|executing|waiting_human|draining|terminal), parameter_bindings (immutable JSON), baselines_pinned (version), target_snapshot, execution_principal_fingerprint, execution_toggles (optional evidence only), trigger_source (server-owned), agent_run_id? (provenance), verification_run_id? (aggregation), idempotency_key (unique), started_at, finished_at, resume_token, error_code, runner_version.

live_execution_binding_ref? plus live_execution_binding_snapshot? are the only persisted live-I/O coordinates. The snapshot is an exact allowlist: binding ref; environment/release/query-model, execution-principal and RLS/security fingerprints; browser-safe checkpoint/action refs; and evidence owner/ref policy. It contains no credential, client, cookie, raw browser context, callable principal, or capture bytes. LiveExecutionCompositionRoot resolves this immutable identity server-side at startup; a missing provider is typed unavailable, and any exact-snapshot mismatch is typed non-pass before I/O.

The deployment-owned settings.scenario_live_execution_bindings[] record stores enabled, binding_snapshot, and query_model_snapshot only. Startup resolves its Environment credentials server-side to build the existing SupersetClient, then registers the exact model/storage tuple. Browser and screenshot providers are process-local registrations and are never serialized in this record; absent registrations yield stable configured-unavailable outcomes.

If the selected revision contains a human step, it is derived as manual_run_only=true: it may start only from the authenticated analyst manual-run route. Scheduler, deploy/release/ETL/API trigger and any background runner or recovery worker are ineligible. The eligibility check occurs before idempotency, ScenarioRun, ActionApprovalGate, notification, queue insertion and dispatcher CAS; no HumanCheckpoint may be created, skipped, defaulted or converted into an automated approval path.

For a PROD request, the service atomically creates ScenarioRun(status=pending_approval) and ActionApprovalGate(owner_type=scenario_run, owner_id=run_id, operation=scenario_execution). Approval transitions only pending_approval → queued; denial/expiry transitions to blocked. Schedulers and external triggers therefore receive a durable run/intent, never an unusable 403. trigger_source is set only by the trusted entry point (manual route, scheduler, deploy connector, or API key), never by a bearer client field. Idempotency uses canonical execution-request hash: same key + same hash returns the existing run; same key + different hash returns 409 IDEMPOTENCY_KEY_REUSED.

EnvironmentExecutionPolicy is resolved server-side from the configured Environment record (stage=PROD or is_production=true) before RBAC, idempotency, run, gate, queue, notification, or adapter work. Any compatibility is_prod request field is non-authoritative and cannot upgrade or downgrade that class; an unknown environment fails closed without creating a run or gate.

ParameterBinding, ExecutionPrincipal, and TargetSnapshot

ParameterBinding: parameter_name, resolved_value, source, resolved_at, validation_fingerprint. It is derived at start from the 038 ParameterDefinition and launch input; it is never embedded in ScenarioRevision or its content_hash.

ExecutionPrincipal: auth_mode, actor_id?, service_identity?, impersonated_user?, effective_roles_hash, rls_context_hash. Its immutable fingerprint is stored on the run and used for every Superset request.

TargetSnapshot: environment_id, dashboard_release_id, dashboard_fingerprint, dataset_lineage_fingerprint, captured_at. It is mandatory even when no release was selected, so environment_id=preprod is never treated as an immutable target.

AnalyticsContextKey is a server-derived SHA-256 over environment_class + compatibility_family + baseline_family + dashboard_release_id + dashboard_fingerprint + dataset_lineage_fingerprint + execution_principal_fingerprint. It is captured on the run and every step result; analytics never groups runs merely by environment or revision id.

Statuses: queued, running, waiting_human, blocked, cancel_requested, cancelled, passed, failed, inconclusive.

Terminal failed, inconclusive and blocked outcomes emit a 036 InvestigationSignal with this immutable execution snapshot. 047 deterministically creates/updates the Queue/Episode from that signal. Signal delivery never changes run status and never auto-starts an agent run; an analyst opens the case explicitly from 045/047.

ScenarioStepRun

Fields: id, run_id FK, logical_step_id (immutable UUID, from #8), step_position (mutable), step_content_hash (mutable), attempt, status (queued|running|waiting_human|passed|failed|inconclusive|blocked|skipped), started_at, finished_at, inputs_snapshot (JSON, no secrets), outputs (JSON), artifact_refs (JSON), error_code, progress, timeout_ms, side_effect_key (nullable), step_outcome (StepOutcome).

StepOutcome { status, reason_codes[], deterministic_evidence_refs[], agent_evaluation_ids[], decision_policy_id?, decision_policy_version?, decided_at } is the authoritative result of one logical step. It is distinct from executor output and from a model verdict.

AgentEvaluation and DecisionPolicy

AgentEvaluation { evaluation_id, scenario_run_id, logical_step_id, attempt, provider_id, model_id, model_version, prompt_template_id, prompt_template_version, input_manifest_hash, evidence_refs, verdict, confidence, findings, reason_codes, raw_response_artifact_ref, started_at, finished_at } is immutable evidence generated only for a declared 038 AgentEvaluationSpec. Its tool/evidence access is bounded by that spec; it cannot mutate program content, invoke mutation actions, change run state or choose downstream scheduling.

DecisionPolicy { policy_id, version, deterministic_hard_failure, high_confidence_failure, low_confidence, disagreement, missing_evidence } is a versioned deterministic mapper. Defaults: deterministic hard failure→failed; high-confidence policy-qualified agent failure→failed; low confidence→inconclusive; evaluation/evidence disagreement→waiting_human via a HumanCheckpoint only for manual runs; missing required evidence→blocked or inconclusive. Scenario aggregation consumes StepOutcome, not AgentEvaluation verdicts directly.

RunnerPlan — deterministic derivation from revision (#2)

RunnerPlan is derived deterministically at run start from the selected immutable ScenarioRevision/Verification Program, NOT read from a stored runner.plan.json. Fields: scenario_revision_id, scenario_content_hash, verification_program_hash, action_registry_version, action_registry_hash, env targets, resolved params, pinned baselines, topological order, and one immutable ActionExecutionDescriptor per step. A descriptor contains the exact {tool, action}, typed input/output contracts, idempotency, retry safety, side-effect-key policy, timeout, mutation contract/risk. The runner persists the descriptor snapshot in both steps and executor_mapping; it derives leases/recovery/retry policy only from that snapshot. A missing, altered, unknown, version/hash-mismatched descriptor, invalid input/output shape, or mutating action without its required contract rejects before run/lease/I/O. Run refuses if its revision/program/action-registry hashes differ from the selected revision.

@INVARIANT RunnerPlan descriptor resolution is exact and version/hash pinned; tool is never a dispatch or retry-policy fallback. @REJECTED A universal idempotent=true, retry_safe=true claim based on a non-human tool was rejected because it can repeat unsafe effects.

The materialized runner.plan.json in git is a reference artifact, never the runtime source of truth; it may be regenerated from any revision.

ScenarioExecutionContext (#11)

Fields: run_id, browser_session (ref, not raw cookies), page/context ref, auth_context ref, current_dashboard, current_filters, artifact_namespace (owner_type=scenario_run), environment client. Browser secrets/cookies live in a separate secure context, NEVER in inputs_snapshot JSON (keeps reproducibility snapshot secret-free). Browser workers resume only by deterministic replay from the last browser-safe checkpoint (dashboard_open, filters_applied, etc.); replay records reconstruction_replay=true and never reclassifies already completed logical steps as rerun. API/XLSX/pure assertion steps may resume directly only when their executor declares retry-safe/idempotent. Browser session/context refs are valid only on the single application-owned provider event loop (loop_binding); they are never shared across runs and never cross loops (see 044 ProviderRuntime contract).

Artifact ownership (#4)

Artifacts use a generic owner: Artifact { id, owner_type: agent_run|scenario_run|verification_run|load_run, owner_id, kind, sha256, content_ref, retention_class, ... }. ScenarioRun evidence/screenshots/report/xlsx use owner_type=scenario_run; no artificial AgentRun is created. Retention_class ties to 046 tiers. Playwright trace/video dumps are not Artifact kinds: an operator debug flag may keep them locally, and they are never registered as Artifact or EvidenceReceipt.

Decision gates (#3)

  • ActionApprovalGate — authorization approval for PROD execution, baseline approval, repository mutation. Generalized 036 gate with owner_type + owner_id.
  • HumanCheckpoint — checkpoint_id, run_id, logical_step_id, checkpoint_type, decision_policy, status, created_at, expires_at, eligible_role?, eligible_actor_ids?, assigned_to?, evidence_refs, decision_version, decided_by?, decided_at?, disposition?, comment?. Status is pending|decided|expired|cancelled; decision is CAS on decision_version, stale/concurrent decision returns 409. finding_review maps confirm→failed, false_positive→passed, inconclusive→inconclusive; manual_assertion maps pass→passed, fail→failed, inconclusive→inconclusive. It is NOT a 036 ApprovalGate decision.

A HumanCheckpoint is never delegated to the agent: it is a manual-run-only analyst decision inside a currently executing run. The agent may explain the evidence in its workspace but cannot consume the checkpoint or convert it into an automated result.

Worker semantics — at-least-once execution (#6)

Runtime primitives: worker lease, heartbeat, run claim, step claim, lease expiration, idempotency key, recovery scheduler. Each executor declares: idempotent? | retry-safe? | side_effect_key? | external_request_id?. POST /scenario-runs requires Idempotency-Key (unique) to prevent double-run on double-click. A crashed step with an external side effect is only re-run if idempotent/retry-safe or keyed.

Retry semantics (#14)

Retry of a failed step invalidates its downstream closure (descendants depending on its output) and re-runs them; retry after the run advanced beyond the step is rejected unless the whole closure re-runs. Bounded attempts per policy.

Result aggregation truth table (#16)

  • any step failed → scenario failed
  • blocked descendants counted as blocked (NOT failed)
  • skipped does not count against pass
  • inconclusive → scenario inconclusive unless a failed also present (then failed)
  • warning is an evidence-level qualifier, not an execution status (045 maps from evidence)

Lifecycle

pending_approval → queued → running → waiting_human | blocked → passed | failed | inconclusive; cancel_requested → cancelled. Human decision atomically consumes the checkpoint and resumes internally. Public /resume is reserved for a recoverable infrastructure pause and requires a typed resume token/reason; it cannot consume a HumanCheckpoint. Cancel drains in-flight within a bounded window.

ScenarioExecutorRegistry and BrowserExecutor

The registry resolves ActionExecutionDescriptor -> executor; the following list states each descriptor's tool family, not a tool-only fallback:

  • browser -> BrowserExecutor → version-pinned 038 ActionRegistry → Playwright/session infrastructure
  • superset_api -> 037 metric_executor_async / SupersetClient.ChartData.Execute
  • sql_evidence -> SqlEvidenceExecutor → Superset SQL Lab backend/API → configured database connection (no credentials exposed to agent)
  • transform -> version-pinned bounded 038 TransformSpec DSL executor
  • xlsx -> xlsx parser + 037 normalization
  • assertion -> 037 comparison.py + 038 ComparisonSpec/AssertionSpec executor
  • agent_evaluation -> bounded provider adapter executing a declared 038 AgentEvaluationSpec and emitting AgentEvaluation; DecisionPolicy owns StepOutcome
  • screenshot -> 038 capture.py + ScreenshotService (owner_type=scenario_run)
  • report -> report-template render + artifact (owner_type=scenario_run)
  • artifact -> generic artifact register (owner_type=scenario_run)
  • human -> EXCLUDED (runner-lifecycle HumanCheckpoint control, not an executor)

BrowserExecutor implements only registered actions: open_dashboard, navigate_tab, apply_native_filter, inspect_filter_state, apply_table_filter, extract_table, scroll_to, inspect_columns, click, select_rows, edit_row, bulk_edit, download, refresh, wait_for_state. ScreenshotService is evidence infrastructure only, not the browser action executor. It validates the registry-declared inputs/outputs/risk/timeout/idempotency before dispatch.

Provider runtime protocol

Every provider implements the following logical records. They are persisted or emitted as immutable operation/evidence receipts; a Python callback alone is not a sufficient provider implementation.

ProviderExecutionContext: run_id, logical_step_id, attempt, descriptor_fingerprint, live_binding_ref, execution_principal_fingerprint, deadline_at, idempotency_key, capacity_lease_id, trace_id, cancellation_requested, cancellation_deadline_at.

ProviderOperationReceipt: operation_id, context identity, provider_id, provider_version, started_at, finished_at, effect_state, external_request_id?, resource_refs, cleanup_state, reconciliation_state, result_digest?. Receipts are append-only; late responses are historical records and cannot mutate the winning attempt.

ProviderEvidenceReceipt: evidence_ref, owner_type=scenario_run, owner_id, run_id, logical_step_id, attempt, operation_id, provider_id/version, descriptor_fingerprint, binding_ref, execution_principal_fingerprint, content_type, byte_length, sha256, retention_class, created_at. The durable store must atomically commit the content and receipt or return no usable ref.

Provider state is registered|health_checked|capacity_claimed|admitted|running|completed|failed| cancel_requested|finalized|reconciliation_required. A provider exposes separate liveness, readiness and dependency-health checks and a redacted capability/limit document. Readiness gates new admissions only; it does not remove active leases or evidence.

Cancellation calls cancel(operation_id) and records stopped|completed|unknown. unknown means the provider cannot prove whether an external effect happened: retry is forbidden until reconcile(operation_id, idempotency_key) returns a terminal receipt. If reconciliation is unavailable before the bounded recovery deadline, the runner terminalizes inconclusive or blocked according to effect risk. Capacity is released only after receipt and cleanup state are durable.

Provider-specific schemas and obligations

Provider Input boundary Required result/evidence Resource and recovery rules
Browser 038 action descriptor, server-owned auth/session binding, typed action input operation receipt, safe checkpoint, page/action provenance and durable evidence where declared isolated context, bounded pages/downloads, cleanup on finalization, safe-checkpoint replay only
Superset API exact 037 query envelope and pinned model/principal/RLS request/response fingerprints, raw response SHA-256 and durable ref request/response limits; 429/5xx bounded retry; unknown timeout reconciles external request
SQL evidence immutable 038 SqlEvidenceSpec plus typed ParameterBindings DB identity, model/principal/RLS fingerprints, raw response digest/ref read-only; caller SQL rejected; reconcile before retry
XLSX verified server-owned artifact ref source ownership/digest and canonical normalized output archive/cell/formula limits; local parse is idempotent
Assertion canonical values plus pinned 037 ComparisonPolicy deterministic input manifest, policy/version and StepOutcome pure, no external resources, retry-safe
Transform bounded 038 TransformSpec DSL source manifest, transform fingerprint and canonical output digest hard operation/row/column/output limits; no code/SQL/network
Screenshot bound capture session/profile and durable evidence policy atomic evidence receipt with media type, length and SHA-256 temp cleanup; unknown capture reconciles before retry
Report pinned template and immutable evidence manifest template/manifest/output digest and durable ref bounded deterministic renderer; partial writes reconcile or discard
Artifact authorized producer ref or verified source artifact producer receipt, owner tuple, digest and retention never accept caller digest alone; registration idempotent
AgentEvaluation immutable 038 spec, evidence allowlist, provider/model/prompt versions immutable AgentEvaluation and raw response artifact receipt token/cost/time limits; verdict cannot directly set StepOutcome

BrowserProvider action contract

Each browser step persists the exact 038 action descriptor plus input_schema_version, output_schema_version, risk, checkpoint_policy, idempotency, retry_safe, limits and mutation contract fingerprint. Defaults are context/authentication timeout 120000ms, action timeout 30000ms, maximum 3 pages, maximum download 26214400 bytes and maximum screenshot 10485760 bytes. A descriptor may lower a limit but cannot raise it from scenario payload.

open_dashboard, navigate_tab, filter application, inspection, scroll, extraction, refresh and wait are read-only/evidence actions; click is rejected unless its effect class is explicit; edit_row and bulk_edit are mutation actions; select_rows is selection-only; download produces a server-owned artifact ref, never a local path. extract_table is bounded to 10,000 rows, 100 columns and 10 MiB output.

Browser resource state is lease_claimed|context_created|authenticated|action_running|checkpointed| effect_recorded|cleanup_started|receipt_finalized|context_closed. One active context is allowed per (run_id, capacity_lease_id) and cleanup failure prevents PASS. Concurrency defaults are 2 concurrent browser contexts for DEV/PREPROD and 1 for PROD per environment, admitted by the shared CapacityManager; screenshot capture shares the same lease accounting, and capacity exhaustion keeps the run queued with CAPACITY_BLOCKED. Recovery never attaches to a dead context; it reconstructs from a safe checkpoint and records reconstruction_replay=true. Unknown mutation effects are held for reconciliation and cannot retry.

Readiness requires browser executable/version, auth binding, allowed-origin, evidence-storage, cleanup-worker and cancellation-capability checks. Deployment requires a read-only PREPROD canary and forced timeout/ cleanup canary before enabling; mutation requires a separate fixture-lease canary.

AgentEvaluation and DecisionPolicy integration

The AgentEvaluation provider claims capacity as workload_class=agent_evaluation, creates an immutable input manifest from deterministic evidence refs, executes within the declared budget, and stores the raw response before publishing a verdict. Provider/model/prompt versions and manifest hash are mandatory. DecisionPolicy(policy_id, version) consumes only the immutable evaluation record plus deterministic evidence and emits StepOutcome. Low confidence, missing evidence or disagreement follow the pinned policy table; the model cannot select a policy, mutate the graph, invoke a provider action, consume a HumanCheckpoint or schedule downstream work.

ExecutionCapacityManager integration

Before any provider I/O, the dispatcher calls the shared allocator with environment, workload_class, provider_id, priority, requested_units, quota, reserved_capacity, run/step and deadline. The atomic result is a durable CapacityLease; no lease means queued or blocked according to policy and no provider invocation. Heartbeat, expiry, release and forced reconciliation are idempotent. Retries claim new leases, while active unknown operations retain their lease until reconciliation or terminal closure.

Verification thresholds

  • scenario_content_hash, verification_program_hash, descriptor_fingerprint, query_model_fingerprint, execution_principal_fingerprint, rls_security_fingerprint and evidence sha256 are lowercase or uppercase hexadecimal SHA-256 values with exactly 64 characters.
  • attempt starts at 1 and increases by exactly 1 for each new logical-step attempt. Historical attempts are immutable; only one attempt may be the active projection for a logical step.
  • byte_length is a positive integer equal to the durable content length. A zero-byte evidence object is never eligible for PASS.
  • deadline_at is later than started_at; a result received after the deadline cannot become the winning outcome. The runner records the late result as historical operation data.
  • CapacityLease validity is [claimed_at, expires_at]; provider I/O requires now < expires_at and a matching run/step/attempt/provider context. Lease release is idempotent: repeated release changes no terminal state or accounting total.
  • Cancellation finalization must occur no later than cancel_drain_deadline_at + 5 seconds under the default scheduler interval. Any exception is a release failure and keeps the run non-GO.

Mutation authorization and execution policy

Authorization is evaluated per run and step as environment_class + scenario risk profile + action mutation profile. Read-only PROD steps require the ScenarioExecution approval policy. Every mutation requires the immutable 038 mutation contract, scoped target keys, precondition evidence, side-effect identity, cleanup/reconciliation outcome and retry_safe=false unless the registry proves an idempotent compensating action. Mutating browser steps in PROD are prohibited. Non-PROD test-data mutation may be delegated only inside an authorized fixture lease; this mutation policy is separate from PROD execution approval.

Immutable Execution Snapshot

A ScenarioRun pins scenario_revision_id + scenario_content_hash at start. Results carry provenance: scenario_revision_id, runner_version, template_version, baseline_revision, target_snapshot, parameter_bindings, execution_principal_fingerprint, query fingerprints. Later edits never alter completed runs.

ExecutionCapacityManager

All AgentRun, VerificationRun, LoadRun, and ScenarioRun claims pass through one environment-scoped capacity manager: environment, workload_class, priority, quota, reserved_capacity. A 046 scenario policy is a consumer of this global allocator, not an independent PROD/PREPROD concurrency limit.

#endregion ScenarioExecution.DataModel


CONTRACTS — Module & Function Contracts

Source: contracts/modules.md

#region ScenarioExecution.Modules [C:5] [TYPE ADR] [SEMANTICS scenario,execution,contracts,modules,runner] @BRIEF Module contracts for the Scenario Execution Engine (044). @defgroup ScenarioExecution Deterministic DAG runner dispatching steps to typed executors. @RELATION DEPENDS_ON -> [ScenarioGraph.Models] @RELATION DEPENDS_ON -> [BaselineEngine.Comparison] @RELATION DEPENDS_ON -> [ScenarioGraph.Capture] @RELATION DEPENDS_ON -> [Services.AgentRuns.Evidence] @RELATION DEPENDS_ON -> [Services.AgentRuns.HitlGate] @RATIONALE Execution is deterministic, persisted, recoverable, and reuses existing executors; human is a lifecycle control. @REJECTED Agent-per-step orchestration; human as executor; reusing AgentRun/VerificationRun as the run.

#region ScenarioExecution.RunnerPlan.Derive [C:4] [TYPE Function] [SEMANTICS scenario,execution,runnerplan,derive]

@ingroup ScenarioExecution

@BRIEF Deterministically derive the execution plan from a ScenarioRevision at run start.

@PRE scenario revision selected and pinned.

@POST returns RunnerPlan with env targets, resolved params, pinned baselines, topological order,

exact ActionExecutionDescriptor mapping plus registry version/hash; refuses on revision mismatch.

@INVARIANT runner.plan.json in git is a reference artifact, never the runtime source of truth.

@INVARIANT Each persisted descriptor is resolved against the pinned 038 registry before a run,

lease or adapter exists; descriptor fields alone derive retry/idempotency/timeout policy.

@REJECTED tool-only executor fallback or universal retry-safe metadata.

@TEST_EDGE missing revision->reject; revision mismatch->reject before run.

def derive_runner_plan(db, scenario_id, revision_id): ...

#endregion ScenarioExecution.RunnerPlan.Derive

@{ ScenarioExecution.RunPreflight [C:5] [TYPE Function]

@BRIEF Validate the server-owned launch snapshot before RunnerPlan derivation or dispatch. @RELATION DEPENDS_ON -> [ScenarioRegistry.RevisionChain] @RELATION CALLS -> [ScenarioExecution.RunnerPlan.Derive]

@PRE ScenarioRevision is persisted and selected; typed ParameterBindings, target, policy, baseline and execution-principal inputs are supplied by the server; an Idempotency-Key is present. @POST Returns a RunPreflightHandle with a server-computed digest bound to scenario/revision/content hash, target snapshot, bindings, policy and action registry fingerprints, or a typed blocking result. @SIDE_EFFECT Persists the immutable preflight result only after all checks pass; blocking validation creates no dispatch, lease, provider I/O, or evidence. @INVARIANT needs_context, needs_selector, and needs_baseline are never guessed or converted into launch values. Client graph, caller digest, runner.plan.json, and client environment classification are non-authoritative. @INVARIANT A PROD result remains pending_approval until the durable gate is consumed; approval and launch continuations use CAS/idempotency semantics. @TEST_EDGE missing_binding->blocked; stale_revision->blocked; policy_changed_before_io->POLICY_CHANGED; same_request_replay->existing run. def run_preflight(db, scenario_id, revision_id, bindings, target, actor): ...

@} ScenarioExecution.RunPreflight

#region ScenarioExecution.Start [C:5] [TYPE Function] [SEMANTICS scenario,execution,start,run]

@ingroup ScenarioExecution

@BRIEF Create and start a ScenarioRun pinned to an immutable revision.

@PRE caller has scenario:run; PROD creates ActionApprovalGate; RunPreflight validates bindings/target; Idempotency-Key supplied.

@POST same idempotency key + canonical request returns existing ScenarioRun; differing request returns IDEMPOTENCY_KEY_REUSED; otherwise run queued/pending_approval and RunnerPlan derived.

@SIDE_EFFECT DB write; enqueue; audit.

@INVARIANT revision snapshot immutable; PROD gated; idempotency identity is key + canonical execution-request hash.

@TEST_EDGE prod_without_approval->pending_approval; stale_revision->blocked; same_idempotency_replay->existing run; changed_request->409.

async def start_run(db, scenario_id, revision_id, params, env, actor, idempotency_key): ...

#endregion ScenarioExecution.Start

#region ScenarioExecution.Dispatch [C:5] [TYPE Function] [SEMANTICS scenario,execution,dispatch,step,executor]

@ingroup ScenarioExecution

@BRIEF Dispatch a descriptor-validated step to its typed executor; record a ScenarioStepRun.

@PRE step dependencies satisfied; {tool, action} exists in the revision-pinned 038 ActionRegistry; mutation policy preflight passed.

@POST returns step outcome; output refs bound; ScenarioStepRun persisted.

@SIDE_EFFECT executor side effects (browser/superset/xlsx/evidence); DB write.

@INVARIANT human not dispatched here; deterministic per registered {tool, action}; unregistered action is rejected before executor I/O.

@TEST_EDGE unknown_tool->rejected; step_fail->dependents blocked; ref_binding->dependent reads producer output.

async def dispatch_step(run, step): ...

#endregion ScenarioExecution.Dispatch

#region ScenarioExecution.HumanCheckpoint [C:4] [TYPE Function] [SEMANTICS scenario,execution,human,suspend,checkpoint]

@ingroup ScenarioExecution

@BRIEF Pause the run at a human step and create a HumanCheckpoint (not a 036 ApprovalGate).

@PRE step.tool == human.

@POST run.waiting_human; step paused; HumanCheckpoint created with evidence; DAG stops.

@SIDE_EFFECT HumanCheckpoint entity; DB write.

@REJECTED using 036 ApprovalGate for test-result disposition — it is authorization approval, not observation disposition.

def suspend_for_human(run, step, evidence): ...

#endregion ScenarioExecution.HumanCheckpoint

#region ScenarioExecution.Resume [C:4] [TYPE Function] [SEMANTICS scenario,execution,resume,run]

@ingroup ScenarioExecution

@BRIEF Resume a paused run from its resume token, continuing dependents without rerunning completed steps.

@PRE checkpoint resolved (confirm/false_positive/inconclusive); resume_token valid; run not terminal.

@POST run.running; continues from resume point.

@SIDE_EFFECT continues executor dispatch.

@INVARIANT completed steps are never rerun.

async def resume_run(db, run_id, disposition): ...

#endregion ScenarioExecution.Resume

#region ScenarioExecution.ActionApprovalGate [C:3] [TYPE Function] [SEMANTICS scenario,execution,approval,gate,prod]

@ingroup ScenarioExecution

@BRIEF Require an ActionApprovalGate for PROD execution (authorization approval).

@PRE environment classified PROD.

@POST gate required before any step dispatch; denial dispatches nothing.

@REJECTED conflating PROD authorization with HumanCheckpoint observation disposition.

def require_action_gate(run): ...

#endregion ScenarioExecution.ActionApprovalGate

#region ScenarioExecution.Cancel [C:4] [TYPE Function] [SEMANTICS scenario,execution,cancel,run]

@ingroup ScenarioExecution

@BRIEF Request cancellation; drain in-flight within a bounded window; terminate cancelled.

@PRE run not terminal.

@POST cancel_requested -> cancelled; in-flight completes or times out.

@SIDE_EFFECT stops new dispatches; bounded drain.

def cancel_run(db, run_id): ...

#endregion ScenarioExecution.Cancel

#region ScenarioExecution.RetryStep [C:4] [TYPE Function] [SEMANTICS scenario,execution,retry,step]

@ingroup ScenarioExecution

@BRIEF Retry a failed step; invalidates its downstream closure and re-runs descendants.

@PRE bounded attempts not exceeded; closure re-run allowed.

@POST increments attempt; re-executes step + downstream closure; rejected after run advanced beyond.

@INVARIANT outputs of dependents are never left stale relative to a re-run producer.

def retry_step(db, run_id, logical_step_id): ...

#endregion ScenarioExecution.RetryStep

#region ScenarioExecution.ClaimStep [C:4] [TYPE Function] [SEMANTICS scenario,execution,worker,lease,claim]

@ingroup ScenarioExecution

@BRIEF Claim a step for a worker with lease/heartbeat for crash recovery.

@PRE worker lease valid; step not already claimed by another live worker.

@POST step claimed; lease+heartbeat started; idempotency/side-effect key checked.

@SIDE_EFFECT lease record; DB write.

@INVARIANT at-least-once: a step with a completed external side effect is not re-executed unless idempotent/retry-safe.

def claim_step(run_id, logical_step_id, worker_id): ...

#endregion ScenarioExecution.ClaimStep

#region ScenarioExecution.ExecutorRegistry [C:3] [TYPE Module] [SEMANTICS scenario,execution,registry,executor]

@ingroup ScenarioExecution

@BRIEF ActionExecutionDescriptor -> executor mapping; BrowserExecutor resolves only exact 038

ActionRegistry actions; human excluded; descriptor declares idempotency/retry-safety.

@REJECTED human as executor (HumanCheckpoint lifecycle control instead).

EXECUTOR_MAP = {browser: BrowserExecutor(ActionRegistry), superset_api, xlsx, assertion, screenshot, report, artifact}

@INVARIANT descriptor, not executor/tool default, declares {idempotent, retry_safe, side_effect_key, timeout, mutation risk}.

@INVARIANT mutation actions require immutable mutation_contract; PROD mutation is rejected; mutation retries default false.

#endregion ScenarioExecution.ExecutorRegistry

#region ScenarioExecution.ProviderProtocol [C:5] [TYPE ADR] [SEMANTICS scenario,execution,provider,protocol,lifecycle,ownership,observability]

@ingroup ScenarioExecution

@BRIEF Common production contract for every external or resource-owning scenario provider.

@DATA_CONTRACT ProviderExecutionContext -> ProviderExecutionResult -> ScenarioStepRun/Artifact

@PRE Provider is registered by trusted startup composition, declares a stable provider_id,

provider_version, supported action descriptors, resource limits, health status, and

cancellation capabilities before accepting a run.

@POST Every invocation is bound to run_id, logical_step_id, attempt, descriptor fingerprint,

execution principal fingerprint, deadline, idempotency key, capacity lease, and trace_id;

result is persisted exactly once for that attempt or is reconciled before retry.

@SIDE_EFFECT External I/O, provider-owned resources, durable evidence, metrics, structured events.

@INVARIANT A provider cannot select authority, environment, principal, action policy, retry policy,

or capacity outside the pinned descriptor and server-owned composition.

@INVARIANT A provider result is accepted only when ownership proof binds every output/evidence ref

to run_id, step_id, attempt, provider_id/version, descriptor fingerprint and digest.

@INVARIANT Provider timeout/cancellation never implies rollback; unknown external effect state is

reconciled or terminalized non-pass before another attempt may start.

@INVARIANT Health failure prevents new claims but does not erase active leases or evidence history.

@REJECTED A generic callable-only provider contract was rejected — it hides resources, ownership,

cancellation, reconciliation and deployability obligations.

ProviderExecutionContext = { "run_id": "UUID", "logical_step_id": "UUID", "attempt": "positive integer", "descriptor_fingerprint": "sha256", "binding_ref": "opaque server-owned ref", "execution_principal_fingerprint": "sha256", "deadline_at": "RFC3339 timestamp", "idempotency_key": "opaque stable key", "capacity_lease_id": "opaque lease ref", "trace_id": "opaque trace ref", "cancellation": {"requested": "bool", "deadline_at": "RFC3339|null"}, }

ProviderExecutionResult = { "status": "passed|failed|inconclusive|blocked", "reason_code": "stable taxonomy code", "provider_id": "stable id", "provider_version": "immutable version", "operation_id": "provider operation ref", "output_refs": "owned output refs", "evidence": "owned evidence refs + sha256 + media/type metadata", "effect_state": "none|completed|not_started|unknown|requires_reconciliation", "retry_disposition": "never|safe|after_reconciliation|manual_only", "observability": "trace/span/metric dimensions", }

Provider lifecycle is: registered -> health_checked -> capacity_claimed -> admitted -> running ->

completed|failed|cancel_requested -> finalized|reconciliation_required. A provider MUST release

capacity and ephemeral resources in finalized/reconciliation_required, while durable evidence and

operation receipts remain immutable. A provider MUST expose readiness (can accept new work), liveness

(process responsive), and dependency health (external system reachable/authenticated) separately.

Cancellation is cooperative first: runner marks cancellation_requested and invokes provider cancel

with operation_id. The provider acknowledges stopped|completed|unknown by the deadline. Unknown means

no retry and no PASS until reconcile(operation_id, idempotency_key) returns a terminal receipt or the

run is terminalized inconclusive/blocked. Retry is a new attempt with a new attempt id but the same

logical side-effect identity where the descriptor requires it.

#endregion ScenarioExecution.ProviderProtocol

#region ScenarioExecution.ProviderRuntime [C:5] [TYPE ADR] [SEMANTICS scenario,execution,provider,runtime,event-loop,concurrency,trace]

@ingroup ScenarioExecution

@BRIEF Pin the async/sync bridge, environment concurrency defaults and debug-capture policy for all providers.

@RELATION IMPLEMENTS -> [ScenarioExecution.ProviderProtocol]

@RELATION DEPENDS_ON -> [ScenarioExecution.CapacityManager]

@PRE The application owns exactly one provider event-loop thread started by startup composition;

concurrency policy and provider limits are server-owned configuration, never per-request input.

@POST Sync executors submit provider coroutines to that loop via run_coroutine_threadsafe and block on

the result under the step deadline; browser/context/session objects live only on that loop and

remain usable across every step of one run without re-authentication.

@INVARIANT Async provider transports (Playwright and successors) are bound to exactly one long-lived

loop; contexts, pages and sessions are never shared across runs and never cross loops.

@INVARIANT All provider concurrency is admitted through the CapacityManager. Pinned defaults: 2

concurrent browser contexts for DEV/PREPROD, 1 for PROD, per environment; screenshot capture

is admitted through the same lease accounting. Exhausted capacity leaves the run queued with

the stable CAPACITY_BLOCKED/CAPACITY_UNAVAILABLE reason; the submit queue is bounded —

overflow waits under the step deadline and never starts I/O.

@REJECTED Per-step asyncio.run was rejected — a fresh loop per call cannot hold live context objects and

forces re-authentication per step. A dedicated browser worker process is deferred, not

rejected: the LiveExecutionAdapter Protocol must survive its later introduction. Warm context

pools are rejected for Phase 1 — cold start must fit the 120s context/auth limit.

@REJECTED Playwright trace/video as durable evidence was rejected — non-deterministic heavyweight bytes

outside descriptor-declared evidence. An operator debug flag may enable trace with local

retention only; such output is never registered as an Artifact or EvidenceReceipt.

@RATIONALE A single in-process loop preserves lease-scoped isolation (one context per run/lease) and

step-to-step session continuity at the lowest operational cost; loop replacement by a worker

process later changes composition only, not the adapter contract.

#endregion ScenarioExecution.ProviderRuntime

#region ScenarioExecution.ProviderCatalog [C:5] [TYPE ADR] [SEMANTICS scenario,execution,provider,catalog,browser,superset,xlsx,screenshot,agent]

@ingroup ScenarioExecution

@BRIEF Per-provider obligations and production acceptance criteria.

@INVARIANT Every catalog entry has input/output schemas, limits, ownership proof, failure taxonomy,

cancel/reconcile semantics, health checks, deployment registration and falsifiable tests.

Provider obligations:

  • BrowserProvider: owns an isolated browser context for one run/lease; authenticates only through the server-owned auth binding; accepts only registry-pinned actions and typed inputs; emits an operation receipt, safe-checkpoint ref, page/dashboard provenance and durable evidence refs where applicable; closes context on finalization. It MUST reject unsafe recovery without a pinned safe checkpoint, report mutation effects as completed|unknown, and reconcile by operation id before retry.
  • SupersetProvider: invokes only the exact 037 query envelope/model/principal/RLS binding; enforces request, response, timeout and row/byte limits; records request/response fingerprints and raw-byte evidence ownership. 403/422 are non-retryable policy/input failures, 429/5xx are bounded retryable transport failures, and timeout/cancel with unknown server execution requires reconciliation by the external request id before retry.
  • SqlEvidenceProvider: is read-only and executes an immutable 038 SqlEvidenceSpec; runtime may bind typed parameters only. It MUST prove database identity, query-model fingerprint, principal/RLS fingerprint, raw response digest and durable evidence ref. Caller SQL or metadata never qualifies.
  • XlsxProvider: consumes a server-owned immutable artifact ref, not unbounded caller bytes; validates content digest, workbook type, archive/resource limits, allowed sheets/formulas/external links and canonical normalization. Parsing is local and idempotent; source artifact ownership is retained.
  • AssertionProvider: is pure and deterministic over canonical inputs; invokes 037 comparison with a pinned policy/version, has no external resources, and emits no PASS without complete inputs.
  • TransformProvider: executes only the bounded 038 TransformSpec DSL with hard operation, row, column and output limits; no code, SQL, imports or network. It is deterministic and retry-safe.
  • ScreenshotProvider: captures through a server-owned browser/capture binding and writes evidence atomically to durable storage. Each ref must contain an ownership receipt for run/step/attempt, binding, capture profile, media type, byte length and SHA-256. Temp files are cleaned after commit or failure; a path or raw bytes alone are never evidence.
  • ReportProvider: renders a version-pinned template from an immutable evidence manifest; output is deterministic for the same manifest/template, size-bounded, durably stored and digest-verified. Partial render/write is cleaned or marked reconciliation-required.
  • ArtifactProvider: registers only content refs created by an authorized producer or verified source artifact. It must not mint evidence from caller-supplied digest strings; ownership, digest, retention and active/inactive projection are immutable and audit-visible.
  • AgentEvaluationProvider: executes only a declared immutable 038 AgentEvaluationSpec, with explicit provider/model/prompt versions, input manifest hash, token/cost/time limits and evidence allowlist. It emits immutable AgentEvaluation; it cannot mutate the graph, invoke actions, consume checkpoints, schedule work or set ScenarioResult. DecisionPolicy alone maps evaluation plus deterministic evidence to StepOutcome.

@REJECTED Treating provider availability as successful execution was rejected — unavailable,

unhealthy, missing-capacity and missing-evidence outcomes are explicit non-pass states.

#endregion ScenarioExecution.ProviderCatalog

#region ScenarioExecution.BrowserProvider [C:5] [TYPE Module] [SEMANTICS scenario,execution,browser,playwright,actions,checkpoint,mutation]

@ingroup ScenarioExecution

@BRIEF Production contract for isolated BrowserProvider execution through the 038 ActionRegistry.

@RELATION IMPLEMENTS -> [ScenarioExecution.ProviderProtocol]

@RELATION DEPENDS_ON -> [ScenarioExecution.CapacityManager]

@RELATION DEPENDS_ON -> [ScenarioExecution.ProviderOperations]

@PRE Startup registration supplies provider/version, browser versions, catalog/binding fingerprints,

limits, health and cancellation capability. A valid CapacityLease and operation receipt exist

before browser/context creation.

@POST Every action returns typed output and an immutable operation receipt; declared evidence has an

EvidenceReceipt; context/pages/downloads/temp files are closed or accounted before finalization.

@SIDE_EFFECT Browser process/context, authenticated session, navigation, downloads, evidence and only

descriptor-authorized mutation effects.

@INVARIANT Authority comes only from the exact persisted binding and 038 descriptor; URLs, IDs, sessions,

cookies and request metadata cannot select credentials or principal.

@INVARIANT Mutation requires mutation_contract, fixture lease, target keys, precondition hash, cleanup

policy and retry_safe=false; PROD mutation is rejected before browser invocation.

@INVARIANT Crash recovery never revives a dead context; it reconstructs from a safe checkpoint. Unknown

mutation effects require reconciliation before retry.

@INVARIANT Browser PASS requires action receipt, output-shape validation and binding/principal/evidence

provenance; page/session metadata alone cannot produce PASS.

@REJECTED Process-global Playwright pages were rejected because they permit cross-run cookies and target

leakage. Dead-context replay was rejected because external effect state is unknowable.

@RATIONALE Mutation surface decision (2026-09-01, test-stand evidence): Superset table charts are

read-only, so edit_row/bulk_edit execute as SQL-Lab-mediated data mutations performed

by the authenticated browser session (page-context fetch to the CSRF-protected SQL Lab

execute endpoint, DML-gated by the target database's allow_dml policy). Preconditions are

the hash of a pre-mutation SELECT over the contract's target keys; cleanup_policy=restore

re-applies the recorded pre-mutation values through the same path; reconciliation is a

post-mutation SELECT comparison. UI-DOM editing was rejected — no editable widget exists

in stock Superset, and a custom plugin is a separate deployment capability.

BrowserProviderActionContract = { "action": "exact registered 038 action", "input_schema": "versioned typed JSON schema", "output_schema": "versioned typed JSON schema", "risk": "read_only|evidence|mutation", "context_timeout_ms": 120000, "action_timeout_ms": 30000, "max_pages": 3, "max_download_bytes": 26214400, "max_screenshot_bytes": 10485760, "checkpoint_policy": "none|safe_reconstruction_point|required", "idempotency": "idempotent|keyed|non_idempotent", "retry_safe": "bool", "mutation_contract": "required for risk=mutation", }

Per-action policy:

  • open_dashboard: typed dashboard/release/binding input; output canonical URL, dashboard fingerprint and dashboard_open checkpoint; read-only.
  • navigate_tab: allowlisted tab/route token only; output selected tab and tab_navigated checkpoint; external origins are rejected.
  • apply_native_filter: typed query-model filter input; output normalized state and filters_applied checkpoint; saving a filter is a separate mutation action.
  • inspect_filter_state, inspect_columns, scroll_to, wait_for_state: typed bounded read-only actions; timeout is inconclusive and cannot manufacture a checkpoint.
  • apply_table_filter: typed normalized table filter; it cannot persist dashboard state.
  • extract_table: bounded selector/projection; max 10,000 rows, 100 columns and 10 MiB output; declared evidence requires an EvidenceReceipt.
  • click: locator plus explicit effect class; submit/save/delete/edit behavior is mutation, and ambiguous clicks are rejected.
  • select_rows: exact row keys/locator; selection-only is read-only and never proves mutation completion.
  • edit_row: exact fixture/record keys, field allowlist and precondition hash; fixture lease, cleanup and reconciliation are mandatory; prohibited in PROD.
  • bulk_edit: same mutation requirements, max 100 record keys and explicit rollback/cleanup plan; prohibited in PROD.
  • download: allowlisted artifact type and max bytes; output is a server-owned artifact ref/digest, never a local path.
  • refresh: bounded read-only render operation; it is not evidence of data correctness without a separate evidence action.

Resource lifecycle is lease_claimed -> context_created -> authenticated -> action_running -> checkpointed|effect_recorded -> cleanup_started -> receipt_finalized -> context_closed. At most one active context exists per (run_id, capacity_lease_id), with max three pages, 25 MiB downloads and 10 MiB screenshots. Context/authentication has a 120-second limit; each action has a 30-second limit unless the descriptor pins a lower value. Cleanup runs on success, failure, timeout and cancellation; cleanup failure prevents PASS and marks mutation reconciliation required.

cancel(operation_id) stops new actions, closes the page/context by the persisted drain deadline and returns stopped|completed|unknown with checkpoint/effect state. Read-only work may retry after clean close; unknown mutation work cannot retry until reconcile(operation_id, idempotency_key) verifies target state. Recovery never attaches to the old process: it reconstructs from dashboard_open, filters_applied, tab_navigated or another descriptor-declared safe checkpoint and records reconstruction_replay=true.

Readiness is ready only when browser executable/version, auth binding, allowed origins, evidence storage, cleanup worker and cancellation capability checks pass. Health exposes provider/version, catalog/binding fingerprints, origin hash, limits and dependency status, excluding cookies, credentials, secret URLs and page content. Deployment requires one read-only PREPROD canary and one forced timeout/cleanup canary before enabling the provider; mutation requires a separate fixture-lease canary.

Acceptance profile: 100% pass for schema rejection, cross-run isolation, unauthorized-origin rejection, binding/lease failures, timeout, cancellation, duplicate/late response, safe recovery, unsafe mutation recovery, cleanup failure, ownership mismatch, size limits, health redaction and startup registration. A real PREPROD canary must produce one owned evidence receipt and one reconstruction trace.

#endregion ScenarioExecution.BrowserProvider

#region ScenarioExecution.CapacityManager [C:5] [TYPE Module] [SEMANTICS scenario,execution,capacity,quota,lease,concurrency]

@ingroup ScenarioExecution

@BRIEF Environment-scoped allocator shared by ScenarioRun and neighboring workloads.

@PRE Claim contains environment, workload_class, priority, requested units, provider_id, run/step,

deadline and idempotency context; capacity policy is server-owned.

@POST Atomic claim returns a lease or CAPACITY_UNAVAILABLE; lease heartbeat, expiry, release and

forced reconciliation are durable and idempotent.

@INVARIANT No provider starts external I/O before a valid capacity lease; retries claim separately.

@INVARIANT Capacity exhaustion creates queued/blocked state according to policy, never an unbounded

provider call or a client-controlled quota increase.

@SIDE_EFFECT Durable lease/quota records, admission metrics, provider concurrency accounting.

@REJECTED Independent per-feature concurrency limits were rejected — they permit cross-workload

starvation and bypass the environment-scoped allocator.

def claim_capacity(context: ProviderExecutionContext, request): ... def heartbeat_capacity(lease_id: str): ... def release_capacity(lease_id: str): ...

#endregion ScenarioExecution.CapacityManager

#region ScenarioExecution.ProviderOperations [C:5] [TYPE Module] [SEMANTICS scenario,execution,provider,operation,cancel,retry,reconcile]

@ingroup ScenarioExecution

@BRIEF Durable operation receipts and recovery protocol for provider side effects.

@PRE Operation receipt is created before external I/O with operation_id, idempotency_key,

descriptor/provider fingerprints and capacity lease.

@POST Every operation reaches completed, failed, cancelled, or reconciliation_required; no receipt

is overwritten, and a late provider response is stored as historical evidence only.

@INVARIANT Retry cannot begin while effect_state is unknown or reconciliation_required.

@SIDE_EFFECT Operation receipt, cancellation request, reconciliation record and audit events.

def request_provider_cancel(operation_id: str, deadline_at: str): ... def reconcile_provider_operation(operation_id: str, idempotency_key: str): ...

#endregion ScenarioExecution.ProviderOperations

#region ScenarioExecution.ProviderOperations.Observability [C:4] [TYPE Module] [SEMANTICS scenario,execution,provider,health,metrics,tracing]

@ingroup ScenarioExecution

@BRIEF Required provider health, readiness, trace and metric surface.

@POST Provider exposes liveness, readiness, dependency health, registration fingerprint and metrics

for admission, latency, timeout, cancellation, retry, reconciliation, capacity and evidence.

@INVARIANT Health output never includes credentials, cookies, raw SQL, capture bytes or secret payloads.

@SIDE_EFFECT Structured events and metrics correlated by trace_id/run_id/step_id/operation_id.

@TEST_EDGE dependency_down->not_ready; auth_failure->unavailable; stale_registration->not_ready;

health_payload_redaction->no_secret_leak.

#endregion ScenarioExecution.ProviderOperations.Observability

#region ScenarioExecution.ReleaseEvidence [C:5] [TYPE ADR] [SEMANTICS scenario,execution,release,verification,evidence,thresholds]

@ingroup ScenarioExecution

@BRIEF Define the measurable evidence required to declare 044 production-ready.

@POST Release evidence contains command output or deployment records for SC-001..SC-011, the full

available backend profile, provider contract profile, PostgreSQL migration verification and

prototype validation.

@INVARIANT A local unit profile cannot close a live-provider, PostgreSQL, scheduler or cross-spec gate.

@INVARIANT Production GO requires 100% pass for all mandatory profiles, zero unresolved P0/P1 rows in

044 and execution-critical dependencies, and no enabled provider without readiness evidence.

@TEST_EDGE missing_profile->NO_GO; stale_profile->NO_GO; provider_not_ready->NO_GO;

unresolved_dependency_P1->NO_GO.

#endregion ScenarioExecution.ReleaseEvidence

@{ ScenarioExecution.AuthoringAdmission [C:5] [TYPE Function]

@BRIEF Reject non-promoted authoring artifacts at production execution admission. @RELATION DEPENDS_ON -> [ScenarioExecution.RunPreflight] @RELATION DEPENDS_ON -> [ScenarioRegistry.AgentAuthoringPromotion]

@PRE Request references a persisted 042 ScenarioRevision, not a workspace/artifact/code payload. @POST Only a valid promoted revision can produce a RunPreflightHandle; sandbox output produces no run, lease, provider I/O or evidence. @INVARIANT Raw code, URL, cookie, secret, filesystem path and caller digest are never authority; future code-backed execution remains separately gated and unimplemented. @TEST_EDGE raw_authoring_payload->rejected; candidate_without_promotion->blocked; exploration_trace->no_run. def admit_promoted_revision(db, scenario_id, revision_id, actor): ...

@} ScenarioExecution.AuthoringAdmission

#endregion ScenarioExecution.Modules


OPENAPI — REST/Event API Contract

Source: contracts/openapi.yaml

openapi: 3.1.0 info: title: Scenario Execution Engine API version: 0.1.0 description: Start and manage execution of dashboard test scenario runs (044). paths: /api/scenario-runs: post: operationId: scenarioRun.start summary: Start a scenario run pinned to an immutable revision security: [{ bearerAuth: [] }] parameters: - { name: Idempotency-Key, in: header, required: true, schema: { type: string } } requestBody: required: true content: application/json: schema: type: object required: [scenario_id, revision_id, environment_id] properties: scenario_id: { type: string } revision_id: { type: string } environment_id: { type: string } params: { type: object, description: "Launch values; server validates against 038 ParameterDefinition and persists immutable ParameterBinding[]" } requested_target_reference: { type: object, description: "User-selected release/target; server verifies it matches the captured TargetSnapshot" } baseline_set: { type: string } execution_toggles: { type: object, description: "Optional evidence only (diagnostic screenshots, verbose logs, optional VLM). Mandatory steps cannot be disabled." } responses: "201": { description: ScenarioRun created (queued or pending_approval), content: { application/json: { schema: { $ref: "#/components/schemas/ScenarioRun" } } } } "403": { description: Permission denied; PROD approval creates pending_approval run rather than returning 403 } "409": { description: Stale revision or IDEMPOTENCY_KEY_REUSED with a different canonical request } /api/scenarios/{scenario_id}/runs: get: operationId: scenarioRun.history summary: Scenario run history (for 045 RunHistoryList) security: [{ bearerAuth: [] }] parameters: [{ name: scenario_id, in: path, required: true, schema: { type: string } }] responses: "200": description: Run history content: application/json: schema: { type: array, items: { $ref: "#/components/schemas/ScenarioRun" } } /api/scenario-runs/{run_id}/result: get: operationId: scenarioRun.result summary: Final result with aggregation + provenance security: [{ bearerAuth: [] }] parameters: [{ name: run_id, in: path, required: true, schema: { type: string } }] responses: { "200": { description: ScenarioExecutionResult, content: { application/json: { schema: { $ref: "#/components/schemas/ScenarioExecutionResult" } } } } } /api/scenario-runs/compare: get: operationId: scenarioRun.compare summary: Compare two runs (per-step deltas, revision-diff warning) security: [{ bearerAuth: [] }] parameters: - { name: a, in: query, required: true, schema: { type: string } } - { name: b, in: query, required: true, schema: { type: string } } responses: { "200": { description: RunComparison, content: { application/json: { schema: { $ref: "#/components/schemas/RunComparison" } } } } } /api/scenario-runs/{run_id}/steps/{logical_step_id}/retry: post: operationId: scenarioRun.retryStep summary: Retry a failed step + invalidated downstream closure security: [{ bearerAuth: [] }] parameters: - { name: run_id, in: path, required: true, schema: { type: string } } - { name: logical_step_id, in: path, required: true, schema: { type: string } } responses: "200": { description: Retried; downstream closure re-run } "409": { description: Run advanced beyond step; closure re-run rejected } /api/scenario-runs/{run_id}: get: operationId: scenarioRun.detail summary: Get a run by id (recoverable) security: [{ bearerAuth: [] }] parameters: [{ name: run_id, in: path, required: true, schema: { type: string } }] responses: { "200": { description: Run detail, content: { application/json: { schema: { $ref: "#/components/schemas/ScenarioRun" } } } } } /api/scenario-runs/{run_id}/events: get: operationId: scenarioRun.events summary: SSE stream of step events (replayable via Last-Event-ID) security: [{ bearerAuth: [] }] parameters: - { name: run_id, in: path, required: true, schema: { type: string } } - { name: Last-Event-ID, in: header, schema: { type: integer }, description: "sequence to replay from" } responses: "200": content: { text/event-stream: { schema: { $ref: "#/components/schemas/ScenarioRunEvent" } } } /api/scenario-runs/{run_id}/cancel: post: operationId: scenarioRun.cancel summary: Cancel a run (bounded drain) security: [{ bearerAuth: [] }] parameters: [{ name: run_id, in: path, required: true, schema: { type: string } }] responses: "202": { description: Cancellation requested; persisted drain deadline returned, content: { application/json: { schema: { $ref: "#/components/schemas/ScenarioRun" } } } } "200": { description: Cancellation already finalized, content: { application/json: { schema: { $ref: "#/components/schemas/ScenarioRun" } } } } "409": { description: Run cannot be cancelled from its current terminal state } /api/scenario-runs/{run_id}/resume: post: operationId: scenarioRun.resume summary: Resume a paused run from resume token security: [{ bearerAuth: [] }] parameters: [{ name: run_id, in: path, required: true, schema: { type: string } }] requestBody: required: true content: application/json: schema: type: object required: [resume_token, resume_reason] properties: resume_token: { type: string } resume_reason: { type: string, enum: [worker_recovered, infrastructure_pause_resolved] } responses: { "200": { description: Infrastructure pause resumed }, "409": { description: Cannot resume HumanCheckpoint or stale token } } /api/scenario-runs/{run_id}/human/decision: post: operationId: scenarioRun.humanDecision summary: Resolve a HumanCheckpoint (observation disposition) security: [{ bearerAuth: [] }] parameters: [{ name: run_id, in: path, required: true, schema: { type: string } }] requestBody: required: true content: application/json: schema: type: object required: [checkpoint_id, decision_version, disposition] properties: checkpoint_id: { type: string } decision_version: { type: integer } disposition: { type: string, enum: [confirm, false_positive, pass, fail, inconclusive] } comment: { type: string } responses: { "200": { description: Checkpoint consumed; step outcome recorded; runner resumed internally }, "409": { description: Checkpoint already decided, expired, or version conflict } } components: securitySchemes: bearerAuth: { type: http, scheme: bearer } schemas: ScenarioRun: type: object required: [id, scenario_id, scenario_revision_id, environment_id, status, parameter_bindings, target_snapshot, execution_principal_fingerprint, verification_program_hash, action_registry_version, step_runs] properties: id: { type: string } scenario_id: { type: string } scenario_revision_id: { type: string } environment_id: { type: string } status: { type: string, enum: [pending_approval, queued, running, waiting_human, blocked, cancel_requested, cancelled, passed, failed, inconclusive] } parameter_bindings: { type: array, items: { $ref: "#/components/schemas/ParameterBinding" } } target_snapshot: { $ref: "#/components/schemas/TargetSnapshot" } execution_principal_fingerprint: { type: string } analytics_context_key: { type: string, description: "Server-derived immutable analytics grouping key" } verification_program_hash: { type: string } action_registry_version: { type: string } step_runs: { type: array, items: { $ref: "#/components/schemas/ScenarioStepRun" } } ScenarioStepRun: type: object required: [id, logical_step_id, step_position, attempt, status, outputs, artifact_refs, error_code, step_outcome] properties: id: { type: string } logical_step_id: { type: string } step_position: { type: integer } attempt: { type: integer } status: { type: string, enum: [queued, running, waiting_human, passed, failed, inconclusive, blocked, skipped, cancelled] } outputs: { type: object } artifact_refs: { type: array, items: { type: string } } error_code: { type: string, nullable: true } step_outcome: { $ref: "#/components/schemas/StepOutcome" } agent_evaluations: { type: array, items: { $ref: "#/components/schemas/AgentEvaluationSummary" } } ParameterBinding: type: object required: [parameter_name, resolved_value, source, resolved_at] properties: parameter_name: { type: string } resolved_value: {} source: { type: string, enum: [launch_input, default, schedule, trigger] } resolved_at: { type: string, format: date-time } TargetSnapshot: type: object required: [environment_id, dashboard_fingerprint, dataset_lineage_fingerprint, captured_at] properties: environment_id: { type: string } dashboard_release_id: { type: [string, "null"] } dashboard_fingerprint: { type: string } dataset_lineage_fingerprint: { type: string } captured_at: { type: string, format: date-time } ScenarioExecutionResult: type: object required: [run_id, status, step_counts, provenance, failures] properties: run_id: { type: string } status: { type: string } step_counts: { type: object, additionalProperties: { type: integer } } failures: { type: array, items: { type: object } } provenance: { $ref: "#/components/schemas/ExecutionProvenance" } analytics_context_key: { type: string } step_outcomes: { type: array, items: { $ref: "#/components/schemas/StepOutcome" } } agent_evaluation_summary: { type: object } StepOutcome: type: object required: [status, reason_codes, deterministic_evidence_refs] properties: status: { type: string, enum: [passed, failed, inconclusive, blocked, waiting_human] } reason_codes: { type: array, items: { type: string } } deterministic_evidence_refs: { type: array, items: { type: string } } agent_evaluation_ids: { type: array, items: { type: string } } decision_policy_id: { type: [string, "null"] } decision_policy_version: { type: [string, "null"] } ExecutionProvenance: type: object required: [scenario_revision_id, scenario_content_hash, verification_program_hash, runner_version, target_snapshot, parameter_bindings, execution_principal_fingerprint, analytics_context_key] properties: scenario_revision_id: { type: string } scenario_content_hash: { type: string, pattern: '^[0-9a-fA-F]{64}$' } verification_program_hash: { type: string, pattern: '^[0-9a-fA-F]{64}$' } runner_version: { type: string, minLength: 1 } target_snapshot: { $ref: "#/components/schemas/TargetSnapshot" } parameter_bindings: { type: array, items: { $ref: "#/components/schemas/ParameterBinding" } } execution_principal_fingerprint: { type: string, minLength: 1 } analytics_context_key: { type: string, pattern: '^[0-9a-fA-F]{64}$' } evidence_receipts: { type: array, items: { $ref: "#/components/schemas/EvidenceReceipt" } } EvidenceReceipt: type: object required: [evidence_ref, owner_type, owner_id, run_id, logical_step_id, attempt, operation_id, provider_id, provider_version, descriptor_fingerprint, content_type, byte_length, sha256] properties: evidence_ref: { type: string, minLength: 1 } owner_type: { type: string, const: scenario_run } owner_id: { type: string } run_id: { type: string } logical_step_id: { type: string } attempt: { type: integer, minimum: 1 } operation_id: { type: string, minLength: 1 } provider_id: { type: string, minLength: 1 } provider_version: { type: string, minLength: 1 } descriptor_fingerprint: { type: string, pattern: '^[0-9a-fA-F]{64}$' } content_type: { type: string, minLength: 1 } byte_length: { type: integer, minimum: 1 } sha256: { type: string, pattern: '^[0-9a-fA-F]{64}$' } AgentEvaluationSummary: type: object required: [evaluation_id, logical_step_id, verdict, confidence, model_id, prompt_template_version] properties: evaluation_id: { type: string } logical_step_id: { type: string } verdict: { type: string, enum: [pass, fail, inconclusive] } confidence: { type: number, minimum: 0, maximum: 1 } model_id: { type: string } prompt_template_version: { type: string } reason_codes: { type: array, items: { type: string } } evidence_refs: { type: array, items: { type: string } } AgentEvaluation: allOf: - $ref: "#/components/schemas/AgentEvaluationSummary" - type: object required: [scenario_run_id, attempt, provider_id, model_version, prompt_template_id, input_manifest_hash, findings, raw_response_artifact_ref, started_at, finished_at] properties: scenario_run_id: { type: string } attempt: { type: integer, minimum: 1 } provider_id: { type: string } model_version: { type: string } prompt_template_id: { type: string } input_manifest_hash: { type: string } findings: { type: array, items: { type: object } } raw_response_artifact_ref: { type: string } started_at: { type: string, format: date-time } finished_at: { type: string, format: date-time } DecisionPolicy: type: object required: [policy_id, version, deterministic_hard_failure, high_confidence_failure, low_confidence, disagreement, missing_evidence] properties: policy_id: { type: string } version: { type: string } deterministic_hard_failure: { type: string, enum: [failed] } high_confidence_failure: { type: string, enum: [failed, inconclusive] } low_confidence: { type: string, enum: [inconclusive] } disagreement: { type: string, enum: [waiting_human, inconclusive] } missing_evidence: { type: string, enum: [blocked, inconclusive] } RunComparison: type: object required: [run_a, run_b, step_deltas, compatibility] properties: run_a: { type: string } run_b: { type: string } compatibility: { type: object } step_deltas: { type: array, items: { type: object } } ScenarioRunEvent: oneOf: - { $ref: "#/components/schemas/RunStartedEvent" } - { $ref: "#/components/schemas/ApprovalRequiredEvent" } - { $ref: "#/components/schemas/RunQueuedEvent" } - { $ref: "#/components/schemas/StepStartedEvent" } - { $ref: "#/components/schemas/StepProgressEvent" } - { $ref: "#/components/schemas/EvidenceCreatedEvent" } - { $ref: "#/components/schemas/AgentEvaluationStartedEvent" } - { $ref: "#/components/schemas/AgentEvaluationCompletedEvent" } - { $ref: "#/components/schemas/StepCompletedEvent" } - { $ref: "#/components/schemas/CheckpointCreatedEvent" } - { $ref: "#/components/schemas/RunCompletedEvent" } - { $ref: "#/components/schemas/RunFailedEvent" } - { $ref: "#/components/schemas/RunCancelledEvent" } RunStartedEvent: type: object required: [id, sequence, event_type, run_id, occurred_at] properties: { id: { type: string }, sequence: { type: integer }, event_type: { const: run_started }, run_id: { type: string }, occurred_at: { type: string, format: date-time }, payload: { type: object } } ApprovalRequiredEvent: type: object required: [id, sequence, event_type, run_id, occurred_at] properties: { id: { type: string }, sequence: { type: integer }, event_type: { const: approval_required }, run_id: { type: string }, occurred_at: { type: string, format: date-time }, payload: { type: object, required: [approval_gate_id] } } RunQueuedEvent: type: object required: [id, sequence, event_type, run_id, occurred_at] properties: { id: { type: string }, sequence: { type: integer }, event_type: { const: run_queued }, run_id: { type: string }, occurred_at: { type: string, format: date-time }, payload: { type: object } } StepStartedEvent: type: object required: [id, sequence, event_type, run_id, logical_step_id, occurred_at] properties: { id: { type: string }, sequence: { type: integer }, event_type: { const: step_started }, run_id: { type: string }, logical_step_id: { type: string }, occurred_at: { type: string, format: date-time }, payload: { type: object } } StepProgressEvent: type: object required: [id, sequence, event_type, run_id, logical_step_id, occurred_at] properties: { id: { type: string }, sequence: { type: integer }, event_type: { const: step_progress }, run_id: { type: string }, logical_step_id: { type: string }, occurred_at: { type: string, format: date-time }, payload: { type: object, required: [progress] } } EvidenceCreatedEvent: type: object required: [id, sequence, event_type, run_id, logical_step_id, occurred_at] properties: { id: { type: string }, sequence: { type: integer }, event_type: { const: evidence_created }, run_id: { type: string }, logical_step_id: { type: string }, occurred_at: { type: string, format: date-time }, payload: { type: object, required: [evidence_ref] } } AgentEvaluationStartedEvent: type: object required: [id, sequence, event_type, run_id, logical_step_id, occurred_at] properties: { id: { type: string }, sequence: { type: integer }, event_type: { const: agent_evaluation_started }, run_id: { type: string }, logical_step_id: { type: string }, occurred_at: { type: string, format: date-time }, payload: { type: object } } AgentEvaluationCompletedEvent: type: object required: [id, sequence, event_type, run_id, logical_step_id, occurred_at] properties: { id: { type: string }, sequence: { type: integer }, event_type: { const: agent_evaluation_completed }, run_id: { type: string }, logical_step_id: { type: string }, occurred_at: { type: string, format: date-time }, payload: { $ref: "#/components/schemas/AgentEvaluationSummary" } } StepCompletedEvent: type: object required: [id, sequence, event_type, run_id, logical_step_id, occurred_at] properties: { id: { type: string }, sequence: { type: integer }, event_type: { const: step_completed }, run_id: { type: string }, logical_step_id: { type: string }, occurred_at: { type: string, format: date-time }, payload: { type: object } } CheckpointCreatedEvent: type: object required: [id, sequence, event_type, run_id, occurred_at] properties: { id: { type: string }, sequence: { type: integer }, event_type: { const: checkpoint_created }, run_id: { type: string }, occurred_at: { type: string, format: date-time }, payload: { type: object } } RunCompletedEvent: type: object required: [id, sequence, event_type, run_id, occurred_at] properties: { id: { type: string }, sequence: { type: integer }, event_type: { const: run_completed }, run_id: { type: string }, occurred_at: { type: string, format: date-time }, payload: { type: object } } RunFailedEvent: type: object required: [id, sequence, event_type, run_id, occurred_at] properties: { id: { type: string }, sequence: { type: integer }, event_type: { const: run_failed }, run_id: { type: string }, occurred_at: { type: string, format: date-time }, payload: { type: object } } RunCancelledEvent: type: object required: [id, sequence, event_type, run_id, occurred_at] properties: { id: { type: string }, sequence: { type: integer }, event_type: { const: run_cancelled }, run_id: { type: string }, occurred_at: { type: string, format: date-time }, payload: { type: object } }


QUICKSTART — Dev Onboarding

Source: quickstart.md

Quickstart: Scenario Execution Engine (044)

Verification status (2026-08-24): the available SQLite/unit profile passes, but this is not production evidence. Browser and Screenshot providers remain unavailable by default. Production GO additionally requires provider contract tests, real PostgreSQL migration checks and live deployment proof.

Prerequisites

  • PostgreSQL reachable through a real DATABASE_URL.
  • backend/.venv activated.
  • SERVICE_JWT set when using Docker Compose.
  • 042 registry and 044 migrations applied.
  • Enabled live bindings configured only through server-owned startup settings.

Available local verification

cd backend
source .venv/bin/activate
python -m pytest -q \
  tests/services/dashboard_testing/registry/test_scenario_*.py \
  tests/api/test_scenario_runs_api.py \
  tests/api/test_scenario_automation_api.py \
  tests/api/test_scenario_analytics_api.py
python -m pytest -q \
  tests/services/dashboard_testing/registry/test_scenario_executors.py \
  tests/services/dashboard_testing/registry/test_live_execution_binding.py \
  tests/services/dashboard_testing/registry/test_scenario_cancel_timeout.py \
  tests/services/dashboard_testing/registry/test_scenario_crash_recovery.py \
  tests/services/dashboard_testing/registry/test_scenario_worker.py \
  tests/services/dashboard_testing/registry/test_scenario_queued_dispatch.py
python -m ruff check src/services/dashboard_testing/execution src/api/routes/dashboard_testing/scenario_runs.py
python -m compileall -q src/services/dashboard_testing/execution src/api/routes/dashboard_testing/scenario_runs.py
cd ..
python3 specs/044-dashboard-scenario-execution/prototype/validate_static.py

Expected current local evidence:

  • full available 044 profile: 246 passed;
  • provider/lifecycle edge profile: 59 passed;
  • prototype static validation: passed;
  • scoped Ruff and compile: passed.

Production verification

cd backend
source .venv/bin/activate
alembic check
alembic upgrade head
python -m pytest -q --run-integration tests/integration/

Run the provider contract profile only after T028-T042 exists:

python -m pytest -q tests/services/dashboard_testing/registry/test_provider_contract.py
python -m pytest -q tests/services/dashboard_testing/registry/test_provider_*.py

Measurable exit gates

  1. SC-001..011 each has a passing named test or deployment evidence record.
  2. 100/100 cancellation trials terminate by cancel_drain_deadline_at + 5 seconds.
  3. 100% of provider I/O attempts have a valid CapacityLease and operation receipt.
  4. 100% of accepted evidence refs have an ownership receipt and verified SHA-256.
  5. 100% of unknown external effects are reconciled or terminalized non-pass before retry.
  6. Every enabled provider has passing liveness, readiness and dependency-health checks.
  7. Browser and Screenshot perform one real authorized deployment run with durable evidence.
  8. No unresolved P0/P1 traceability row remains in 044 or execution-critical dependencies 036, 037, 038, 041, 042, 046 and 047.

Current boundary

The local profile proves fail-closed adapters, exact Superset binding behavior, lifecycle closure and prototype state coverage. It does not prove Browser/Screenshot live composition, shared provider capacity, AgentEvaluation runtime, real PostgreSQL migration validity, scheduler deployment behavior or 047 case ingestion.


TRACEABILITY — Requirements Matrix

Source: traceability.md

Traceability: Scenario Execution Engine (044)

Story Requirement Model API operationId Contract Task Test Actual status / gap
US1 Start SCEX-FR-001/008 ScenarioRun scenarioRun.start Execution.Start, Execution.Runner.QueuedDispatch T006-T008, T024 test_runner, test_scenario_queued_dispatch, test_scenario_scheduler_callbacks, test_scenario_runs_api, test_scenario_automation_api, test_live_execution_binding [~] HTTP/automation start and replay persist queued/pending rows without request-time dispatch; only the scheduler composition's durable queued->running CAS walks its winner. Fixed scheduler callback registration, database-edge containment, and repeat-tick terminal side-effect idempotency are unit-proven. Automated human plans are rejected before they reach CAS/walker; a manual human graph reaches HumanCheckpoint only after that dispatcher claim. Approval-to-real live dispatch remains unproven.
US2 Dispatch SCEX-FR-002/006/009 ScenarioStepRun, LiveExecutionBinding scenarioRun.step Execution.Dispatch, Execution.LiveCompositionRoot T009-T011, T024 test_dispatch, test_scenario_executors, test_live_execution_binding [~] Browser/Superset/Screenshot use fail-safe typed adapter boundaries: no explicit adapter success means no PASS; invalid evidence digest/ref remains inconclusive. Lifespan bootstraps settings.scenario_live_execution_bindings through the existing SupersetClient, exact model and durable storage, so configured Superset dispatch invokes 037 and stores the exact raw-byte digest/ref; mismatched/unavailable providers make no I/O call. Browser safe-checkpoint and Screenshot durable-evidence registration are supported but no provider is deployed, so enabled bindings return stable configured-unavailable codes.
BrowserProvider SCEX-FR-016..023 BrowserProviderActionContract, BrowserOperationReceipt, BrowserEvidenceReceipt scenarioRun.step ScenarioExecution.BrowserProvider T028-T034, T040-T042, T042b test_provider_browser_*, Browser PREPROD canaries [~] Contract completeness is 90/100: all actions, risk classes, limits, lifecycle, checkpoint/recovery, ownership, cancellation/reconciliation, readiness and canary gates are specified. Runtime provider, shared capacity integration, deployment registration and real canary evidence remain open.
US3 Human SCEX-FR-004/010 ScenarioRun(waiting_human) scenarioRun.humanDecision Execution.SuspendForHuman, Execution.Resume T012-T014, T025 test_human_resume, test_scenario_runs_api [~] persisted HumanCheckpoint and infrastructure-resume continuations advance only the missing DAG frontier; completed steps are not re-run. Full live-composition closure remains pending.
Manual-only boundary SCEX-FR-004a ScenarioRun, HumanCheckpoint — Execution.Runner.Start, RunnerPlan.Derive T027 test_scenario_manual_run_only, test_scenario_automation_api [x] Trusted scheduled/deploy/release/ETL/API origins reject persisted human revisions before idempotency or any run/gate/notification/queue side effect. Manual origin remains eligible; HumanCheckpoint is not an approval gate.
Automated human prohibition SCEX-FR-004a/013, SCAUTO-FR-001/007 RunnerPlan.manual_run_only — Execution.Runner.TriggerSource, Automation.Trigger T027, 046 T016/T019 test_scenario_manual_run_only, test_scenario_automation_api, test_scenario_automation_trigger [x] Every trusted non-manual origin, including scheduled, deploy, release, ETL, API and background recovery, is rejected before idempotency/run/gate/notification/queue/dispatch side effects. Only authenticated manual origin may create a human-containing run.
US4 Lifecycle SCEX-FR-005/006 ScenarioRun(status, cancel deadline), ScenarioArtifact(active projection) scenarioRun.cancel, scenarioRun.retry Execution.Lifecycle.Cancel, Execution.Lifecycle.Timeout, Execution.Lifecycle.Retry, Execution.Runner.ContinueAfterRetry T014g, T015-T016, T025 test_scenario_lifecycle, test_scenario_cancel_timeout, test_scenario_retry_closure [~] Strict eligible retry invalidates the persisted target/descendant closure, archives old step attempts, and retains artifact rows only as inactive historical provenance. Cancellation pins a durable bounded drain deadline; deadline finalization expires leases and retires active projections without deleting audit evidence. A timeout during adapter I/O wins over a late PASS and materializes every otherwise-lazy pinned-plan descendant as blocked, so continuation cannot dispatch it; it emits one idempotent inconclusive signal. Full live-I/O composition remains fail-closed/unproven.
US5 Snapshot/API SCEX-FR-003/007 ScenarioExecutionResult scenarioRun.detail, scenarioRun.events Execution.RunnerPlan, Execution.Runner.CrashRecovery T017-T020, T025 test_result, test_api, test_scenario_runner_walker, test_scenario_crash_recovery [~] API/projection and real-digest artifact integrity checks exist. Server-driven crash recovery uses only the persisted run/RunnerPlan and expired lease: completed work is not rerun, safe claims get a new attempt with active evidence retired to history, unsafe claims require reconciliation, and browser recovery without a pinned safe checkpoint is non-pass without adapter I/O. Rejected terminal contexts reuse their idempotent signal. Production live-executor provenance remains unproven.
Terminal signals SCEX-FR-011 InvestigationQueueItem — Execution.Runner.TerminalSignal T026 test_scenario_terminal_signals [x] Failed/blocked/inconclusive terminal runs emit one idempotent immutable 047 queue input with run/artifact provenance; passed runs emit none. Producer-only ingestion starts no case, AgentRun, chat, remediation action, or recurrence classification.
Gate/RBAC SCEX-FR-008 ScenarioRun, ActionApprovalGate scenarioRun.start Execution.EnvironmentPolicy, Execution.Runner.Start T019 test_scenario_runner, test_scenario_runs_api, test_scenario_automation_api, test_scenario_automation_trigger [~] Server ConfigManager classifies every target before persistence: client flags cannot select PROD, unknown targets create no run-side effect, and every trusted source enters the same durable pending_approval gate boundary. HTTP/trigger paths remain persistence-only and dispatcher excludes pending gates. A dedicated real APScheduler scheduled-PROD integration test remains coverage debt; this is not a dispatch bypass.
Provider protocol SCEX-FR-017..023 ProviderExecutionContext, ProviderOperationReceipt, ProviderEvidenceReceipt — ScenarioExecution.ProviderProtocol, ProviderOperations, ProviderOperations.Observability T028-T031, T040-T042 test_provider_contract, test_provider_health, deployment health checks [ ] New production gate: common context/result, operation receipts, ownership proof, cancellation/reconciliation, health/readiness and startup registration are specified but not implemented.
Capacity SCEX-FR-024 CapacityLease, ExecutionCapacityManager — ScenarioExecution.CapacityManager T032, T041-T042 test_provider_capacity, scheduler/worker integration [ ] New production gate: shared environment-scoped capacity admission is specified but current 046 checks do not close it.
Agent evaluation SCEX-FR-025 AgentEvaluation, DecisionPolicy — ScenarioExecution.ProviderCatalog, data-model AgentEvaluation and DecisionPolicy T039, T041-T042 test_provider_agent_evaluation [ ] New production gate: bounded provider, immutable evaluation evidence and deterministic policy mapping are specified but absent.
Success criteria SC-001 RunnerPlan, ScenarioStepRun — ScenarioExecution.RunnerPlan.Derive, ScenarioExecution.Dispatch T004-T011, T041 canonical fixture order and duplicate-dispatch assertions [~] Existing fixture/lifecycle coverage passes; exact 100% canonical dispatch evidence remains a release gate.
Success criteria SC-002 HumanCheckpoint, ScenarioRun scenarioRun.humanDecision ScenarioExecution.HumanCheckpoint, ScenarioExecution.Resume T012-T014, T025 human CAS/frontier tests [x] Human checkpoint CAS and missing-frontier resume are verified; live provider closure remains separate.
Success criteria SC-003 ScenarioRun.cancel_drain_deadline_at scenarioRun.cancel ScenarioExecution.Cancel T015-T016, T025, T041 100 cancellation trials plus scheduler finalizer [~] Bounded drain is unit-proven; required 100-trial production evidence is open.
Success criteria SC-004 RunnerPlan, worker lease, operation receipt scenarioRun.detail ScenarioExecution.Runner.CrashRecovery, ScenarioExecution.ProviderOperations T014c, T025, T030, T041 safe/unsafe recovery and reconciliation tests [~] Safe/unsafe persisted recovery is verified; provider operation reconciliation is not implemented.
Success criteria SC-005 immutable run provenance scenarioRun.result Execution.RunnerPlan, Execution.Result T017, T020, T041 revision-edit immutability and evidence receipt tests [~] Snapshot fields exist; complete receipt-level byte identity evidence is open.
Success criteria SC-006 ActionApprovalGate, ExecutorRegistry scenarioRun.start Execution.ActionApprovalGate, ScenarioExecution.ExecutorRegistry T014d, T019, T041 PROD no-call and human-not-executor tests [x] Current no-call and registry rejection tests pass.
Success criteria SC-007..010 Provider protocol/health/capacity — ScenarioExecution.ProviderProtocol, ProviderOperations.Observability, CapacityManager T028-T042 common/provider-specific contract and deployment profiles [ ] New production gate not implemented.
Success criteria SC-011 release evidence record — ScenarioExecution.ProviderProtocol T022, T041-T042 full profiles, PostgreSQL and semantic audit outputs [ ] Release evidence package is incomplete.

N/A: Registry (042), Editor (043), Monitor UX (045), Automation (046), Analytics (047).

Cross-Spec Pipeline Traceability

Stage / requirement 044 authority Handoff / ownership Evidence
persisted revision is the only launch program input ScenarioExecution.RunPreflight 042 ScenarioRevision -> preflight handle test_scenario_runner, test_scenario_automation_api
server-owned preflight digest and failure boundary ScenarioExecution.RunPreflight blocks before plan, lease, dispatch or provider I/O test_scenario_runner, policy/binding tests
deterministic plan and plan digest ScenarioExecution.RunnerPlan.Derive preflight -> pinned RunnerPlan -> run test_scenario_runner_plan, test_scenario_dispatch
queued/pending approval continuation ScenarioExecution.Start, ScenarioExecution.ActionApprovalGate approved pending run -> queued; denial/expiry -> blocked test_scenario_queued_dispatch, test_scenario_automation_api
ScenarioRun snapshot and idempotency ScenarioExecution.Start / ScenarioRun model exact revision/content/target/principal/binding snapshot test_scenario_runs_api, test_scenario_scheduler_callbacks
evidence ownership and authoritative StepOutcome ScenarioExecution.ProviderProtocol, ScenarioStepRun provider receipt -> Artifact(owner_type=scenario_run) -> StepOutcome test_provider_contract, test_scenario_terminal_signals
terminal signal to 047 ScenarioExecution.TerminalSignal failed/blocked/inconclusive -> idempotent 047 queue input; passed -> none test_scenario_terminal_signals
no reference artifact authority ScenarioExecution.RunnerPlan.Derive runner.plan.json is diagnostic/reference only test_scenario_runner_plan, crash-recovery tests
authoring promotion admission ScenarioExecution.AuthoringAdmission 042 promoted revision -> 044 preflight raw-artifact rejection, candidate-without-promotion, no-run sandbox tests

All rows are required together with the 038/042/050 matrices. Provider, capacity, PostgreSQL, or live-composition gaps keep the coordinated release gate NO-GO, even when lifecycle unit tests pass.

Authoring E2E is additionally NO-GO until persistent workspace exploration, typed proposal conversion, user diff review, 042 handle-based save, and 044 promoted-revision-only admission are evidenced. Sandbox security and code-backed provider readiness are separate gates and remain unimplemented unless explicitly proven.

Provider production boundary: Existing typed unavailable outcomes and exact Superset binding tests prove only the fail-closed boundary. Production readiness additionally requires T028-T042 and real startup/dependency health checks for every enabled live provider.

BrowserProvider boundary: The BrowserProvider is contract-ready at 90/100 but implementation-ready only after T034, T040-T042 and T042b. A typed unavailable result remains the correct behavior until the PREPROD read-only/timeout/cleanup/reconstruction canaries produce retained operation and evidence receipts.

Cross-spec production boundary: 044 readiness depends on 036 authority/evidence, 037 query and baseline provenance, 038 executable graph identity, 041 lineage target state, 042 registry revisions, 046 automation dispatch and 047 terminal-signal/case ingestion. 039/043/045 are operator continuity dependencies: they do not authorize execution, but incomplete typed API/SSE/revision flows prevent a complete production workflow. Current dependency-weighted aggregate for 036-047 is approximately 67/100 and remains NO-GO.

Full production tool boundary: There is no reduced preview target. Browser/Screenshot, controlled non-PROD mutation, AgentEvaluation/DecisionPolicy, automated schedules/triggers, case investigation, analytics and remediation are mandatory capability gates. Missing or unproven capability blocks GO; only human-containing automated revisions remain prohibited and PROD mutation remains policy-forbidden.

Verification boundary (2026-08-24): the available local profile is 246 backend tests, 59 provider/ lifecycle edge tests, scoped Ruff/compile and prototype validation. Real PostgreSQL migration checks, provider contract T028-T042, live Browser/Screenshot composition, scheduler deployment and Axiom index rebuild remain open.


TASKS — Implementation Tasks

Source: tasks.md

#region ScenarioExecution.Tasks [C:3] [TYPE ADR] [SEMANTICS tasks,scenario,execution,implementation] @BRIEF Ordered TDD backlog for the Scenario Execution Engine (044). Tests FIRST for every C3+ contract.

Prerequisites: plan.md, spec.md; contracts/modules.md, traceability.md.

Format: - [ ] T### [P] [USx] Description with exact file path

Factual audit 2026-08-20: [x] means code plus relevant evidence; [~] means partial implementation; [ ] means absent integration or unperformed verification.

Phase 1 — Setup

  • T001 Create ScenarioRun/ScenarioStepRun models + alembic migration in backend/src/models/scenario_run.py
  • T002 [P] Create canonical fixtures under specs/044-dashboard-scenario-execution/fixtures/ (graph, runner plan, run states)
  • T003 [P] Materialize fixtures to backend/tests/fixtures/scenario_execution/

Phase 2 — RunnerPlan (promote stub to contract)

  • T004 [US2] Write failing RunnerPlan derivation/validation tests in backend/tests/services/dashboard_testing/registry/test_scenario_runner_plan.py
  • T005 [US2] Implement derive_runner_plan in backend/src/services/dashboard_testing/execution/runner_plan.py @POST: env targets, resolved params, pinned baselines, topological order, executor mapping @TEST_EDGE: missing plan->reject; unsafe plan->reject

Phase 3 — US1 Start + Executors

  • T006 [US1] Write failing start_run tests in backend/tests/services/dashboard_testing/registry/test_scenario_runner.py @TEST_EDGE: prod_without_approval->blocked; stale_revision->blocked
  • [~] T007 [US1] Implement ScenarioExecutorRegistry + typed dispatch boundary in execution/executor_registry.py
  • [~] T008 [US1] Implement start_run in execution/runner.py @POST: run pinned to immutable revision; queued->running; RunnerPlan deterministically derived

Phase 4 — US2 Deterministic Dispatch

  • T009 [US2] Write failing dispatch_step tests in backend/tests/services/dashboard_testing/registry/test_scenario_dispatch.py
  • [~] T010 [US2] Implement dispatch_step in execution/dispatch.py @INVARIANT: human not dispatched here; deterministic per tool; ref binding @TEST_EDGE: unknown_tool->rejected; step_fail->dependents blocked; ref_binding->dependent reads producer output
  • [~] T011 [US2] Implement dependency readiness + failure propagation in execution/dispatch.py

Phase 5 — US3 Human Suspend/Resume

  • T012 [US3] Write failing suspend/resume contracts in backend/tests/services/dashboard_testing/registry/test_scenario_lifecycle.py
  • T013 [US3] Implement suspend_for_human + checkpoint CAS decision in execution/lifecycle.py @POST: run.waiting_human; step paused; HumanCheckpoint created (confirm/false_positive/inconclusive); resume from token without rerun @REJECTED: 036 ApprovalGate for observation disposition
  • T014 [US3] Add POST /scenario-runs/{id}/human/decision (disposition enum)

Phase 5b — Worker semantics + action gate (P0 #3/#6)

  • T014b [P] Write worker claim/lease/idempotency contracts in backend/tests/services/dashboard_testing/registry/test_scenario_worker.py
  • T014c [P] Implement claim_step (lease/heartbeat/idempotency) in execution/worker.py @INVARIANT: at-least-once; persisted lease/idempotency checks are proven; an external side effect is not re-run unless idempotent/retry-safe.
  • T014d [P] Implement require_action_gate (PROD authorization) + duplicate Idempotency-Key rejection in execution/runner.py

Phase 5c — RunnerPlan derivation + artifacts (P0 #2/#4)

  • [~] T014e [P] Implement derive_runner_plan (deterministic derivation from revision) in execution/runner_plan.py @INVARIANT: runner.plan.json is reference, never runtime source of truth; revision mismatch rejected
  • [~] T014f [P] Generalize artifact owner (owner_type=scenario_run) for evidence/screenshot/report in execution/artifacts.py

Phase 5d — Retry closure + result aggregation (P0 #14/#16)

  • T014g [P] Implement retry downstream-closure invalidation in execution/lifecycle.py @INVARIANT: strict eligible retry retires the active target/descendant closure, archives prior step attempts, and preserves artifact rows as inactive historical provenance. A reterminalized attempt has a new immutable terminal signal; replay of the same terminal context is idempotent.
  • [~] T014h [P] Implement result aggregation truth table in execution/result.py

Phase 6 — US4 Cancel/Retry/Timeout

  • T015 [US4] Write failing cancel/retry contracts in backend/tests/services/dashboard_testing/registry/test_scenario_lifecycle.py
  • T016 [US4] Implement cancel_run + retry_step with downstream closure invalidation in execution/lifecycle.py @POST: cancel->cancelled with bounded drain; retry increments attempt; timeout->failed/inconclusive

Phase 7 — US5 Snapshot + API + PROD gate

  • [~] T017 [US5] Implement immutable execution snapshot + provenance in execution/result.py
  • [~] T018 [US5] Add POST /scenario-runs (Idempotency-Key), cancel, resume, GET /{id}, GET /{id}/events (SSE), GET /scenarios/{id}/runs, GET /scenario-runs/{id}/result, GET /scenario-runs/compare, POST /scenario-runs/{id}/steps/{logical_step_id}/retry in api/routes/dashboard_testing/scenario_runs.py
  • [~] T019 [US5] PROD approval gate (036) + RBAC scenario:run / scenario:run:prod tests. Server ConfigManager policy now ignores client is_prod compatibility flags, fails closed for unknown environments, and gives manual/API/event/scheduled origins the same durable pending_approval gate before dispatcher CAS. A dedicated real APScheduler scheduled-PROD integration test remains coverage debt; it does not authorize dispatch around the gate.
  • T020 [P] Frontend DTOs (ScenarioRun, ScenarioStepRun, ScenarioExecutionResult) in frontend/src/types/scenario-run.ts

Phase 8 — Polish

  • T021 [P] Belief-runtime instrumentation for C5 runner/dispatch
  • T022 Run quickstart-equivalent full scenario backend scope, scoped Ruff, ATTN/orphan static audit, real PostgreSQL alembic check/upgrade head, and semantic rebuild. Completion requires command output attached to the release record and zero unresolved P0/P1 rows.
  • T023 Prototype validation: every declared @UX_STATE is reachable via prototype/index.html. Proof: python specs/044-dashboard-scenario-execution/prototype/validate_static.py.

Phase 9 — Provider Production Contract

  • T028 [P] [US2] Implement common ProviderExecutionContext, operation receipts and typed ProviderExecutionResult schemas in backend/src/services/dashboard_testing/execution/provider_protocol.py.
  • T029 [P] [US2] Add provider ownership receipts and atomic evidence commit contract in backend/src/services/dashboard_testing/execution/provider_evidence.py.
  • T030 [P] [US2] Add provider operation lifecycle, cancellation and reconciliation contracts in backend/src/services/dashboard_testing/execution/provider_operations.py.
  • T031 [P] [US2] Add provider liveness/readiness/dependency health and redacted telemetry contract in backend/src/services/dashboard_testing/execution/provider_health.py.
  • T032 [US2] Integrate atomic shared ExecutionCapacityManager admission, heartbeat, expiry, release and reconciliation with dispatcher claims in backend/src/services/dashboard_testing/execution/capacity.py. @INVARIANT: no provider I/O without a capacity lease; retries claim a new lease. Include the ProviderRuntime contract (provider_runtime.py): one application-owned long-lived event-loop thread for async transports, bounded submissions, and the pinned concurrency defaults (2 browser contexts DEV/PREPROD, 1 PROD; screenshot capture on the same lease accounting).
  • T033 [P] [US2] Add per-provider input/output schemas, limits and error taxonomy in backend/src/services/dashboard_testing/execution/provider_contracts.py for browser, Superset API, SQL evidence, XLSX, assertion, transform, screenshot, report and artifact. Include bounded verification response bytes/rows/cells/canonicalization time: oversized output is RESULT_TOO_LARGE + inconclusive and can never be truncated into PASS or a baseline update.
  • T034 [US2] Implement BrowserProvider resource ownership, safe-checkpoint replay, cancel and reconciliation adapter in backend/src/services/dashboard_testing/execution/providers/browser.py. Required action catalog: open_dashboard, navigate_tab, apply_native_filter, inspect_filter_state, apply_table_filter, extract_table, scroll_to, inspect_columns, click, select_rows, edit_row, bulk_edit, download, refresh, wait_for_state. Required limits: 120s context/auth, 30s action, 3 pages, 25 MiB downloads, 10 MiB screenshots; mutation actions require fixture lease and cleanup. Context/session objects are bound to the shared provider event loop and stay live across all steps of one run; no per-step loop creation and no warm context pool.
  • T035 [US2] Implement ScreenshotProvider atomic durable evidence adapter with cleanup and ownership receipts in backend/src/services/dashboard_testing/execution/providers/screenshot.py, wrapping the existing llm_analysis Playwright capture stack; capture runs on the shared provider event loop under the same capacity admission; trace/video output is never registered as evidence.
  • T036 [US2] Harden SupersetProvider and SqlEvidenceProvider request limits, error taxonomy, external request reconciliation and principal/RLS evidence in backend/src/services/dashboard_testing/execution/providers/superset.py.
  • T037 [P] [US2] Implement server-owned XLSX artifact intake and resource limits in backend/src/services/dashboard_testing/execution/providers/xlsx.py.
  • T038 [P] [US2] Implement deterministic ReportProvider and ArtifactProvider durable registration with manifest/digest ownership in backend/src/services/dashboard_testing/execution/providers/artifacts.py.
  • T039 [US2] Implement AgentEvaluationProvider and deterministic DecisionPolicy integration in backend/src/services/dashboard_testing/execution/providers/agent_evaluation.py. @INVARIANT: model verdict never directly sets ScenarioResult or consumes HumanCheckpoint.
  • T040 [US2] Add startup deployment registration and readiness preflight for all provider capabilities in backend/src/services/dashboard_testing/execution/provider_bootstrap.py.
  • T041 [US2] Add common provider contract tests for unavailable/dependency failure/capacity exhaustion/timeout/cancel/duplicate/late response/ownership mismatch/malformed result/cleanup and reconciliation in backend/tests/services/dashboard_testing/registry/test_provider_contract.py.
  • T042 [US2] Add provider-specific contract tests in backend/tests/services/dashboard_testing/registry/test_provider_*.py and real deployment health checks for Superset, Browser and Screenshot bindings.
  • T042c [US2] Add dispatcher policy/binding revalidation tests in backend/tests/services/dashboard_testing/registry/test_provider_contract.py: reclassification to PROD or a changed provider/security fingerprint after approval prevents all provider I/O and returns typed POLICY_CHANGED or BINDING_CHANGED.
  • T042b [US2] Run BrowserProvider PREPROD canaries and retain evidence in specs/044-dashboard-scenario-execution/evidence/browser-provider/: read-only action canary, forced timeout/cleanup canary, safe-checkpoint reconstruction trace and readiness/health payload. GO requires all BrowserProvider acceptance vectors to pass and one owned evidence receipt. Contract target before runtime enablement: 90/100; runtime score remains below target until T034, T040-T042 and this canary task are complete.

Requirement evidence rules

  • A task claiming provider PASS MUST assert a valid CapacityLease, operation receipt and ownership receipt; mocked callback success alone is insufficient.
  • A cancellation/retry task MUST include unknown external effect, late response and duplicate request cases, with the expected durable state asserted after the scheduler interval.
  • A deployment task MUST record provider/version, capability fingerprint, readiness result and redacted dependency diagnostics for every enabled binding.
  • A cross-spec task is complete only when its named dependency row in traceability.md is [x]; local unit tests cannot close an unresolved 036/037/041/042/046/047 runtime boundary.

Audit Follow-ups (2026-08-20)

  • [~] T024 Replace synthetic default executor outcomes with typed adapter boundaries for required 037/038/036 and existing browser/XLSX infrastructure; add an import/compile test for the executor module. Assertion uses 037 compare_values; browser/Superset/Screenshot never synthesize PASS without explicit typed adapter success, and invalid evidence digest/ref is inconclusive. The Superset binding slice persists an immutable identity snapshot and invokes the existing 037 query envelope only after an exact composition-owned resolver match, storing the exact raw-byte digest/ref. This is fail-safe partial closure, not proof of real live composition.
  • [~] T025 Add falsifiable DAG tests for actual executor output, failures, descendants, artifact writes, checkpoint resume, timeout/cancel drain, recovery and approval-to-dispatch. Failed assertion now blocks descendants and xlsx registers an artifact (test_failed_assertion_blocks_descendants_and_xlsx_registers_artifact). Cancellation now persists a one-time drain deadline: queued work is skipped, a claimed adapter may finish only during the window, and an expired sweep retires leases/current projections while preserving immutable evidence history. A timeout during claimed adapter I/O wins over a late PASS, materializes every otherwise-lazy declared descendant as a durable blocked row from the pinned RunnerPlan, and emits one inconclusive signal (proof: test_scenario_cancel_timeout.py). Human and infrastructure resume advance only the missing DAG frontier, so completed steps are not re-run. Server-driven crash recovery loads only the persisted run/RunnerPlan and expired lease: safe frontier attempts are archived/retried once, unsafe effects require reconciliation, and browser recovery requires a pinned safe checkpoint (test_scenario_crash_recovery.py). Retired evidence remains historical and rejected terminal contexts reuse their idempotent signal. Persisted artifact evidence needs a real valid digest/ref. Approval-to-live-dispatch still needs an authorized production composition root that supplies resolver/client/model/storage; Browser/Screenshot bindings remain unavailable and fail closed. Queued HTTP/automation starts remain non-dispatching; the existing scheduler claims eligible rows through durable queued->running CAS before walking them, so repeated ticks do not repeat a side effect. A manual human run reaches its checkpoint only after that claim, while an automated human plan is rejected before a row reaches CAS/walker (test_scenario_queued_dispatch.py, test_scenario_manual_run_only.py). Fixed queue/cancel callbacks are unit-proven to retain their exact five-second singleton/ coalescing registration, contain database-edge errors, and preserve terminal side effects on repeated ticks (test_scenario_scheduler_callbacks.py); this does not prove a live scheduler process, browser composition, or T022 closure.
  • T026 [P] Emit failed/blocked/inconclusive terminal ScenarioRuns as one idempotent 047 queue producer signal with immutable run/artifact provenance; passed runs emit none. The producer never opens a case, AgentRun, chat, remediation action, or recurrence classification. Proof: test_scenario_terminal_signals.py.
  • T027 [P] Derive manual_run_only from a persisted human graph and reject every trusted 046 automation origin before idempotency or run/gate/notification/queue/dispatch side effects. The rejected key remains valid for a manual start; HumanCheckpoint is not ActionApprovalGate. Proof: test_scenario_manual_run_only.py, test_scenario_automation_api.py.

Verified profile (2026-08-20): the targeted service suite independently reverified 35 passes. The exact API file has a 29-pass result outside this sandbox; here FastAPI TestClient is blocked by AnyIO self-pipe EPERM, which is a sandbox limitation rather than an application failure.

Dependencies

Setup → RunnerPlan; US1 (start+executors) → US2 (dispatch); US3 (human) depends on US2; US4 (lifecycle) depends on US2; US5 (snapshot/API/gate) depends on US3/4. 045 monitor consumes run/step results; 042 provides persisted scenario+revision.

#endregion ScenarioExecution.Tasks


PROTOTYPE — State/Manifest

Source: prototype/manifest.md

#region ScenarioExecution.PrototypeManifest [C:3] [TYPE ADR] [SEMANTICS prototype,manifest,scenario,execution] @defgroup Prototype Interactive HTML prototype manifest for Scenario Execution Engine (engine-level states). @RELATION DEPENDS_ON -> [ScenarioExecution.PrototypeRun] @RELATION DEPENDS_ON -> [Test.ScenarioExecution.PrototypeStatic] @RATIONALE The manifest distinguishes declared source-level reachability from browser, assistive-technology, and production-runtime evidence so the prototype cannot overstate its proof. @REJECTED Treating static source checks as browser validation or production execution evidence was rejected — neither rendering nor runtime integrations execute in this verifier.

Prototype Metadata

Factual audit 2026-08-21: validate_static.py proves source-level lifecycle names, persistent live-region attributes, native state controls, 44px sizing and reduced-motion declarations. It does not prove browser rendering, keyboard operation, assistive-technology announcements, runner execution, executor outcomes, artifacts or recovery in production code.

  • Feature: 044 Scenario Execution Engine
  • Source contracts: ux_reference.md, contracts/modules.md
  • Screens represented: 1 (Scenario Run engine view)
  • Total states: 6 (running, waiting_human, resumed, cancelled, failed, passed)
  • Accessibility (static): native buttons, persistent role=status/polite atomic live region, ≥44px target declarations, prefers-reduced-motion override
  • Responsive: 375px, 900px

State Coverage

@UX_STATE Prototype State Reachable? Recovery
running running ✅ statebar —
waiting_human waiting_human ✅ initial + statebar confirm/false-positive/inconclusive
resumed resumed ✅ statebar/confirm continues from resume token
cancelled cancelled ✅ statebar/stop —
failed failed ✅ statebar/non-pass decision retry / triage (047)
passed passed ✅ statebar —

Screen ↔ Story Traceability

Story Prototype Feature Intended acceptance coverage
US1 Start run header + snapshot pinned revision
US2 Dispatch tool→executor row, step statuses deterministic dispatch
US3 Human WAITING_FOR_HUMAN box + resume suspend/resume
US4 Lifecycle cancel/stop cancel
US5 Snapshot snapshot row immutable revision + provenance
#endregion ScenarioExecution.PrototypeManifest

PROTOTYPE — Interactive HTML

Source: prototype/index.html

<!doctype html>

<html lang="ru"> <head> </head>
Superset Tools · BI testing
СценарииЗапускиАвтоматизацияКачество Ручной запуск
Scenario run SR-1842 · revision r18

XLSX reconciliation waiting_human

PREPROD · Target: rc-17 · Анна · параметры сохранены в snapshot

Остановить запуск

Ход проверки 4 из 6

1
Открыть dashboard
BrowserExecutor · 6 сек
Готово
2
Применить filters
Контрагент = ACME · дата = 10.08.2026
Готово
3
Скачать XLSX
XlsxExecutor · file: fi-0080.xlsx
Готово
4
Сравнить с baseline
AssertionExecutor · 2 различия обнаружены
Проверка
5
Ручная проверка evidence
Manual assertion · доступна только в ручном запуске
Ожидает вас
6
Сформировать отчёт
Начнётся после решения
Ожидает
State: Running Waiting humanResumedFailedPassed Cancelled
<script src="../../prototype-ui.js"></script> <script> let checkpointDecision = null; const stateLabels = { "running": "running · выполняется", "waiting_human": "waiting_human · ожидает проверки", "resumed": "resumed · зависимые шаги продолжаются", "cancelled": "cancelled · запуск остановлен", "failed": "failed · non-pass результат", "passed": "passed · завершён успешно", }; const decisionLabels = { confirm: "Решение confirm сохранено.", false_positive: "Решение false_positive сохранено; PASS не заявлен.", inconclusive: "Решение inconclusive сохранено; PASS не заявлен.", }; function decide(v) { checkpointDecision = v; document.getElementById("checkpoint").innerHTML = 'Решение сохранено

' + { confirm: "Соответствует", false_positive: "Не соответствует", inconclusive: "Недостаточно данных", }[v] + '

Runner продолжает только зависимые шаги. Решение, evidence и автор зафиксированы.

'; setProtoState(v === "confirm" ? "resumed" : "failed"); } protoState("waiting_human", (s) => { let b = document.getElementById("run-status"), c = document.getElementById("checkpoint"), announcement = document.getElementById("run-announcement"), detail = document.getElementById("state-detail"); document.documentElement.dataset.protoState = s; b.className = "badge " + (s === "passed" ? "ok" : s === "cancelled" || s === "failed" ? "danger" : s === "waiting_human" ? "warn" : "info"); b.textContent = s; c.classList.toggle("hidden", s !== "waiting_human"); detail.textContent = stateLabels[s]; announcement.textContent = checkpointDecision ? stateLabels[s] + ". " + decisionLabels[checkpointDecision] : stateLabels[s]; document.querySelectorAll("[data-proto-state]").forEach((control) => { control.setAttribute("aria-pressed", String(control.dataset.protoState === s)); }); }); </script> </html>

SESSION_STATE.md

Source: SESSION_STATE.md

044 Scenario Execution — Session State

Updated: 2026-08-21 17:12 +03:00 Purpose: Durable handoff for the current implementation/review session. This is a decision and verification ledger; tasks.md and traceability.md remain the canonical feature backlog and requirement matrix.

Current Objective

Bring the 044 Scenario Execution Engine into material conformance with its contracts while preserving fail-closed live I/O and GRACE-Poly invariants. Every code change is followed by an independent verifier pass and semantic curation.

Readiness Assessment

Production readiness is weighted toward real external execution because local fail-closed tests cannot prove that a configured Browser/Superset/Screenshot/DB path performs authorized I/O and produces durable evidence.

Category Weight Score Weighted contribution Evidence / limiting factor
Real live composition and external evidence 50% 25/100 12.5 Exact injected 037 Superset binding is proven; Browser/Screenshot providers remain unavailable; no real deployment run
DAG, RunnerPlan and executor behavior 15% 87/100 13.05 Pinned descriptors, DAG closure, typed executors, read-only sql_evidence and bounded transform pass; required browser/screenshot integrations remain partial
Security and fail-closed policy 15% 86/100 12.9 Server-owned PROD gate, CAS decisions, descriptor allowlist and no-I/O rejection paths are covered
Retry, cancel, timeout and crash recovery 10% 84/100 8.4 Persisted lifecycle and terminal cleanup profiles pass
Deployment, PostgreSQL migrations and operations 5% 62/100 3.1 One Alembic head; real PostgreSQL upgrade/check not run because environment has placeholder DATABASE_URL
Tests and independent verification 5% 82/100 4.1 242 backend 044 tests, scoped Ruff/compile pass, prototype validation passes; semantic rebuild unavailable
Weighted total 100% 54.05/100 NO-GO

Hard production gate: live composition must score at least 70/100. The current live score is 25/100, so the feature remains NO-GO for full production, regardless of the 53/100 weighted total. It is suitable only for internal fail-safe preview/shadow mode with unavailable providers explicitly surfaced as non-pass.

Architecture Decisions Already Implemented

  • HTTP and 046 start paths are persistence-only. The scheduler-owned queued dispatcher is the sole initial execution authority, claiming queued -> running with a durable CAS.
  • Environment classification is server-owned (ConfigManager), not request-owned. A configured PROD run always becomes pending_approval with an ActionApprovalGate; an unknown environment fails before any run/gate/notification/queue side effect.
  • LiveExecutionBinding persists only immutable identity/fingerprint data. Application startup builds LiveExecutionCompositionRoot from trusted configuration. Exact Superset bindings may call the existing 037 query envelope; unavailable/mismatched browser, Superset, or screenshot bindings are typed non-pass and perform no I/O.
  • Artifact evidence needs a non-zero 64-hex digest and durable matching ref. Historical evidence is retained for audit but retired from current result/signal projections on retry/cancel/recovery.
  • Retry, timeout, cancel and crash recovery operate only on persisted run state and the pinned RunnerPlan. Unsafe effects require reconciliation; browser recovery needs a pinned safe checkpoint.
  • Failed/blocked/inconclusive terminal results publish one immutable, idempotent 047 queue input; they do not create a Case, AgentRun, chat or remediation action.

Independently Verified Closures

  • Server-owned PROD gate: external 044/046 verification profile passed 76 tests.
  • Queued dispatcher API + scheduler profile: 94 tests passed; scheduler callback profile: 124 tests passed. FastAPI TestClient tests run outside this sandbox because AnyIO self-pipe gets EPERM inside it.
  • Live composition bootstrap: 57 tests passed; exact configured Superset binding produces the expected durable draft:{run}:{sha256} ref/digest. Browser/Screenshot providers still default to typed unavailable.
  • Lifecycle closure: retry/cancel/timeout/recovery profiles were independently verified. Timeout materializes otherwise-lazy declared descendants as durable blocked rows.

Orthogonal Review Findings

The independent review found these material gaps. The first is fixed; the second is active.

  1. P0 — fixed: Client-controlled is_prod could bypass ActionApprovalGate. The new server-owned environment policy is verified and semantically curated.
  2. P0 — fixed (2026-08-21): RunnerPlan persists version/hash-pinned exact {tool, action} ActionExecutionDescriptor snapshots. Tool-only fallback and universal non-human retry-safe metadata are rejected; malformed legacy plans terminalize before I/O, and server-owned PROD blocks browser mutation before provider invocation.
  3. P1 — fixed (2026-08-21): HumanCheckpoint and ActionApprovalGate decisions now consume pending state through atomic SQL compare-and-set predicates. Stale/concurrent decisions fail after exactly one winner; existing cancellation expiry remains terminal and non-reactivating.
  4. P1 — fixed (2026-08-21): Queued-dispatch exceptions now terminalize through a shared cleanup path that retires current evidence projections, marks active steps inconclusive, expires leases, and emits the existing idempotent terminal side effects.
  5. P1 — partial: sql_evidence now has a version-pinned read-only action descriptor and exact 037 binding adapter with durable raw-response evidence. Bounded transform, AgentEvaluation/DecisionPolicy and global capacity management remain absent despite the complete 044 target contract. Bounded transform now has a version-pinned transform_result action with only select_fields, sort_rows, and limit_rows, hard row/operation limits and canonical normalization; arbitrary expressions/code remain rejected. Existing 046 candidate capacity/dedup checks do not constitute the 044 dispatcher capacity boundary.
  6. P2 — pending: Quickstart paths/counts and a stale automation migration-head test must be aligned with the current linear Alembic chain.

Invariants That Must Not Regress

  • No client payload controls PROD classification, authority, RLS, action identity or retry safety.
  • No external adapter returns PASS from metadata, identifiers, missing binding, invalid status, missing digest or mismatched evidence ref.
  • HumanCheckpoint is never an executor or an ActionApprovalGate.
  • Only an ActionRegistry-pinned descriptor may reach executor I/O; unknown action must fail before lease/I/O.
  • A terminal run has no active lease/current evidence projection/running step.
  • Repeated requests, scheduler ticks, retries and terminal projections are idempotent within their immutable context.

Verification Boundaries / Deployment Debt

  • Browser-safe action provider and ScreenshotService-to-durable-evidence provider are not yet registered in deployment; both correctly fail closed.
  • A real PostgreSQL alembic upgrade head remains required. SQLite cannot run an earlier unrelated migration using drop_constraint.
  • Axiom MCP is unavailable in this session, so semantic index rebuild is not claimed.
  • Legacy ScreenshotService is intentionally not registered as a 044 provider: its output paths and environment credentials do not prove principal/RLS-bound durable evidence. The provider registration hook remains startup-owned and fail-closed until an explicit lawful adapter exists.

Active Work Contract

Task: Replace tool-only dispatch with version-pinned {tool, action} descriptors from the 038 ActionRegistry; make descriptor data the only source of executor policy, side-effect key, idempotency, retry safety, timeout and mutation guard.

Completion evidence required: preflight and legacy-plan malformed cases, safe vs unsafe recovery/retry, server-owned PROD mutation no-call, queued error closure, targeted 038/044 tests, linear Alembic head, lint/compile/diff checks, independent verifier result, then semantic curation.

Latest Verification

  • Targeted lifecycle/approval/queued-dispatch/scheduler profile: 32 passed.
  • RunnerPlan, worker, retry, terminal-signal and timeout profile: 19 passed.
  • python -m compileall and git diff --check: passed.
  • Retry API now explicitly rejects legacy unsafe tool-only plans with 409 RETRY_CONFLICT; the stale API expectation was aligned with the pinned-descriptor contract.
  • recover_run was decomposed into a bounded running-step recovery helper; scoped Ruff for execution and scenario-run API modules passes.
  • Full available 044 backend profile: 241 passed (test_scenario_*.py registry suite plus scenario run/automation/analytics API suites).
  • Live binding and recovery profile: 40 passed; frontend profile previously verified 227 files / 3930 tests; prototype static validation now passes.
  • alembic heads remains a single linear head (e3f4a5b6c7d8), but alembic check and real upgrade are blocked in this environment because DATABASE_URL is the placeholder __MUST_SET_DATABASE_URL__; Docker Compose is also blocked by missing SERVICE_JWT.
  • Screenshot evidence validation now accepts one or more provider-issued refs only when every ref has a valid non-zero SHA-256 and the first ref agrees with the declared sha256; focused live executor/composition profile: 36 passed.
  • Revalidated after commit ffa4d6a8: full available 044 backend profile is 242 passed; scoped Ruff and targeted compile pass; validate_static.py passes; alembic heads reports the single head e3f4a5b6c7d8.
  • The subsequent auth/UI commit 1a5c1473 and current unrelated frontend working-tree changes are excluded from this 044 readiness score.
  • Added sql_evidence as a registered capture_sql_evidence action. It reuses only the exact server-owned 037 binding and persists raw response evidence; caller SQL, static result metadata, missing digest, or missing durable ref cannot pass. Full available 044 profile is now 244 passed after updating the registry fingerprint fixture.
  • Added bounded transform_result action and executor. It applies only declarative field selection, deterministic string sorting and row limiting to completed table results, then canonicalizes via the existing 037 normalizer. No eval, exec, arbitrary imports, SQL or unbounded output path exists. Full available 044 profile is now 246 passed; scoped Ruff and compile pass.

Contract Hardening Update (2026-08-21)

The specification was expanded from a fail-closed adapter boundary into a production provider contract. The new normative requirements are SCEX-FR-017..025 and tasks T028..T042.

  • Every provider now has a common execution context/result, immutable operation receipt, effect state, ownership/evidence receipt, cancellation and reconciliation protocol.
  • Provider-specific obligations are explicit for Browser, Superset API, SQL evidence, XLSX, Assertion, Transform, Screenshot, Report, Artifact and AgentEvaluation.
  • Provider lifecycle now includes resource ownership, server-owned limits, capacity leases, cleanup, liveness/readiness/dependency health, startup registration and redacted observability.
  • AgentEvaluation is bounded by immutable 038 spec and deterministic DecisionPolicy; it cannot directly set StepOutcome, mutate the graph, consume HumanCheckpoint or schedule work.
  • ExecutionCapacityManager is a shared environment-scoped admission boundary; no provider I/O is lawful without an atomic lease.

This hardening is specification-only. T028-T042 remain unimplemented, so production status remains NO-GO. Existing typed-unavailable and exact Superset tests prove only the fail-closed boundary, not the new provider production gate. The Axiom semantic index was not rebuilt in this session.

Orthogonal Edge-Case Verification (2026-08-21 18:13 +03:00)

  • Full available 044 registry/API profile: 246 passed.
  • Provider and lifecycle edge profile (executors, live binding, cancel/timeout, crash recovery, worker, queued dispatch): 59 passed.
  • Full registry scenario edge profile: 186 passed.
  • Scoped Ruff and compile checks for execution/API modules: passed.
  • Prototype state validation: passed.
  • The first edge command referenced a nonexistent test_scenario_live_execution_binding.py; it was corrected to the existing test_live_execution_binding.py before the successful 59-test run.
  • No new provider protocol, capacity, AgentEvaluation, operation receipt, health or deployment tests exist yet; therefore these results do not close T028-T042.

Cross-Spec Production Readiness: 036-047 (2026-08-21)

044 production readiness is not a single-feature property. The execution chain is:

036 authority/evidence/gates -> 037 query/baseline -> 038 executable scenario model -> 042 registry -> 043 editor -> 044 runner/providers -> 046 automation -> 047 triage/analytics -> 045 monitor, with 039 scenario authoring UI consuming 036/037/038 and feeding 042/043.

The following scores are categorical implementation-readiness estimates, not claims of completed verification. A score reflects contract completeness, runtime closure, external integration evidence and operator/deployment readiness. A neighboring spec may score well in isolation and still block 044 when its missing boundary is on the execution critical path.

Spec Domain Contract Runtime External/deploy Operational Readiness Effect on 044
036 Agent runs, evidence, gates 90 82 70 80 81/100 Strong substrate; 044 still needs provider-specific ownership receipts and live deployment proof
037 Superset baseline/query engine 88 82 68 78 79/100 Exact query envelope is reusable; migration/API and real PostgreSQL/live checks remain limiting
038 Scenario graph/compiler/model 92 88 75 84 85/100 Executable graph and pinned specs are a strong source of truth; AgentEvaluation runtime remains open in 044
039 Scenario authoring UI 78 65 45 58 62/100 Non-blocking for backend execution, but incomplete API binding weakens authoring-to-run continuity
040 Load testing/capacity patterns 88 78 65 76 77/100 Useful capacity precedent, but not a substitute for shared 044 CapacityManager
041 Dataset lineage/blast radius 88 76 60 72 74/100 Required for target/staleness/blast-radius provenance; fan-out and production sync evidence remain material
042 Scenario registry/revisions 86 72 55 68 70/100 Direct blocker: runtime 037/041 staleness and 036 signal integration remain open
043 Scenario editor 82 68 50 62 66/100 Not required for dispatcher execution, but immutable revision production flow depends on 042 and open E2E/policy checks
044 Scenario execution/providers 86 58 25 62 54/100 Current hard NO-GO; provider contract T028-T042 and live composition are open
045 Run monitor/results UX 78 65 40 58 61/100 Operator-facing production blocker: typed launch alignment and 047 handoff are incomplete
046 Scenario automation/operations 82 58 40 55 59/100 Direct production blocker for scheduled/triggered runs; startup reload, event dispatch and persisted scheduler semantics are open
047 Investigation queue/analytics 82 60 35 58 59/100 Direct terminal-flow blocker: canonical signal ingestion, immutable case evidence and conformant closure are open

Dependency-weighted aggregate

For a full 036-047 production surface, the practical aggregate is approximately 67/100. The execution-critical subset {036,037,038,041,042,044,046,047} is approximately 68/100. Both remain below a production GO because 044 live composition, 046 automation, and 047 terminal ingestion are hard gates rather than optional UX debt.

Category impact on the 044 production score

Cross-spec category Current contribution to 044 Required closure Expected 044 effect
Authority, RBAC, approval and evidence substrate (036) 80/100 Preserve 036 gate/CAS semantics and bind provider receipts to 036 evidence +2 to +4 weighted
Superset query, normalization and baseline provenance (037) 72/100 Real PostgreSQL/live query verification; exact binding remains mandatory +3 to +5
Executable graph and immutable program identity (038) 82/100 Close AgentEvaluationSpec/DecisionPolicy runtime boundary +2 to +4
Target lineage and staleness (041) 55/100 Production sync/fan-out and pinned blast-radius snapshot +3 to +6
Registry/revision source of truth (042) 58/100 Staleness subscription, health signal and final independent verification +4 to +7
Operator authoring/editor continuity (039/043) 45/100 API binding, policy/E2E and revision save-to-run proof +1 to +3; not a live-I/O gate
Monitor and human operations (045) 48/100 Typed 044 launch/SSE contract and real evidence/047 handoff +2 to +4
Automation and scheduling (046) 42/100 Event dispatch, startup schedule reload, persisted APScheduler behavior and shared capacity +5 to +9
Triage, cases and analytics (047) 35/100 Canonical signal ingestion, immutable evidence, case CAS and closure +3 to +6
Provider runtime contracts (new T028-T042) 25/100 Receipts, ownership, cancellation/reconciliation, health, deployment, capacity +15 to +25

Hard gates for full production GO

The following are conjunctive, not averaged away:

  1. 044 Browser and Screenshot providers perform authorized live I/O and produce durable owned evidence.
  2. 044 shared CapacityManager admits every provider operation before I/O and survives retry/crash paths.
  3. 046 persisted schedules/triggers execute through 044 queued CAS with startup reload and real scheduler verification.
  4. 047 ingests one immutable terminal signal, preserves run/evidence provenance, and opens/closes cases through CAS without changing run truth.
  5. 042 registry revision/staleness state is production-connected to 037/041 and emits canonical 036 signals.
  6. 036/037/041 PostgreSQL migration and deployment checks pass against a real PostgreSQL environment.
  7. 045/039/043 operator paths use typed contracts end-to-end; no agent prose or compatibility payload substitutes for API truth.

Interpretation

Implementing only the new 044 provider contracts would likely raise 044 from 54/100 to 80-88/100, but the complete 036-047 product would remain around 67/100 until 042, 046 and 047 close their cross-spec runtime boundaries. Conversely, completing 045/039 UI polish without 046/047 and live 044 composition would not materially change the production decision.

Requirements Clarification Review (2026-08-24)

The documentation package was reviewed for ambiguous, non-measurable and cross-document requirements. The following consistency changes were applied:

  • SC-001..011 now define exact thresholds for dispatch order, checkpoint uniqueness, 100 cancellation trials, recovery identity, provenance immutability, provider readiness and release evidence.
  • SCEX-FR-017..025 now have explicit MUST semantics, named evidence expectations and measurable ownership, capacity, cancellation, health and deployment rules.
  • OpenAPI now requires core run/step fields, strict step status enum, typed execution provenance and evidence receipts; cancel responses distinguish 202 request, 200 already-finalized and 409 invalid state.
  • Quickstart commands now match repository paths and current verified counts: 246 full 044 tests and 59 provider/lifecycle edge tests. PostgreSQL and provider contract profiles are explicitly separate release gates.
  • Checklist expanded from FR-001..010 to FR-001..025 and SC-001..011, including per-provider and cross- spec production gates.
  • UX reference defines observable deadlines, event replay, capacity-unavailable, reconciliation-required, ownership failure and redaction behavior.
  • Plan, data model, tasks and traceability now use the same 5-second scheduler tolerance, 64-hex digest, one-receipt-per-operation and zero-open-P0/P1 release rules.

Verification after documentation update: OpenAPI YAML parsed, GRACE anchors balanced, git diff --check passed, and the available 044 regression profile passed 246 tests. This remains documentation clarification, not implementation of T028-T042 or a production GO.

Automated Human Checkpoint Clarification (2026-08-24)

Confirmed invariant: a HumanCheckpoint MUST NOT exist in an automated scheduled or triggered run. manual_run_only=true is rejected before idempotency lookup and before creation of any ScenarioRun, ActionApprovalGate, notification, queue item or dispatcher claim for scheduled, deploy, release, ETL, API and background-recovery origins. Only the authenticated manual route may create a human-containing run and later reach waiting_human.

This is covered by the existing manual-only tests and is now stated identically in 044 spec.md, data-model.md, 046 spec.md, 046 quickstart/checklist/traceability and the 044 cross-spec traceability. PROD ActionApprovalGate remains valid only for eligible automated revisions without human checkpoints; it never substitutes for or bypasses HumanCheckpoint.

Production Scope Assessment (2026-08-24)

This assessment distinguishes the full feature roadmap from a production-usable scope. The full 036-047 surface remains approximately 67/100 and NO-GO. A narrower production rollout is possible only as an explicitly constrained release profile; it must not be presented as the complete automated scenario product.

Orthogonal scope categories

Category Included capability Current readiness Production scope decision
Scenario authoring Agent creates typed 038 graph/draft; validator rejects unsafe graph; analyst reviews 78/100 Include after revision-save E2E; agent cannot activate or execute directly
Registry and revision identity 042 immutable revision/current activation and run pinning 70/100 Include only with stale-state policy and activation CAS; no git-file-only truth
Read-only manual execution Manual analyst start, exact RunnerPlan, assertion/transform/XLSX/Superset read paths 72/100 Candidate for controlled production/shadow rollout after real PostgreSQL and Superset deployment proof
Browser execution Browser navigation/read/inspection with safe checkpoint and owned evidence 90/100 contract, 28/100 runtime Contract now includes per-action schemas/risk, limits, lifecycle, receipts, recovery and canaries; exclude until implementation and PREPROD proof
Screenshot evidence Durable principal/RLS-bound capture and receipt 20/100 Exclude; legacy ScreenshotService paths are not evidence
Mutation execution Browser/test-data mutation and cleanup/reconciliation 15/100 Exclude entirely from initial production scope; PROD mutation remains prohibited
AgentEvaluation Bounded model call plus immutable evaluation and DecisionPolicy 15/100 Exclude; no model-backed execution decisions in initial rollout
Human checkpoints Manual-only analyst disposition 85/100 Include only for manual runs; prohibited in scheduled/triggered/background runs
Retry/recovery Persisted lifecycle, safe retry, unsafe reconciliation block, timeout closure 78/100 Include for pure/read-only and explicitly retry-safe actions; unsafe external effects remain blocked
Scheduled automation 046 cron/deploy/release/ETL/API triggers 42/100 Exclude from initial production; 046 scheduler/subscriber lifecycle is not closed
Notifications Completion/failure/stale/triage domain events 35/100 Exclude as a production promise until lifecycle wiring is verified
Run monitor 045 typed SSE, provenance, human action and recovery UI 58/100 Include as operator surface only after typed 044 API alignment; not evidence of backend readiness
Investigation queue 047 queue ingestion and explicit analyst case opening 45/100 Include only read-only queue projection if signal ingestion is proven; agent case closure remains excluded
Analytics/health Flakiness, health, trends and recurring failures 40/100 Exclude from release acceptance; derived analytics cannot be operational truth yet
Capacity/admission Shared 044 CapacityManager across workloads 25/100 Hard blocker; no production provider I/O without this boundary
Security/authority Server-owned environment class, PROD gate, RBAC, exact bindings 86/100 Include; mandatory invariant for every rollout profile
Deployment/operations PostgreSQL migrations, readiness, health, observability, rollback 45/100 Hard blocker until real deployment evidence exists

The only defensible near-term production scope is:

  1. Agent-authored, validator-approved, immutable scenario revisions.
  2. Authenticated manual starts only; no schedules, triggers or background execution.
  3. PREPROD or explicitly non-production environments only until real PROD evidence is complete.
  4. Read-only Superset/SQL evidence through the exact 037 binding, bounded transform, assertion and server-owned XLSX artifact intake.
  5. HumanCheckpoint allowed only for manual runs.
  6. Retry limited to pure/idempotent read-only actions; unknown external effects become non-pass and are never replayed automatically.
  7. 045 monitor for typed status/provenance and 047 queue projection only where the corresponding APIs are verified.
  8. Explicitly surface unavailable Browser/Screenshot/AgentEvaluation/automation capabilities as disabled or non-pass, never as successful execution.

Estimated readiness of this constrained profile: 72-78/100, conditional on closing PostgreSQL, CapacityManager and deployment-readiness gates. It is a preview/shadow production profile, not full production automation.

Explicitly excluded from production scope

  • unattended scheduled, deploy, release, ETL or API-triggered execution;
  • any human-containing revision in an automated origin;
  • BrowserProvider and ScreenshotProvider until real owned evidence is proven;
  • browser or data mutation, including fixture mutation and cleanup side effects;
  • AgentEvaluation as a runtime decision source;
  • automatic retry after unknown external effects;
  • 047 agent-led remediation and unresolved case closure;
  • analytics-derived health as an authoritative deployment or release decision.

Full production scope gate

The complete feature may be called production-ready only when the constrained profile additionally has: real Browser/Screenshot providers, common provider receipts, shared CapacityManager, AgentEvaluation and DecisionPolicy, closed 046 scheduler/event lifecycle, canonical 047 signal/case closure, 042 staleness integration, real PostgreSQL migration checks and end-to-end agent-authoring-to-execution-to-investigation evidence. Until then, the correct release label is manual-readonly-preview, not production-ready.

BrowserProvider contract reassessment

BrowserProvider is now 90/100 contract-complete and 28/100 runtime-complete. The contract closes the prior ambiguity around every action, read-only versus mutation risk, limits, isolated context lifecycle, safe checkpoint/reconstruction, evidence ownership, cancellation, reconciliation, health and deployment. The score must not be interpreted as production enablement: T034, T040-T042, T042b and a real PREPROD canary remain release blockers. Until those pass, BrowserProvider stays typed unavailable and remains outside manual-readonly-preview.

Contract Package Reassessment: 042-047 (2026-08-24)

This reassessment separates two questions:

  • Contract/system description score: how completely the documents define a production system, including inputs, outputs, invariants, failure behavior, recovery, security, operations and measurable release evidence.
  • Production proof score: how much of that described system is implemented and externally verified.
Spec Contract/system description Production proof Key contract gap Production blocker
042 Registry 88/100 70/100 Staleness/health signal timing and cross-spec subscription semantics need stricter operational evidence 037/041 runtime staleness integration and final independent verification
043 Editor 84/100 66/100 Save/activate/revision API and policy are clear, but end-to-end conflict and accessibility evidence are incomplete Depends on 042 activation/staleness and 044 typed launch contract
044 Execution 91/100 54/100 Provider contract is now detailed; common runtime protocol/capacity/live providers remain unimplemented Browser/Screenshot runtime, receipts, CapacityManager, PostgreSQL, deployment canaries
045 Monitor 83/100 61/100 Typed event/provenance UX is strong, but API schema alignment and event replay evidence remain incomplete Depends on 044 runtime truth and 047 investigation ingestion
046 Automation 84/100 59/100 Trigger, schedule, dedup, retention and manual-human prohibition are clear; scheduler operational semantics need live proof Persisted scheduler reload, event subscribers, notifications, capacity and real cron/trigger run
047 Analytics 86/100 59/100 Queue/case/health/flakiness contracts are clear; evidence snapshot and closure invariants are not runtime-closed Canonical signal ingestion, immutable case evidence, verification/reconciliation closure

Functional contract categories

Category Contract completeness Production proof Assessment
Scenario identity, revision pinning and stale policy 88/100 69/100 042/043 define the model; runtime staleness propagation remains incomplete
Agent-authored graph/editor workflow 84/100 65/100 Draft/validation/save/activate are described; full agent-to-activation E2E is not proven
Deterministic execution and provider boundary 93/100 54/100 044 is contractually strongest; live provider protocol is the largest implementation gap
Browser/evidence ownership 90/100 28/100 BrowserProvider contract is detailed, but no deployed provider/canary receipt exists
Human checkpoint and manual-only policy 94/100 85/100 Strongly specified and tested; automated human revisions are rejected pre-create
Retry, timeout, cancel and crash recovery 88/100 78/100 Runner lifecycle is strong; provider-level reconciliation is absent
Scheduled/event automation 84/100 42/100 046 requirements are clear; operational scheduler/subscriber lifecycle is not demonstrated
Monitoring, SSE and operator recovery 83/100 58/100 045 renders typed state, but depends on incomplete 044/047 boundaries
Investigation/case/triage workflow 86/100 45/100 047 producer signal exists; durable case evidence and closure are missing
Analytics, health and recurrence 82/100 40/100 Deterministic formulas are described; end-to-end historical aggregation and 042 feedback are open
Security, authority and RBAC 89/100 80/100 Server-owned gates and object-level access are well covered; deployment audit remains required
Deployment, health, observability and rollback 76/100 38/100 Requirements exist but real PostgreSQL, provider readiness, scheduler and rollback evidence are open

Recalculated package scores

Using contract completeness as a descriptive index, the 042-047 package scores 86/100. This means the documents describe most of the intended production system with clear boundaries and measurable requirements.

Using production proof and hard-gate weighting, the package scores 58/100. The score is lower than the previous 67/100 estimate for 036-047 because this narrower 042-047 package gives greater weight to the unclosed execution, automation and case-closure boundaries rather than to the mature 036/037 substrate.

The complete package remains NO-GO. Hard gates are conjunctive:

  1. 042 revision/staleness state is connected to 037/041 and emits canonical signal evidence.
  2. 043 save/activate flow is verified end to end against 042 and 044.
  3. 044 provider runtime passes common receipts, capacity, cancellation/reconciliation and BrowserProvider PREPROD canaries; BrowserProvider contract completeness is 90/100 but runtime remains 28/100.
  4. 045 consumes typed 044 events with gap-free replay and renders evidence/provenance without compatibility payloads.
  5. 046 persisted scheduler reload, event subscribers, dedup, notifications and trigger-to-044 execution are verified in a real deployment; human-containing revisions remain pre-create rejected for automation.
  6. 047 ingests canonical signals, persists immutable evidence snapshots, enforces verification/reconciliation before resolution and preserves immutable run truth.
  7. Real PostgreSQL migration/deployment checks, readiness, observability and rollback evidence pass.

Production interpretation

The package defines a complete production tool, not a reduced profile. BrowserProvider, Screenshot, mutation, AgentEvaluation, unattended automation and agent-led remediation are mandatory capability gates; missing proof keeps the system NO-GO rather than creating a preview release.

Superseding Product Scope Decision (2026-08-24)

The earlier manual-readonly-preview recommendation is superseded and must not be used as a release target. The product target is the complete production tool: agent authoring and activation, all trusted manual and automated trigger origins, Browser/Screenshot providers, controlled non-PROD fixture mutation, AgentEvaluation/DecisionPolicy, monitoring, automation, investigation cases, analytics and remediation.

These are mandatory capability gates, not permanent exclusions. A missing or unproven capability blocks production GO. The only policy exclusions are: HumanCheckpoint-containing revisions cannot be automated and PROD mutation is prohibited. Unknown external effects cannot retry before reconciliation. This decision overrides earlier preview/shadow wording in this historical session ledger.

================================================================================ FEATURE: 045-dashboard-run-monitor Files: 14


SPEC — Feature Specification

Source: spec.md

#region ScenarioRunMonitor.Spec [C:3] [TYPE ADR] [SEMANTICS spec,requirements,ux,scenario,run,monitor,result,history] @BRIEF User-facing live Scenario Run Monitor and Results UX: run configuration, live step timeline with evidence, human-checkpoint actions, final result with provenance, run history and comparison. A standalone surface, not a small panel. @RELATION DEPENDS_ON -> [Doc.Adr.ADR0001] @RELATION DEPENDS_ON -> [Doc.Adr.ADR0006] @RELATION DEPENDS_ON -> [ScenarioExecution.Spec] @RELATION DEPENDS_ON -> [ScenarioRegistry.Spec] @RELATION DEPENDS_ON -> [DashboardLoadTesting.Spec] @RATIONALE After creating a scenario the user needs to run it and understand results; 042/044 provide the backend run/result model, 045 renders it live with recovery, evidence, human actions, history, and comparison (reusing 040 LoadRunComparison UX ideas). Without this, execution is a headless API. @REJECTED A tiny inline panel — rejected because live monitoring, evidence review, human checkpoint, and comparison each need dedicated surfaces; the run is a primary workflow, not a widget. @REJECTED Naming this "Verification history" — rejected because 037 VerificationRun is release-pipeline verification; Scenario runs must be labeled distinctly. @RATIONALE After the 050 MCP drift, RUNMON-FR-011 "Investigate with agent" transitions into an external MCP client session over the same InvestigationCase; the monitor keeps rendering typed events only.

Navigation (DSA Indexer keywords)

@SEMANTICS: spec, requirements, feature, ux, scenario, run, monitor, result, history, compare

Feature Branch: 045-dashboard-run-monitor Created: 2026-08-07 | Status: Partially implemented — factual audit pending remediation Input: "Provide a live Scenario Run Monitor and Results UX: persistent run configuration panel, live step timeline with per-step status/duration/logs/evidence, human checkpoint actions inside the monitor, final result/provenance, comparison, and an entry to agent-led investigation."

User Scenarios

Story 1 — Configure and Launch a Run (P1)

Why P1: Run is not a single button; it needs environment, revision, baseline set, parameters, and toggles.

Independent Test: Open the run configuration for a scenario and verify it pre-populates environment, revision, baseline set, parameters, and per-category toggles; launching starts a run.

Acceptance:

  1. Given a scenario detail When "Run" is clicked Then a persistent configuration panel shows Environment, Revision (current), target/release information, Baseline set, and Parameters. It separates mandatory graph checks (for example XLSX comparison) from optional diagnostic enrichment (screenshots, verbose logs, VLM commentary), which alone may be toggled.
  2. Given the user picks PROD When launch is requested Then a 036 approval gate appears (concurrency ceiling, volume, blast-radius dependents, reason).
  3. Given the user confirms When submitted Then a ScenarioRun starts and the monitor opens.

Story 2 — Live Run Monitor (P1)

Why P1: Users must watch progress, step status, and evidence as the run executes.

Independent Test: Feed a running run's events and verify the monitor renders progress, per-step status/duration, logs, and evidence links.

Acceptance:

  1. Given a running run When the monitor renders Then it shows elapsed time, progress (N/M), and per-step status/duration/attempts with icons.
  2. Given a step completes When inspected Then inputs, outputs, refs, error, logs, and artifacts/evidence are visible.
  3. Given the browser disconnects When the user reopens the run id Then live status and step timeline are recovered server-side.

Story 3 — Human Checkpoint Inside the Monitor (P1)

Why P1: Pause/resume must be actionable from the monitor, not after the fact.

Independent Test: Reach a human step and verify the monitor shows a WAITING FOR HUMAN panel with evidence and confirm/false-positive/inconclusive actions that resume the run.

Acceptance:

  1. Given a run is waiting at a human step When the monitor renders Then a WAITING FOR HUMAN panel shows the evidence (screenshot, VLM finding, confidence) and step context.
  2. Given the user chooses a disposition When submitted Then the step records the outcome and the run resumes.
  3. Given a VLM finding When disposition changes Then it is auditable and never alters scenario graph structure.

Story 4 — Final Result (P2)

Why P2: The run's outcome and reproducibility are the business value.

Independent Test: Complete a run and verify the result view shows summary, pass/fail counts, failures with expected/actual/delta, evidence, and full provenance.

Acceptance:

  1. Given a completed run When the result view renders Then it shows status, check counts (passed/failed/warning/inconclusive), duration, revision, release, environment, baseline set.
  2. Given failures exist When expanded Then each shows expected vs actual vs delta with evidence links.
  3. Given the user inspects provenance When expanded Then scenario_revision_id, content_hash, runner_version, template_version, baseline_revision, environment, parameter snapshot, query fingerprints are shown.

Story 5 — Run History and Comparison (P2)

Why P2: Analysts compare runs and spot drift/flakiness.

Independent Test: Load run history for a scenario and compare two runs; verify per-step deltas and consistency findings.

Acceptance:

  1. Given a scenario When the Runs tab renders Then a chronological list shows status, environment, revision, duration, and timestamp.
  2. Given two runs are selected When compared Then per-step latency/value deltas and consistency findings are shown (reuse 040 LoadRunComparison ideas).
  3. Given a run references a different revision When compared Then the revision difference is surfaced to avoid false comparisons.

Story 6 — Global Run Operations Center (P2) (#10)

Why P2: With automation (046) and many scenarios, operators need a cross-scenario view of all runs, not per-scenario navigation.

Independent Test: Open the global runs center and verify it lists all runs with filters and a "Waiting for me" view.

Acceptance:

  1. Given runs exist across scenarios When the Global Run Operations Center opens Then it lists active, queued, waiting-human, failed, and recently completed runs with scenario, environment, trigger, and owner columns and filters.
  2. Given human checkpoints are pending When the "Waiting for me" filter is applied Then only runs awaiting the current user's disposition are shown, actionable inline.
  3. Given a run is selected When clicked Then it navigates to that run's live monitor.

Edge & Failure Cases

# Scenario Expected Behavior Recovery
E1 PROD gate denied No dispatch; denial recorded Adjust / retry
E2 Run failed Failure detail + provenance + Investigation Queue item Open 047 case with agent
E3 Disconnect mid-run Recover by run_id Reopen monitor
E4 Run terminated mid-step In-flight completes/times out View partial results
E5 Compare different revisions Revision diff surfaced Note before compare
E6 RBAC result denied 403 permission_denied Contact admin

Requirements

Functional

  • RUNMON-FR-001: Run configuration MUST include environment, revision, release, baseline set, parameters, and execution toggles; PROD requires a 036 approval gate.
  • RUNMON-FR-002: The live monitor MUST render elapsed time, progress (N/M), and per-step status, duration, attempts, inputs/outputs/refs, error, logs, and evidence links.
  • RUNMON-FR-003: The monitor MUST be recoverable by scenario_run_id after disconnect (server-side run).
  • RUNMON-FR-004: Human checkpoints MUST be actionable inside the monitor with confirm/false-positive/inconclusive; disposition auditable, never alters graph.
  • RUNMON-FR-005: The result view MUST show summary, check counts, failures (expected/actual/delta), evidence, and full provenance (revision, runner/template/baseline versions, env, param snapshot, query fingerprints).
  • RUNMON-FR-006: Run history MUST be a chronological list; comparison MUST show per-step deltas and consistency findings, surfacing revision differences.
  • RUNMON-FR-007: Scenario runs MUST be labeled distinctly from 037 release verification and 040 load tests in all UI.
  • RUNMON-FR-008: All UI MUST follow Svelte 5 runes/model-first conventions and be keyboard-accessible.
  • RUNMON-FR-009: A Global Run Operations Center MUST list all runs (active/queued/waiting-human/failed/recent) with filters and a "Waiting for me" view for pending human checkpoints.
  • RUNMON-FR-010: Run Configuration MUST match the 044 start contract (environment, revision, release, baseline_set, parameters, execution_toggles for optional evidence only); mandatory graph steps MUST NOT be toggleable off.
  • RUNMON-FR-011: Failed, blocked and inconclusive results MUST expose their Investigation Queue item and an explicit "Investigate with agent" transition into the persistent 047 case workspace. Opening it never mutates run truth or auto-executes tools.
  • RUNMON-FR-012: Live Monitor MUST render only typed 044 ScenarioRunEvent payloads (including approval, evidence, agent-evaluation, checkpoint and terminal events), never parse assistant prose. Result and step inspection MUST render StepOutcome, EvidenceReference and AgentEvaluationSummary separately.
  • RUNMON-FR-013: An AgentEvaluation display MUST show declared prompt/model version, evidence manifest, verdict, confidence, reason codes and the DecisionPolicy-derived StepOutcome. It must never present the raw model verdict as ScenarioResult.
  • RUNMON-FR-012: Run configuration, gate decisions, conflict recovery and investigation entry MUST use persistent pages/panels or inline cards; modal/dialog interaction MUST NOT be required.

Key Entities

  • RunConfiguration: Launch parameters (env, revision, baseline set, params, toggles, gate linkage).
  • RunMonitorModel: Frontend screen model (.svelte.ts) holding run state, step timeline, and human-action actions; binds to SSE events.
  • ScenarioResultView: Rendered final result with summary, failures, evidence, and provenance.
  • RunComparison: Delta view of two runs (per-step latency/value + consistency).

Success Criteria

  • SC-001: A fixture run renders live progress and per-step detail from events only (no prose parsing).
  • SC-002: Human checkpoint resolves and resumes the run from the monitor.
  • SC-003: Reopening a run after disconnect recovers live status by run id.
  • SC-004: Final result shows full provenance and reproduces expected/actual/delta for failures.
  • SC-005: Comparison surfaces revision differences and per-step deltas without false cross-revision conclusions.
  • SC-006: PROD runs are gated; denial dispatches nothing.

Clarifications

Session 2026-08-07

  • Q: Is 045 a separate surface? → A: Yes, standalone Run Monitor + Results, not an extension of the 039 create workspace or 043 editor.
  • Q: How does it differ from 039 VerificationHistoryList? → A: 039 renders 037 VerificationRun (release pipeline); 045 renders ScenarioRun (user-created scenario execution). Distinct entities, distinct labels.
  • Q: Reuse of 040? → A: Reuse the LoadRunComparison UX concept for run comparison.

Implementation Status & MVP Debt (factual audit 2026-08-20)

Monitor model, timeline, checkpoint/result/history/compare components and scenario-run routes exist.

  • [~] RunMonitorModel launches runs, but release, baseline set and execution toggles are embedded in params.launch_config rather than represented by the 044 typed start contract.
  • [~] The UI can render typed events, but its runtime truth is limited by incomplete 044 execution.
  • [ ] Failed/blocked/inconclusive investigation entry is not end-to-end because 047 does not ingest production signals into the queue.
  • [ ] Current browser, reconnect and accessibility evidence has not been retained.

Drift Amendment — MCP Interface (2026-08-24)

  • The monitor remains a pure web surface over typed events. Its "Investigate with agent" button targets a HandoffSurface (connection hint + case context) instead of an in-product chat route.

Status (2026-09-02): done — реализовано в рамках 050: инструменты и гейты (specs/050-mcp-interface/tasks.md T012–T028 [x]), handoff-поверхность (050 T030–T033), демонтаж чата и сервиса agent/ (050 T040–T041, чекпоинты specs/WORKSTATE-043-047.md).

#endregion ScenarioRunMonitor.Spec


UX REFERENCE — Interaction Narrative

Source: ux_reference.md

#region ScenarioRunMonitor.UxReference [C:3] [TYPE ADR] [SEMANTICS ux,reference,scenario,run,monitor,result] @BRIEF UX interaction reference for the Scenario Run Monitor & Results (045).

Feature Branch: 045-dashboard-run-monitor | Created: 2026-08-07

1. User Persona & Context

  • User: BI analyst / quality engineer running a dashboard test scenario.
  • Goal: Configure, launch, watch, and interpret a scenario run; handle human checkpoints; compare runs.
  • Context: Browser; from scenario detail "Run" or automation trigger.

2. Happy Path

Analyst opens the persistent launch panel, picks PREPROD + r17 + v31 baseline + params, and selects optional diagnostics. The monitor streams progress; a manual-only run may show an inline HumanCheckpoint. A failed/blocked/inconclusive result links to Investigation Queue, where the analyst may explicitly open an agent case. The result and comparison retain full provenance.

3. Screens & States

Screen: Run Configuration Panel

  • Layout: Environment select, Revision (current), Release, Baseline set, Parameters, Execution toggles.
  • @UX_STATE: idle, validating, prod_gate, launching, error.
  • @UX_RECOVERY: prod gate → approve/deny; 403 → no confirm.

Screen: Live Run Monitor

  • Layout: Header (run id, status badge, elapsed, progress) + step timeline + step inspector + human panel (conditional).
  • @UX_STATE: connecting, running, waiting_human, failed, cancelled, disconnected, reconnected.
  • @UX_RECOVERY: disconnect → reconnect by run_id (server-side).

Screen: Final Result

  • Layout: Summary card (counts, duration, revision, env, baseline) + failures (expected/actual/delta) + evidence + provenance + persistent Investigation Queue entry when qualifying evidence exists.
  • @UX_STATE: loaded, failed, provenance_expanded.

Screen: Runs History + Compare

  • Layout: chronological table + compare view (per-step deltas, revision-diff warning).
  • @UX_STATE: idle, loaded, cross_revision_warning.

4. Error Experience

  • PROD gate denied → no dispatch; record denial.
  • Run failed → failure detail + Investigation Queue entry (047); “Investigate with agent” opens a persistent case.
  • Queue is a navigation destination from the result and Operations Center; it does not become a modal or start case work before analyst selection.
  • Compare different revisions → warning surfaced.

5. Tone & Voice

  • Style: Concise, technical. Terminology: "Scenario run" distinct from "Release verification" and "Load tests".

Edge & Failure Matrix (feed to prototype)

NET_01/02/03, VAL_01/02, AUTH_01/02, NF_01, CONF_01/02, 422, 429, 5XX, STALE, PARTIAL, EMPTY, MALFORMED, A11Y, RESP.

#endregion ScenarioRunMonitor.UxReference


CHECKLISTS — Requirements Quality — requirements.md

Source: checklists/requirements.md

Requirements Checklist: Scenario Run Monitor & Results UX (045)

Purpose: Verify RUNMON-FR-001..008 completeness. | Created: 2026-08-07

Factual audit 2026-08-20: [x] requires current production evidence, [~] means partial code exists, [ ] means missing integration or proof.

Configure & Launch (FR-001)

  • [~] CHK001 Run config includes environment, revision, release, baseline set, parameters, toggles
  • CHK002 PROD requires 036 approval gate; denial dispatches nothing

Live Monitor (FR-002/003)

  • CHK003 Renders elapsed, progress (N/M), per-step status/duration/attempts
  • CHK004 Step inspector shows inputs/outputs/refs/error/logs/evidence
  • CHK005 Recoverable by scenario_run_id after disconnect (server-side run)
  • CHK006 Renders from events only, never parses prose

Human Checkpoint (FR-004)

  • CHK007 WAITING FOR HUMAN panel with evidence + confirm/false-positive/inconclusive
  • CHK008 Disposition resumes run; auditable; never alters graph

Final Result (FR-005)

  • CHK009 Summary + counts + failures (expected/actual/delta) + evidence
  • CHK010 Full provenance (revision, runner/template/baseline versions, env, param snapshot, query fingerprints)

History & Compare (FR-006)

  • CHK011 Chronological run history list
  • CHK012 Comparison shows per-step deltas + consistency; surfaces revision differences

Naming / UX (FR-007/008)

  • CHK013 Scenario runs labeled distinctly from 037/040
  • CHK014 Svelte 5 runes, keyboard-accessible

Success Criteria

  • CHK015 SC-001..006 verified

UX DECISIONS — Final Choices

Source: contracts/ux/decisions.md

#region ScenarioRunMonitor.Ux.Decisions [C:3] [TYPE ADR] [SEMANTICS scenario,run,monitor,ux,decisions] @BRIEF Final UX decisions for the Scenario Run Monitor & Results (045). @RELATION DEPENDS_ON -> [ScenarioRunMonitor.Spec]

Decision 1 — Standalone surface

Run Monitor + Results is a dedicated route, not a panel; distinct from 039 create workspace and 043 editor.

Decision 2 — Events-only rendering

Live timeline renders from 044 SSE events; never parses prose.

Decision 3 — Human actions in monitor

WAITING FOR HUMAN panel with confirm/false-positive/inconclusive → 044 human/decision; run resumes.

Decision 4 — Distinct naming

Scenario runs labeled "Scenario runs", distinct from "Release verification" (037) and "Load tests" (040).

Decision 5 — Comparison reuse

Reuse 040 LoadRunComparison UX concept; surface revision differences to avoid false cross-revision conclusions. #endregion ScenarioRunMonitor.Ux.Decisions


PLAN — Implementation Plan

Source: plan.md

Implementation Plan: Scenario Run Monitor & Results UX

Branch: 045-dashboard-run-monitor | Date: 2026-08-07 | Spec: spec.md | Status: Partially implemented — factual audit pending remediation

Implementation audit, 2026-08-20: monitor UI exists, but typed launch-contract alignment and end-to-end evidence depend on unresolved 044 execution and 047 investigation ingestion.

Summary

A standalone live Scenario Run Monitor and Results surface: persistent run configuration panel, live step timeline bound to 044 SSE events, human-checkpoint actions inside the monitor, final result/provenance, run history/comparison and Investigation Queue entry. Consumes 044 run/step API plus 047 queue/case DTOs.

Technical Context

Language/Version: TypeScript + Svelte 5 runes (frontend); Python DTOs (backend 044) Primary Dependencies: SvelteKit 5, Vite, Tailwind; SSE consumption; existing 044 run API, 040 comparison pattern Storage: none new (reads 044 run/step state) Testing: vitest L1 model + L2 UX (@testing-library/svelte) Frontend Architecture: RunMonitorModel.svelte.ts, model-first, runes-only Performance Goals: live update < 200ms; result render < 200ms Constraints: no prose parsing; scenario runs labeled distinctly from 037/040; PROD gated Scale: single run at a time per monitor; run history paginated

Constitution Check

Principle Result
VI. Svelte 5 Runes Only PASS
VII. Test-Driven PASS — L1 model then L2 UX
II. Decision Memory PASS — research R1-R4
VIII. Attention-Optimized PASS

Project Structure

specs/045-dashboard-run-monitor/
├── spec.md / data-model.md / research.md / plan.md / tasks.md / traceability.md / quickstart.md / ux_reference.md
├── checklists/requirements.md
├── contracts/modules.md, contracts/ux/
└── prototype/index.html + manifest.md

frontend/src/lib/models/RunMonitorModel.svelte.ts
frontend/src/lib/components/scenario-run/ (RunConfigurationPanel, RunTimeline, StepInspector, HumanCheckpointPanel, ScenarioResultView, RunComparison, RunHistoryList, InvestigationEntry)
frontend/src/routes/dashboard-testing/scenarios/[id]/runs/[runId]/+page.svelte

Delivery Phases

  1. Persistent RunConfiguration panel + DTOs.
  2. RunMonitorModel + SSE binding (L1 tests first).
  3. Live timeline + step inspector.
  4. Human checkpoint panel + disposition.
  5. Final result view + provenance.
  6. Run history + comparison.
  7. Routes integration, polish, regression gates.

Traceability

traceability.md maps Story → model → action → contract → task → test.

Cross-Spec Boundary

  • Consumes 044 run/step API + SSE; reads scenarios from 042.
  • Reuses 040 LoadRunComparison UX concept.
  • Distinct labels from 037 release verification.

Complexity Tracking

No exception planned. Monitor model is bounded C3-C4; comparison and result view model-first.


RESEARCH — Technical Decisions

Source: research.md

Scenario Run Monitor & Results — Phase 0/1 Research (045)

Branch: 045-dashboard-run-monitor | Date: 2026-08-07 | Spec: spec.md

R1. Standalone surface

Decision: Dedicated Run Monitor + Results UI (run config, live timeline, evidence, human actions, result, history, compare). Not a panel.

Rationale: The run is a primary workflow; 042/044 provide the runtime. A tiny panel would be a facade.

Alternatives: inline panel (rejected); reusing 039 VerificationHistoryList for scenario runs (rejected: different entity).

Impact: new routes/components + RunMonitorModel.svelte.ts.

R2. Event binding

Decision: Monitor consumes 044 GET /scenario-runs/{id}/events (SSE) for live updates; recovers full state by run id on reconnect.

Rationale: Server-side run recovery (pattern from 036/040).

Impact: SSE binding in model; no prose parsing.

R3. Human actions in monitor

Decision: WAITING FOR HUMAN panel inside the monitor with confirm/false-positive/inconclusive → 044 human/decision.

Impact: human disposition actionable from live view.

R4. Comparison reuse

Decision: Reuse 040 LoadRunComparison UX concept for ScenarioRun comparison; surface revision differences.

Impact: per-step delta + consistency comparison.

Contracts & API

  • contracts/modules.md — frontend model + view contracts.
  • Consumes 044 API; UI-only feature (no new backend entities beyond DTOs).

Constitution Check

Principle Result
VI. Svelte 5 Runes Only PASS — RunMonitorModel .svelte.ts
VII. Test-Driven PASS — L1 model + L2 UX tests first
II. Decision Memory PASS — R1-R4
VIII. Attention-Optimized PASS

DATA MODEL — Entities & Relations

Source: data-model.md

#region ScenarioRunMonitor.DataModel [C:4] [TYPE ADR] [SEMANTICS data-model,scenario,run,monitor,result,compare] @BRIEF Run configuration, monitor screen model, result view, and run-comparison models for 045. @RELATION DEPENDS_ON -> [ScenarioRunMonitor.Research] @RATIONALE The monitor binds to 044 run/step DTOs + SSE events; typed configuration and result/provenance models make launches reproducible and comparisons reliable. @REJECTED Reusing 037 VerificationRun for scenario-run UI — distinct concepts, distinct entities.

RunConfiguration

Fields: scenario_id, revision_id, environment_id, target_preview, baseline_set (version), parameters (map), required_checks (read-only graph steps), execution_toggles (optional evidence only — diagnostic screenshots, verbose logs, optional VLM; mandatory graph steps such as XLSX are NOT toggleable), prod_gate (ActionApprovalGate ref when PROD). This is pre-run data and is loaded into a persistent launch panel from GET /scenarios/{scenario_id}/run-configuration; validated before launch and matches 044 start contract.

RunCenterModel

Frontend .svelte.ts model for the Global Run Operations Center: run rows with scenario/environment/trigger/owner columns, status/dashboard/env/trigger/owner filters, "Waiting for me" view (runs awaiting the current user's HumanCheckpoint disposition). @ACTION: filter, list, openRun, waitingForMe.

RunMonitorModel

Frontend .svelte.ts model: run (ScenarioRun), steps (ScenarioStepRun[]), elapsed, progress (derived), human_action (active gate), selected_step (inspector), events (SSE binding). @STATE: idle, configuring, running, waiting_human, failed, passed, cancelled. @ACTION: launch, cancel, dispose(human), openStep, reconnect.

ScenarioResultView

Fields: overall status, counts (passed/failed/warning/inconclusive), duration, revision_id, content_hash, release, environment, baseline_set, failures[] (step, expected, actual, delta, evidence), provenance (scenario_revision_id, content_hash, runner_version, template_version, baseline_revision, environment, param_snapshot, query_fingerprints).

RunComparison

Two runs pinned to the same scenario (optionally same revision). Per-step latency/value deltas; consistency findings (identical step diverged); revision difference surfaced when revisions differ.

Evidence & Disposition

Human checkpoint evidence (screenshot, VLM finding) bound to a gate; disposition (confirm/false-positive/inconclusive) is auditable and never alters the graph.

Failure/blocked/inconclusive results show a linked Investigation Queue item when one exists. Investigate with agent opens the persistent 047 case workspace; it is not a modal, does not change the run, and never consumes a HumanCheckpoint.

AgentEvaluationPanel is a result subprojection, not a chat surface: evaluation id, logical_step_id, model/prompt version, evidence manifest, typed verdict/confidence/reason codes and DecisionPolicy-derived StepOutcome. The UI reads this only from typed 044 result/SSE schemas.

#endregion ScenarioRunMonitor.DataModel


CONTRACTS — Module & Function Contracts

Source: contracts/modules.md

#region ScenarioRunMonitor.Modules [C:4] [TYPE ADR] [SEMANTICS scenario,run,monitor,contracts,modules,result] @BRIEF Frontend model/view contracts for the Scenario Run Monitor & Results (045). @defgroup ScenarioRunMonitor Live run monitoring, results, history, comparison. @RELATION DEPENDS_ON -> [ScenarioExecution.Api] @RELATION DEPENDS_ON -> [ScenarioRegistry.Api] @RELATION DEPENDS_ON -> [DashboardLoadTesting.Api] @RATIONALE The monitor binds to 044 run/step DTOs + SSE; reuse 040 comparison UX; scenario runs labeled distinctly from 037. @REJECTED Rendering scenario runs via 037 VerificationHistoryList.

#region RunMonitor.Launch [C:3] [TYPE Action] [SEMANTICS scenario,run,monitor,launch,config]

@ingroup ScenarioRunMonitor

@BRIEF Open and validate run configuration; launch via 044.

@ACTION launch(params): POST /scenario-runs

@POST opens monitor bound to run id; PROD gate shown when env is PROD.

def launch(config): ...

#endregion RunMonitor.Launch

#region RunMonitor.BindEvents [C:4] [TYPE Action] [SEMANTICS scenario,run,monitor,events,sse]

@ingroup ScenarioRunMonitor

@BRIEF Bind to 044 SSE events and update step timeline.

@ACTION bindEvents(runId): GET /scenario-runs/{id}/events

@POST live step status/duration/progress update; never parse prose.

def bind_events(run_id): ...

#endregion RunMonitor.BindEvents

#region RunMonitor.HumanAction [C:4] [TYPE Action] [SEMANTICS scenario,run,monitor,human,disposition]

@ingroup ScenarioRunMonitor

@BRIEF Resolve a human checkpoint from the monitor.

@ACTION dispose(gateId, decision, comment): POST /scenario-runs/{id}/human/decision

@POST step outcome recorded; run resumes; disposition auditable, never alters graph.

def dispose(gate_id, decision, comment): ...

#endregion RunMonitor.HumanAction

#region RunMonitor.ResultView [C:3] [TYPE Component] [SEMANTICS scenario,run,monitor,result,provenance]

@ingroup ScenarioRunMonitor

@BRIEF Render final result with summary, failures (expected/actual/delta), evidence, provenance.

@UX_STATE idle, loading, loaded, failed, error, provenance_expanded.

def result_view(run): ...

#endregion RunMonitor.ResultView

#region RunMonitor.Compare [C:4] [TYPE Component] [SEMANTICS scenario,run,monitor,compare,delta]

@ingroup ScenarioRunMonitor

@BRIEF Compare two scenario runs; surface per-step deltas and revision differences.

@UX_STATE idle, loaded, cross_revision_warning.

def compare(run_a, run_b): ...

#endregion RunMonitor.Compare

#endregion ScenarioRunMonitor.Modules


OPENAPI — REST/Event API Contract

Source: contracts/openapi.yaml

openapi: 3.1.0 info: title: Scenario Run Monitor API (view contract) version: 0.1.0 description: Frontend-facing surface for 045 — global run center and run configuration. Execution state/events come from 044. paths: /api/scenario-runs: get: operationId: runCenter.list summary: Global Run Operations Center — all runs with filters security: [{ bearerAuth: [] }] parameters: - { name: status, in: query, schema: { type: string, enum: [pending_approval, queued, running, waiting_human, blocked, cancel_requested, failed, passed, cancelled, inconclusive] } } - { name: environment_id, in: query, schema: { type: string } } - { name: scenario_id, in: query, schema: { type: string } } - { name: dashboard_id, in: query, schema: { type: string } } - { name: owner, in: query, schema: { type: string } } - { name: trigger, in: query, schema: { type: string } } - { name: waiting_for_me, in: query, schema: { type: boolean } } - { name: page, in: query, schema: { type: integer, default: 1 } } - { name: page_size, in: query, schema: { type: integer, default: 50 } } responses: "200": description: Paged scenario runs content: application/json: schema: type: object properties: items: type: array items: { $ref: "#/components/schemas/ScenarioRunRow" } total: { type: integer } /api/scenarios/{scenario_id}/run-configuration: get: operationId: runCenter.configuration summary: Pre-populated run configuration for a scenario (env/revision/release/baseline/params/toggles) security: [{ bearerAuth: [] }] parameters: - { name: scenario_id, in: path, required: true, schema: { type: string } } responses: "200": description: RunConfiguration template content: application/json: schema: { $ref: "#/components/schemas/RunConfiguration" } components: securitySchemes: bearerAuth: { type: http, scheme: bearer } schemas: ScenarioRunRow: type: object properties: run_id: { type: string } scenario_id: { type: string } scenario_name: { type: string } environment_id: { type: string } status: { type: string } trigger: { type: string } owner: { type: string } started_at: { type: string, format: date-time } waiting_checkpoint_id: { type: string, nullable: true } RunConfiguration: type: object properties: scenario_id: { type: string } revision_id: { type: string } environment_id: { type: string } release: { type: string, nullable: true } baseline_set: { type: string } params: { type: object } execution_toggles: type: object description: "Optional evidence only (diagnostic screenshots, verbose logs, optional VLM). Mandatory graph steps cannot be disabled."


QUICKSTART — Dev Onboarding

Source: quickstart.md

Quickstart: Scenario Run Monitor & Results UX (045)

Factual audit 2026-08-20: pending verification checklist only; launch-contract and investigation handoff scenarios depend on remediation in 044 and 047.

Prereqs

  • 044 execution API live (scenario-runs, events, human/decision)
  • Frontend deps installed

Commands

cd frontend && npm run test -- RunMonitor
npm run lint

Exit Gates

  • Persistent run configuration panel opens with env/revision/baseline/params/toggles; inline PROD gate appears
  • Live timeline renders from events only; step inspector shows inputs/outputs/evidence
  • Reconnect recovers run by scenario_run_id
  • WAITING FOR HUMAN panel resolves and resumes
  • Final result shows counts + failures + full provenance
  • Failed/blocked/inconclusive result exposes Investigation Queue entry without auto-starting agent work
  • Comparison surfaces per-step deltas + revision-diff warning
  • Scenario runs labeled distinctly from 037/040

TRACEABILITY — Requirements Matrix

Source: traceability.md

Traceability: Scenario Run Monitor & Results (045)

Factual audit 2026-08-20: rows identify code/test ownership only. The launch contract currently transports release/baseline/toggles through params.launch_config; complete 044/047 runtime proof is open.

Story Requirement Model API operationId Contract Task Test
US1 Launch RUNMON-FR-001 RunConfiguration scenarioRun.start RunMonitor.Launch T003-T005 RunMonitorModel.test, run.ux.test
US2 Live RUNMON-FR-002/003 RunMonitorModel scenarioRun.events RunMonitor.BindEvents T006-T008 run.ux.test
US3 Human RUNMON-FR-004 HumanCheckpointPanel scenarioRun.humanDecision RunMonitor.HumanAction T009-T010 run.ux.test
US4 Result RUNMON-FR-005 ScenarioResultView scenarioRun.detail RunMonitor.ResultView T011-T012 run.ux.test
US5 Compare RUNMON-FR-006 RunComparison — RunMonitor.Compare T013-T014 compare.ux.test
Naming RUNMON-FR-007/008 — — — T015 run.ux.test

N/A: Registry (042), Editor (043), Execution backend (044), Automation (046), Analytics (047).


TASKS — Implementation Tasks

Source: tasks.md

#region ScenarioRunMonitor.Tasks [C:3] [TYPE ADR] [SEMANTICS tasks,scenario,run,monitor,implementation] @BRIEF Ordered TDD backlog for Scenario Run Monitor & Results UX (045). Tests FIRST.

Prerequisites: plan.md, spec.md; contracts/modules.md, traceability.md.

Format: - [ ] T### [P] [USx] Description with exact file path

Factual audit 2026-08-20: [x] means code plus relevant evidence; [~] means partial implementation; [ ] means absent integration or unperformed verification.

Phase 1 — Setup

  • T001 Define frontend DTOs (RunConfiguration, ScenarioResultView, RunComparison) in frontend/src/types/scenario-run.ts
  • T002 [P] Create canonical fixtures in specs/045-dashboard-run-monitor/fixtures/ (run states, step timeline, result) Present: specs/045-dashboard-run-monitor/fixtures/scenario-runs.json

Phase 2 — US1 Configure and Launch

  • T003 [US1] L1 model test for RunMonitorModel launch in frontend/src/lib/models/__tests__/RunMonitorModel.test.ts
  • T004 [US1] Build persistent RunConfigurationPanel.svelte (env, revision, baseline, params, toggles, inline PROD gate)
  • T005 [US1] L2 UX test for PROD gate Test: frontend/src/lib/components/scenario-run/__tests__/RunConfigurationPanel.test.ts

Phase 3 — US2 Live Monitor

  • T006 [US2] Implement RunMonitorModel.svelte.ts + SSE binding (bindEvents) @POST: live step status/duration/progress; never parse prose
  • T007 [US2] Build typed RunTimeline.svelte (StepInspector detail remains incremental)
  • T008 [US2] L2 UX test for step timeline from events Test: frontend/src/lib/components/scenario-run/__tests__/RunTimeline.test.ts (typed status/progress + SSE-derived)

Phase 4 — US3 Human Checkpoint

  • T009 [US3] Build HumanCheckpointPanel.svelte (evidence + confirm/false-positive/inconclusive) @POST: dispose -> 044 human/decision; run resumes; disposition auditable
  • T010 [US3] L2 UX test for WAITING_FOR_HUMAN + resume Test: frontend/src/lib/components/scenario-run/__tests__/RunMonitorViews.test.ts (disposition dispatch) + SSE-driven resume in RunTimeline.test.ts/RunMonitorModel.test.ts

Phase 5 — US4 Final Result

  • T011 [US4] Build ScenarioResultView.svelte (summary, counts, provenance)
  • T012 [US4] L2 UX test for result + provenance render Test: frontend/src/lib/components/scenario-run/__tests__/RunMonitorViews.test.ts

Phase 6 — US5 History + Compare

  • T013 [US5] Build RunHistoryList.svelte (chronological list)
  • T014 [US5] Build RunComparison.svelte (per-step deltas + revision-diff warning)

Phase 6b — Global Run Operations Center (P0 #10)

  • T014b [P] L1 model test for Global Run Center in frontend/src/lib/models/__tests__/RunCenterModel.test.ts
  • T014c [P] Build GlobalRunCenterModel.svelte.ts + /dashboard-testing/runs route with filters (status/dashboard/env/trigger/owner) + "Waiting for me" view
  • T014d [P] Build WaitingForMeView.svelte (actionable human checkpoints inline)
  • T014e [P] L2 UX test for global center + waiting-for-me filter Test: frontend/src/lib/components/scenario-run/__tests__/WaitingForMeView.test.ts

Phase 6c — Run Configuration binding (P0 #12/#17)

  • [~] T014f [P] Bind RunConfigurationPanel.svelte to 044 start contract (release/baseline_set/toggles); mandatory steps not toggleable
  • T014g [P] L2 UX test for mandatory-step toggle protection Test: frontend/src/lib/components/scenario-run/__tests__/RunConfigurationPanel.test.ts

Phase 7 — Polish

  • T015 [P] Distinct labeling from 037/040 in all UI (copy + aria) Present in: RunConfigurationPanel, ScenarioResultView, GlobalRunCenter route header
  • T016 Run quickstart-equivalent monitor checks, full frontend tests, ATTN static audit and semantic rebuild.
  • T017 Prototype validation: every declared @UX_STATE is reachable via prototype/index.html.

Audit Follow-ups (2026-08-20)

  • T018 Extend the 044 start contract with typed release, baseline set and optional evidence toggles; remove params.launch_config transport and prove mandatory steps remain non-toggleable.
  • T019 Add monitor E2E evidence using real 044 outcomes and 047 queue entries after their runtime contracts are implemented.

Dependencies

Setup → US1; US2 depends on 044 SSE; US3 depends on US2; US4 depends on US2; US5 depends on US4.

#endregion ScenarioRunMonitor.Tasks


PROTOTYPE — State/Manifest

Source: prototype/manifest.md

#region ScenarioRunMonitor.PrototypeManifest [C:3] [TYPE ADR] [SEMANTICS prototype,manifest,scenario,run,monitor] @defgroup Prototype Interactive HTML prototype manifest for Scenario Run Monitor & Results.

Prototype Metadata

Factual audit 2026-08-20: prototype state coverage is not evidence that the 044 launch/event contract or 047 investigation handoff works end-to-end.

  • Feature: 045 Scenario Run Monitor & Results UX
  • Source contracts: ux_reference.md, contracts/modules.md
  • Screens represented: 4 (Config, Live Monitor, Result, History/Compare) + Investigation Queue handoff → 047
  • Total states: 7 (config, running, waiting_human, result, history, compare, disconnected/reconnect)
  • Accessibility: keyboard nav, focus-visible, aria-live, ≥44px, prefers-reduced-motion
  • Responsive: 375px, 900px

State Coverage

@UX_STATE Prototype State Reachable? Recovery
configuring config ✅ —
running running ✅ —
waiting_human waiting_human ✅ confirm/false-positive/inconclusive
failed result ✅ triage (047)
loaded history ✅ —
loaded / cross_revision compare ✅ revision-diff warning
disconnected disconnected/reconnect ✅ recover by run_id

Screen ↔ Story Traceability

Story Prototype Feature Intended acceptance coverage
US1 Launch config state env/revision/baseline/toggles
US2 Live running timeline + step rows events-only rendering
US3 Human WAITING FOR HUMAN panel resume from monitor
US4 Result result tab + provenance counts/failures/provenance
US5 Compare compare tab + history deltas + revision check
Failed result Queue handoff analyst explicitly opens 047 case; no auto-chat
#endregion ScenarioRunMonitor.PrototypeManifest

PROTOTYPE — Interactive HTML

Source: prototype/index.html

<!doctype html>

<html lang="ru"> <head> </head>
Superset Tools · BI testing
СценарииЗапускиАвтоматизацияКачествоРасследования 2 3 активных
Operations Center

Запуски сценариев

Ручные и автоматические проверки, evidence и результаты.

Требуют моей проверки: 1
Все статусы Выполняется Ожидает проверки Ошибка Все environments PREPROD PRODПрименить
Запуск Сценарий Environment Источник Статус
SR-1842
сегодня, 10:42
Комментарии по строкам
FI-0080
PREPROD Ручной Ожидает вас Открыть
SR-1841 XLSX reconciliation PREPROD Schedule Выполняется Открыть
SR-1839 Фильтры и метрики PREPROD Deploy FAIL Результат
← К запускам

SR-1842 Ожидает вашей проверки

Ручной ScenarioRun · шаг 5 из 6 · checkpoint создан 2 минуты назад

Проверка: комментарий виден у нужной строки

Система уже сохранила evidence. Проверьте, что «Проверено аналитиком» принадлежит строке ACME-184.

Screenshot · выбранная строка: ACME-184
Submitted comment: «Проверено аналитиком»
✓ Соответствует ✕ Не соответствует Неясно
← К запускам

SR-1839 FAIL

Фильтры и метрики · PREPROD · deploy trigger

Разобрать причину
Шаги
4 / 5
пройдено
Отклонение
+18
строк в XLSX
Target
rc-17
dashboard fingerprint pinned

Failure evidence

Шаг «Сравнить row set» · expected 124, actual 142.

Открыть XLSX Сравнить с SR-1828

Конфигурация запуска

PREPROD PROD Revision r18

Mandatory steps включены и недоступны для отключения.

Запустить сценарий
Scenario run выполняется

SR-1841 · live timeline

1
Open dashboard
progress 100%
Готово
2
Compare XLSX
progress 64%
Running
Cross-revision comparison. r17 и r18 имеют различающиеся execution snapshots.

Step deltas

Шаг SR-1828 SR-1839
compare-row-set PASS FAIL
Соединение с events потеряно.
Run восстанавливается по run_id; prose не используется для статусов.
Переподключиться
State: ConfigRunning WaitingResultHistoryCompare Disconnected
<script src="../../prototype-ui.js"></script> <script> function finish(out) { alert( "Решение «" + out + "» сохранено. Runner продолжает зависимые шаги.", ); setProtoState("history"); } protoState("history", (s) => [ "config", "running", "waiting_human", "result", "history", "compare", "disconnected", ].forEach((id) => document.getElementById(id).classList.toggle("hidden", id !== s), ), ); </script> </html>

================================================================================ FEATURE: 046-dashboard-scenario-automation Files: 14


SPEC — Feature Specification

Source: spec.md

#region ScenarioAutomation.Spec [C:3] [TYPE ADR] [SEMANTICS spec,requirements,scenario,automation,schedule,trigger,notification,operations] @BRIEF Scenario Automation & Operations: schedules and triggers for running scenarios automatically, notifications, concurrency policies, retention, and operational metrics. Reuses 037 trigger semantics in a common framework. @RELATION DEPENDS_ON -> [Doc.Adr.ADR0001] @RELATION DEPENDS_ON -> [Doc.Adr.ADR0005] @RELATION DEPENDS_ON -> [ScenarioExecution.Spec] @RELATION DEPENDS_ON -> [ScenarioRegistry.Spec] @RELATION DEPENDS_ON -> [SupersetBaselineEngine.Spec] @RATIONALE After manual runs, users want scheduled/triggered runs after each deploy, ETL, or release, plus notifications and retention. 037 already has trigger semantics for VerificationRun; 046 generalizes them for scenario runs. @REJECTED Reimplementing a separate scheduler — rejected; reuse the 037 trigger framework and existing APScheduler infrastructure. @REJECTED Writing scenario run results into the 037 baseline catalog or verification pipeline — rejected; scenario runs are a distinct entity (044). @RATIONALE SCAUTO-FR-013 delegated schedule management is exercisable through MCP tools after the 050 drift; scheduler dispatch, eligibility and deduplication remain deterministic server behavior for every actor.

Navigation (DSA Indexer keywords)

@SEMANTICS: spec, requirements, feature, scenario, automation, schedule, trigger, notification, operations

Feature Branch: 046-dashboard-scenario-automation Created: 2026-08-07 | Status: Not production-complete — factual audit pending remediation Input: "Provide Scenario Automation & Operations: schedule and trigger scenario runs (manual, on PREPROD deploy, on release created, after ETL, scheduled, API), notifications for completion/failure/blocked/stale, concurrency policies, retention, deduplication, and operational metrics — reusing the 037 trigger framework."

User Scenarios

Story 1 — Schedule and Trigger Scenario Runs (P1)

Why P1: Scenarios must run automatically, not only manually.

Independent Test: Configure a schedule and a deploy/release trigger for a scenario and verify runs are created on the trigger with the correct environment and revision.

Acceptance:

  1. Given a scenario When a schedule is configured Then runs are created per cron/daily schedule with the pinned eligible revision or the atomically activated current revision and environment.
  2. Given a PREPROD deploy or release-created event fires When matched Then a scenario run is triggered automatically (reuse 037 trigger framework).
  3. Given an ETL completion event When matched Then a scenario run is triggered.

Story 2 — Notifications (P2)

Why P2: Users must not sit on the page for a long run.

Independent Test: Trigger a completed and a failed run and verify the corresponding notification domain events are emitted.

Acceptance:

  1. Given a run completes When the run ends Then a "Scenario completed" notification domain event is emitted.
  2. Given a run fails or blocks When the run ends Then a "Scenario failed"/"Scenario blocked" event is emitted.
  3. Given a scenario revision contains a human step When automation is configured or dispatch preflight runs Then it is rejected as manual-run-only; no scheduled run and no false-PASS path exists.
  4. Given a scenario becomes stale When detected Then a "Scenario became stale" event is emitted.

Story 3 — Concurrency, Retention, Dedup (P2)

Why P2: Automation must not flood environments or accumulate unbounded runs.

Acceptance:

  1. Given multiple triggers fire for the same scenario When runs are created Then concurrency policies and deduplication prevent redundant parallel runs on the same environment+revision.
  2. Given a schedule overlaps a deployment window When evaluated Then the policy blocks or warning-gates the overlap.
  3. Given runs accumulate When retention runs Then old runs are pruned per policy (results/artifacts retained per retention rules).

Story 4 — Operational Metrics (P3)

Why P3: Ops needs visibility into automation health.

Acceptance:

  1. Given automation is active When metrics are queried Then scheduled runs, success rate, and trigger distribution are shown.
  2. Given a scheduled run fails repeatedly When detected Then an alerting signal is emitted (repeated flaky failure).

Story 5 — Manage Automation (P2) (#10/#20/#21)

Why P2: Schedules/triggers/policies must be manageable (create/update/disable), not just creatable.

Acceptance:

  1. Given automation is configured When the Automation Management UI opens Then schedules, trigger rules, and policies are listed and editable (create/update/delete/enable).
  2. Given an external system needs to trigger a run When it calls POST /scenarios/{id}/trigger Then a run starts (distinct from configuring a trigger rule).
  3. Given a PROD external trigger When invoked Then an ActionApprovalGate is required before dispatch.
  4. Given an authenticated external MCP client with automation scopes When it manages schedules, trigger rules or policies Then it receives the same bounded projections, validation, RBAC and gate outcomes as the Automation Management UI/API.

Edge & Failure Cases

# Scenario Expected Behavior Recovery
E1 Trigger event arrives while a run is active Dedup / queue per policy —
E2 Overlap with deployment window Block or warning-gate Reschedule
E3 PROD scheduled run Requires approval gate Approve / skip
E4 Scheduler down Missed runs flagged; no silent loss Manual run
E5 Retention prunes referenced artifacts Retain per policy; never break provenance —

Requirements

Functional

  • SCAUTO-FR-001: Scenario runs MUST be triggerable by manual, PREPROD deploy, release-created, ETL-completed, scheduled, and API triggers, reusing the 037 trigger framework. This applies only to revisions whose persisted RunnerPlan has manual_run_only=false.
  • SCAUTO-FR-002: Scheduled runs MUST pin the scenario revision and environment at trigger time. A newly saved candidate MUST NOT be adopted until 042 activation; a pinned rule MUST reject an ineligible revision.
  • SCAUTO-FR-003: The system MUST emit domain notification events: completed, failed, blocked, scenario-stale, repeated-flaky-failure. human-action-required is excluded because HumanCheckpoint scenarios are manual-run-only.
  • SCAUTO-FR-004: Concurrency and deduplication policies MUST prevent redundant parallel runs on the same environment+revision.
  • SCAUTO-FR-005: Schedule/deployment-window overlap MUST be blocked or warning-gated per policy.
  • SCAUTO-FR-006: Retention MUST prune old runs per policy while preserving provenance and referenced artifacts.
  • SCAUTO-FR-007: A PROD-classified automated run for an eligible revision (manual_run_only=false) MUST create an ActionApprovalGate before dispatch. A revision with manual_run_only=true MUST be rejected before idempotency lookup, ScenarioRun creation, ActionApprovalGate creation, notification, queue insertion or dispatcher CAS for every non-manual origin, including scheduled, deploy, release, ETL, API and background recovery. Approval MUST NOT convert a human checkpoint into an automated run.
  • SCAUTO-FR-008: Operational metrics MUST include scheduled run counts, success rate, trigger distribution, and repeated-failure alerts.
  • SCAUTO-FR-009: Scenario automation MUST NOT write into the 037 baseline catalog or release verification pipeline; scenario runs remain distinct.
  • SCAUTO-FR-010: Schedules, trigger rules, and policies MUST be fully manageable (CRUD + enable/disable) via an Automation Management UI.
  • SCAUTO-FR-011: An external API run trigger MUST exist as POST /scenarios/{id}/trigger, distinct from configuring a trigger rule; PROD requires an ActionApprovalGate.
  • SCAUTO-FR-012: Failed, blocked, stale and repeated-failure automation events MUST emit idempotent 036 InvestigationSignals with trigger/run provenance; 047 creates/updates the Queue item. They MUST NOT automatically start an agent conversation or action.
  • SCAUTO-FR-013: An analyst-opened case MAY let the agent create, edit, pause or resume schedules and trigger rules when delegated policy permits. Scheduler dispatch, eligibility, deduplication, capacity and trigger provenance MUST remain deterministic; non-delegated mutations require an inline ActionApprovalGate.
  • SCAUTO-FR-014: Automation management and policy decisions MUST be completed in persistent pages, agent cases or inline cards; modal/dialog interaction MUST NOT be required.
  • SCAUTO-FR-015: Scheduler semantics MUST be explicit (timezone, DST, misfire_grace_time, coalesce, max_instances, missed-execution policy, scheduler-restart handling).
  • SCAUTO-FR-016: Retention MUST use layered tiers; the analytics minimum history window (047) MUST be guaranteed independent of run retention.
  • SCAUTO-FR-017: MCP MUST expose curated operations to list, upsert and delete schedules and trigger rules; get/upsert automation policy; and get automation metrics. Each mutation MUST accept an idempotency key, explicit active revision and environment, reuse the 046 services, and return the same typed validation/conflict/approval outcomes as REST. Generic scheduler-job invocation is forbidden.
  • SCAUTO-FR-018: An MCP schedule mutation for a revision containing HumanSteps MUST be rejected before schedule, run, queue, notification or approval-gate mutation. A scheduled PROD due event for an eligible revision MUST create at most one durable gate before dispatch; denial or expiry leaves dispatch unstarted.

Key Entities

  • ScenarioSchedule: Cron/daily schedule binding a scenario, revision, environment, and policy.
  • ScenarioTriggerRule: Event-based trigger (deploy/release/ETL) mapping to a scenario run.
  • NotificationEvent: Domain event for automation outcomes (completed/failed/blocked/stale/flaky). human-action-required MUST NOT be emitted by automation because human-containing revisions are ineligible before run creation.
  • AutomationPolicy: Concurrency cap, deduplication window, overlap rule, retention, PROD gating.

Success Criteria

  • SC-001: A fixture schedule and each trigger type produce a scenario run with pinned revision+env in tests.
  • SC-002: 100% of required notification events are emitted on the matching run/state transitions.
  • SC-003: No redundant parallel runs on the same environment+revision under concurrency/dedup policy.
  • SC-004: Overlap with deployment windows is blocked or warning-gated.
  • SC-005: Zero writes to the 037 baseline catalog by automation.
  • SC-006: 100/100 automated start attempts for a manual_run_only=true revision across scheduled, deploy, release, ETL, API and background origins produce no ScenarioRun, gate, notification, queue or dispatcher side effect; 100/100 eligible PROD automated attempts create a gate before dispatch.

Clarifications

Session 2026-08-07

  • Q: New scheduler? → A: No. Reuse the 037 trigger framework + existing APScheduler.
  • Q: Can a scheduled run pause at HumanCheckpoint? → A: No. Human-containing revisions are manual-run-only and are rejected before any durable automation side effect. The PROD approval gate is an authorization gate for eligible automated revisions, never a substitute for HumanCheckpoint.
  • Q: Where do results go? → A: Scenario runs (044), never into 037 baseline/verification.

Implementation Status & MVP Debt (factual audit 2026-08-20)

Models, CRUD API, management UI and pure schedule/policy/retention/metrics helpers are present, but the automation workflow is not wired end-to-end.

  • [~] handle_trigger_event() returns policy-checked candidates and dispatch_trigger_event() passes each real server-owned event origin into 044 start_run. A persisted human revision fails manual-run-only pre-create with no run-side effect. HTTP/scheduler trigger paths persist only queued or server-gated pending_approval rows; separate 044 queued->running CAS dispatch is the sole initial adapter authority and excludes pending gates. No long-running event-subscriber proof is yet available.
  • [ ] notify() only appends to a caller-provided list and is not connected to run, stale or repeated failure lifecycle events.
  • [~] Scheduler startup reloads enabled ScenarioSchedule rows and registration forwards their timezone, misfire_grace_time, max_instances, and missed-execution policy (derived coalesce). Independent callback tests also prove the fixed 044 queued-dispatch and cancel-drain jobs register their exact IDs, five-second intervals, singleton/coalescing options, and contain database-edge failures; repeated callbacks preserve their exact terminal side-effect counts. This is not proof of cron scheduled-scenario firing, a live scheduler process, or subscriber composition. Existing scheduler baseline lint debt is separate and remains unclosed.
  • [~] Server-owned EnvironmentPolicy now sends manual/API/event/scheduled origins through the same 044 durable pending_approval gate and excludes pending rows from dispatcher execution; client flags and unknown environments cannot bypass it. A dedicated real APScheduler scheduled-PROD integration test is still required, alongside the remaining retention/runtime workflow tests, before feature closure.

Drift Amendment — MCP Interface (2026-08-24)

  • Automation management UI stays; MCP decision tools may drive the same CRUD under SCAUTO-FR-013 policy. Signals never auto-start client activity (pull-only).

Status (2026-09-02): done — реализовано в рамках 050: инструменты и гейты (specs/050-mcp-interface/tasks.md T012–T028 [x]), handoff-поверхность (050 T030–T033), демонтаж чата и сервиса agent/ (050 T040–T041, чекпоинты specs/WORKSTATE-043-047.md).

#endregion ScenarioAutomation.Spec


UX REFERENCE — Interaction Narrative

Source: ux_reference.md

#region ScenarioAutomation.UxReference [C:3] [TYPE ADR] [SEMANTICS ux,reference,scenario,automation] @BRIEF UX interaction reference for Scenario Automation & Operations (046).

Feature Branch: 046-dashboard-scenario-automation | Created: 2026-08-07

1. User Persona & Context

  • User: BI analyst / ops engineer configuring automated scenario runs.
  • Goal: Run scenarios on a schedule or after deploy/ETL, get notified, keep environments bounded.
  • Context: Browser; scenario detail → Automation tab.

2. Happy Path

Analyst configures "Daily at 07:00" + "Every PREPROD deployment" for a scenario, sets a concurrency cap and PROD gate, enables notifications. After a deploy, the scenario auto-runs; a notification is emitted on completion.

3. Screens & States

Screen: Automation Config

  • Layout: Triggers section (checkboxes: Every PREPROD deployment / Daily at 07:00 / ETL completed), policy fields (concurrency, dedup, overlap rule, retention, PROD gate), notifications toggles.
  • @UX_STATE: idle, saving, saved, error, overlap_warning.
  • @UX_RECOVERY: overlap → block/warn; 403 → no confirm.

4. Error Experience

  • Overlap with deployment window → warning-gated.
  • PROD scheduled run → approval gate.
  • Scheduler down → missed runs flagged.

5. Tone & Voice

  • Style: Concise, technical. Terminology: schedule/trigger/policy/notification.

Edge & Failure Matrix (feed to prototype)

NET, VAL, AUTH, CONF, 429, 5XX, EMPTY, A11Y, RESP.

#endregion ScenarioAutomation.UxReference


CHECKLISTS — Requirements Quality — requirements.md

Source: checklists/requirements.md

Requirements Checklist: Scenario Automation & Operations (046)

Purpose: Verify SCAUTO-FR-001..009 completeness. | Created: 2026-08-07

Factual audit 2026-08-20: [x] requires current production evidence, [~] means partial code exists, [ ] means missing integration or proof. Human checkpoints are manual-run-only.

Schedule & Trigger (FR-001/002)

  • CHK001 Runs triggerable by manual, PREPROD deploy, release-created, ETL, scheduled, API only when persisted manual_run_only=false
  • CHK002 Reuses 037 trigger framework + APScheduler (no parallel scheduler)
  • CHK003 Scheduled/triggered runs pin revision + env at trigger time

Notifications (FR-003)

  • CHK004 Domain events: completed, failed, blocked, scenario-stale, repeated-flaky-failure

Policy (FR-004/005/006/007)

  • CHK005 Concurrency/dedup prevent redundant parallel runs (same env+revision)
  • CHK006 Schedule/deployment-window overlap blocked or warning-gated
  • CHK007 Retention prunes while preserving provenance/artifacts
  • CHK008 Eligible PROD scheduled runs gated (036); human-containing revisions rejected before any run/gate/queue/notification/dispatch side effect

Metrics & Boundary (FR-008/009)

  • [~] CHK009 Operational metrics: scheduled counts, success rate, trigger distribution, repeated-failure alerts
  • CHK010 Zero writes to 037 baseline catalog / verification pipeline

Success Criteria

  • CHK011 SC-001..006 verified

UX DECISIONS — Final Choices

Source: contracts/ux/decisions.md

#region ScenarioAutomation.Ux.Decisions [C:3] [TYPE ADR] [SEMANTICS scenario,automation,ux,decisions] @BRIEF Final UX decisions for Scenario Automation & Operations (046). @RELATION DEPENDS_ON -> [ScenarioAutomation.Spec]

Decision 1 — Trigger-first config

Automation config centers on triggers (deploy/daily/release/ETL) with policy fields below.

Decision 2 — Overlap/prod awareness

Overlap with deployment windows → block/warn banner; PROD scheduled runs → approval gate.

Decision 3 — Notification toggles

Grouped notification toggles (completed, failed, blocked, human-action, stale, repeated-flaky).

Decision 4 — Ops metrics visible

Operational metrics (scheduled runs, success rate, trigger distribution, repeated-failure alert) surfaced in the same surface. #endregion ScenarioAutomation.Ux.Decisions


PLAN — Implementation Plan

Source: plan.md

Implementation Plan: Scenario Automation & Operations

Branch: 046-dashboard-scenario-automation | Date: 2026-08-07 | Spec: spec.md | Status: Not production-complete — factual audit pending remediation

Implementation audit, 2026-08-20: management and pure helpers exist; event dispatch, lifecycle notifications, startup schedule reload and persisted APScheduler semantics remain open.

Summary

Automate scenario runs: schedules and triggers (manual, PREPROD deploy, release-created, ETL, scheduled, API) reusing the 037 trigger framework + APScheduler, typed notification events, concurrency/dedup/overlap/retention policies, PROD gating, and operational metrics. Scenario results stay in 044 — never written to 037 baseline/verification.

Technical Context

Language/Version: Python 3.13+ (backend) Primary Dependencies: FastAPI, APScheduler (existing), 037 trigger framework, 044 run API Storage: new tables scenario_schedules, scenario_trigger_rules, notification_events Testing: pytest (trigger mapping, policy enforcement, dedup, retention, notification) Performance Goals: trigger → run start < 1s; policy evaluation < 100ms Constraints: reuse 037 triggers/APScheduler; PROD gated; no 037 baseline writes; retention preserves provenance Scale: hundreds of schedules, per-env concurrency bounded

Constitution Check

Principle Result
I. Semantic Contract First PASS
II. Decision Memory PASS — research R1-R4
V. RBAC Enforcement PASS — automation scopes, PROD gate
VII. Test-Driven C3+ PASS — trigger/policy/dedup tests first
VIII. Attention-Optimized PASS

Project Structure

specs/046-dashboard-scenario-automation/
├── spec.md / data-model.md / research.md / plan.md / tasks.md / traceability.md / quickstart.md / ux_reference.md
├── checklists/requirements.md
├── contracts/modules.md, contracts/openapi.yaml, contracts/ux/
└── prototype/index.html + manifest.md

backend/src/models/scenario_automation.py
backend/src/services/dashboard_testing/automation/ (schedule.py, trigger.py, policy.py, notify.py, retention.py, metrics.py)
backend/src/api/routes/dashboard_testing/scenario_automation.py

Delivery Phases

  1. Automation models + migration + fixtures.
  2. Schedule (APScheduler binding).
  3. Trigger rules + event mapping (037 framework reuse).
  4. Policy (concurrency/dedup/overlap/PROD gate).
  5. Notifications.
  6. Retention + operational metrics.
  7. REST routes + regression gates.

Traceability

traceability.md maps Story → model → operationId → contract → task → test.

Cross-Spec Boundary

  • Triggers runs via 044; reads scenario+revision from 042; reuses 037 trigger framework + scheduler.
  • Never writes to 037 baseline/verification.

Complexity Tracking

No exception planned. Automation is bounded C3-C4; policy/trigger decomposed.


RESEARCH — Technical Decisions

Source: research.md

Scenario Automation & Operations — Phase 0/1 Research (046)

Branch: 046-dashboard-scenario-automation | Date: 2026-08-07 | Spec: spec.md

R1. Reuse 037 trigger framework + APScheduler

Decision: Generalize the 037 trigger semantics (release_create/scheduled) into a shared trigger framework; use existing APScheduler for cron schedules.

Rationale: 037 already has deploy/release/scheduled trigger machinery; a parallel scheduler would duplicate ~200 lines and diverge.

Alternatives: new scheduler (rejected); per-feature trigger code (rejected: drift).

Impact: ScenarioSchedule, ScenarioTriggerRule bound to 037 hooks.

R2. Notification domain events

Decision: Emit typed NotificationEvent domain events for completed/failed/blocked/human/stale/flaky. Channel delivery (email/WS) left to infrastructure.

Rationale: Decouples domain semantics from delivery; 040/036 already have event patterns.

Impact: notification event contracts.

R3. Concurrency/dedup/retention policies

Decision: AutomationPolicy with max_concurrent_per_env, dedup window, overlap rule, retention, PROD gating, repeated-failure alerting.

Impact: policy enforcement on run creation.

R4. Boundary — never touch 037 baseline

Decision: Scenario automation writes runs to 044 only; zero writes to 037 baseline catalog or verification pipeline.

Rationale: Scenario runs are a distinct entity (040 also forbids baseline writes).

Impact: enforcement + test.

Contracts & API

  • contracts/modules.md — Automation.Schedule, Automation.Trigger, Automation.Notify, Automation.ApplyPolicy, Automation.Retention.
  • OpenAPI: POST /scenario-schedules, POST /scenario-trigger-rules, GET /automation/metrics, GET /notifications.

Constitution Check

Principle Result
I. Semantic Contract First PASS
II. Decision Memory PASS — R1-R4
V. RBAC PASS — automation scopes, PROD gate
VII. Test-Driven C3+ PASS — trigger/policy/dedup tests first
VIII. Attention-Optimized PASS

DATA MODEL — Entities & Relations

Source: data-model.md

#region ScenarioAutomation.DataModel [C:4] [TYPE ADR] [SEMANTICS data-model,scenario,automation,schedule,trigger,policy] @BRIEF Schedule, trigger-rule, notification, and automation-policy models for 046. @RELATION DEPENDS_ON -> [ScenarioAutomation.Research] @RATIONALE Automation binds a scenario+revision to schedules/triggers with explicit policies; reuses 037 trigger semantics rather than a parallel scheduler. @REJECTED Reimplementing scheduler; writing scenario results into 037 baseline/verification.

ScenarioSchedule

Fields: id, scenario_id, revision_policy (current|pinned), revision_id (required when pinned), environment_id, cron_expr, timezone, missed_execution_policy, enabled, policy_id, created_by, created_at. Reuses existing APScheduler.

Schedule/trigger creation and every current-policy dispatch preflight require automation_eligible=true and an activated 042 current revision. A newly saved candidate is never silently adopted by a current schedule; it becomes eligible only through 042 atomic activation. A revision containing any 044 human step is manual_run_only; the service rejects the rule/dispatch with AUTOMATION_INELIGIBLE_HUMAN_STEP, rather than skipping the step or emitting a false PASS.

Scheduler semantics (#22)

Explicit APScheduler configuration: timezone (per schedule, IANA), DST policy (cron still fires on wall-clock), misfire_grace_time (default 300s), and max_instances (1 — no overlapping instances of the same schedule). coalesce is derived, never independently configured: skip ignores missed occurrences, run_latest sets coalesce=true, and queue_all sets coalesce=false. Scheduler restart applies that derived policy.

ScenarioTriggerRule

Fields: id, scenario_id, environment_id, trigger (deploy_to_preprod | release_created | etl_completed | api), revision_policy, revision_id (required when pinned), enabled, policy_id. Mapped onto the 037 trigger framework (release_create/scheduled hooks).

NotificationEvent

Fields: id, type (completed|failed|blocked|scenario_stale|repeated_flaky_failure), scenario_id, run_id?, severity, payload, emitted_at, investigation_signal_id?, investigation_queue_item_id?. Channel delivery is infrastructure-owned; qualifying attention events emit the canonical 036 InvestigationSignal; 047 creates/updates the Queue item from it rather than auto-starting agent work. human_action_required is not an automation event because human-step revisions are manual-run-only.

AutomationPolicy

Fields: id, name, enabled, workload_class, max_concurrent_per_env, dedup_window_seconds, overlap_rule (block|warn), retention_days, prod_gate_required (bool), on_repeated_failure (alert|disable). Applies to schedule/trigger runs through the global 044 ExecutionCapacityManager; it cannot reserve capacity owned by another workload class.

An agent may create, modify, pause or resume schedules/trigger rules under delegated policy and may investigate a repeated failure after an analyst opens its queue item. Scheduler dispatch, eligibility checks, deduplication, capacity, retention and trigger provenance stay deterministic. A policy-gated automation mutation is represented by an inline ActionApprovalGate, never a modal.

Retention tiers (#23)

Layered retention independent of the analytics minimum history window: run metadata 180d, triage/audit 365d (or policy), step metrics 90d, heavy artifacts 30d, screenshots 30d, raw VLM 7d. The analytics minimum history window (047) is guaranteed even when aggressive run retention prunes.

RunBinding

A triggered run pins scenario_id + revision_id + immutable Verification Program/content hash + environment_id + target snapshot + execution principal + server-owned trigger source (from 044). Concurrency bucket (for capacity) may be environment+workload class; dedup identity is canonical_execution_request_hash, or for an event trigger (source_type, source_event_id, scenario_id). External API triggers require an Idempotency-Key; same key/hash returns the same run, while a changed canonical request returns 409. Automation cannot supply or rewrite runtime SQL/DSL/program content.

#endregion ScenarioAutomation.DataModel


CONTRACTS — Module & Function Contracts

Source: contracts/modules.md

#region ScenarioAutomation.Modules [C:4] [TYPE ADR] [SEMANTICS scenario,automation,contracts,modules,schedule,trigger] @BRIEF Module contracts for Scenario Automation & Operations (046). @defgroup ScenarioAutomation Schedules, triggers, notifications, policies for scenario runs. @RELATION DEPENDS_ON -> [ScenarioExecution.Modules] @RELATION DEPENDS_ON -> [BaselineEngine.TriggerFramework] @RELATION DEPENDS_ON -> [Services.Scheduler] @RATIONALE Automation reuses 037 triggers + APScheduler; scenario results stay in 044, never in 037 baseline. @REJECTED Parallel scheduler; writing scenario results into 037 baseline/verification.

#region Automation.Schedule [C:4] [TYPE Function] [SEMANTICS scenario,automation,schedule,cron]

@ingroup ScenarioAutomation

@BRIEF Create/enable a cron schedule binding scenario+revision+env to APScheduler.

@PRE caller has automation scope; policy validated.

@POST schedule registered; runs created on cron with pinned revision+env.

@SIDE_EFFECT APScheduler job; DB write.

def upsert_schedule(db, schedule): ...

#endregion Automation.Schedule

#region Automation.Trigger [C:4] [TYPE Function] [SEMANTICS scenario,automation,trigger,event]

@ingroup ScenarioAutomation

@BRIEF Map an event (deploy/release/etl) to scenario trigger rules and start runs.

@PRE event classified by 037 trigger framework; rules enabled.

@POST matched rules start runs (via 044) with policy applied.

@SIDE_EFFECT run creation; dedup/overlap evaluation.

@INVARIANT PROD runs gated; no redundant parallel runs per policy.

@TEST_EDGE event->run; dedup suppresses second; overlap blocks/warns; PROD gated.

async def handle_trigger_event(db, event): ...

#endregion Automation.Trigger

#region Automation.Notify [C:3] [TYPE Function] [SEMANTICS scenario,automation,notify,event]

@ingroup ScenarioAutomation

@BRIEF Emit a NotificationEvent for a run/state outcome.

@POST emits typed event (completed/failed/blocked/human/stale/flaky).

def notify(db, type, scenario_id, run_id, severity): ...

#endregion Automation.Notify

#region Automation.ApplyPolicy [C:4] [TYPE Function] [SEMANTICS scenario,automation,policy,concurrency]

@ingroup ScenarioAutomation

@BRIEF Enforce concurrency/dedup/overlap/PROD-gate before starting an automated run.

@PRE policy loaded; run candidate.

@POST permits or suppresses run; PROD requires gate.

@INVARIANT never exceeds max_concurrent_per_env; dedup window honored.

def apply_policy(db, policy, run_candidate): ...

#endregion Automation.ApplyPolicy

#region Automation.Retention [C:3] [TYPE Function] [SEMANTICS scenario,automation,retention,prune]

@ingroup ScenarioAutomation

@BRIEF Prune old runs per retention policy while preserving provenance/artifacts.

@POST prunes beyond retention; referenced artifacts retained per policy.

def run_retention(db, policy): ...

#endregion Automation.Retention

#region Automation.Metrics [C:3] [TYPE Function] [SEMANTICS scenario,automation,metrics,ops]

@ingroup ScenarioAutomation

@BRIEF Report scheduled-run counts, success rate, trigger distribution, repeated-failure alerts.

def automation_metrics(db): ...

#endregion Automation.Metrics

#region Automation.McpSchedule [C:5] [TYPE Function] [SEMANTICS scenario,automation,mcp,schedule,policy]

@ingroup ScenarioAutomation

@BRIEF Expose schedule, trigger-rule, policy and automation-metric management through curated MCP operations.

@PRE Caller has the same automation scopes as the REST surface; each write supplies an idempotency key and an explicit active revision, environment and policy input.

@POST List/get operations return bounded server-owned projections. Upsert/delete operations return the durable schedule/rule/policy projection or a typed approval_required envelope; no duplicate scheduler job is registered on replay.

@SIDE_EFFECT Schedule/rule/policy writes may update APScheduler registration; a PROD due event may create one approval gate but never dispatch before its decision.

@DATA_CONTRACT McpScheduleRequest -> ScenarioSchedule | ScenarioTriggerRule | AutomationPolicy | AutomationMetricsProjection

@INVARIANT A revision containing HumanSteps is rejected before schedule, run, queue, notification or approval-gate mutation. MCP wrappers reuse Automation.Schedule/Trigger/ApplyPolicy and never implement a second scheduler.

@RELATION CALLS -> [Automation.Schedule]

@RELATION CALLS -> [Automation.Trigger]

@RELATION CALLS -> [Automation.ApplyPolicy]

@RELATION CALLS -> [Automation.Metrics]

@RATIONALE The 050 external-client boundary must have the same policy-governed automation capability as the 046 UI/API; hiding it behind a browser surface breaks MCP parity.

@REJECTED A generic MCP payload that can invoke arbitrary scheduler jobs or bypass revision eligibility, deduplication, environment policy or PROD gates.

@TEST_EDGE human-step revision -> no side effect; replay -> same schedule ID; stale revision -> typed conflict; PROD due event -> one pending gate before dispatch.

def manage_automation_over_mcp(db, request, actor): ...

#endregion Automation.McpSchedule

#endregion ScenarioAutomation.Modules


OPENAPI — REST/Event API Contract

Source: contracts/openapi.yaml

openapi: 3.1.0 info: title: Scenario Automation & Operations API version: 0.1.0 description: Schedules, triggers, notifications, policies, and metrics for scenario runs (046). paths: /api/scenario-schedules: post: operationId: automation.schedule.create summary: Create/enable a cron schedule binding a scenario to an eligible current/pinned revision and environment security: [{ bearerAuth: [] }] requestBody: required: true content: application/json: schema: $ref: "#/components/schemas/ScenarioSchedule" responses: { "201": { description: Schedule created }, "422": { description: AUTOMATION_INELIGIBLE_HUMAN_STEP, candidate/not-activated revision, or preflight failure } } get: operationId: automation.schedule.list summary: List schedules security: [{ bearerAuth: [] }] responses: { "200": { description: "Schedule[]" } } /api/scenario-schedules/{schedule_id}: patch: operationId: automation.schedule.update summary: Update a schedule (cron, revision policy, enabled) security: [{ bearerAuth: [] }] parameters: [{ name: schedule_id, in: path, required: true, schema: { type: string } }] responses: { "200": { description: Updated } } delete: operationId: automation.schedule.delete summary: Delete a schedule security: [{ bearerAuth: [] }] parameters: [{ name: schedule_id, in: path, required: true, schema: { type: string } }] responses: { "204": { description: Deleted } } /api/scenario-trigger-rules: post: operationId: automation.trigger.create summary: Create a trigger rule using an eligible current/pinned revision (deploy/release/ETL) security: [{ bearerAuth: [] }] requestBody: required: true content: application/json: schema: $ref: "#/components/schemas/ScenarioTriggerRule" responses: { "201": { description: Rule created }, "422": { description: AUTOMATION_INELIGIBLE_HUMAN_STEP, candidate/not-activated revision, or preflight failure } } get: operationId: automation.trigger.list summary: List trigger rules security: [{ bearerAuth: [] }] responses: { "200": { description: "TriggerRule[]" } } /api/scenario-trigger-rules/{rule_id}: patch: operationId: automation.trigger.update summary: Update a trigger rule security: [{ bearerAuth: [] }] parameters: [{ name: rule_id, in: path, required: true, schema: { type: string } }] responses: { "200": { description: Updated } } delete: operationId: automation.trigger.delete summary: Delete a trigger rule security: [{ bearerAuth: [] }] parameters: [{ name: rule_id, in: path, required: true, schema: { type: string } }] responses: { "204": { description: Deleted } } /api/automation-policies: post: operationId: automation.policy.create summary: Create an automation policy security: [{ bearerAuth: [] }] requestBody: { required: true, content: { application/json: { schema: { $ref: "#/components/schemas/AutomationPolicy" } } } } responses: { "201": { description: Policy created } } get: operationId: automation.policy.list summary: List automation policies security: [{ bearerAuth: [] }] responses: { "200": { description: "AutomationPolicy[]" } } /api/automation-policies/{policy_id}: patch: operationId: automation.policy.update summary: Update or enable/disable an automation policy security: [{ bearerAuth: [] }] parameters: [{ name: policy_id, in: path, required: true, schema: { type: string } }] requestBody: { required: true, content: { application/json: { schema: { $ref: "#/components/schemas/AutomationPolicy" } } } } responses: { "200": { description: Updated } } delete: operationId: automation.policy.delete summary: Delete an unused automation policy security: [{ bearerAuth: [] }] parameters: [{ name: policy_id, in: path, required: true, schema: { type: string } }] responses: { "204": { description: Deleted } } /api/scenarios/{scenario_id}/trigger: post: operationId: automation.triggerApi summary: External API run trigger (distinct from creating a trigger rule) security: [{ bearerAuth: [] }] parameters: - { name: scenario_id, in: path, required: true, schema: { type: string } } - { name: Idempotency-Key, in: header, required: true, schema: { type: string } } requestBody: required: true content: application/json: schema: type: object required: [environment_id] properties: environment_id: { type: string } revision_id: { type: string, nullable: true } params: { type: object } baseline_bindings: { type: object } target_reference: { type: object } execution_toggles: { type: object, description: "Optional diagnostics only" } responses: "202": { description: "ScenarioRun accepted: queued or pending_approval; response contains run_id, status, approval_gate_id?" } "403": { description: Permission denied; an authorized PROD request creates pending_approval rather than 403 } /api/automation/metrics: get: operationId: automation.metrics summary: Operational metrics security: [{ bearerAuth: [] }] responses: "200": description: Metrics (scheduled counts, success rate, trigger distribution, alerts) /api/notifications: get: operationId: automation.notifications summary: List notification domain events security: [{ bearerAuth: [] }] responses: "200": description: Notification events content: application/json: schema: { type: array, items: { $ref: "#/components/schemas/NotificationEvent" } } components: securitySchemes: bearerAuth: { type: http, scheme: bearer } schemas: ScenarioSchedule: type: object required: [scenario_id, environment_id, cron_expr] properties: scenario_id: { type: string } environment_id: { type: string } cron_expr: { type: string } revision_policy: { type: string, enum: [current, pinned], default: current } revision_id: { type: string, nullable: true, description: "Required when revision_policy=pinned; must be an activated, automation-eligible revision" } timezone: { type: string, description: "IANA timezone" } policy_id: { type: string } missed_execution_policy: { type: string, enum: [skip, run_latest, queue_all], default: skip } enabled: { type: boolean, default: true } ScenarioTriggerRule: type: object required: [scenario_id, environment_id, trigger] properties: scenario_id: { type: string } environment_id: { type: string } trigger: { type: string, enum: [deploy_to_preprod, release_created, etl_completed, api] } revision_policy: { type: string, enum: [current, pinned], default: current } revision_id: { type: string, nullable: true, description: "Required when pinned; must be an activated, automation-eligible revision" } policy_id: { type: string } enabled: { type: boolean, default: true } AutomationPolicy: type: object required: [name, enabled, workload_class, max_concurrent_per_env] properties: id: { type: string, nullable: true } name: { type: string } enabled: { type: boolean } workload_class: { type: string, enum: [scenario_smoke, scenario_regression] } max_concurrent_per_env: { type: integer, minimum: 1 } dedup_window_seconds: { type: integer, minimum: 0 } overlap_rule: { type: string, enum: [block, warn] } retention_days: { type: integer, minimum: 1 } prod_gate_required: { type: boolean } on_repeated_failure: { type: string, enum: [alert, disable] } NotificationEvent: type: object properties: id: { type: string } type: { type: string, enum: [completed, failed, blocked, scenario_stale, repeated_flaky_failure] } scenario_id: { type: string } run_id: { type: string, nullable: true } severity: { type: string }


QUICKSTART — Dev Onboarding

Source: quickstart.md

Quickstart: Scenario Automation & Operations (046)

Factual audit 2026-08-20: pending verification checklist only; scheduler/event/notification integration is not yet production-complete.

Prereqs

  • 044 run API, 042 registry, DB migrated (scenario_automation tables)
  • APScheduler infrastructure (037) available

Commands

cd backend && source .venv/bin/activate
alembic upgrade head
python -m pytest -v tests/services/dashboard_testing/automation/
python -m pytest -v tests/api/test_scenario_automation.py
python -m ruff check src/services/dashboard_testing/automation/

Exit Gates

  • Each trigger type (deploy/release/ETL/schedule/api) starts an eligible run with pinned revision+env; a manual_run_only=true revision is rejected before any durable side effect
  • Notification events emitted for completed/failed/blocked/stale/flaky; no human-action-required event is emitted by automation
  • Concurrency/dedup prevents redundant parallel runs; overlap blocked/warned
  • Retention prunes preserving provenance/artifacts
  • Eligible PROD automated runs create ActionApprovalGate before dispatch; human-containing revisions are rejected before gate/run/queue/notification creation
  • Zero writes to 037 baseline catalog
  • ruff clean; prototype states covered

TRACEABILITY — Requirements Matrix

Source: traceability.md

Traceability: Scenario Automation & Operations (046)

Story Requirement Model API operationId Contract Task Test Actual status / gap
US1 Schedule SCAUTO-FR-001/002/007 ScenarioSchedule, ScenarioTriggerRule, ActionApprovalGate automation.schedule, automation.trigger Automation.Schedule, Automation.Trigger, Execution.EnvironmentPolicy, Execution.Runner.QueuedDispatch T003-T005, T016, T019 test_trigger, test_scenario_runner, test_scenario_runs_api, test_scenario_automation_api, test_scenario_manual_run_only, test_scenario_queued_dispatch, test_scenario_scheduler_callbacks [~] Any persisted human-containing revision is rejected before idempotency, ScenarioRun, gate, notification, queue or dispatcher CAS for scheduled/deploy/release/ETL/API/background origins; manual origin remains eligible. Eligible PROD automation creates ActionApprovalGate before dispatch. Dedicated real scheduled-PROD callback integration, cron firing, a live scheduler process and long-running subscribers remain open.
US2 Notify SCAUTO-FR-003 NotificationEvent automation.notify Automation.Notify T006-T007 test_notify [ ] helper has no run/staleness lifecycle caller.
US3 Policy SCAUTO-FR-004/005/006/007 AutomationPolicy automation.policy Automation.ApplyPolicy, Automation.Retention T008-T010 test_policy [~] pure helpers exist; persisted runtime enforcement is unproven.
US4 Metrics SCAUTO-FR-008 — automation.metrics Automation.Metrics T011 test_metrics [~] aggregate helper exists; source events/runs are not fully wired.
Boundary SCAUTO-FR-009 — — — T013 test_policy [~] routes/RBAC exist; scheduled PROD runtime proof is open.

N/A: Registry (042), Editor (043), Execution (044), Monitor (045), Analytics (047).


TASKS — Implementation Tasks

Source: tasks.md

#region ScenarioAutomation.Tasks [C:3] [TYPE ADR] [SEMANTICS tasks,scenario,automation,implementation] @BRIEF Ordered TDD backlog for Scenario Automation & Operations (046). Tests FIRST.

Prerequisites: plan.md, spec.md; contracts/modules.md, traceability.md.

Format: - [ ] T### [P] [USx] Description with exact file path

Factual audit 2026-08-20: [x] means code plus relevant evidence; [~] means partial implementation; [ ] means absent integration or unperformed verification.

Phase 1 — Setup

  • T001 Create ScenarioSchedule, ScenarioTriggerRule, NotificationEvent models in backend/src/models/scenario_automation.py
  • T002 [P] Create canonical fixtures in specs/046-dashboard-scenario-automation/fixtures/

Phase 2 — US1 Schedule & Trigger

  • [~] T003 [US1] Write failing schedule/trigger tests (trigger semantics contract coverage: backend/tests/services/dashboard_testing/registry/test_scenario_automation_trigger.py — filename differs from plan; coverage closes the task: event->run, release_create->run, ETL->run, disabled/mismatch skip, capacity/PROG/dedup gates)
  • [~] T004 [US1] Implement upsert_schedule schedule semantics in automation/schedule.py
  • [~] T005 [US1] Implement handle_trigger_event policy-bound event mapping in automation/trigger.py @POST: matched rules start runs via 044; pinned revision+env @TEST_EDGE: event->run; release_create->run; ETL->run

Phase 3 — US2 Notifications

  • [~] T006 [US2] Write notification tests in backend/tests/services/dashboard_testing/registry/test_scenario_automation.py
  • [~] T007 [US2] Implement notify in automation/notify.py

Phase 4 — US3 Concurrency/Dedup/Retention

  • [~] T008 [US3] Write failing policy/retention tests in backend/tests/services/dashboard_testing/registry/test_scenario_automation_policy.py
  • [~] T009 [US3] Implement apply_policy (concurrency/dedup/overlap/PROD gate) in backend/src/services/dashboard_testing/automation/policy.py @INVARIANT: never exceeds max_concurrent_per_env; dedup window honored
  • [~] T010 [US3] Implement layered run_retention in automation/retention.py

Phase 5 — US4 Operational Metrics

  • [~] T011 [US4] Implement automation_metrics in automation/metrics.py (verified by registry/test_scenario_automation_semantics.py Metrics section + API test)

Phase 6 — API + Polish

  • T012 Add routes in backend/src/api/routes/dashboard_testing/scenario_automation.py (schedules, trigger-rules, metrics)
  • T013 [P] RBAC automation scopes + PROD gate tests (verified by backend/tests/api/test_scenario_automation_api.py Rbac section: missing_manage_scope 403, missing_trigger_scope 403, missing_prod_scope 403, direct trigger + idempotency 202/409)

Phase 6b — Management UI + API trigger + semantics (P0 #10/#20/#21/#22/#23)

  • T013b [P] Automation CRUD/API trigger Schedule/rule/policy CRUD + direct POST /scenarios/{id}/trigger (idempotency + PROD gate) implemented and covered by test_scenario_automation_api.py. Only the UI direct-trigger button is not yet exposed (tracked in T013c panel extension).
  • T013c [P] Automation Management UI mounted: AutomationPanel.svelte (schedules/triggers/ policies list + edit) at frontend/src/routes/dashboard-testing/automation/ (route +page.svelte + route test automation_page.ux.test.ts; run center links to it)
  • [~] T013d [P] APScheduler semantics config (timezone/DST/misfire_grace_time/coalesce/ max_instances/missed-policy) in automation/schedule.py (verified by registry/test_scenario_automation_semantics.py Schedule section)
  • T013e [P] Layered retention tiers (run metadata/triage/step metrics/artifacts/screenshots/ raw VLM) in automation/retention.py (verified by registry/test_scenario_automation_semantics.py Retention section + API GET /retention)

Phase 7 — Polish

  • T014 Run scoped backend tests, ruff, frontend tests and build; record current command-level evidence.
  • T015 Prototype validation: every declared @UX_STATE is reachable via prototype/index.html.

Audit Follow-ups (2026-08-20)

  • [~] T016 Wire deploy/release/ETL/scheduled events through policy/dedup into 044 start_run; verify pinning, PROD gates and idempotency at the production boundary. The dispatcher now forwards server-owned event provenance before creation, and human revisions reject manual-run-only without a side effect. HTTP/scheduler callbacks persist only queued or server-gated pending_approval rows; the separate 044 queued->running CAS owns initial adapter dispatch and excludes pending gates. Fixed 044 queue/cancel due callbacks now have unit proof for registration options, database-edge containment, and repeat-tick exact side-effect idempotency (test_scenario_scheduler_callbacks.py); cron scheduled-scenario due firing and subscriber integration remain open. Add a dedicated real scheduled-PROD integration test proving server policy produces pending_approval plus one durable gate before dispatcher CAS; this remaining coverage is not a gate bypass.
  • T017 Emit persisted lifecycle notification events and canonical InvestigationSignals for required run/stale/repeated-failure transitions; prove no auto-started agent work.
  • [~] T018 Load enabled ScenarioSchedule rows at scheduler startup and map timezone, misfire grace, coalesce and max instances into APScheduler. Startup reload/registration code exists; restart and cron scheduled-scenario due-job behavior still need independent proof. The fixed 044 queue/cancel maintenance callbacks are separately unit-proven, not live-process proof.

Dependencies

Setup → US1; US2 depends on US1; US3 depends on US1; US4 depends on US3. Requires 044 run API + 037 trigger framework.

#endregion ScenarioAutomation.Tasks


PROTOTYPE — State/Manifest

Source: prototype/manifest.md

#region ScenarioAutomation.PrototypeManifest [C:3] [TYPE ADR] [SEMANTICS prototype,manifest,scenario,automation] @defgroup Prototype Interactive HTML prototype manifest for Scenario Automation & Operations.

Prototype Metadata

Factual audit 2026-08-20: prototype coverage is not evidence of event dispatch, scheduler restart/reload, notification persistence, or PROD runtime behavior.

  • Feature: 046 Scenario Automation & Operations
  • Source contracts: ux_reference.md, contracts/modules.md
  • Screens represented: 1 (Automation Config)
  • Total states: 5 (idle, saved, overlap_warning, prod_gate, metrics)
  • Accessibility: keyboard nav, focus-visible, aria-live, ≥44px, prefers-reduced-motion
  • Responsive: 375px, 900px

State Coverage

@UX_STATE Prototype State Reachable? Recovery
idle idle ✅ —
saved saved ✅ —
overlap_warning overlap_warning ✅ block/warn
prod_gate prod_gate ✅ approve (036)
metrics metrics ✅ —

Screen ↔ Story Traceability

Story Prototype Feature Intended acceptance coverage
US1 Schedule trigger checkboxes deploy/daily/release/ETL
US2 Notify notify toggle + metrics flaky alert notification events
US3 Policy policy grid + overlap/prod banners concurrency/dedup/retention/gate
US4 Metrics metrics card ops visibility
#endregion ScenarioAutomation.PrototypeManifest

PROTOTYPE — Interactive HTML

Source: prototype/index.html

<!doctype html>

<html lang="ru"> <head> </head>
Superset Tools · BI testing
СценарииЗапускиАвтоматизацияКачество Новое правило
Automation management

Расписания и правила запуска

Аналитик сам настраивает автоматические проверки. В список попадают только полностью автоматизируемые revisions.

Расписания

Сценарий Когда Revision Environment Статус
XLSX reconciliation Каждый день · 09:00 Europe/Simferopol Current PREPROD Включено
Фильтры и метрики Пн–Пт · 08:30 r12 pinned PREPROD Включено

Trigger rules

Сценарий Событие Policy
XLSX reconciliation Deploy to PREPROD Smoke / max 2 Изменить
Фильтры и метрики ETL completed Regression / max 1 Изменить
← К правилам

Новое расписание

Сначала проверим, может ли выбранная revision выполняться без человека.

Что запускать

Сценарий XLSX reconciliation · fully automated Комментарии по строкам · manual-run-only Revision policy Current revision Pin r18РасписаниеEnvironment PREPROD PROD
Отмена Сохранить правило
Правило включено

XLSX reconciliation будет запускаться ежедневно в 09:00

Dedup использует canonical execution request, а capacity контролируется отдельно для PREPROD.

К правилам
Overlap policy blocked a duplicate run.
PREPROD capacity 2/2; событие сохранено без нового запуска.
К правилам
PROD approval required

Daily XLSX reconciliation

Blast radius: 4 dashboards · revision r18 · concurrency 1.

Отклонить Одобрить
Schedules
12
Success rate
96%
Deduplicated
8
State: IdleNewSaved Overlap PROD gateMetrics
<script src="../../prototype-ui.js"></script> <script> function checkScenario(v) { const e = document.getElementById("eligibility"); const b = document.getElementById("save"); if (v === "comments") { e.className = "notice danger"; e.innerHTML = "Нельзя автоматизировать
Revision содержит RunHumanCheckpoint. Создайте полностью автоматизируемую revision."; b.disabled = true; } else { e.className = "notice info"; e.innerHTML = "Можно автоматизировать
Нет HumanSteps; registry action contracts и target policy валидны."; b.disabled = false; } } protoState("idle", (s) => [ "idle", "new", "saved", "overlap_warning", "prod_gate", "metrics", ].forEach((id) => document.getElementById(id).classList.toggle("hidden", id !== s), ), ); </script> </html>

================================================================================ FEATURE: 047-dashboard-scenario-analytics Files: 14


SPEC — Feature Specification

Source: spec.md

#region ScenarioAnalytics.Spec [C:3] [TYPE ADR] [SEMANTICS spec,requirements,scenario,triage,flakiness,analytics,health] @BRIEF Investigation Queue and agent-led case workspaces over deterministic scenario analytics, so failures become evidence-led remediation work rather than isolated forms. @RELATION DEPENDS_ON -> [Doc.Adr.ADR0001] @RELATION DEPENDS_ON -> [Doc.Adr.ADR0005] @RELATION DEPENDS_ON -> [ScenarioExecution.Spec] @RELATION DEPENDS_ON -> [ScenarioRegistry.Spec] @RELATION DEPENDS_ON -> [ScenarioRunMonitor.Spec] @RATIONALE 036 supplies durable agent work and 044 supplies immutable run evidence, but neither provides an analyst-controlled queue and long-lived investigation case. Health and failure identity remain deterministic inputs to that work. @REJECTED Ending at pass/fail/inconclusive — rejected because operational workflow begins there. @REJECTED Opening an agent chat for every failure — rejected because queue deduplication and analyst intent are needed to avoid noise. @REJECTED Server-rendered case chat as the only conversation medium — superseded 2026-08-24: durable case threads continue in external MCP clients per specs/050-mcp-interface/spec.md; evidence snapshot, AgentAction timeline and linked runs remain server-owned records.

Navigation (DSA Indexer keywords)

@SEMANTICS: spec, requirements, feature, scenario, triage, flakiness, analytics, health, trend

Feature Branch: 047-dashboard-scenario-analytics Created: 2026-08-07 | Status: Not production-complete — factual audit pending remediation Input: "Provide an Investigation Queue and agentic case workspace for failed/stale/load/automation evidence, backed by deterministic flakiness, health, trends and recurring-failure analytics."

User Scenarios

Story 1 — Open an Agent-Led Investigation (P1)

Why P1: A failed test must enter a deliberate evidence-led investigation, not a classification form.

Independent Test: Feed a failed run into the queue, open it explicitly, and verify the case creates an agent thread with immutable evidence and an audited disposition projection.

Acceptance:

  1. Given a failed/inconclusive/blocked run When its deterministic evidence is produced Then an Investigation Queue item is created or updated; no agent chat or tool run starts automatically.
  2. Given an analyst selects “Investigate with agent” When the item opens Then a durable case has chat, source/evidence snapshots, tool timeline and linked AgentRuns.
  3. Given the case reaches a disposition When it is persisted Then a versioned, audited TriageRecord projection updates the run/step view without changing historical run truth.

Story 2 — Detect Flakiness and Compute Health (P1)

Why P1: Distinguish a broken dashboard from a broken test.

Independent Test: Feed a run history with repeated step failures and verify flakiness and contextual health are calculated deterministically before any agent decision.

Acceptance:

  1. Given a step fails intermittently When analytics compute Then it is classified flaky with a flakiness ratio.
  2. Given run history When health is derived Then 30d success rate, flaky ratio, infra-failure ratio, and most-unstable-step are shown.
  3. Given a scenario health changes When thresholds are crossed Then the registry health badge updates (042), while an untriaged failure remains attention with unknown cause.

Story 3 — Analyze Trends and Recurring Failures (P2)

Why P2: Repeated failures and trends need surfacing.

Independent Test: Feed multiple runs and verify trends and recurring-failure grouping render.

Acceptance:

  1. Given run history When trends are rendered Then success-rate trend and failure-classification distribution are shown.
  2. Given the same step fails repeatedly When aggregated Then recurring failures are grouped with a count and first/last occurrence.
  3. Given a failure recurs after its FailureEpisode was resolved When aggregated Then a new alertable episode opens; only duplicates inside an active episode are suppressed.

Edge & Failure Cases

# Scenario Expected Behavior Recovery
E1 No run history Health unknown, no flaky signal —
E2 Case disposition concurrent edit 409 conflict panel Reload
E3 Case/agent access denied 403 permission_denied Contact admin
E4 Infra failure spike Classified infra, not product regression Investigate infra

Requirements

Functional

  • SCAN-FR-001: Failed/inconclusive/blocked runs, staleness, baseline immutability, load and repeated automation findings MUST emit the canonical 036 InvestigationSignal; 047 MUST create/update a deduplicated Investigation Queue item from that idempotent envelope. The queue MUST NOT auto-start an agent chat, AgentRun or tool action.
  • SCAN-FR-002: An analyst MUST be able to open a queue item into a durable InvestigationCase with chat, immutable evidence snapshot, hypotheses, AgentAction timeline, linked AgentRuns, approvals, verification and final disposition. The agent may execute only delegated/policy-authorized actions.
  • SCAN-FR-003: Triage MUST be a CAS/audited compact projection of a case disposition; the historical RunResult stays immutable truth.
  • SCAN-FR-004: The system MUST detect flakiness per step using same environment class, logical_step_id, compatibility_family, baseline family and comparable context. It requires a post-failure pass and two result transitions; infra/cancelled/inconclusive outcomes are excluded from the pass/fail denominator.
  • SCAN-FR-005: Health MUST be contextual and computed separately for product, scenario-test, infrastructure and agent-evaluation health plus overall attention. Agent verdict variance MUST NOT be classified as deterministic flakiness.
  • SCAN-FR-006: Health MUST feed the 042 registry health badge.
  • SCAN-FR-007: The system MUST render trends and group recurring failures by immutable compatibility-scoped fingerprint, never including triage classification.
  • SCAN-FR-008: A matching failure inside an active FailureEpisode is deduplicated; a matching failure after resolution MUST open a new alertable FailureEpisode and Queue item.
  • SCAN-FR-009: Object-level result/evidence access and agent-action authority MUST enforce ACL separately from scenario-result:view and scenario-result:triage.
  • SCAN-FR-010: Health/trends/recurring MUST aggregate scenario history, not a single run.
  • SCAN-FR-011: Investigation, approval, conflict recovery and closure MUST use persistent case workspaces and inline cards; modal/dialog interaction MUST NOT be required.
  • SCAN-FR-012: A case may become resolved only with verification evidence and completed reconciliation. accepted requires an analyst-recorded accepted-risk/won't-fix/duplicate rationale. A new matching active signal MUST reopen the work (or create a new case for a new FailureEpisode).
  • SCAN-FR-013: Health and flakiness eligibility MUST use the exact 044 AnalyticsContextKey; grouping by environment or revision alone is forbidden.
  • SCAN-FR-014: Analytics MUST distinguish deterministic failure, agent-evaluation disagreement, low confidence, model instability, infrastructure failure, scenario-definition defect and product regression. Recurring fingerprints remain immutable raw evidence fingerprints and MUST NOT include a human/agent classification.

Key Entities

  • InvestigationQueueItem: Deduplicated attention item with source/evidence context; opening it is explicit.
  • InvestigationCase: Agent-led durable workstream with chat, tools, approvals, verification and closure.
  • TriageRecord: Compact audited projection of case disposition for a run/step.
  • FlakinessSignal: Per-step intermittent-failure metric (ratio, window).
  • ScenarioHealth: 30d success rate, flaky ratio, infra ratio, most unstable step.
  • RecurringFailureGroup: Grouped identical failures with count, first/last occurrence, current triage.

Success Criteria

  • SC-001: 100% of qualifying fixture events create/update a deduplicated queue item without starting agent work; an opened case has immutable evidence and an audited disposition projection.
  • SC-002: Flaky steps are detected and health derived from run history with deterministic output.
  • SC-003: Health badge in 042 updates when thresholds cross.
  • SC-004: Trends and recurring failures render; recurrence after a resolved episode creates a new alertable episode.
  • SC-005: RBAC enforces view vs triage.

Clarifications

Session 2026-08-07

  • Q: New data? → A: Reuses 044/037/040/041 evidence; adds Queue, Case, AgentAction projection and deterministic analytics; feeds 042 health.
  • Q: Does a case change graph/result? → A: No. A case may create a separately immutable revision or policy-bound action, but it never rewrites historical run truth.
  • Q: Does every failure open chat? → A: No. It enters a deduplicated Investigation Queue; the analyst explicitly opens the case.

Implementation Status & MVP Debt (factual audit 2026-08-20)

Queue/case, flakiness/health/trend/recurring primitives, API routes and UI routes are present. The required evidence-led investigation workflow is not production-complete.

  • [~] 044 terminal failed/blocked/inconclusive runs now call canonical queue ingestion with an idempotent immutable run/artifact signal; no case/chat/AgentRun/remediation is started. Staleness, baseline, load and automation producers plus compatibility-scoped recurring-episode classification remain outside this bounded signal path.
  • [ ] InvestigationCase lacks the required immutable evidence snapshot and durable chat/linked- AgentRun representation.
  • [ ] set_disposition() unconditionally marks a case resolved; it does not enforce verification evidence/reconciliation or accepted-risk rationale required by SCAN-FR-012.
  • [~] Deterministic aggregation/UI primitives exist but need end-to-end history and signal-ingestion proof; T013 remains partial.

Drift Amendment — MCP Interface (2026-08-24)

  • InvestigationCase chat (open T016 item) is re-scoped: the server persists evidence snapshot, hypotheses, tool timeline and linked runs; the conversational thread lives in the analyst's external MCP client. Queue ingestion and "no auto-start" semantics are unchanged.

Status (2026-09-02): done — реализовано в рамках 050: инструменты и гейты (specs/050-mcp-interface/tasks.md T012–T028 [x]), handoff-поверхность (050 T030–T033), демонтаж чата и сервиса agent/ (050 T040–T041, чекпоинты specs/WORKSTATE-043-047.md).

#endregion ScenarioAnalytics.Spec


UX REFERENCE — Interaction Narrative

Source: ux_reference.md

#region ScenarioAnalytics.UxReference [C:3] [TYPE ADR] [SEMANTICS ux,reference,scenario,triage,analytics] @BRIEF UX interaction reference for Investigation Queue and agent-led case workspaces (047).

Feature Branch: 047-dashboard-scenario-analytics | Created: 2026-08-07

1. User Persona & Context

  • User: BI analyst / quality engineer investigating scenario failures.
  • Goal: Open selected attention items into evidence-led agent investigations, understand flakiness vs real regression, and monitor scenario health.
  • Context: Browser; Investigation Queue in 042/045/047, then persistent case workspace.

2. Happy Path

A run fails and updates an Investigation Queue item. The analyst explicitly opens it with the agent, which receives evidence and may run delegated tools. The analyst records the final case disposition; a compact triage projection is updated without changing run truth. A recurring failure after resolution opens a new alertable episode and a new queue item.

3. Screens & States

Screen: Investigation Queue

  • Layout: severity, source, recurrence, evidence summary, suggested next action, “Investigate with agent”.
  • @UX_STATE: loading, loaded, empty, error.

Screen: Investigation Case

  • Layout: evidence, target/RLS provenance, agent chat, hypotheses, tool timeline, inline ActionApprovalGate cards, verification and final disposition.
  • @UX_STATE: open, investigating, awaiting_approval, awaiting_external_change, verifying, resolved, accepted, reopened.
  • @UX_RECOVERY: 409 → persistent conflict panel; 403 → permission state.

Screen: Scenario Health / Analytics

  • Layout: health card (success rate, flaky ratio, infra ratio, most unstable step), trends chart, recurring failures list.
  • @UX_STATE: idle, loading, loaded, empty, error.

4. Error Experience

  • Concurrent case disposition → 409 persistent conflict panel.
  • RBAC denied → permission_denied.

5. Tone & Voice

  • Style: Concise, evidence-led. Terminology: queue, case, evidence, action, disposition, flakiness, health.

Edge & Failure Matrix (feed to prototype)

NET, VAL, AUTH, CONF, 429, 5XX, EMPTY, A11Y, RESP.

#endregion ScenarioAnalytics.UxReference


CHECKLISTS — Requirements Quality — requirements.md

Source: checklists/requirements.md

Requirements Checklist: Investigation Queue & Scenario Analytics (047)

Purpose: Verify SCAN-FR-001..011 completeness. | Created: 2026-08-07

Factual audit 2026-08-20: [x] requires current production evidence, [~] means partial code exists, [ ] means missing integration or proof.

Queue and Case (FR-001..003/009/011)

  • CHK001 Qualifying events create/update a deduplicated queue item and never auto-start agent work
  • CHK002 Analyst explicitly opens Queue item into Case with evidence/chat/actions
  • CHK003 Case disposition projects triage through CAS/audit; never alters graph/result/baseline
  • CHK004 Object ACL plus scenario-result:view vs :triage enforced

Flakiness & Health (FR-004..006)

  • [~] CHK005 Per-step flaky detection requires post-failure pass and two transitions; excluded outcomes are not in denominator
  • CHK006 Contextual health: product/test/infra/overall plus confidence
  • CHK007 Health feeds 042 registry badge on threshold cross
  • [~] CHK008 Success-rate trend + failure-classification distribution
  • [~] CHK009 Recurring failures grouped (count, first/last occurrence)
  • CHK010 Matching recurrence after resolved episode opens a new alertable FailureEpisode

Success Criteria

  • CHK011 SC-001..005 verified

UX DECISIONS — Final Choices

Source: contracts/ux/decisions.md

#region ScenarioAnalytics.Ux.Decisions [C:3] [TYPE ADR] [SEMANTICS scenario,analytics,ux,decisions] @BRIEF Final UX decisions for Investigation Queue & Scenario Analytics (047). @RELATION DEPENDS_ON -> [ScenarioAnalytics.Spec]

Decision 1 — Analyst-opened case, orthogonal disposition

Qualifying evidence enters Investigation Queue; it never auto-starts agent work. The analyst opens a persistent case with chat/evidence/actions, and its compact triage projection never alters graph, result truth or baselines.

Decision 2 — Health feeds registry

Scenario health (success rate, flaky ratio, infra ratio, most unstable step) derived and feeds the 042 registry health badge.

Decision 3 — Recurring no re-alert

Recurring failures are grouped; only active episodes are deduplicated, while recurrence after resolution is newly alertable.

Decision 4 — Object ACL and disposition

scenario-result:view grants only object-authorized read; scenario-result:triage permits case disposition. Agent actions separately obey delegated policy and tool ACL.

Decision 5 — No modal workflow

Queue, case, conflict recovery and approvals are persistent work surfaces or inline cards; no modal/dialog is required to complete investigation. #endregion ScenarioAnalytics.Ux.Decisions


PLAN — Implementation Plan

Source: plan.md

Implementation Plan: Investigation Queue & Scenario Analytics

Branch: 047-dashboard-scenario-analytics | Date: 2026-08-07 | Spec: spec.md | Status: Not production-complete — factual audit pending remediation

Implementation audit, 2026-08-20: analytics primitives/UI exist; canonical signal ingestion, immutable case evidence and conformant case closure remain open.

Summary

Investigation Queue and analyst-opened agentic cases over scenario evidence, backed by compact audited triage projection, deterministic flakiness detection, health feeding 042, trends, and recurring-failure grouping.

Technical Context

Language/Version: Python 3.13+ (backend), TypeScript + Svelte 5 runes (frontend) Primary Dependencies: FastAPI, SQLAlchemy; 044 run/step results, 042 registry health Storage: queue/case/action/triage projection tables; derived analytics (flakiness, health, trends) computed, not stored as truth Testing: pytest (queue/case, flakiness, health, trends), vitest (L1/L2 for queue/case UI) Frontend Architecture: model-first .svelte.ts for persistent queue/case/health views Performance Goals: health/flakiness derivation < 500ms for fixture histories; trends bounded Constraints: case/triage never alters graph/result/baseline; analyst explicitly opens agent work; RBAC/object ACL; derived health feeds 042 Scale: hundreds of runs per scenario; windows up to 90d

Constitution Check

Principle Result
I. Semantic Contract First PASS
II. Decision Memory PASS — research R1-R3
V. RBAC PASS — scenario-result:view / :triage
VII. Test-Driven C3+ PASS — flakiness/health/trend tests first
VIII. Attention-Optimized PASS

Project Structure

specs/047-dashboard-scenario-analytics/
├── spec.md / data-model.md / research.md / plan.md / tasks.md / traceability.md / quickstart.md / ux_reference.md
├── checklists/requirements.md
├── contracts/modules.md, contracts/openapi.yaml, contracts/ux/
└── prototype/index.html + manifest.md

backend/src/models/scenario_investigation.py
backend/src/services/dashboard_testing/analytics/ (investigation.py, flakiness.py, health.py, trends.py, recurring.py)
backend/src/api/routes/dashboard_testing/scenario_analytics.py
frontend/src/lib/models/InvestigationQueueModel.svelte.ts, InvestigationCaseModel.svelte.ts
frontend/src/lib/components/scenario-analytics/ (QueueList, CaseWorkspace, AgentActionTimeline, ScenarioHealthCard, TrendsChart, RecurringFailuresList)

Delivery Phases

  1. Queue/Case/AgentAction/TriageRecord models + migration + fixtures.
  2. Queue projection, explicit case open, disposition CAS + RBAC/object ACL.
  3. Flakiness detection.
  4. Health derivation → 042 badge.
  5. Trends + recurring failures.
  6. Frontend model/components.
  7. Polish, regression gates.

Traceability

traceability.md maps Story → model → operationId → contract → task → test.

Cross-Spec Boundary

  • Consumes 044 run/step results; feeds 042 health badge.
  • Case/analytics never rewrite 037 baselines or historical run truth.
  • Queue is surfaced in 042 Registry and 045 Results; the persistent case workspace is 047.

Complexity Tracking

No exception planned. Queue/case orchestration is C4/C5; deterministic derivations stay decomposed.


RESEARCH — Technical Decisions

Source: research.md

Investigation Queue & Scenario Analytics — Phase 0/1 Research (047)

Branch: 047-dashboard-scenario-analytics | Date: 2026-08-07 | Spec: spec.md

R1. Case disposition as orthogonal projection

Decision: An analyst-opened InvestigationCase is the primary workstream; TriageRecord is its compact orthogonal projection over 044 run/step results. Neither alters graph, result truth or baselines.

Rationale: Triage is analyst judgment; changing result truth on triage would corrupt reproducibility.

Alternatives: triage mutates result (rejected); no triage (rejected: graveyard).

Impact: InvestigationQueueItem, InvestigationCase, AgentAction and TriageRecord projection; object ACL plus RBAC view vs disposition.

R2. Flakiness + health derived, not stored-truth

Decision: FlakinessSignal and ScenarioHealth are derived from run history (window), feeding the 042 health badge. Deterministic.

Rationale: Derived avoids stale truth; matches 042 health derivation.

Impact: derivation service + thresholds.

R3. Recurring-failure grouping

Decision: Group identical failures by immutable fingerprint (logical_step_id + error_code + normalized_error_signature + assertion_kind + affected_ref) with counts. Classification is triage metadata, not fingerprint input. A new matching occurrence after an episode is resolved opens a new episode and alerts.

Impact: recurring group + FailureEpisode + active-episode deduplication.

Contracts & API

  • contracts/modules.md — Analytics.Triage, Analytics.Classify, Analytics.Flakiness, Analytics.Health, Analytics.Trends, Analytics.Recurring.
  • OpenAPI: queue list/open-case/disposition, GET /scenarios/{id}/health, GET /scenarios/{id}/trends, GET /scenarios/{id}/recurring-failures.

Constitution Check

Principle Result
I. Semantic Contract First PASS
II. Decision Memory PASS — R1-R3
V. RBAC PASS — view vs triage
VII. Test-Driven C3+ PASS — flakiness/health/trend tests first
VIII. Attention-Optimized PASS

DATA MODEL — Entities & Relations

Source: data-model.md

#region ScenarioAnalytics.DataModel [C:5] [TYPE ADR] [SEMANTICS data-model,scenario,investigation,queue,agent,analytics,health] @BRIEF Investigation Queue/Case, compact triage projection, and deterministic analytics models for 047. @RELATION DEPENDS_ON -> [AgentInvestigation.Cases] @RATIONALE Failed runs need an evidence-led agentic workstream, while historical execution truth and calculated analytics must remain independently reproducible. @REJECTED Treating triage as a standalone form — rejected because analysts investigate through an agent thread with evidence and tools, not isolated classification fields. @REJECTED Letting agent judgment rewrite RunResult, health calculations, baseline truth or recurring identity — rejected because those are deterministic historical facts.

InvestigationQueueItem — attention, not automatic chat

Queue entries are generated only by consuming the canonical 036 InvestigationSignal from failed/inconclusive/blocked ScenarioRuns, staleness signals, baseline immutability violations, load circuit-breaker/consistency findings and repeated automation failures. Fields: id, source_type, source_id, scenario_id?, run_id?, logical_step_id?, severity, fingerprint?, active_episode_id?, evidence_summary, target_snapshot, execution_principal_fingerprint?, suggested_next_action, state, count, first_seen_at, last_seen_at, case_id?. The signal idempotency identity is preserved, so a producer cannot manufacture duplicate Queue work by retry.

State: new | acknowledged | case_opened | suppressed | resolved. A matching occurrence updates one item only inside an active FailureEpisode; a matching occurrence after its resolution creates a new item. Queue creation does not start AgentRun, tool calls or a chat.

InvestigationCase and AgentAction — agentic workstream

An analyst explicitly opens a queue item into InvestigationCase { id, queue_item_id, status, source_snapshot, evidence_snapshot, owner_actor_id, agent_thread_id, opened_at, resolved_at?, final_disposition?, resolution_summary? }.

Status: open | investigating | awaiting_approval | awaiting_external_change | verifying | resolved | accepted | reopened.

The case owns chat, hypotheses, tool timeline and linked AgentRuns. Every tool call is the shared 036 AgentAction record: canonical inputs, risk/policy decision, target/key scope, side-effect identity, pre/postcondition evidence, cleanup/reconciliation, approvals and actor/agent/tool provenance. The agent may autonomously read, diagnose, run policy-permitted diagnostics, mutate authorized fixture data, and save validated scenario revisions. It never bypasses ACL, deterministic validation, ActionRegistry mutation contracts, capacity, immutable revision creation, or a required ActionApprovalGate. Failed cleanup prevents case resolution.

Closure policy: resolved requires non-empty verification evidence proving the stated acceptance condition and no unresolved cleanup/reconciliation action. accepted requires an analyst-confirmed accepted-risk/won't-fix/duplicate rationale; it is not a claim that the product passed. awaiting_external_change is used while an external remediation or cleanup is pending. A new matching active signal after a terminal disposition reopens the case (or opens a new case if the episode is new); it never silently remains resolved.

TriageRecord — compact audited case projection

RunResult is immutable historical truth. TriageRecord projects the current case disposition onto run_id + logical_step_id? for Registry, Run Monitor and analytics: investigation_status (new|investigating|resolved), classification (product_regression|data_regression|baseline_stale|scenario_bug|infrastructure_failure|not_confirmed), resolution (fixed|accepted_risk|duplicate|wont_fix|false_positive), comment, case_id, decision_version, actor_id, timestamps.

The projection is versioned/CAS and append-only audited. It is never a free-standing modal form, never changes a RunResult, scenario graph or baseline, and is derived/updated only by an authenticated case decision.

FlakinessSignal — deterministic rules

Window: last N eligible runs (default 30). Eligibility per deterministic step: same environment class, same logical_step_id, same compatibility_family, same baseline family, and comparable target/principal context. Infrastructure outcomes, cancelled and inconclusive runs are excluded from the deterministic pass/fail denominator. AgentEvaluation verdict variation is never counted as deterministic flakiness. It is separately aggregated as evaluation disagreement, low-confidence rate and model instability keyed by AgentEvaluationSpec/model/prompt version. A deterministic step is flaky only when pass and fail are both observed, a pass occurs after the first failure, at least two pass/fail state transitions occur, and failure ratio is within (X,Y) (default 5%–50%). A one-way PASS→FAIL change is a regression signal, never flaky.

Fields: scenario_id, logical_step_id, compatibility_family, context_key, window, eligible_runs, total_runs, failures, flaky_runs, ratio, is_flaky, excluded_outcomes.

ScenarioHealth — deterministic, contextual

AnalyticsContextKey is exactly the server-derived 044 SHA-256 over environment_class + compatibility_family + baseline_family + dashboard_release_id + dashboard_fingerprint + dataset_lineage_fingerprint + execution_principal_fingerprint. Fields: scenario_id, environment_class, context_key (=AnalyticsContextKey), window, product_health, scenario_test_health, infrastructure_health, agent_evaluation_health, overall_attention, success_rate, flaky_ratio, infra_failure_ratio, inconclusive_ratio, agent_disagreement_ratio, low_confidence_ratio, model_instability_ratio, most_unstable_step, generated_at, confidence.

Product health consumes product/data regressions; scenario-test health consumes scenario bugs, stale baselines and deterministic flaky signals; infrastructure health consumes typed infrastructure outcomes; agent-evaluation health consumes disagreement/low-confidence/model-instability. Untriaged failures raise overall_attention=attention with cause unknown, but do not fabricate a classification. 042 displays only overall_attention, or unknown if analytics is unavailable.

RecurringFailureGroup and FailureEpisode — immutable identity

Group identity is (scenario_id, compatibility_family, logical_step_id, error_code, normalized_error_signature, assertion_kind, affected_ref). Triage data is excluded. RecurringFailureGroup has counts and first/last occurrence; FailureEpisode { id, group_id, opened_at, resolved_at?, resolution? } is active iff resolved_at is null. A matching occurrence in an active episode increments it and suppresses a duplicate queue item; a matching occurrence after resolution opens a new alertable episode and a new queue item.

Boundary

047 owns queue, case, triage projection and analytics. It consumes 044 execution, 042 staleness, 037 baseline and 040 load evidence; it may request actions through 036 but never replaces their deterministic execution or ownership.

#endregion ScenarioAnalytics.DataModel


CONTRACTS — Module & Function Contracts

Source: contracts/modules.md

#region ScenarioAnalytics.Modules [C:4] [TYPE ADR] [SEMANTICS scenario,analytics,contracts,modules,triage,flakiness] @BRIEF Module contracts for Investigation Queue/Case and deterministic scenario analytics (047). @defgroup ScenarioAnalytics Queue/case, compact disposition projection, flakiness, health, trends for scenario runs. @RELATION DEPENDS_ON -> [ScenarioExecution.Modules] @RELATION DEPENDS_ON -> [ScenarioRegistry.Modules] @RATIONALE Queue/case makes red runs actionable through an analyst-controlled agent workstream, while deterministic analytics distinguish product/data regressions from flaky/infra failures; never mutate graph/result/baseline. @REJECTED Ending at pass/fail; triage altering result truth.

#region Analytics.OpenCase [C:4] [TYPE Function] [SEMANTICS scenario,analytics,queue,case,agent]

@ingroup ScenarioAnalytics

@BRIEF Explicitly open a queued signal into an InvestigationCase and agent thread.

@PRE caller has source-object/evidence access; queue item is active.

@POST one durable case returned/created; no RunResult or baseline changed.

@SIDE_EFFECT DB write; audit; agent thread provision only.

@INVARIANT a queue event alone never starts the case or tool action.

@TEST_EDGE duplicate-open->existing case; denied->403; suppressed->409.

def open_case(db, queue_item_id, actor): ...

#endregion Analytics.OpenCase

#region Analytics.Disposition [C:4] [TYPE Function] [SEMANTICS scenario,analytics,case,triage,status]

@ingroup ScenarioAnalytics

@BRIEF Persist an analyst-confirmed case disposition as compact TriageRecord projection.

@PRE caller has scenario-result:triage; case decision version matches.

@POST TriageRecord updated/audited; RunResult unchanged (immutable truth).

@SIDE_EFFECT DB write; audit.

@INVARIANT case disposition is orthogonal metadata; does not change run result or baselines.

@TEST_EDGE concurrent->409; denied->403; audit recorded; run stays FAILED.

def set_disposition(db, case_id, decision_version, disposition, actor): ...

#endregion Analytics.Disposition

#region Analytics.Flakiness [C:4] [TYPE Function] [SEMANTICS scenario,analytics,flakiness,detect]

@ingroup ScenarioAnalytics

@BRIEF Detect per-step flakiness from run history.

@POST returns FlakinessSignal (ratio, is_flaky); deterministic.

def detect_flakiness(db, scenario_id): ...

#endregion Analytics.Flakiness

#region Analytics.Health [C:4] [TYPE Function] [SEMANTICS scenario,analytics,health,derive]

@ingroup ScenarioAnalytics

@BRIEF Derive scenario health and feed 042 badge.

@POST returns ScenarioHealth; updates 042 registry badge on threshold cross.

@SIDE_EFFECT updates 042 health.

def derive_health(db, scenario_id): ...

#endregion Analytics.Health

#region Analytics.Trends [C:3] [TYPE Function] [SEMANTICS scenario,analytics,trends,series]

@ingroup ScenarioAnalytics

@BRIEF Render success-rate trend and failure-classification distribution.

@POST returns time series + distribution.

def trends(db, scenario_id): ...

#endregion Analytics.Trends

#region Analytics.Recurring [C:4] [TYPE Function] [SEMANTICS scenario,analytics,recurring,group]

@ingroup ScenarioAnalytics

@BRIEF Group recurring failures by fingerprint; deduplicate only an active FailureEpisode and alert a recurrence after resolution.

@POST returns RecurringFailureGroup[] with counts and first/last occurrence.

def recurring(db, scenario_id): ...

#endregion Analytics.Recurring

#endregion ScenarioAnalytics.Modules


OPENAPI — REST/Event API Contract

Source: contracts/openapi.yaml

openapi: 3.1.0 info: title: Investigation Queue & Scenario Analytics API version: 0.3.0 description: Analyst-opened agentic investigation cases over deterministic scenario health, trends and recurring failures. paths: /api/internal/investigation-signals: post: operationId: investigations.ingestSignal summary: Idempotently ingest a deterministic producer signal into the Investigation Queue description: Internal producer route. Ingestion creates/updates Queue/Episode only and never starts an agent chat or AgentRun. security: [{ serviceAndUser: [] }] requestBody: required: true content: application/json: schema: { $ref: "#/components/schemas/InvestigationSignal" } responses: "202": { description: Signal accepted; Queue item created or updated deterministically, content: { application/json: { schema: { $ref: "#/components/schemas/InvestigationQueueItem" } } } } "409": { description: Signal identity was reused with non-canonical content } /api/investigation-queue: get: operationId: investigations.listQueue summary: List deduplicated attention items; listing never starts an agent run security: [{ bearerAuth: [] }] parameters: - { name: state, in: query, schema: { type: string, enum: [new, acknowledged, case_opened, suppressed, resolved] } } - { name: scenario_id, in: query, schema: { type: string } } - { name: severity, in: query, schema: { type: string, enum: [info, warning, critical] } } responses: "200": description: InvestigationQueueItem[] content: { application/json: { schema: { type: array, items: { $ref: "#/components/schemas/InvestigationQueueItem" } } } } /api/investigation-queue/{queue_item_id}/open-case: post: operationId: investigations.openCase summary: Explicitly open a persistent InvestigationCase and its agent thread security: [{ bearerAuth: [] }] parameters: [{ name: queue_item_id, in: path, required: true, schema: { type: string, format: uuid } }] responses: "201": description: Case opened content: { application/json: { schema: { $ref: "#/components/schemas/InvestigationCase" } } } "200": description: Existing case for the active queue item content: { application/json: { schema: { $ref: "#/components/schemas/InvestigationCase" } } } "403": { description: Requires source-object and evidence access } "409": { description: Queue item is suppressed/resolved or has incompatible case state } /api/investigation-cases/{case_id}: get: operationId: investigations.getCase summary: Get persistent case, evidence, compact triage and action timeline security: [{ bearerAuth: [] }] parameters: [{ name: case_id, in: path, required: true, schema: { type: string, format: uuid } }] responses: "200": { description: InvestigationCase, content: { application/json: { schema: { $ref: "#/components/schemas/InvestigationCase" } } } } "403": { description: Requires source-object and evidence access } "404": { description: Not found } /api/investigation-cases/{case_id}/disposition: post: operationId: investigations.setDisposition summary: Record an analyst-confirmed case disposition and update compact triage projection security: [{ bearerAuth: [] }] parameters: [{ name: case_id, in: path, required: true, schema: { type: string, format: uuid } }] requestBody: required: true content: application/json: schema: { $ref: "#/components/schemas/CaseDispositionRequest" } responses: "200": { description: Updated case and audited TriageRecord projection, content: { application/json: { schema: { $ref: "#/components/schemas/InvestigationCase" } } } } "403": { description: Requires scenario-result:triage and object access } "409": { description: Stale decision_version or terminal-case conflict } "422": { description: Invalid state/disposition combination } /api/scenarios/{scenario_id}/health: get: operationId: analytics.health summary: Contextual deterministic health feeding 042 badge security: [{ bearerAuth: [] }] parameters: - { name: scenario_id, in: path, required: true, schema: { type: string } } - { name: environment_class, in: query, schema: { type: string } } - { name: context_key, in: query, schema: { type: string } } responses: "200": { description: ScenarioHealth, content: { application/json: { schema: { $ref: "#/components/schemas/ScenarioHealth" } } } } /api/scenarios/{scenario_id}/trends: get: operationId: analytics.trends summary: Deterministic success and classified/disposition trend series security: [{ bearerAuth: [] }] parameters: [{ name: scenario_id, in: path, required: true, schema: { type: string } }] responses: "200": { description: Trends, content: { application/json: { schema: { $ref: "#/components/schemas/Trends" } } } } /api/scenarios/{scenario_id}/recurring-failures: get: operationId: analytics.recurring summary: Compatibility-scoped recurring groups and episodes security: [{ bearerAuth: [] }] parameters: [{ name: scenario_id, in: path, required: true, schema: { type: string } }] responses: "200": { description: "RecurringFailureGroup[]", content: { application/json: { schema: { type: array, items: { $ref: "#/components/schemas/RecurringFailureGroup" } } } } } components: securitySchemes: bearerAuth: { type: http, scheme: bearer } serviceAndUser: { type: http, scheme: bearer, description: "Trusted producer service with originating user/provenance context" } schemas: InvestigationQueueItem: type: object required: [id, source_type, severity, state, count, first_seen_at, last_seen_at] properties: id: { type: string, format: uuid } source_type: { type: string, enum: [scenario_run, staleness_signal, baseline_immutability, load_finding, automation_failure] } scenario_id: { type: [string, 'null'] } run_id: { type: [string, 'null'], format: uuid } logical_step_id: { type: [string, 'null'], format: uuid } severity: { type: string, enum: [info, warning, critical] } fingerprint: { type: [string, 'null'] } active_episode_id: { type: [string, 'null'], format: uuid } evidence_summary: { type: object, additionalProperties: true } target_snapshot: { type: object, additionalProperties: true } suggested_next_action: { type: [string, 'null'] } state: { type: string, enum: [new, acknowledged, case_opened, suppressed, resolved] } count: { type: integer, minimum: 1 } first_seen_at: { type: string, format: date-time } last_seen_at: { type: string, format: date-time } case_id: { type: [string, 'null'], format: uuid } InvestigationSignal: type: object required: [source_type, source_id, severity, evidence_refs, occurred_at] properties: source_type: { type: string, enum: [scenario_run, staleness_signal, baseline_immutability, load_finding, automation_failure] } source_id: { type: string } scenario_id: { type: [string, 'null'] } run_id: { type: [string, 'null'], format: uuid } logical_step_id: { type: [string, 'null'], format: uuid } severity: { type: string, enum: [info, warning, critical] } canonical_fingerprint: { type: [string, 'null'] } evidence_refs: { type: array, items: { type: string, format: uuid } } target_snapshot: { type: [object, 'null'], additionalProperties: true } execution_principal_fingerprint: { type: [string, 'null'] } occurred_at: { type: string, format: date-time } InvestigationCase: type: object required: [id, queue_item_id, status, source_snapshot, evidence_snapshot, owner_actor_id, agent_thread_id, opened_at] properties: id: { type: string, format: uuid } queue_item_id: { type: string, format: uuid } status: { type: string, enum: [open, investigating, awaiting_approval, awaiting_external_change, verifying, resolved, accepted, reopened] } source_snapshot: { type: object, additionalProperties: true } evidence_snapshot: { type: object, additionalProperties: true } owner_actor_id: { type: string } agent_thread_id: { type: string } opened_at: { type: string, format: date-time } resolved_at: { type: [string, 'null'], format: date-time } final_disposition: { type: [string, 'null'], enum: [fixed, accepted_risk, duplicate, wont_fix, false_positive] } resolution_summary: { type: [string, 'null'] } triage: { $ref: "#/components/schemas/TriageRecord" } actions: { type: array, items: { $ref: "#/components/schemas/AgentAction" } } CaseDispositionRequest: type: object required: [decision_version, investigation_status] properties: decision_version: { type: integer, minimum: 1 } investigation_status: { type: string, enum: [new, investigating, resolved] } classification: { type: string, enum: [product_regression, data_regression, baseline_stale, scenario_bug, infrastructure_failure, not_confirmed] } resolution: { type: string, enum: [fixed, accepted_risk, duplicate, wont_fix, false_positive] } comment: { type: string, maxLength: 4000 } verification_evidence_refs: { type: array, items: { type: string, format: uuid }, description: "Required to transition a case to resolved" } acceptance_rationale: { type: string, maxLength: 4000, description: "Required for accepted-risk/wont-fix terminal acceptance" } TriageRecord: type: object properties: investigation_status: { type: string } classification: { type: [string, 'null'] } resolution: { type: [string, 'null'] } comment: { type: [string, 'null'] } decision_version: { type: integer } actor_id: { type: string } AgentAction: type: object required: [id, intent, risk_class, policy_decision, status] properties: id: { type: string, format: uuid } intent: { type: string } risk_class: { type: string, enum: [read, diagnostic_run, controlled_test_data_mutation, draft_write, scenario_revision_write, activate_current_revision, baseline_approval, automation_policy_write, prod_mutation] } policy_decision: { type: string, enum: [delegated, approval_required, denied] } status: { type: string, enum: [planned, running, completed, failed, awaiting_reconciliation] } evidence_refs: { type: array, items: { type: string, format: uuid } } approval_gate_id: { type: [string, 'null'], format: uuid } ScenarioHealth: type: object required: [scenario_id, overall_attention, generated_at, confidence] properties: scenario_id: { type: string } environment_class: { type: string } context_key: { type: string, description: "044 AnalyticsContextKey" } product_health: { type: string, enum: [healthy, attention, unknown] } scenario_test_health: { type: string, enum: [healthy, attention, unknown] } infrastructure_health: { type: string, enum: [healthy, attention, unknown] } agent_evaluation_health: { type: string, enum: [healthy, attention, unknown] } overall_attention: { type: string, enum: [healthy, attention, unknown] } success_rate: { type: number } flaky_ratio: { type: number } infra_failure_ratio: { type: number } inconclusive_ratio: { type: number } agent_disagreement_ratio: { type: number } low_confidence_ratio: { type: number } model_instability_ratio: { type: number } generated_at: { type: string, format: date-time } confidence: { type: string, enum: [insufficient_history, low, high] } Trends: type: object required: [generated_at, success_rate_series, disposition_distribution] properties: generated_at: { type: string, format: date-time } success_rate_series: { type: array, items: { type: object } } disposition_distribution: { type: object, additionalProperties: { type: integer } } RecurringFailureGroup: type: object required: [id, fingerprint, compatibility_family, count, first_occurred_at, last_occurred_at, episodes] properties: id: { type: string, format: uuid } fingerprint: { type: string, description: "logical_step_id + error_code + normalized_error_signature + assertion_kind + affected_ref" } compatibility_family: { type: string } count: { type: integer } first_occurred_at: { type: string, format: date-time } last_occurred_at: { type: string, format: date-time } episodes: { type: array, items: { type: object, properties: { id: { type: string, format: uuid }, opened_at: { type: string, format: date-time }, resolved_at: { type: [string, 'null'], format: date-time } } } }


QUICKSTART — Dev Onboarding

Source: quickstart.md

Quickstart: Investigation Queue & Scenario Analytics (047)

Factual audit 2026-08-20: pending verification checklist only; qualifying production signals do not currently enter the queue and case closure is not conformant.

Prereqs

  • 044 run/step results, 042 registry, 036 AgentAction contract, DB migrated (investigation tables)

Commands

cd backend && source .venv/bin/activate
alembic upgrade head
python -m pytest -v tests/services/dashboard_testing/analytics/
python -m pytest -v tests/api/test_scenario_analytics.py
python -m ruff check src/services/dashboard_testing/analytics/

cd frontend && npm run test -- ScenarioHealth
npm run lint

Exit Gates

  • Qualifying events queue without auto-starting agent work; analyst-opened cases retain evidence/actions/disposition audit
  • Flaky steps detected; health derived and feeds 042 badge
  • Trends + recurring failures render; recurrence after resolution opens a new alertable episode
  • Object ACL + RBAC view vs disposition enforced
  • Case/triage never alters graph/result/baseline
  • ruff clean; prototype states covered

TRACEABILITY — Requirements Matrix

Source: traceability.md

Traceability: Investigation Queue & Scenario Analytics (047)

Story Requirement Model API operationId Contract Task Test Actual status / gap
US1 Queue/Case SCAN-FR-001/002/003/009/011 InvestigationQueueItem, InvestigationCase, TriageRecord investigations.listQueue/openCase/setDisposition Analytics.OpenCase, Analytics.Disposition T003-T005, T015 test_investigation, test_scenario_terminal_signals [~] 044 terminal failed/blocked/inconclusive runs now produce one idempotent immutable queue signal with provenance and no automatic case/AgentRun/chat/action. This bounded producer does not classify recurrence; immutable case evidence/chat linkage and SCAN-FR-012 closure remain open.
US2 Flakiness/Health SCAN-FR-004/005/006 FlakinessSignal, ScenarioHealth analytics.health Analytics.Flakiness, Analytics.Health T006-T008 test_flakiness [~] primitive exists; real history/signal-ingestion proof is pending.
US3 Trends/Recurring SCAN-FR-007/008/010 RecurringFailureGroup analytics.trends, analytics.recurring Analytics.Trends, Analytics.Recurring T009-T010 test_trends [~] aggregate primitives exist; episode/case lifecycle needs E2E proof.
Frontend SCAN-FR-011 InvestigationQueueModel, InvestigationCaseModel — — T011-T012 investigation.ux.test [~] UI exists but cannot prove an unconnected queue workflow.

N/A: Registry (042), Editor (043), Execution (044), Monitor (045), Automation (046).


TASKS — Implementation Tasks

Source: tasks.md

#region ScenarioAnalytics.Tasks [C:3] [TYPE ADR] [SEMANTICS tasks,scenario,analytics,implementation] @BRIEF Ordered TDD backlog for Investigation Queue/Case and deterministic quality analytics (047). Tests FIRST.

Prerequisites: plan.md, spec.md; contracts/modules.md, traceability.md.

Format: - [ ] T### [P] [USx] Description with exact file path

Factual audit 2026-08-20: [x] means code plus relevant evidence; [~] means partial implementation; [ ] means absent integration or unperformed verification.

Phase 1 — Setup

  • T001 Create InvestigationQueueItem, InvestigationCase, AgentAction projection, and migration in backend/src/models/scenario_investigation.py
  • T002 [P] Create canonical fixtures in specs/047-dashboard-scenario-analytics/fixtures/ Proof: fixtures/analytics.json (287 lines, schema_version: 1, fixture_version: 047.1.0, references canonical 042 registry fixture, RunResult rows declared immutable).

Phase 2 — US1 Queue and Agentic Case

  • T003 [US1] Write failing queue/case tests in backend/tests/services/dashboard_testing/registry/test_scenario_investigation.py @TEST_EDGE: event queues but does not start agent; duplicate active episode updates count; explicit open is idempotent; disposition CAS->409
  • [~] T004 [US1] Implement queue projection, explicit open_case, AgentAction linkage and compact set_disposition in analytics/investigation.py @INVARIANT: case/triage orthogonal; RunResult immutable; never alters graph/baseline
  • [~] T005 [US1] Add queue/case/disposition API + object/RBAC tests

Phase 3 — US2 Flakiness + Health

  • T006 [US2] Write failing contextual flakiness/health tests in backend/tests/services/dashboard_testing/registry/test_scenario_analytics.py @TEST_EDGE: one-way PASS→FAIL -> regression (NOT flaky); post-failure PASS + two transitions -> flaky; infra/cancelled/inconclusive excluded
  • [~] T007 [US2] Implement detect_flakiness + derive_health in backend/src/services/dashboard_testing/analytics/flakiness.py @POST: deterministic signals; feeds 042 badge on threshold cross
  • T008 [US2] Add scenario health endpoints in registry and scenario_analytics.py
  • T009 [US3] Write failing trend/recurring tests in backend/tests/services/dashboard_testing/registry/test_scenario_trends.py @TEST_EDGE: classification change does NOT change fingerprint (immutable group identity)
  • [~] T010 [US3] Implement deterministic trends/fingerprint summary and recurring episode persistence in analytics services @POST: compatibility-scoped immutable fingerprint; a matching occurrence after a resolved episode opens a new alertable episode and queue item

Phase 5 — Frontend + Polish

  • [~] T011 [P] Build analytics models/components InvestigationQueueModel, InvestigationCaseModel, ScenarioAnalyticsModel, HealthCard and queue/case/trends components implemented on disk (frontend/src/lib/models/*, frontend/src/lib/components/scenario-analytics/); full queue/case/trends UI remains.
  • [~] T012 [P] L1/L2 model + UX tests for queue/case/health primitives
  • [~] T013 Run quickstart, scoped/full backend + frontend tests, ruff; ATTN_1-4; semantic rebuild Partial: scoped frontend vitest (26 passed) + npm run build (✔ built) + scoped backend API/RBAC (23 passed) verified. Full backend suite, ruff, quickstart and semantic rebuild not yet run in this session.
  • T014 Prototype validation: every declared @UX_STATE is reachable via prototype/index.html.

Audit Follow-ups (2026-08-20)

  • T015 Wire canonical 036/run/automation signals to idempotent queue_signal; prove qualifying events create/update queue items without starting agent work. ingest_investigation_signal + auto_queue_failed_run + registry staleness emission; 044 terminal failed/blocked/inconclusive runs additionally emit one exact immutable provenance signal (passed emits none). This producer does not create a case/AgentRun/action or classify recurring episodes.
  • [~] T016 Persist immutable evidence snapshot, chat/AgentRun linkage and object ACL on InvestigationCase. Evidence snapshot, linked_run_ids and owner_id ACL persist (can_access_case / GET /cases/{id} 403). Chat/AgentRun workspace remains open.
  • T017 Replace unconditional resolved disposition with a CAS state machine enforcing verification evidence/reconciliation for resolved and accepted-risk rationale for accepted; test recurrence reopening. Recurrence after a resolved episode opens a new episode and queue item (test_recurrence_after_resolved_episode_opens_new_queue_item).
  • T018 Run independent end-to-end analytics verification and update T013 only with current evidence.

Dependencies

Setup → US1; US2 depends on US1 + 044 results; US3 depends on US2; US5 frontend depends on backend analytics.

#endregion ScenarioAnalytics.Tasks


PROTOTYPE — State/Manifest

Source: prototype/manifest.md

#region ScenarioAnalytics.PrototypeManifest [C:3] [TYPE ADR] [SEMANTICS prototype,manifest,scenario,analytics] @defgroup Prototype Interactive HTML prototype manifest for Failure Triage & Quality Analytics.

Prototype Metadata

Factual audit 2026-08-20: prototype coverage is not evidence of canonical signal ingestion, immutable case evidence, or conformant disposition closure.

  • Feature: 047 Failure Triage & Quality Analytics
  • Source contracts: ux_reference.md, contracts/modules.md
  • Screens represented: 2 (Investigation Queue and persistent agent-led Case; health summary embedded in Queue)
  • Total states: 4 (idle, saved, conflict, empty)
  • Accessibility: keyboard nav, focus-visible, aria-live, ≥44px, prefers-reduced-motion
  • Responsive: 375px, 900px

State Coverage

@UX_STATE Prototype State Reachable? Recovery
idle (triage) idle ✅ —
saved saved ✅ audited
conflict (409) conflict ✅ reload triage
empty empty ✅ —

Screen ↔ Story Traceability

Story Prototype Feature Intended acceptance coverage
US1 Investigation queue item → explicit case, evidence, agent tools, disposition case/triage projection persisted + audited
US2 Flakiness/Health health card + unstable step flaky detection + health
US3 Recurring recurring list + episode state grouping + alert after resolution
#endregion ScenarioAnalytics.PrototypeManifest

PROTOTYPE — Interactive HTML

Source: prototype/index.html

<!doctype html>

<html lang="ru"> <head> </head>
Superset Tools · BI testing
СценарииЗапускиАвтоматизацияРасследования 2 требуют внимания
Investigation Queue

Очередь расследований

Сигналы не запускают агента сами. Аналитик открывает только нужный case.

Серьёзность Источник и evidence Повторяемость Следующее действие
Critical XLSX reconciliation · SR-1839
ROW_SET_CHANGED · target rc-17 · RLS analyst
Новая episode #2
18 строк
Проверить ETL-1421 Разобрать с агентом
Warning Фильтры и метрики
browser timeout · PREPROD
3 occurrences Сравнить с прошлым run Открыть case
Здоровье продукта
Attention
2 regression runs
Здоровье теста
Stable
flaky 0%
Инфраструктура
Healthy
0 infra failures
← К очереди
Case IC-204 · opened from SR-1839

Расхождение строк после ETL

Run truth остаётся FAILED. Ниже — расследование и управляемые действия.

Investigating

Agent chat

Agent
18 дополнительных строк появились после ETL-1421. Сначала сопоставлю lineage, предыдущий XLSX и release rc-17.
Tool timeline
✓ Сравнил SR-1839 и SR-1812
✓ Получил lineage snapshot
• Проверяю ETL event и affected datasets
Показать evidence Запустить diagnostic run
Disposition сохранён.
Case audit обновлён; immutable RunResult остаётся FAILED.
К очереди
Конфликт решения (409).
Другой аналитик обновил decision_version. Перезагрузите triage projection.
Перезагрузить case

Очередь пуста

Новых alertable episodes нет. Исторические cases доступны в архиве.

Показать фикстуру очереди
State: IdleCaseSaved ConflictEmpty
<script src="../../prototype-ui.js"></script> <script> function showAction(v) { let e = document.getElementById("action-result"); e.classList.remove("hidden"); e.innerHTML = v === "run" ? "Diagnostic run SR-1845 started.
Действие delegated: read-only scenario run, quota проверена." : v === "resolve" ? "Disposition сохранён.
Case audit и compact triage projection обновлены; historical run не изменён." : "Evidence готов.
Сравнение, lineage и target snapshot закреплены за case."; } protoState("idle", (s) => ["idle", "case", "saved", "conflict", "empty"].forEach((id) => document.getElementById(id).classList.toggle("hidden", id !== s), ), ); </script> </html>

================================================================================ FEATURE: 048-rls-management-workspace Files: 13


SPEC — Feature Specification

Source: spec.md

#region Std.Specify.FeatureSpec [C:3] [TYPE ADR] [SEMANTICS spec,requirements,feature,rls,superset,idm] @BRIEF Feature specification — WHAT the user needs and WHY. Implementation-free. Survives HCA 128× via @SEMANTICS grouping.

Navigation (DSA Indexer keywords)

@SEMANTICS: spec, requirements, feature, rls, superset, idm, audit

Feature Branch: 048-rls-management-workspace Created: 2026-08-04 | Status: Draft Input: "Версионирование SQL-скрипта корпоративной RLS-системы Apache Superset (rls_t) с помощью в разработке, аудит доступа пользователей к датасетам (user→unit_id), аудит актуальности BI-пользователей через IDM (уволенные/неактивные), конструктор кастомных RLS-правил на основе справочников целевой БД (фильтрация IN/LIKE/regexp, привязка пользователей по AD-логину, AD-группам и оргструктуре)"

Clarifications

Session 2026-08-04 (обсуждение до спецификации)

  • Q: Что значит «аналитика для аккаунтов пользователей»? → A: Компонент аудита: сверка внесённых в RLS BI-пользователей (rls_bi_users) с данными IDM — уволенные, неактивные, в отпуске.
  • Q: Инструмент обновляет rls.rls_t сам? → A: Нет. Обновление остаётся за внешним SQL-скриптом (TRUNCATE+INSERT). superset-tools версионирует скрипт, помогает в его разработке и контролирует актуальность таблицы (возраст данных).
  • Q: Как фильтруются справочники для кастомных правил? → A: Полностью на целевой БД — SQL-выражения (WHERE IN / LIKE / regexp) исполняются на DWH, как их будет исполнять финальный скрипт. Предпросмотр результата обязателен.
  • Q: Как устроен жизненный цикл скриптов? → A: Версии скриптов живут в репозитории superset-tools (branch 048), затем push в корпоративный репозиторий, оттуда — развёртывание в прод DWH внешним контуром.
  • Q: К каким атрибутам применяются фильтры правил? → A: К любым атрибутам справочника; unit_id — одна выделенная колонка справочника.
  • Q: Как привязываются пользователи к ролям? → A: Три способа: по AD-логину (вручную), массовый импорт всех аккаунтов из AD-группы, назначение кусками оргструктуры (весь отдел / БЕ) через поиск по IDM.
  • Q: Откуда superset-tools получает данные rls_t/rls_bi_users? → A: Гибрид: снапшот rls_t/rls_bi_users/rls_bi_roles в собственную БД superset-tools (обновление по расписанию + вручную) для аудита; предпросмотр правил — живой SELECT на целевой БД.
  • Q: Как связать датасет Superset с unit_id для аудита? → A: Гибрид-каталог с учётом rls_type: автоопределение колонки-носителя единицы по имени + ручное подтверждение/переопределение; привязка хранит тип единицы (unit_balance_code / plant_code), матрица аудита пересекает rls_type пользователя с типом единицы датасета. Архитектурно rls_t НЕ изменяется — новые типы RLS добавляются значениями rls_type (текстовая колонка), маппинг датасетов живёт в superset-tools.
  • Q: Какие RBAC-роли работают с разделом RLS? → A: Две операционные роли: rls_operator — привязка пользователей к ролям, ведение правил, аудит bi_users, применение INSERT; rls_script_dev — разработка и версионирование скрипта rls_t (diff, push). Глобальный admin наследует обе.
  • Q: Как применяются INSERT в bi_users/bi_roles? → A: Через интерфейс: INSERT ролей и пользователей исполняется сразу на целевой БД (live-запись, идемпотентно), а сгенерированный SQL сохраняется в репозиторий как миграция для истории; деактивация пользователей — тоже из интерфейса.
  • Q: Как пушить версии скрипта в корпоративный репозиторий? → A: Второй git remote (corp): корпоративный репозиторий подключается как дополнительный remote в git-контуре, push — стандартным git-механизмом; credentials через существующий шифрованный механизм проекта.

User Scenarios

Каждая story — независимо тестируемая единица. Приоритеты P1 (MVP) → P2. Все stories разделяют доменные ключевые слова: rls, superset, idm, audit.

Story 1 — Версионирование и разработка RLS-скрипта (P1)

Why P1: Это фундамент: без версий скрипта нельзя безопасно развивать остальные компоненты (правила, аудит), а «помощь в разработке» — прямая потребность DWH-команды.

Independent Test: Открыть раздел RLS → список версий скрипта rls_t main insert.sql, создать новую версию из текущей, увидеть diff, пометить версию как активную, увидеть статус свежести таблицы.

Acceptance:

  1. Given пользователь в разделе RLS-скриптов, When открывает список версий, Then видит все сохранённые версии с датой, автором и комментарием
  2. Given есть текущая активная версия, When пользователь создаёт новую версию из активной, Then создаётся копия SQL для редактирования, активная версия не меняется
  3. Given две версии скрипта, When пользователь открывает diff, Then изменения показаны построчно (added/removed/modified)
  4. Given пользователь завершил редактирование версии, When сохраняет и помечает как активную, Then версия фиксируется в git-контуре репозитория и готова к push в корпоративный репозиторий
  5. Given настроено подключение к целевой БД, When пользователь открывает статус актуальности, Then видит дату последнего обновления rls_t, возраст данных (dt_created/dttm_updated), количество строк и предупреждение при устаревании (порог настраивается)
  6. Given SQL-скрипт содержит синтаксическую ошибку, When пользователь сохраняет версию, Then система показывает предупреждение о синтаксисе и не блокирует сохранение (скрипт исполняется внешним контуром)
  7. Given пользователь нажимает «Push в корпоративный репозиторий», When изменения отсутствуют или push невозможен, Then система показывает понятную ошибку с описанием причины

Story 2 — Аудит доступа к датасетам (P1)

Why P1: Прямой ответ на вопрос «какой пользователь какие данные видит» — контрольная функция RLS, ценность для безопасности без новых источников (только rls_t + Superset API).

Independent Test: Открыть раздел аудита → ввести AD-логин → увидеть список датасетов Superset и видимые unit_id; выбрать датасет → увидеть список пользователей с доступом к каждому unit_id.

Acceptance:

  1. Given пользователь вводит AD-логин в поиске аудита, When система находит записи, Then показывает: список датасетов, к которым у пользователя есть доступ, и видимые unit_id (с учётом source_system, origin_role и rls_type)
  2. Given пользователь выбирает датасет, When система строит обратную матрицу, Then показывает список пользователей и их unit_id по этому датасету
  3. Given пользователь ищет логин без записей в rls_t, When поиск завершён, Then показывается пустое состояние с пояснением «пользователь не найден в RLS» и подсказкой проверить регистр (user_id хранится в нижнем регистре)
  4. Given rls_t недоступна или снапшот устарел, When пользователь открывает аудит, Then система показывает дату снапшота и предупреждение о его возрасте
  5. Given датасет не привязан к RLS-единицам, When он попадает в матрицу, Then он отображается как «без RLS-фильтрации» отдельной группой

Story 3 — Аудит актуальности BI-пользователей через IDM (P1)

Why P1: Самый дешёвый компонент (мок IDM уже есть) и немедленная ценность: вычистка уволенных/неактивных из RLS — риск информационной безопасности.

Independent Test: Открыть раздел аудита пользователей → увидеть список rls_bi_users, обогащённый статусами IDM (активен/уволен/отпуск/аккаунт отключён, риск-скор, VIP, adminRights) → отфильтровать «уволенные» → увидеть рекомендацию по деактивации.

Acceptance:

  1. Given в системе есть rls_bi_users, When пользователь открывает аудит пользователей, Then каждая строка обогащена данными IDM: статус сотрудника (dateOut/leaveType), статус AD-аккаунта (enabled), риск-скор, VIP, adminRights
  2. Given пользователь включает фильтр «уволенные / неактивные», When список обновляется, Then показываются только кандидаты на вычистку с обоснованием по каждому
  3. Given IDM недоступен, When аудит открыт, Then статусы помечаются как «недоступно», список продолжает работать на данных bi_users, появляется предупреждение о деградации
  4. Given пользователь открывает карточку сотрудника в аудите, When карточка загружена, Then видны данные IDM (ФИО, отдел, должность, компания, даты) и все аккаунты/приложения с OU и adminRights
  5. Given пользователь выбрал одного или нескольких неактивных пользователей, When формирует деактивацию, Then система применяет деактивацию (is_deleted=true) на целевой БД из интерфейса и сохраняет SQL-миграцию в репозиторий для истории
  6. Given пользователь ищет по AD-логину, When логин не найден ни в bi_users, ни в IDM, Then показывается явное «не найден» с разграничением «нет в RLS» / «нет в IDM»

Story 4 — Конструктор кастомных RLS-правил (P2)

Why P2: Самая сложная story, требует доступа к справочникам DWH; но именно она превращает ведение кастомного RLS из ручных миграций в поддерживаемый процесс (пример: ЦО-БЕ-ПФМ).

Independent Test: Создать правило «ЦО-БЕ-ПФМ»: выбрать справочник, колонку unit_id (PFM), фильтр по центру ответственности → предпросмотр списка unit_id на целевой БД → привязать пользователей (по логину / AD-группе / отделу) к определению правила → сохранить правило и увидеть effect_at следующего запуска скрипта.

Acceptance:

  1. Given пользователь создаёт правило, When выбирает справочник (витрину) и колонку unit_id, Then система показывает доступные атрибуты справочника и типы данных для фильтрации
  2. Given пользователь задал фильтр (WHERE IN / LIKE / regexp по любому атрибуту), When нажимает «Предпросмотр», Then фильтрация выполняется на целевой БД и показывается количество и список unit_id, попадающих в правило
  3. Given фильтр вернул 0 строк, When предпросмотр завершён, Then система предупреждает о пустом результате с подсказкой проверить фильтр (правило сохранить можно — ролей не будет, пока справочник не расширится)
  4. Given правило сохранено, When пользователь привязывает пользователей, Then доступны три способа: ручной ввод AD-логинов (с проверкой через IDM), импорт всех аккаунтов AD-группы (поиск по OU/группе в IDM), назначение куском оргструктуры (отдел/БЕ через поиск IDM)
  5. Given правило с привязками готово, When пользователь нажимает «Сохранить правило», Then определение UPSERT-ится в rls.rls_roles_filter, а привязки пользователей — в rls.rls_rule_users (идемпотентно по rule_id + user_id); скрипт v13 динамически вычисляет итоговые пары user_id → unit_id; UI показывает индикатор «эффект при запуске скрипта 06:00»
  6. Given пользователь изменил фильтр существующего правила, When повторяет предпросмотр, Then показывается новый ожидаемый срез (счётчик unit_id) до сохранения — скрипт пересчитает роли автоматически на следующем запуске
  7. Given правило использует regexp-фильтр, When система валидирует выражение, Then некорректный паттерн подсвечивается с сообщением об ошибке (как в rls_roles_mask)
  8. Given пользователь открывает раздел правил, When список загружен, Then виден каталог правил: название, справочник, колонка unit_id, тип фильтра, количество unit_id, версия сгенерированного SQL

Edge Cases

  • rls_t пустая или скрипт никогда не запускался → аудит показывает пустое состояние + явное указание «таблица не заполнена»
  • Пользователь IDM имеет несколько AD-аккаунтов в разных доменах (TEST.LOCAL, IE.CORP) → в аудите показываются все аккаунты; сопоставление с rls_t по логину без домена, регистр не важен
  • IDM возвращает невалидные идентификаторы (accountId с пробелами/мусором, как в моке) → все ID трактуются как непрозрачные строки, без UUID-валидации
  • Справочник содержит >100 000 строк → предпросмотр лимитирован (например, 500 значений + счётчик), поиск/фильтрация остаётся на стороне БД
  • В справочнике появилась новая БЕ → следующий запуск скрипта автоматически добавит роли новой БЕ (динамическое правило, R13) — без ручных действий
  • БЕ удалена из справочника → следующий запуск скрипта перестанет включать её роли автоматически (доступ отзывается); между удалением и запуском скрипта действует латентность, показываемая в UI («эффект при запуске 06:00»)
  • Правило сохранено, но скрипт ещё не запускался → роли не действуют; индикатор effect_at объясняет ожидание
  • Два пользователя параллельно редактируют одно правило → версии правил конфликтуют: показывается modal «перезагрузить или отменить» (последний сохранённый выигрывает)
  • Уволенный пользователь всё ещё фигурирует в rls_t (SAP-часть) → аудит bi_users показывает только bi_users; расхождение rls_t vs bi_users фиксируется в отчёте
  • Датасет содержит единицы обоих типов (unit_balance_code и plant_code) → в каталоге несколько привязок; в матрице пользователь сопоставляется с датасетом по своему rls_type (оба типа — оба списка unit_id)
  • INSERT применён повторно (двойной клик/повтор операции) → идемпотентность: существующие пары не дублируются (RLS-FR-023)
  • Сбой INSERT на целевой БД в середине операции → операция выполняется по одной паре (user, role) с ошибкой на конкретной строке; уже применённые пары остаются, ошибка показывается с указанием строки
  • AD-группа/OU не найдена в IDM → ошибка с пояснением, привязка не создаётся
  • Специальные символы в фильтрах (кавычки, %, _) → экранирование при генерации SQL, предпросмотр безопасен
  • Реальный IDM строже мока (поиск только по UID) → клиент пробует несколько идентификаторов, деградация с сообщением

Requirements

Functional (IDs survive HCA 128× via hierarchical naming)

  • RLS-FR-001: Хранение версий SQL-скрипта rls_t с метаданными (дата, автор, комментарий, активная версия) в git-контуре репозитория
  • RLS-FR-002: Построчный diff между любыми двумя версиями скрипта
  • RLS-FR-003: Мониторинг актуальности rls_t: возраст данных (dt_created/dttm_updated), количество строк, настраиваемый порог устаревания с предупреждением; работает по снапшоту
  • RLS-FR-004: Push подготовленных версий в корпоративный репозиторий через второй git remote (corp); credentials — существующий шифрованный механизм проекта; при недоступности — понятная ошибка без потери версии
  • RLS-FR-005: Матрица аудита «пользователь × датасет × unit_id» на основе снапшота rls_t и каталога датасетов; сопоставление учитывает rls_type: пользователь с rls_type=user_to_unit_balance_code сопоставляется с датасетами, привязанными к unit_balance_code, с user_to_plant_code — к plant_code
  • RLS-FR-006: Обратная матрица «датасет → пользователи с доступом»
  • RLS-FR-007: Обогащение rls_bi_users данными IDM: статус сотрудника, статус аккаунта, риск-скор, VIP, adminRights, карточка (ФИО, отдел, должность, компания, даты, аккаунты/приложения)
  • RLS-FR-008: Деактивация пользователей (is_deleted=true) применяется из интерфейса на целевой БД; SQL-миграция сохраняется в репозиторий для истории
  • RLS-FR-009: Определение кастомного правила: справочник, колонка unit_id, фильтры (WHERE IN / LIKE / regexp) по произвольным атрибутам
  • RLS-FR-010: Предпросмотр результата правила — живой SELECT на целевой БД (количество и выборка unit_id, лимит на выборку)
  • RLS-FR-011: Привязка пользователей к ролям тремя способами: AD-логин, AD-группа (все аккаунты группы/OU), кусок оргструктуры (отдел/БЕ)
  • RLS-FR-012: Сохранение динамического правила: определение UPSERT в rls_roles_filter + пользовательские привязки UPSERT в rls_rule_users; скрипт v13 разворачивает rule_id → unit_id и rule_id → user_id в итоговые пары (user_id, unit_id)
  • RLS-FR-013: Diff между поколениями SQL одного правила до сохранения
  • RLS-FR-014: Валидация regexp-фильтров с человекочитаемыми ошибками
  • RLS-FR-015: Каталог правил: список с метаданными и количеством unit_id
  • RLS-FR-016: Расширение мока IDM: поиск пользователей по AD-группе/OU и по оргструктуре (отдел/БЕ) для массовой привязки
  • RLS-FR-017: Деградация при недоступности IDM (аудит работает на данных bi_users, статусы помечаются «недоступно»)
  • RLS-FR-018: RBAC: две операционные роли — rls_operator (ведение правил, привязка пользователей к ролям, аудит bi_users, применение INSERT на DWH, каталог датасетов) и rls_script_dev (версии скрипта rls_t, diff, редактор SQL, push в корпоративный репозиторий); глобальный admin наследует обе; персональные данные IDM (ФИО, риск-скор, VIP) видят rls_operator и admin; просмотр аудитов — обе роли + admin
  • RLS-FR-019: Деградация при недоступности DWH: чтение из снапшота rls_t, пометка возраста данных
  • RLS-FR-020: Снапшот rls_t/rls_bi_users/rls_bi_roles в собственную БД superset-tools: обновление по расписанию и вручную (кнопка «Обновить»); аудит работает по снапшоту, предпросмотр правил — живой SELECT на DWH
  • RLS-FR-021: Каталог датасетов (dataset_rls_bindings): автоопределение колонки-носителя единицы по имени (unit_id/unit_balance_code/plant_code) + ручное подтверждение/переопределение; датасет может иметь несколько привязок (разные типы единиц); rls_t архитектурно не изменяется
  • RLS-FR-022: Новые типы RLS (например, user_to_center_responsibility) вводятся значениями rls_type без ALTER корпоративной таблицы
  • RLS-FR-023: Идемпотентность привязок и определений: повторное сохранение правила и повторный UPSERT привязок (rule_id, user_id) в rls_rule_users не создают дубликатов
  • RLS-FR-024: Динамические кастомные правила (R13): определение хранится в rls_roles_filter, пользователи правила — в rls_rule_users; скрипт v13 вычисляет unit_id из справочника при каждом запуске и формирует user_id → unit_id — новые единицы добавляются, удалённые исчезают без materialized-ролей
  • RLS-FR-025: rls_roles_mask, rls_bi_roles и rls_bi_users сохраняют статическую семантику; динамические правила используют отдельные DISTRIBUTED REPLICATED таблицы rls_roles_filter и rls_rule_users; деактивация — is_deleted=true; rls_t флаг не получает

Все маркеры неясностей закрыты в фазе clarify (сессия 2026-08-04).

Key Entities

  • RlsScriptVersion: Версия SQL-скрипта (rls_t main insert.sql): содержимое, номер/дата, автор, комментарий, флаг активной версии, статус push
  • RlsRule: Определение кастомного правила: название, справочник (витрина), колонка unit_id, набор фильтров, rls_type, версия сгенерированного SQL
  • RlsRuleFilter: Атрибут справочника + тип фильтра (IN / LIKE / regexp) + выражение
  • RlsRoleBinding: Привязка пользователя к роли: AD-логин (или источник: группа/OU/оргструктура), роль, способ привязки
  • RlsSnapshot: Снапшот rls_t и rls_bi_users/rls_bi_roles для аудита: данные, дата снятия, источник
  • RlsDatasetBinding: Привязка датасета Superset к RLS: датасет → колонка-носитель единицы + тип единицы (unit_balance_code / plant_code); несколько привязок на датасет; способ определения (авто/вручную)
  • RlsOperator: Роль «простой администратор» — ведение правил, привязка пользователей, аудит, применение INSERT
  • RlsScriptDeveloper: Роль разработчика — версии скрипта rls_t, diff, push в корпоративный репозиторий
  • IdmProfile: Карточка сотрудника из IDM: ФИО, отдел, должность, компания, даты приёма/увольнения, статусы, риск-скор, VIP, аккаунты с OU и adminRights
  • RlsAuditReport: Отчёт аудита: расхождения bi_users с IDM, устаревшие записи, кандидаты на деактивацию

Success Criteria

  • SC-001: Создание новой кастомной роли (правило + привязка + SQL) занимает не более 5 минут и не требует написания SQL вручную
  • SC-002: Аудит bi_users выявляет 100% уволенных/неактивных пользователей при доступности IDM (сверка по dateOut/leaveType/enabled)
  • SC-003: Для любого AD-логина ответ «какие датасеты и unit_id видит пользователь» строится за ≤ 5 секунд на снапшоте
  • SC-004: Все версии скрипта rls_t доступны в git-истории; diff между произвольными версиями отображается без обращения к внешним сервисам
  • SC-005: Сгенерированный SQL правила проходит валидацию синтаксиса и всегда содержит пары (user_id, unit_id) — без исключений

#endregion Std.Specify.FeatureSpec


UX REFERENCE — Interaction Narrative

Source: ux_reference.md

#region Std.Specify.UxReference [C:3] [TYPE ADR] [SEMANTICS ux,reference,rls,superset,idm] @BRIEF UX interaction reference — persona, flows, states, recovery paths, and edge/failure matrix coverage. Drives @UX_* contract tags in Phase 1 and feeds prototype + OpenAPI generation.

Feature Branch: 048-rls-management-workspace Created: 2026-08-04 | Status: Draft

1. User Persona & Context

  • Who is the user?: Две операционные роли: (1) администратор RLS (rls_operator) — ведёт правила, привязывает пользователей к ролям, аудирует bi_users, применяет изменения на DWH; (2) разработчик RLS-скрипта (rls_script_dev) — версионирует и редактирует SQL rls_t main insert.sql, пушит в корпоративный репозиторий. Второстепенная аудитория: глобальный admin (наследует обе роли) и специалист по безопасности (просмотр аудитов).
  • What is their goal?: Поддерживать RLS актуальной: развивать скрипт rls_t, видеть кто что видит, вычищать уволенных, заводить новые кастомные правила без ручного SQL.
  • Context: Браузер, рабочий стол, корпоративный контур. Источники: целевая БД DWH (Greenplum/ClickHouse), IDM (в dev — мок на localhost:3215), Superset API. Раздел в навигации «RLS» (рядом с migration/datasets).

2. The "Happy Path" Narrative

Администратор открывает раздел «RLS» и видит дашборд: актуальность rls_t (зелёный бейдж «обновлено вчера»), список версий скрипта и предупреждения аудита. Он открывает аудит пользователей — система сама подсвечивает двух уволенных; он формирует SQL-миграцию деактивации одним кликом. Затем переходит в конструктор правил, выбирает витрину «ЦО-БЕ-ПФМ», задаёт фильтр по центру ответственности через пару кликов, видит предпросмотр: 34 unit_id попадут в роль. Привязывает отдел разработки через поиск оргструктуры — 12 человек. Жмёт «Применить» — INSERT выполнен на целевой БД, SQL-миграция сохранена в репозиторий, diff показан. Всё за три минуты, без единой строчки SQL вручную.

3. Interface Mockups

UI Layout & Flow

Экран 1: RLS Dashboard (обзор)

  • Layout: Три карточки сверху: «Актуальность rls_t» (возраст данных, строки, порог), «Версии скрипта» (активная версия, последний push), «Аудит» (счётчики: уволенные, неактивные, расхождения). Ниже — таблица последних версий скрипта.
  • Key Elements:
    • Карточка актуальности: бейдж статуса (green/amber/red), текст «Обновлено: 03.08.2026, 214 503 строки». Порог устаревания настраивается в настройках.
    • Кнопка «Новая версия из активной»: создаёт копию SQL для редактирования.
    • Кнопка «Push в корпоративный репозиторий»: отправляет активную версию внешнему контуру.
  • Contract Mapping:
    • @UX_STATE: loading / loaded / stale (rls_t устарела) / push-failed
    • @UX_FEEDBACK: toast при сохранении версии, бейдж статуса push
    • @UX_RECOVERY: при недоступности DWH — снапшот + баннер «данные от 02.08.2026»
    • @UX_REACTIVITY: загрузка данных через SvelteKit load(); карточки — $derived от состояния модели RlsDashboardModel
  • States:
    • Idle/Default: все карточки заполнены, бейджи зелёные.
    • Loading: скелетоны в карточках.
    • Success: toast «Версия v12 сохранена и помечена активной».
    • Error/Degraded: баннер «DWH недоступна, показан снапшот от <дата>» + кнопка «Обновить».

Экран 2: Аудит доступа к датасетам

  • Layout: Поисковая строка (AD-логин) + переключатель направления: «Пользователь → данные» / «Датасет → пользователи». Результат — матрица: строки датасеты, колонки unit_id (или наоборот), ячейки с source_system/origin_role.
  • Key Elements:
    • Поиск пользователя: autocomplete по rls_t (нижний регистр, без учёта регистра), подсказка «user_id хранится в нижнем регистре».
    • Выбор датасета: поиск по метаданным Superset API.
    • Группа «Без RLS-фильтрации»: датасеты, не связанные с unit_id — отдельным блоком.
  • Contract Mapping:
    • @UX_STATE: idle / searching / loaded / empty / degraded (устаревший снапшот)
    • @UX_FEEDBACK: счётчик найденных записей, бейдж возраста снапшота
    • @UX_RECOVERY: пустой результат — подсказка «проверьте регистр логина или наличие записей в rls_t»
    • @UX_REACTIVITY: модель RlsAuditMatrixModel (.svelte.ts) — $state для направления/поиска, $derived для матрицы
  • States:
    • Idle/Default: пустая матрица с подсказкой.
    • Loading: спиннер в области матрицы.
    • Success: матрица со счётчиком «42 датасета, 5 unit_id».
    • Error/Degraded: баннер возраста данных + данные снапшота.

Экран 3: Аудит пользователей (bi_users × IDM)

  • Layout: Таблица rls_bi_users, каждая строка обогащена статусами IDM: колонки AD-логин, роль, статус сотрудника (активен/уволен/отпуск), статус AD-аккаунта (enabled/disabled), риск-скор, VIP, adminRights, дата обновления. Фильтры-чипы сверху: «Уволенные», «Неактивные», «В отпуске», «adminRights». Строка раскрывается в карточку сотрудника.
  • Key Elements:
    • Фильтр «Уволенные»: подсвечивает кандидатов на вычистку красным.
    • Чекбоксы выбора: для группового формирования рекомендации деактивации.
    • Кнопка «Сформировать SQL деактивации»: генерирует INSERT (is_deleted=true) в версионируемый контур.
    • Карточка сотрудника: ФИО, отдел, должность, компания, даты, все аккаунты/приложения с OU и adminRights (вкладки по приложениям).
  • Contract Mapping:
    • @UX_STATE: loading / loaded / idm-degraded / empty
    • @UX_FEEDBACK: чипы-фильтры со счётчиками, бейдж «IDM недоступен» на строках
    • @UX_RECOVERY: при недоступности IDM — статусы «недоступно», список работает
    • @UX_REACTIVITY: модель RlsUserAuditModel — $state фильтры/выбор, $derived отфильтрованный список
  • States:
    • Idle/Default: таблица со всеми bi_users.
    • Loading: скелетон таблицы.
    • Success: строка с зелёным «активен» / красным «уволен».
    • Error/Degraded: бейджи «IDM недоступен», данные bi_users на месте.

Экран 4: Конструктор правил (пошаговый)

  • Layout: Слева — каталог правил (список с метаданными). Справа — мастер из 4 шагов: (1) Справочник и колонка unit_id; (2) Фильтры (таблица «атрибут × тип × выражение», добавление строк: WHERE IN — список значений, LIKE — паттерн, regexp — выражение с валидацией); (3) Предпросмотр: счётчик «34 unit_id», таблица первых 500 значений, кнопка «Обновить счётчик»; (4) Привязка пользователей: вкладки «По логину» (автокомплит через IDM), «AD-группа/OU» (поиск группы, показ найденных аккаунтов), «Оргструктура» (дерево отделов/БЕ, мультивыбор, показ количества сотрудников). Финал: «Сгенерировать SQL» → модал с SQL + diff vs предыдущая версия + «Сохранить версию».
  • Key Elements:
    • Шаг 1: селектор справочника (из доступных на DWH), селектор колонки unit_id.
    • Шаг 2: строки фильтров, кнопка «+ фильтр»; regexp-поле с live-валидацией.
    • Шаг 3: «Предпросмотр на целевой БД» — выполняется SQL на DWH; предупреждение при 0 строк (нельзя сохранить).
    • Шаг 4: три вкладки привязки; счётчик выбранных пользователей.
    • Кнопка «Применить»: только при сохранённом правиле и выбранных пользователях; исполняет INSERT на целевой БД (идемпотентно) и сохраняет SQL-миграцию в репозиторий.
  • Contract Mapping:
    • @UX_STATE: step-1 … step-4, preview-loading, preview-empty, saving, saved
    • @UX_FEEDBACK: валидация regexp под полем, счётчики unit_id/пользователей, модал diff
    • @UX_RECOVERY: ошибка БД при предпросмотре — баннер с текстом ошибки DWH + «Повторить»
    • @UX_REACTIVITY: модель RlsRuleBuilderModel — $state шаг/фильтры/ruleUsers/effectAt, $derived готовность к сохранению; save — @ACTION
  • States:
    • Idle/Default: пустой мастер, шаг 1.
    • Loading: счётчик предпросмотра.
    • Success: «Правило сохранено, SQL v1 сгенерирован».
    • Error/Degraded: ошибка подключения к DWH — баннер + повтор.

4. The "Error" Experience

Philosophy: Не просто сообщать об ошибке — вести к исправлению. Каждый путь ошибки ниже мапится на @UX_RECOVERY / @UX_FEEDBACK в контрактах компонентов.

Edge & Failure State Matrix Reference

State Class Trigger Applicable? Visual/Feedback Recovery
NET_01 — Offline Сеть недоступна Да (все экраны) Оффлайн-баннер Автоповтор при reconnect
NET_02 — Timeout >30s нет ответа DWH/IDM Да Toast + счётчик Retry (3 попытки)
NET_03 — Retry exhausted 3 неудачные попытки Да Устойчивый баннер Ручной retry
VAL_01 — Field validation Некорректный regexp/логин Да Красная рамка + текст Исправить выражение
VAL_02 — Cross-field validation Колонка unit_id = колонка фильтра / пустой справочник Да Сводный баннер Исправить и повторить
AUTH_01 — 401 Unauthorized Истёк токен Да Редирект на login Login → назад
AUTH_02 — 403 Forbidden Роль viewer открывает RLS Да Полностраничное объяснение Навигация на дашборды
NF_01 — 404 Not Found Правило/версия удалены Да Страница not found Навигация к списку
CONF_01 — 409 Concurrent edit Параллельное редактирование правила Да Модал «Перезагрузить?» Reload или отмена
422 — Server validation Бизнес-правило (пустая роль) Да Toast с деталями Исправить + повтор
429 — Rate limited Слишком много запросов к IDM Да Countdown Ждать Retry-After
5XX — Server error Backend/DWH/IDM упал Да Error section + retry Retry
STALE — Background update Новый снапшот rls_t Да Refresh banner Клик «Обновить»
PARTIAL — Partial load Часть пользователей не найдена в IDM Да Плейсхолдер строки «не найдено в IDM» Retry по строке
DUP_01 — Double submit Двойной клик «Сгенерировать SQL» Да Кнопка disabled Нормальное завершение
DUP_02 — Navigation interrupt Грязный мастер + уход со страницы Да Confirm dialog Остаться или отменить
LARGE — Large dataset Справочник >100k строк Да Счётчик + лимит выборки 500 Уточнение фильтра
EMPTY — No data 0 unit_id в предпросмотре / 0 записей в аудите Да Пустое состояние с пояснением Изменить фильтр
MALFORMED — Bad response IDM вернул мусор (accountId с пробелами) Да Error ID + retry Зафиксировать ID ошибки
A11Y — Screen reader Смена состояния Да aria-live announcements Встроено в переходы
RESP — Responsive Viewport <768px Да Stacked layout Встроено

Scenario A: Пустой предпросмотр правила

  • User Action: Задал фильтр по ЦО, попало 0 unit_id.
  • System Response: (UI) Шаг 3 показывает жёлтый баннер: «Фильтр не вернул ни одной записи. Проверьте атрибуты и выражение» + кнопка «Вернуться к фильтрам». Кнопка «Применить» заблокирована.
  • Recovery: Возврат к шагу 2, изменение выражения, повторный предпросмотр.

Scenario B: IDM недоступен при аудите пользователей

  • System Response: Строки таблицы остаются, колонки статусов показывают «—» и серый бейдж «IDM недоступен». Баннер: «Сервис IDM не отвечает. Статусы сотрудников недоступны, проверка по bi_users продолжается». Кнопка «Повторить».
  • Recovery: Ручной retry или автоповтор через 30s при reconnect.

5. Tone & Voice

  • Style: Технический, спокойный, на русском языке. Терминология без жаргона.
  • Terminology: «RLS-правило» (не «rule config»), «справочник» (не «витрина»), «unit_id» — как в корпоративных справочниках, «деактивация» (не «удаление») для мягкого удаления пользователей, «AD-логин» (не «username»). Сохранять термины RLS-репозитория: rls_t, rls_bi_users, rls_bi_roles, rls_type, source_system.

#endregion Std.Specify.UxReference


CHECKLISTS — Requirements Quality — requirements.md

Source: checklists/requirements.md

[REQUIREMENTS] Checklist: 048-rls-management-workspace

Purpose: Проверка полноты требований фичи RLS Management Workspace перед планированием — каждое требование из spec.md покрыто проверяемым пунктом. Created: 2026-08-04 Feature: spec.md

Data Sources & Connectivity

  • CHK001 Снапшот rls_t/rls_bi_users/rls_bi_roles в собственную БД superset-tools (расписание + ручное обновление) для аудита; предпросмотр правил — живой SELECT на целевой БД ([RLS-FR-010, RLS-FR-020])
  • CHK002 Деградация при недоступности DWH: работа по снапшоту, пометка возраста данных ([RLS-FR-019])
  • CHK003 Интеграция с Superset API для метаданных датасетов (аудит матрицы) ([RLS-FR-005])
  • CHK004 Интеграция с IDM: карточка сотрудника, аккаунты/приложения, статусы (мок на dev, реальный сервис в prod) ([RLS-FR-007])
  • CHK005 Расширение мока IDM: поиск по AD-группе/OU и оргструктуре (отдел/БЕ) ([RLS-FR-016])

Versioning & Script Lifecycle

  • CHK006 Хранение версий SQL-скрипта rls_t в git-контуре с метаданными (дата, автор, комментарий, активная версия) ([RLS-FR-001])
  • CHK007 Построчный diff между произвольными версиями скрипта ([RLS-FR-002])
  • CHK008 Создание новой версии из активной без изменения активной версии (acceptance US1-2)
  • CHK009 Мониторинг актуальности rls_t: возраст dt_created/dttm_updated, количество строк, настраиваемый порог с предупреждением ([RLS-FR-003])
  • CHK010 Push версий в корпоративный репозиторий через второй git remote (corp) с понятной ошибкой при недоступности ([RLS-FR-004])
  • CHK011 Валидация синтаксиса SQL как предупреждение (не блокирует сохранение) (acceptance US1-6)

Audit: Datasets × RLS

  • CHK012 Матрица «пользователь → датасеты × unit_id» по AD-логину ([RLS-FR-005])
  • CHK013 Обратная матрица «датасет → пользователи с доступом» ([RLS-FR-006])
  • CHK014 Группа «датасеты без RLS-фильтрации» отображается отдельно (acceptance US2-5)
  • CHK015 Пустое состояние поиска с подсказкой про регистр логина (acceptance US2-3)
  • CHK016 Отображение source_system и origin_role в ячейках матрицы (acceptance US2-1)
  • CHK017 Время построения матрицы ≤ 5 секунд на снапшоте ([SC-003])

Audit: BI-Users × IDM

  • CHK018 Обогащение rls_bi_users статусами IDM: сотрудник (dateOut/leaveType), аккаунт (enabled), риск-скор, VIP, adminRights ([RLS-FR-007])
  • CHK019 Фильтры-чипы: уволенные / неактивные / в отпуске / adminRights (acceptance US3-2)
  • CHK020 Карточка сотрудника: ФИО, отдел, должность, компания, даты, аккаунты/приложения с OU и adminRights (acceptance US3-4)
  • CHK021 Применение деактивации (is_deleted=true) на целевой БД из интерфейса + SQL-миграция в репозиторий ([RLS-FR-008])
  • CHK022 Разграничение «нет в RLS» vs «нет в IDM» при поиске (acceptance US3-6)
  • CHK023 Деградация при недоступности IDM с пометкой статусов ([RLS-FR-017])
  • CHK024 Полнота выявления уволенных/неактивных при доступном IDM ([SC-002])

Rule Builder (Custom RLS)

  • CHK025 Определение правила: справочник, колонка unit_id, фильтры по произвольным атрибутам ([RLS-FR-009])
  • CHK026 Поддержка фильтров WHERE IN / LIKE / regexp (acceptance US4-2)
  • CHK027 Предпросмотр на целевой БД: счётчик + выборка с лимитом ([RLS-FR-010])
  • CHK028 Блокировка сохранения при 0 unit_id (acceptance US4-3)
  • CHK029 Три способа привязки пользователей: AD-логин / AD-группа / оргструктура ([RLS-FR-011])
  • CHK030 Валидация AD-логинов через IDM при ручной привязке (acceptance US4-4)
  • CHK031 Сохранение правила: preview SELECT DISTINCT + UPSERT определения в rls_roles_filter + UPSERT пользователей в rls_rule_users; effect_at ([RLS-FR-012])
  • CHK032 Инвариант: применяемые пары всегда (user_id, unit_id) ([SC-005])
  • CHK033 Diff между поколениями SQL одного правила до применения ([RLS-FR-013])
  • CHK034 Валидация regexp с человекочитаемой ошибкой ([RLS-FR-014])
  • CHK035 Каталог правил с метаданными и количеством unit_id ([RLS-FR-015])
  • CHK036 Экранирование спецсимволов фильтров при генерации SQL (edge case)
  • CHK037 Идемпотентность INSERT: повторное применение не создаёт дубликатов ([RLS-FR-023])

RBAC & Security

  • CHK038 Две операционные роли: rls_operator (правила, привязки, аудит, INSERT) и rls_script_dev (версии скрипта, diff, push) ([RLS-FR-018])
  • CHK039 Глобальный admin наследует обе роли; просмотр аудитов — обе роли + admin ([RLS-FR-018])
  • CHK040 Персональные данные IDM (ФИО, риск-скор, VIP) — только rls_operator и admin ([RLS-FR-018])
  • CHK041 Мутирующие операции проходят существующий механизм require_role (ADR-0005)

Edge Cases & Degradation

  • CHK040 rls_t пустая / скрипт не запускался → явное пустое состояние (edge case)
  • CHK041 Множественные AD-аккаунты пользователя в разных доменах → показ всех, сопоставление по логину (edge case)
  • CHK042 Невалидные ID из IDM трактуются как непрозрачные строки (edge case)
  • CHK043 Конфликт параллельного редактирования правила → модал reload/discard (CONF_01)
  • CHK044 Двойной сабмит применения → кнопка disabled + идемпотентность (DUP_01)
  • CHK045 Справочник >100k строк → лимит выборки + счётчик (LARGE)
  • CHK046 Сбой INSERT на середине операции → ошибка с указанием строки, применённые пары остаются (edge case)
  • CHK047 Динамические правила (R13): definitions в rls_roles_filter + users в rls_rule_users; скрипт v13 вычисляет user→unit при каждом запуске ([RLS-FR-024])
  • CHK048 Структуры rls_roles_mask/rls_bi_roles/rls_bi_users сохраняются; новые rls_roles_filter + rls_rule_users DISTRIBUTED REPLICATED; soft-delete; rls_t без флага ([RLS-FR-025])
  • CHK049 Сохранение правила = определение + привязки (идемпотентно); preview обязателен до save (422); индикатор effect_at ([RLS-FR-023])

Notes

  • Check items off as completed: [x]
  • Все маркеры [NEEDS CLARIFICATION] закрыты: dwh-connection (RLS-FR-020), dataset-rls-mapping (RLS-FR-021), git-push-target (RLS-FR-004).
  • Инвариант фичи: superset-tools НЕ исполняет финальный скрипт rls_t на проде — только готовит версии; обновление rls_t остаётся за внешним контуром. INSERT в bi_users/bi_roles применяется из интерфейса на целевую БД (это не rls_t).

PLAN — Implementation Plan

Source: plan.md

Implementation Plan: RLS Management Workspace

Branch: 048-rls-management-workspace | Date: 2026-08-04 | Spec: spec.md Input: Feature specification from /specs/048-rls-management-workspace/spec.md

Summary

Корпоративная RLS-система Apache Superset получает полноценный инструментарий в superset-tools (внешний оркестратор, ADR-0003): версионирование SQL-скрипта rls_t в git-контуре репозитория с push в корпоративный remote (corp), аудит «пользователь × датасет × unit_id» на снапшотах rls_t с учётом rls_type, аудит bi_users через IDM (уволенные/неактивные), конструктор кастомных RLS-правил (справочник → фильтры IN/LIKE/regexp → живой предпросмотр на DWH → идемпотентное применение INSERT через интерфейс + SQL-миграция в репозиторий). Две операционные роли: rls_operator (правила/привязки/аудит/применение) и rls_script_dev (версии скрипта/push).

Технический подход (research.md): пакет backend/src/services/rls/ (IdmClient, SnapshotService, ScriptStore, AuditService, RuleBuilder C5, BindingsService, UserAuditService) на существующих Core.DbExecutor (DWH: снапшоты + живой предпросмотр/применение определений), Core.Scheduler (cron), шифрованных credentials; API /api/rls/* с permission-энфорсментом; фронт — 4 страницы /rls/* с Screen Models (.svelte.ts), reuse паттернов DashboardHubModel/WizardModel/SelectionModel и дизайн-атомов $lib/ui; прототип 4 экранов уже валидирован (24 состояния, 57/57 токенов). Кастомные правила — динамические определения в новой DWH-таблице rls.rls_roles_filter, вычисляемые скриптом v13 при каждом запуске (R13): обновления справочников «прокручиваются» автоматически, синхронизация/дрейф исключены по построению; структуры rls_roles_mask/rls_bi_roles не меняются.

Technical Context

Language/Version: Python 3.13+ (backend), TypeScript (frontend Svelte 5 runes-only) Primary Dependencies: FastAPI, SQLAlchemy, APScheduler, httpx, gitpython (backend); SvelteKit 5, Vite, Tailwind CSS (frontend) Storage: PostgreSQL 16 (собственная БД superset-tools: снапшоты rls, правила, каталог датасетов, версии-индекс); целевая DWH (Greenplum/ClickHouse) — только через Core.DbExecutor (asyncpg/clickhouse-connect) Testing: pytest (backend, моки DbExecutor/IDM/git), vitest (L1 модели без рендера + L2 компоненты @testing-library/svelte), Playwright (browser-first на implement) Target Platform: Linux server (Docker), современные браузеры Project Type: web application (FastAPI REST backend, SvelteKit SPA frontend) Frontend Architecture: TypeScript-first, model-first (Screen Models .svelte.ts), runes-only, typed callback props Performance Goals: матрица аудита ≤5с на снапшоте (SC-003); предпросмотр правил — лимит выборки 500 + счётчик (LARGE-кейс); снапшот rls_t ≤ 1M строк — bulk insert Constraints: RBAC permission-энфорсмент; rls_t архитектурно не меняется; определения/пользователи динамических правил UPSERTятся в rls_roles_filter/rls_rule_users; superset-tools не исполняет финальный скрипт rls_t Scale/Scope: ~10-50 пользователей платформы; rls_t до ~250k строк; справочники до ~100k+ строк (предпросмотр лимитирован); 4 страницы, 4 модели, ~20 API-эндпоинтов

Constitution Check

GATE: пройден до Phase 0; перепроверен после Phase 1.

  • I. Semantic Contract First: все планируемые модули имеют GRACE-контракты (contracts/modules.md, ATTN-комплаенс проверен: иерархические ID Rls.*/Api.Rls.*/Test.Rls.*, единый доменный ключ rls, anchor-однострочники) ✅
  • II. Decision Memory: research.md фиксирует @RATIONALE/@REJECTED по каждому решению (гибрид снапшот+live, corp-remote push, rls_type-измерение, отказ от live-only/DB-only/sync-requests) ✅
  • III. External Orchestrator: superset-tools остаётся оркестратором — прод-обновление rls_t за внешним контуром; кастомные правила — определения в rls_roles_filter на DWH, скрипт v13 вычисляет роли динамически (R13) ✅
  • IV. Module Discipline: пакет services/rls декомпозирован на 7 модулей (<400 строк каждый); RuleBuilder C5 вынесен отдельно; Invariants: INV_7 учтены ✅
  • V. RBAC Enforcement: permission-модель has_permission("rls"/"rls_scripts"/"idm", ...) через существующий каталог (rbac_permission_catalog.py); мутации — только rls WRITE; персональные данные — idm READ_PERSONAL ✅
  • VI. Svelte 5 Runes Only: 4 Screen Models .svelte.ts, runes-only, typed callback props, requestApi (не fetch) ✅
  • VII. Test-Driven C3+: 25 канонических фикстур (fixtures/), @TEST_EDGE на каждом C3+ контракте, rejected-path покрытие (SQL-инъекция, пустой предпросмотр, 403, идемпотентность) ✅
  • VIII. Attention-Optimized: ATTN_1-4 соблюдены (проверено в contracts/modules.md) ✅
  • ADR-0001: размещение по каноническим границам (api/services/models/schemas) ✅
  • ADR-0005: auth/RBAC поверх существующей модели ✅
  • ADR-0011: async-first (httpx.AsyncClient, async DbExecutor) ✅
  • ADR-0003: оркестраторный паттерн не нарушен ✅

Блокирующих конфликтов не обнаружено.

Project Structure

Documentation (this feature)

specs/048-rls-management-workspace/
├── plan.md              # This file
├── spec.md              # 4 user stories, 23 FR, 5 SC
├── ux_reference.md      # Персона, 4 экрана, failure-матрица
├── research.md          # Phase 0: R1-R10 (решения + @RATIONALE/@REJECTED)
├── data-model.md        # Phase 1: сущности, DTO-пары, Screen Models
├── quickstart.md        # Phase 1: Makefile-таргеты + мок IDM
├── traceability.md      # Phase 1: RTM (gate: T??? до /speckit.tasks)
├── checklists/requirements.md  # 46 проверок
├── prototype/
│   ├── index.html       # 4 экрана, 24 состояния, интерактивные recovery
│   └── manifest.md      # 19/19 состояний, токен-аудит 57/57
├── contracts/
│   └── modules.md       # Контракты: 7 сервисов, 20 роутов, 4 модели, 4 страницы
└── fixtures/
    ├── manifest.md      # 25 фикстур-индекс
    └── api/             # 25 JSON (preview/apply/deactivate/push/snapshot/binding)

Source Code (repository root)

backend/src/
├── api/routes/rls.py            # /api/rls/* (20 эндпоинтов, permission-декларации)
├── services/rls/
│   ├── __init__.py
│   ├── idm_client.py            # Rls.IdmClient (C4)
│   ├── snapshot_service.py      # Rls.SnapshotService (C4) + cron
│   ├── script_store.py          # Rls.ScriptStore (C4, gitpython + corp remote)
│   ├── audit_service.py         # Rls.AuditService (C4)
│   ├── rule_builder.py          # Rls.RuleBuilder (C5)
│   ├── bindings_service.py      # Rls.BindingsService (C4)
│   └── user_audit_service.py    # Rls.UserAuditService (C4)
├── models/rls.py                # Models.Rls (C1)
└── schemas/rls.py               # Schemas.Rls (C1, cross-stack DTO-пары)

frontend/src/
├── routes/rls/
│   ├── +page.svelte             # Rls.DashboardPage
│   ├── audit-datasets/+page.svelte
│   ├── audit-users/+page.svelte
│   └── rules/+page.svelte       # Rls.RulesPage
├── lib/models/
│   ├── RlsDashboardModel.svelte.ts
│   ├── RlsAuditMatrixModel.svelte.ts
│   ├── RlsUserAuditModel.svelte.ts
│   └── RlsRuleBuilderModel.svelte.ts
├── lib/components/rls/          # страницы-компоненты (тонкие, BINDS_TO)
├── lib/api/rls.ts               # Lib.Api.Rls
├── types/rls.ts                 # TS DTO ↔ Schemas.Rls
└── lib/components/layout/sidebarNavigation.ts  # категория «RLS» (requiredPermission "rls")

rls_scripts/                     # версии SQL-скрипта (git-история + version.json)

Structure Decision: web application, отдельные backend/frontend директории; домен RLS — новый пакет services/rls + собственная папка rls_scripts/ для SQL-версий (по образцу dashboards/). DWH-доступ — только через существующий Core.DbExecutor (никаких новых коннекторов).

Semantic Contract Guidance

См. contracts/modules.md — все контракты прошли Attention Compliance Gate (ATTN_1: однострочные anchors; ATTN_2: иерархические Rls.*/Api.Rls.*; ATTN_3: единый ключ rls (+ idm, snapshot, audit, rule как вторичные); ATTN_4: контракты ≤150 строк, модули ≤400).

Ключевые контракты C4/C5 с полными заголовками: Rls.RuleBuilder.Preview (C5, @TEST_EDGE: zero_rows/dwh_error/sql_injection), Rls.RuleBuilder.SaveDefinition (C5, duplicate_binding/dwh_failure/regexp_invalid), Rls.UserAuditService.Deactivate (C4), Rls.ScriptStore.PushToCorp (C4), Rls.SnapshotService.Refresh (C4), Api.Rls.SaveRule (C5). Cross-stack DTO-пары: Schemas.Rls ↔ Lib.Api.Rls/types/rls.ts (общий @SEMANTICS rls).

Один [NEED_CONTEXT: db-executor-query-rows] в Rls.RuleBuilder.Preview — уточняется на implement (возможность fetch-строк у DbExecutor; при отсутствии — расширение DbExecutor read-методом query()).

Complexity Tracking

Fill ONLY if Constitution Check has violations that must be justified

Нарушений нет — таблица пуста.

ADR Continuity (итог)

  • ADR-0001 (module layout): соблюдён — все новые файлы в канонических границах.
  • ADR-0005 (auth RBAC): расширение permission-каталога значениями (rls READ/WRITE, rls_scripts WRITE, idm READ_PERSONAL) — без изменения модели.
  • ADR-0011 (async backend): все внешние вызовы async (httpx), DWH — через async DbExecutor.
  • ADR-0003 (orchestrator): прод-скрипт rls_t не исполняется из superset-tools — только подготовка версий и push.
  • Micro-ADR (feature-local): rls_scripts/ корневой каталог (аналог dashboards/); снапшот-модель (RlsSnapshotMeta+Rows); permission-based RBAC вместо жёстких ролей (фиксируется в research.md R5/R9).

RESEARCH — Technical Decisions

Source: research.md

#region Rls.Plan.Research [C:4] [TYPE ADR] [SEMANTICS research,plan,rls,superset,idm] @BRIEF Phase 0 research for 048-rls-management-workspace — resolves implementation-shaping unknowns with decision memory. @RELATION DEPENDS_ON -> [Doc.Adr.ADR0001] @RELATION DEPENDS_ON -> [Doc.Adr.ADR0005] @RELATION DEPENDS_ON -> [Doc.Adr.ADR0011]

R1. Module placement (backend)

Decision: New domain package backend/src/services/rls/ (business logic), backend/src/api/routes/rls.py (routes), backend/src/models/rls.py (SQLAlchemy), backend/src/schemas/rls.py (Pydantic). External clients inside the package: services/rls/idm_client.py. DWH access reuses existing Core.DbExecutor + Core.ConnectionService — no new DB layer.

Rationale: ADR-0001 boundaries: routes contain no logic; services hold orchestration; models/schemas stay pure. DbExecutor already supports asyncpg/clickhouse-connect pools with connection resolution via connection_service.get_connection() — the RLS feature needs exactly this for snapshot reads and rule preview/apply on the target DWH.

Alternatives Considered:

  • Plugin (backend/src/plugins/rls/) — rejected: ADR-0004 plugin restrictions and subprocess semantics do not fit a first-class operational domain with own DB models and scheduled jobs.
  • Direct asyncpg client without DbExecutor — rejected: duplicates pool lifecycle, config resolution, and dialect routing already in Core.DbExecutor; violates constitution principle IV (module discipline).

Impact On Contracts / Tasks: New contracts Rls.IdmClient, Rls.SnapshotService, Rls.ScriptStore, Rls.AuditService, Rls.RuleBuilder, Rls.BindingsService, Api.Rls.*. DbExecutor may need a read/query method (see R2).

R2. DWH connectivity mode (clarified hybrid)

Decision: Snapshot rls_t / rls_bi_users / rls_bi_roles into superset-tools own PostgreSQL for audit screens; rule preview and SaveDefinition execute live SQL on DWH (SELECT DISTINCT preview, UPSERT rls_roles_filter + rls_rule_users).

Rationale: Matches clarify session (hybrid). Audit matrix must answer in ≤5s and survive DWH outages; preview must reflect exactly what the final script will compute, so it must run on the DWH (user requirement: «инструмент должен полностью выполняться на целевой БД»).

Alternatives Considered:

  • Live-only reads — rejected: audit degrades when DWH down, conflicts with SC-003 latency.
  • Snapshot-only preview — rejected: preview may diverge from production reference data.
  • File import — rejected: no live preview possible.

Open gap: DbExecutor.execute_sql returns DbExecutionResult (rows_affected/timing). SELECT preview needs fetched rows. Decision: extend DbExecutor with a read-query method query(connection_id, sql, limit) (C3, read-only) OR reuse existing SQL Lab path. Final call in contract Rls.RuleBuilder.Preview — check Core.DbExecutor internals during implementation; if rows are already exposed, no new method needed. Emit [NEED_CONTEXT: db-executor-query-rows] in preview contract.

Impact On Contracts / Tasks: Rls.SnapshotService (C4, scheduler cron via Core.Scheduler APScheduler — CronTrigger pattern at scheduler.py:175), Rls.RuleBuilder (C5) with preview/apply, Api.Rls.SnapshotRefresh.

R3. IDM client & mock extension

Decision: Rls.IdmClient — httpx.AsyncClient, base URL + optional token from config.json block rls (idm_url, idm_token), timeouts 10s, retry ×2, total degradation → IdmUnavailableError with stable envelope. Methods: get_person_info, get_person_extensions (APPS), get_account_info, plus search endpoints search_by_ou(ou) and search_by_department(department) for group/org bindings (US4). The mock in research/rls/idm-mock is extended with the two search endpoints (same response shapes as real service contract).

Rationale: Mock mirrors real API 1:1 (README: structure identical); adding search to both mock and client keeps dev/prod parity. Degradation pattern mirrors Services.SupersetLookupService (success or degraded payload, stable shape) — the established project pattern for external identity lookups.

Alternatives Considered:

  • Sync requests — rejected: project is async-first (ADR-0011).
  • Caching IDM responses in DB — rejected for v1: IDM data is lookup-grade (not snapshot-grade); audit screen already degrades gracefully. Revisit if IDM latency hurts UX.

Impact On Contracts / Tasks: Rls.IdmClient (C4), Rls.UserAudit enrichment (C4), Rls.BindingsService (C4). Mock extension tasks in US3/US4 (test data: OU paths from users.json).

R4. Snapshot data model

Decision: SQLAlchemy models in models/rls.py:

  • RlsSnapshotMeta — id, taken_at, source, row_count, status (pending/success/failed), error
  • RlsSnapshotRow — snapshot_meta_id FK, user_id, unit_id, source_system, origin_role, business_area_type, rls_type
  • RlsBiUserSnapshot — snapshot_meta_id FK, ad_user, role, role_descr, is_deleted
  • RlsBiRoleSnapshot — snapshot_meta_id FK, role, unit_id, rls_type, source_system

Refresh = TRUNCATE previous snapshot rows for the meta + bulk insert (mirrors DWH script's own TRUNCATE+INSERT). Snapshot meta keeps N latest snapshots (default 7, configurable).

Rationale: Audit queries join snapshot rows against dataset catalog; versioned snapshots give «age of data» visibility (SC) and staleness banners. Keeping 7 generations supports time-window comparisons (edge case «расхождение rls_t vs bi_users»).

Alternatives Considered:

  • Overwrite single snapshot — rejected: loses age/trend context needed for «актуальность» monitoring.
  • Storage in ClickHouse — rejected: own PostgreSQL is the project SSOT for operational state.

Impact On Contracts / Tasks: Rls.SnapshotService (C4), models (C1), Api.Rls.Overview reads meta + counts.

R5. Script versioning & corp push

Decision: SQL script versions live as files in repo root rls_scripts/ (v012_rls_t_main_insert.sql + version.json with author/comment/active). Rls.ScriptStore (gitpython Repo on the repo itself) creates versions (copy active → edit), computes diffs, marks active. Push to corporate repo = add corp remote + push with PAT from the project's encrypted credentials mechanism (core/encryption.py), reusing git patterns from Services.Git (_clone_with_auth, push_changes(pat)).

Rationale: Clarify answer A (second git remote). File-based versions give the external DWH contour a plain directory to deploy from; git history is the version store — no separate versions table needed (diff = git diff).

Alternatives Considered:

  • Versions in DB table — rejected: diff/push/review would need a custom git emulation; the corporate pipeline consumes a git repo.
  • Export patches via UI — rejected: manual, no history.

Impact On Contracts / Tasks: Rls.ScriptStore (C4), Api.Rls.Scripts*, Api.Rls.Push. New root dir rls_scripts/ (mirrors dashboards/ convention).

R6. Dataset catalog (audit matrix, clarified)

Decision: RlsDatasetBinding model: dataset_key (Superset dataset UUID), dataset_name, unit_column, unit_type (unit_balance_code | plant_code), detection (auto | manual), confirmed (bool). Auto-detection: Superset API dataset metadata → column names matching unit_id|unit_balance_code|plant_code → proposal rows in auto; admin confirms/overrides in UI (hybrid, clarify Q2). Audit matrix joins snapshot rows to bindings by unit_type × rls_type; datasets without bindings = «без RLS-фильтрации» group.

Rationale: rls_type is the missing dimension (user clarified): a user with user_to_unit_balance_code sees only datasets bound to unit_balance_code. rls_t stays untouched architecturally (FR-021/022).

Alternatives Considered: manual-only catalog (rejected — new datasets require per-dataset manual work); Superset-internal RLS objects (rejected — different mechanism, rls_t unused).

Impact On Contracts / Tasks: Rls.AuditService (C4), Api.Rls.Datasets.Bindings*, Api.Rls.Audit.Datasets*.

R7. Rule builder (custom RLS)

Decision: RlsRule model: name, reference_table (schema.table on DWH), unit_id_column, rls_type (string, extensible), filters (JSON: list of {column, op: in|like|regexp, value}), created_by, updated_at. RlsRuleBinding model: rule_id, ad_user, source_type (manual | ad_group | org_unit), source_ref (login | group/OU dn | department id), role_name (generated BI_<prefix>_<unit_id> or manual). Flow: preview (live SELECT DISTINCT unit_id + row limit 500 + count) → bind users (manual autocomplete via IDM / group search / org tree) → apply: idempotent INSERTs (existence check per (role, unit_id) and (ad_user, role)) executed on DWH via DbExecutor + generated SQL migration saved into rls_scripts/migrations/<rule>/v<N>_<ts>.sql + registered as script version.

Rationale: Mirrors the existing rls_bi_roles migration pattern (SELECT DISTINCT ... FROM dict_dds... — seen in REPOSITORY_ANALYSIS §4.3). Idempotency (FR-023) prevents double-apply duplicates. Generated SQL doubles as history and external-contour artifact.

Alternatives Considered:

  • Only generate SQL without live INSERT — rejected (user: «INSERT ... должен быть доступен из интерфейса»).
  • Rule definition only as SQL files — rejected (clarify: definitions are entities; SQL is generated output).

Impact On Contracts / Tasks: Rls.RuleBuilder (C5, belief markers), Rls.BindingsService (C4), Api.Rls.Rules*, fixtures for preview/apply/regexp-validation.

R8. Frontend topology

Decision: Routes under frontend/src/routes/rls/: +page.svelte (dashboard/overview + script versions), audit-datasets/+page.svelte, audit-users/+page.svelte, rules/+page.svelte (catalog + builder wizard). Screen Models (.svelte.ts, from ux_reference): RlsDashboardModel, RlsAuditMatrixModel, RlsUserAuditModel, RlsRuleBuilderModel. Reuse: list/filter/pagination pattern from DashboardHubModel.svelte.ts; wizard-step pattern from Migration.WizardModel.svelte.ts; selection submodel pattern from Dashboards.SelectionModel.svelte.ts. DTOs in frontend/src/types/rls.ts mirroring schemas/rls.py. API client frontend/src/lib/api/rls.ts using requestApi wrapper (native fetch forbidden).

Rationale: Model-first per semantics-svelte §IIIa — cross-widget invariants (filter resets, step gating, apply disabled on empty preview) live in models, verified without render (L1). Sidebar entry: new category in layout/sidebarNavigation.ts with requiredPermission: "rls".

Alternatives Considered: Component-first with inline $state — rejected (invariants scattered across onclick handlers).

Impact On Contracts / Tasks: 4 models (C4), components components/rls/* (C3, @RELATION BINDS_TO), routes (C2/C3).

R9. RBAC & permissions

Decision: New permissions discovered via the existing has_permission("rls", "READ"|"WRITE"), ("rls_scripts", "WRITE"), ("idm", "READ_PERSONAL") catalog mechanism (rbac_permission_catalog.py auto-discovers from route files). Two operational roles map to permission sets: rls_operator → rls READ+WRITE, idm READ_PERSONAL; rls_script_dev → rls READ, rls_scripts WRITE. Global admin inherits all (existing admin-role logic). Mutating endpoints enforce has_permission (ADR-0005 pattern, not require_role string role — permissions are the granular unit; roles are DB-managed groups).

Rationale: ADR-0005 RBAC; permission-based enforcement keeps role assignment flexible (two clarify-договорённые operational roles can be assigned by admin without code change).

Alternatives Considered: Hardcoded new roles in code — rejected: permission catalog + DB roles is the established extensible path.

Impact On Contracts / Tasks: Api.Rls.* route contracts carry permission declarations; Api.Rls.Push requires rls_scripts WRITE; Api.Rls.Rules.Apply requires rls WRITE.

R10. Test strategy

Decision: Backend pytest: unit tests with patched DbExecutor (fake connection results), mocked IDM transport (httpx MockTransport), temp git repo for ScriptStore (real gitpython, in tmp dir). Frontend vitest: L1 model invariants without render (pattern: models/__tests__/DashboardHubModel.test.ts), L2 component/UX tests with @testing-library/svelte (pattern: components/__tests__/*.ux.test.ts). Rejected-path regression: tests proving double-apply does not duplicate, deactivation is idempotent, IDM-unavailable envelope, preview-on-empty blocks apply.

Rationale: semantics-testing §VI-VII; fixtures in specs/048/fixtures/ are the hardcoded expected values (anti-tautology); @TEST_EDGE per C3+ contract.

Alternatives Considered: Integration tests with real Greenplum/Superset — rejected for v1 (no DWH in CI); covered by mocked DbExecutor + testcontainers path exists for later.

Impact On Contracts / Tasks: Fixtures (canonical JSON), Test.Rls.* modules, Test.Rls.RuleBuilder rejected-path coverage.

R11. INSERT target for custom rules & impact on rls_t script (architecture answer)

Decision: Custom rule INSERTs go exclusively into the existing rls.rls_bi_roles (one row per unit_id: role = role_prefix||unit_id, rls_type, source_system='CUSTOM') and rls.rls_bi_users (ad_user, role, is_deleted=false). No new tables. The rls_t main insert.sql script requires NO changes for the base case: the existing CTE bi_users_dwh (bi_users JOIN bi_roles → unit_id) picks custom rules up automatically on next run; TRUNCATE touches only rls_t, never bi_users/bi_roles. Invariant: rls_t is generated exclusively by the script, always yielding (user_id, unit_id) pairs.

Rationale: Mirrors the existing rls_bi_roles migration pattern (REPOSITORY_ANALYSIS §4.3: SELECT DISTINCT 'BI_bukrs_'||unit_balance_code FROM dict_dds...). Rule builder is a parameterization of that pattern (role_prefix + unit_id from live preview). Deactivation via is_deleted=true is already honored by the script's WHERE is_deleted = false. BI_ALL priority (cross join) is unaffected.

Alternatives Considered:

  • New custom-rule table + script rewrite — rejected: duplicates the bi_users/bi_roles machinery, breaks the external contour's deploy pipeline, requires script surgery for zero functional gain.
  • Direct INSERT into rls_t — rejected: violates the single-generator invariant; next TRUNCATE+INSERT run would erase manual edits.

Optional script changes (only if transparency wanted):

  • source_system='CUSTOM' visibility in rls_t → 1-line change in bi-ветка/финальный INSERT → new script version v13 (diff tracked in rls_scripts/).
  • New rls_type values (e.g. user_to_center_responsibility) → value-only extension, no ALTER (FR-022); Superset-side filtering changes are outside the script.

Impact On Contracts / Tasks: Rls.RuleBuilder.Apply writes exactly these two INSERTs (idempotent, FX_Rls.RuleBuilder.Apply.*); Rls.RuleBuilder.GenerateSql emits the same SQL as migration artifact; Rls.ScriptStore versions any optional script change; audit displays source_system including 'CUSTOM'.

R12. Reference-table drift — REJECTED: materialized rule sync (superseded by R13)

Decision: REJECTED. Earlier design materialized custom rules into rls_bi_roles slices with sync (INSERT + soft-delete + reactivation + confirm + drift indicator). Superseded by R13 dynamic filter-based rules — materialization creates drift by construction, requires sync machinery, confirm flows, and reactivation edge cases, all of which disappear when rules are computed inside the script on each run. Decision memory retained for the exploration path: sync variants considered (one-shot apply, dynamic CTE in script, silent auto-DELETE) and rejected as noted below.

Rationale (retained): Materialization mirrors the existing rls_bi_roles migration pattern but couples rule definitions to a snapshot of reference data; reference updates (БЕ added/removed) then require re-apply to stay correct, and removed units keep granting access until sync runs.

Rejected paths (carried forward):

  • One-shot apply without sync — rejected: drift accumulates, removed БЕ keep granting access.
  • Silent auto-DELETE of stale roles — rejected: destructive without operator visibility.
  • is_deleted flag on rls_t — rejected: ephemeral table (TRUNCATE+INSERT rebuild), flag meaningless after next run.

R13. Dynamic filter-based custom rules (ACCEPTED — replaces R7 materialized apply + R12 sync)

Decision: Custom rules use two new DWH tables: rls.rls_roles_filter (definition: id, reference_table, unit_id_column, role_prefix, rls_type, filters, is_deleted) and rls.rls_rule_users (rule_id, user_id, source_type, source_ref, is_deleted), both DISTRIBUTED REPLICATED. Script v13 computes active rule_id→unit_id from reference data, joins active rule_id→user_id, and emits user_id→unit_id. rls_bi_roles/rls_bi_users retain the static BI-role branch; rls_roles_mask remains SAP regexp masks.

UI/API consequences: «Save» = own-DB SSOT + UPSERT rls_roles_filter + UPSERT rls_rule_users; effect at next script run (effect_at). Preview is the exact live unit slice. Rule/user membership deactivation uses is_deleted=true. No per-rule materialized-role SQL.

Rationale: Fully satisfies «инструмент должен полностью выполняться на целевой БД, как он в итоге будет обновляться скриптом»; correctness by construction; removes the entire sync surface (R7 Apply + R12 Sync, fixtures, drift indicator, confirm dialogs, reactivation). Rule preview doubles as the script's own computation, so preview-apply mismatch is impossible.

Alternatives Considered:

  • Materialized roles in rls_bi_roles + sync (R7/R12) — rejected: drift, sync machinery, confirm flows, reactivation edge cases.
  • Changing rls_roles_mask structure to host rule definitions — rejected: mixes SAP mask semantics with rule definitions, risks SAP branch and backward compatibility; new table rls_roles_filter is semantically clean with same replication profile.
  • Rules stored only in superset-tools DB (no DWH table) — rejected: script cannot read them; DWH table is the script-facing contract.

Impact On Contracts / Tasks: Rls.RuleBuilder.Preview unchanged (live slice); Rls.RuleBuilder.Apply → replaced by Rls.RuleBuilder.SaveDefinition (own DB + UPSERT rls_roles_filter + bi_users bindings); Rls.RuleBuilder.Sync (R12) removed; Api.Rls.Rules.Save replaces Apply/Sync endpoints; fixtures: rule_sync_* replaced by rule_define_* (definition save, definition deactivate, definition reactivate = is_deleted toggle); RlsRule gains is_deleted + latency fields; script v13 = one CTE addition; RlsOverviewResponse drift counter removed (impossible by construction), replaced by rules-active count.

R7. Rule builder (custom RLS) — SUPERSEDED by R13

Decision: SUPERSEDED. Original design materialized rules into rls_bi_roles/rls_bi_users (INSERT + idempotency + per-rule SQL migrations). Replaced by R13 dynamic filter-based rules (definitions in rls.rls_roles_filter, computed inside the script). Retained elements: RlsRule entity shape (name, reference_table, unit_id_column, rls_type, role_prefix, filters JSON), preview flow (live SELECT DISTINCT + limit 500 + count), binding flows (manual/group/org via IDM), regexp validation. Dropped elements: materialized INSERT of roles, per-rule migration files, apply-idempotency against bi_roles, RlsRuleBinding role_name derivation from unit_id.

Rationale (retained): Mirrors the existing rls_bi_roles migration pattern (SELECT DISTINCT ... FROM dict_dds... — REPOSITORY_ANALYSIS §4.3) — R13 moves that same pattern into the script's CTE, so the rule definition IS the migration.

R11. INSERT target for custom rules & impact on rls_t script (SUPERSEDED by R13)

Decision: SUPERSEDED. Original answer: INSERTs into rls_bi_roles + rls_bi_users. R13 replaces role materialization: definitions → rls.rls_roles_filter (UPSERT), bindings → rls_bi_users (INSERT, applied immediately; effect at next script run). rls_bi_roles remains for static BI roles only. Invariant unchanged: rls_t is generated exclusively by the script, always (user_id, unit_id) pairs, and now always reflects current reference state by construction.

Retained invariants: TRUNCATE touches only rls_t; BI_ALL priority unaffected; deactivation of bindings via is_deleted=true honored by script; source_system='CUSTOM' visibility in rls_t remains an optional 1-line script change (v13.1 if wanted).

R14. Explainability layer lives in DWH (user decision) — evidence via script, not reconstruction

Decision: rls_access_evidence (+ rls_rejection_evidence, rls_run) are DWH tables populated BY the main rls_t script (inserts BEFORE the GROUP BY aggregation), not reconstructed in superset-tools. Rationale (user): DWH-side join with datasets for validation/analytics (Dataset Guard) requires evidence co-located with dataset tables; per-request reconstruction in the tool was evaluated and rejected for this reason. superset-tools role: version the script change that writes evidence (v14 via ScriptStore — «помощь в разработке»), snapshot metadata/watermarks, and read evidence on demand for UI explanations («почему есть/нет доступ» in user card) via targeted indexed SELECT on DWH (no full snapshot of evidence — volume ×5-10 of rls_t). Rejections (rls_rejection_evidence) likewise script-generated; tool displays per (user, unit) lookups.

Rationale: Keeps single source of truth (the script) for BOTH rls_t and evidence — no interpreter drift risk. Tool remains orchestrator/observer, never a logic source (ADR-0003). Dataset validation joins happen naturally in DWH.

Alternatives Considered:

  • Evidence reconstruction in superset-tools (R14-predecessor) — rejected: no direct join with dataset tables, interpreter drift risk, duplicates script logic.
  • Full evidence snapshot into tool DB — rejected: volume ×5-10 of rls_t, no analytical value outside DWH.

Impact On Contracts / Tasks: 048 unchanged (v13 = dynamic rules; evidence is v14, a follow-up explainability feature with DWH-team agreement on script change). Explainability = script v14 (evidence inserts) + Rls.EvidenceReader (targeted SELECT by user/unit) + UI «почему» buttons in user card + rejection display. Snapshot meta already carries source_watermarks groundwork (see R2/R4 extension candidate).

#endregion Rls.Plan.Research


DATA MODEL — Entities & Relations

Source: data-model.md

#region Rls.Plan.DataModel [C:3] [TYPE ADR] [SEMANTICS datamodel,rls,snapshot,rule] @BRIEF Data model for 048-rls-management-workspace — SQLAlchemy entities, Pydantic schemas, TypeScript DTOs, Screen Model state. @RELATION DEPENDS_ON -> [Rls.Plan.Research]

1. SQLAlchemy entities (backend/src/models/rls.py)

RlsSnapshotMeta

Field Type Notes
id int PK
taken_at datetime Время снятия снапшота
source str 'dwh' (будущее: 'import')
row_count int Строк rls_t в снапшоте
status str pending / success / failed
error str | None Текст ошибки при failed

RlsSnapshotRow (снапшот rls_t)

Field Type Notes
id int PK
snapshot_meta_id int FK → RlsSnapshotMeta
user_id str AD-логин lowercase (как в rls_t)
unit_id str Код единицы
source_system str SAP / SAP_ISO / FM_COMMENTS_INPUT
origin_role str | None Агрегированные роли
business_area_type str | None SD/FI/MM/LE/BI
rls_type str user_to_unit_balance_code / user_to_plant_code / кастомные

Индексы: (user_id), (unit_id), (rls_type), (snapshot_meta_id). Uniqueness не требуется (снапшот-копия), но дубликаты исключаются на этапе загрузки (как GROUP BY в скрипте).

RlsBiUserSnapshot / RlsBiRoleSnapshot

Поле (users) Тип Поле (roles) Тип
id PK int id PK int
snapshot_meta_id FK int snapshot_meta_id FK int
ad_user str role str
role str unit_id str
role_descr str | None rls_type str
is_deleted bool source_system str | None
is_deleted bool

Примечание (R13): структуры DWH-таблиц rls_bi_roles, rls_bi_users и rls_roles_mask НЕ изменяются. Динамические правила используют две новые DISTRIBUTED REPLICATED таблицы: rls_roles_filter (определения) и rls_rule_users (rule_id, user_id, source_type, source_ref, is_deleted). Скрипт v13 соединяет обе таблицы со справочником и формирует user_id → unit_id. rls_t флаг не получает.

RlsScriptVersion (метаданные версий скрипта)

Файлы живут в rls_scripts/vNNN_*.sql + version.json; в БД — только индекс для UI-скорости:

Field Type Notes
id int PK
version str v012
filename str v012_rls_t_main_insert.sql
author str AD-логин
comment str | None
is_active bool Одна активная
pushed_at datetime | None Последний push в corp
created_at datetime

RlsDatasetBinding

Field Type Notes
id int PK
dataset_key str UUID датасета Superset
dataset_name str
unit_column str Колонка-носитель единицы
unit_type str unit_balance_code / plant_code
detection str auto / manual
confirmed bool Ручное подтверждение
updated_at datetime

Уникальность: (dataset_key, unit_column, unit_type) — датасет может иметь несколько привязок (edge case).

RlsRule

Field Type Notes
id int PK
name str «БЕ по ЦО»
reference_table str schema.table на DWH
unit_id_column str pfm
rls_type str Расширяемое значение
role_prefix str BI_czo_
filters JSON [{column, op: in|like|regexp, value}]
is_deleted bool Мягкая деактивация определения правила
roles_expected int | None Последний preview-счётчик (ожидаемый срез)
effect_at str | None Ближайший запуск скрипта (индикация латентности)
created_by str
created_at / updated_at datetime

RlsRuleBinding / DWH rls_rule_users

Field Type Notes
id int PK
rule_id int FK → RlsRule
user_id str Канонический AD-логин
source_type str manual / ad_group / org_unit
source_ref str | None OU DN / department id / группа
created_at datetime

Уникальность: (rule_id, user_id). source_type/source_ref сохраняют происхождение привязки; is_deleted обеспечивает мягкую деактивацию.

2. Pydantic schemas (backend/src/schemas/rls.py) ↔ TypeScript DTOs (frontend/src/types/rls.ts)

Cross-stack пары (имя класса ↔ интерфейс TS), общий @SEMANTICS ключ rls:

Pydantic TS DTO Поля (кратко)
RlsOverviewResponse RlsOverview snapshot: {taken_at, row_count, age_days, stale: bool}; active_version; audit: {fired, inactive, no_rls}
RlsScriptListItem RlsScriptListItem version, filename, author, comment, is_active, pushed_at
RlsScriptDiffResponse RlsScriptDiff version_a, version_b, diff: string (unified)
RlsPushRequest / Response RlsPushResult version, ok, message
RlsSnapshotRefreshResponse RlsSnapshotRefreshResult meta_id, status, row_count
RlsBiUserRow RlsBiUserRow ad_user, role, idm_status: {employee: active|fired|vacation|unknown, account: enabled|disabled|unknown, risk_score, vip, admin_rights}, updated_at
RlsUserCardResponse RlsUserCard name, department, job, company, date_in, date_out, accounts: [{app, account_name, ou, enabled, admin_rights}]
RlsDeactivateRequest RlsDeactivateRequest ad_users: string[]
RlsDatasetBindingDto RlsDatasetBindingDto dataset_key, dataset_name, unit_column, unit_type, detection, confirmed
RlsAuditMatrixResponse RlsAuditMatrix direction, rows: [{dataset, unit_type, unit_ids[], source_systems[], roles[]}], no_rls_datasets[]
RlsRuleDto RlsRuleDto id, name, reference_table, unit_id_column, rls_type, role_prefix, filters: RlsFilterDto[], unit_count
RlsFilterDto RlsFilterDto column, op: 'in'|'like'|'regexp', value
RlsRulePreviewResponse RlsRulePreview unit_count, sample: [{unit_id, ...attrs}] (limit 500)
RlsRuleBindingRequest RlsRuleBindingRequest rule_id, source_type, ad_users[] | group_ref | org_ref
RlsApplyResponse RlsApplyResponse deactivated_users, skipped_users, migration_path (только аудит пользователей)
RlsRuleSaveResponse RlsRuleSaveResult rule_id, roles_expected, rule_users_saved, duplicates_skipped, effect_at

3. Screen Models (frontend/src/lib/models/, .svelte.ts)

RlsDashboardModel — @STATE idle/loading/loaded/stale/push-failed; атомы snapshot, activeVersion, versions[], auditCounters; @ACTION load(), createVersion(), activate(v), push()

Инварианты: активная версия одна; stale = age > threshold.

RlsAuditMatrixModel — @STATE idle/searching/loaded/empty/degraded; атомы direction ('user'|'dataset'), query, matrix, snapshotAge; @ACTION search(), setDirection(), retry()

Инвариант: смена direction сбрасывает результат; searching при пустом запросе запрещён.

RlsUserAuditModel — @STATE loading/loaded/idm-degraded/empty; атомы rows[], filter ('all'|'fired'|'inactive'|'vacation'|'admin'), selected: Set; @ACTION load(), setFilter(), toggleSelect(), deactivate()

Инварианты: фильтр применяется к уже загруженным строкам (без повторного запроса); кнопка деактивации disabled при пустом selected.

RlsRuleBuilderModel — @STATE step-1..4, preview-loading/preview-empty/preview-loaded, saving/saved/error; атомы rule (draft), filters[], preview, ruleUsers[], step, effectAt; @ACTION next(), setReference(), addFilter(), validateRegexp(), preview(), bindManual(), bindGroup(), bindOrg(), save()

Инварианты: save требует выполненный preview и хотя бы одного rule user; zero-unit preview разрешён с предупреждением; смена справочника сбрасывает фильтры и preview; saved показывает effect_at.

#endregion Rls.Plan.DataModel


CONTRACTS — Module & Function Contracts

Source: contracts/modules.md

#region Rls.Modules [C:5] [TYPE Module] [SEMANTICS rls,modules,contracts] @defgroup Rls Design contracts for 048-rls-management-workspace — backend services, API, frontend models/components. @RELATION DEPENDS_ON -> [Rls.Plan.Research] @RELATION DEPENDS_ON -> [Rls.Plan.DataModel]

Все контракты проходят Attention Compliance Gate (ATTN_1-4): иерархические ID, один доменный ключ rls, anchor на одной строке.


Backend: services/rls (пакет)

#region Rls.IdmClient [C:4] [TYPE Module] [SEMANTICS rls,idm,client] @defgroup RlsServices IDM HTTP client with degradation envelope. @BRIEF Async IDM client (GetPersonInfo/GetPersonExtensionInfo/GetAccountInfo + search by OU/department) with stable degraded response. @RELATION DEPENDS_ON -> [EXT:httpx:AsyncClient] @RATIONALE Mirrors Services.SupersetLookupService degradation pattern (established for external identity lookups). Mock and prod share wire shape (research/rls/idm-mock README). @REJECTED Sync requests — rejected (ADR-0011 async-first). Response caching in DB — rejected for v1 (lookup-grade data, graceful degradation covers outage).

#region Rls.IdmClient.GetPersonInfo [C:4] [TYPE Function] [SEMANTICS rls,idm,card]

@ingroup RlsServices

@BRIEF Fetch employee card by uid/login/accountName.

@PRE client configured with base_url; identifier non-empty.

@POST Returns IdmPersonInfo or degraded None with warning flag.

@SIDE_EFFECT External HTTP call; retry ×2 on timeout.

@RELATION DEPENDS_ON -> [DTO:IdmPersonInfoDto]

@TEST_EDGE: idm_unavailable -> degraded envelope, no raise

@TEST_EDGE: malformed_id -> id treated as opaque string

#endregion Rls.IdmClient.GetPersonInfo

#region Rls.IdmClient.SearchByOu [C:4] [TYPE Function] [SEMANTICS rls,idm,search]

@ingroup RlsServices

@BRIEF Search accounts by AD group/OU (mock extension, US4 mass binding).

@PRE ou non-empty.

@POST List of IdmAccountRef (login, ou, enabled) or degraded [].

@SIDE_EFFECT External HTTP call.

@RELATION DEPENDS_ON -> [DTO:IdmAccountRefDto]

@TEST_EDGE: group_not_found -> empty list + warning, binding not created

#endregion Rls.IdmClient.SearchByOu

#region Rls.IdmClient.SearchByDepartment [C:4] [TYPE Function] [SEMANTICS rls,idm,search]

@ingroup RlsServices

@BRIEF Search employees by org structure (department/БЕ) for org-unit binding.

@PRE department identifier non-empty.

@POST List of IdmPersonRef (uid, login, department) or degraded [].

@SIDE_EFFECT External HTTP call.

@TEST_EDGE: department_unknown -> empty list + warning

#endregion Rls.IdmClient.SearchByDepartment

#endregion Rls.IdmClient

#region Rls.SnapshotService [C:4] [TYPE Module] [SEMANTICS rls,snapshot,dwh] @defgroup RlsServices Snapshot rls_t/bi_users/bi_roles into own PostgreSQL. @BRIEF Scheduled + manual snapshot refresh from target DWH via Core.DbExecutor; TRUNCATE+bulk insert per snapshot generation. @RELATION DEPENDS_ON -> [Core.DbExecutor] @RELATION DEPENDS_ON -> [Core.Scheduler] @RELATION DEPENDS_ON -> [Models.Rls] @RATIONALE Hybrid clarified model: audit reads own snapshots (≤5s, outage-safe), preview stays live. Keeps N generations for staleness/trend views. @REJECTED Live-only audit reads — rejected (SC-003 + DWH outage resilience). Single overwrite snapshot — rejected (loses age context).

#region Rls.SnapshotService.Refresh [C:4] [TYPE Function] [SEMANTICS rls,snapshot,refresh]

@ingroup RlsServices

@BRIEF Execute full snapshot refresh: rls_t + bi_users + bi_roles, one meta generation.

@PRE dwh connection_id configured; rls source tables exist.

@POST New RlsSnapshotMeta (success/failed) with row_count; previous generations retained (default 7).

@SIDE_EFFECT Reads DWH (SELECT via DbExecutor); writes snapshot tables; belief markers REASON/REFLECT.

@DATA_CONTRACT Input: connection_id -> Output: RlsSnapshotRefreshResponse

@RELATION DEPENDS_ON -> [DTO:RlsSnapshotRefreshResponse]

@TEST_EDGE: dwh_unreachable -> meta status=failed, error text, no partial rows

@TEST_EDGE: empty_source -> success with row_count=0 (empty state on UI)

#endregion Rls.SnapshotService.Refresh

#region Rls.SnapshotService.RegisterCron [C:3] [TYPE Function] [SEMANTICS rls,snapshot,cron]

@ingroup RlsServices

@BRIEF Register daily cron (configurable, default 05:00 UTC) via Core.Scheduler CronTrigger.

@PRE scheduler running.

@POST Job registered; manual refresh endpoint also available.

#endregion Rls.SnapshotService.RegisterCron

#endregion Rls.SnapshotService

#region Rls.ScriptStore [C:4] [TYPE Module] [SEMANTICS rls,script,git] @defgroup RlsServices SQL script versioning in repo rls_scripts/ + corp remote push. @BRIEF Version lifecycle for rls_t main insert.sql: create from active, diff, activate, push to corporate repo (remote corp). @RELATION DEPENDS_ON -> [EXT:gitpython:Repo] @RELATION DEPENDS_ON -> [Core.Encryption] @RELATION DEPENDS_ON -> [Models.Rls.RlsScriptVersion] @RATIONALE Clarify answer A: second git remote; file-based versions give the external contour a deploy-ready directory. git history IS the version store (diff = git diff). @REJECTED DB-only versions — rejected (corporate pipeline consumes git). UI patch export — rejected (no history, manual).

#region Rls.ScriptStore.CreateVersion [C:4] [TYPE Function] [SEMANTICS rls,script,version]

@ingroup RlsServices

@BRIEF Copy active script to new version file, record metadata, keep active untouched.

@PRE repo writable; active version exists (or seed from template).

@POST New version file vNNN_*.sql + version.json entry; active flag unchanged.

@SIDE_EFFECT File write + git add/commit on branch; belief markers.

@TEST_EDGE: no_active_version -> seeds from bundled template, first version

@TEST_EDGE: duplicate_version -> version number auto-incremented

#endregion Rls.ScriptStore.CreateVersion

#region Rls.ScriptStore.Diff [C:3] [TYPE Function] [SEMANTICS rls,script,diff]

@ingroup RlsServices

@BRIEF Unified diff between two versions.

@PRE both versions exist.

@POST Diff string (line-based added/removed/modified).

#endregion Rls.ScriptStore.Diff

#region Rls.ScriptStore.Activate [C:3] [TYPE Function] [SEMANTICS rls,script,active]

@ingroup RlsServices

@BRIEF Mark version active (single active invariant), commit.

@POST Exactly one is_active=true.

@SIDE_EFFECT DB flag update + git commit.

#endregion Rls.ScriptStore.Activate

#region Rls.ScriptStore.PushToCorp [C:4] [TYPE Function] [SEMANTICS rls,script,push]

@ingroup RlsServices

@BRIEF Push rls_scripts/ to corporate repo via remote corp with PAT from encrypted credentials.

@PRE corp remote configured; PAT decrypted.

@POST Push attempted; pushed_at updated on success; error envelope on failure without data loss.

@SIDE_EFFECT git push (network); belief markers.

@RELATION DEPENDS_ON -> [Rls.ScriptStore]

@TEST_EDGE: corp_unreachable -> error envelope, local version intact, retry possible

@TEST_EDGE: no_corp_remote -> configuration error with hint

#endregion Rls.ScriptStore.PushToCorp

#endregion Rls.ScriptStore

#region Rls.AuditService [C:4] [TYPE Module] [SEMANTICS rls,audit,datasets] @defgroup RlsServices Dataset × user audit matrix over snapshot + catalog. @BRIEF Build audit matrices (user→datasets×unit_id, dataset→users) joining RlsSnapshotRow with RlsDatasetBinding by unit_type×rls_type; no-RLS group separate. @RELATION DEPENDS_ON -> [Models.Rls] @RELATION DEPENDS_ON -> [Core.SupersetClient] @RATIONALE rls_type is the join dimension (clarify Q2): user's rls_type must match binding unit_type; datasets without bindings surface as «без RLS-фильтрации». @REJECTED Joining by unit_id value alone — rejected (unit_balance_code vs plant_code namespaces collide).

#region Rls.AuditService.MatrixByUser [C:4] [TYPE Function] [SEMANTICS rls,audit,user]

@ingroup RlsServices

@BRIEF Build user→datasets matrix from snapshot.

@PRE snapshot exists; user login normalized lowercase.

@POST Matrix rows (dataset, unit_type, unit_ids, source_systems, roles) + no_rls_datasets; empty result → empty state envelope.

@TEST_EDGE: user_not_found -> empty envelope with hint about lowercase

@TEST_EDGE: stale_snapshot -> age flagged in response (STALE)

#endregion Rls.AuditService.MatrixByUser

#region Rls.AuditService.MatrixByDataset [C:4] [TYPE Function] [SEMANTICS rls,audit,dataset]

@ingroup RlsServices

@BRIEF Inverse matrix dataset→users.

@POST Rows (dataset, unit_type, users[] with unit_ids).

@TEST_EDGE: dataset_not_bound -> dataset listed in no_rls group

#endregion Rls.AuditService.MatrixByDataset

#endregion Rls.AuditService

#region Rls.RuleBuilder [C:5] [TYPE Module] [SEMANTICS rls,rule,builder] @defgroup RlsServices Custom RLS rule lifecycle: preview on DWH, save definition, bind users. @BRIEF Rule definition is UPSERTed to rls_roles_filter and user membership to rls_rule_users; script v13 computes user→unit dynamically — no materialized roles. @RELATION DEPENDS_ON -> [Core.DbExecutor] @RELATION DEPENDS_ON -> [Rls.IdmClient] @RELATION DEPENDS_ON -> [Models.Rls] @RATIONALE R13: dynamic filter-based rules make rls_t correct by construction on every script run; eliminates drift/sync/confirm/reactivation surface. rls_roles_mask structure untouched (SAP masks); new table rls_roles_filter with same replication profile. @REJECTED Materialized roles in rls_bi_roles + sync (R7/R12) — rejected: drift by construction, sync machinery, confirm flows. Changing rls_roles_mask to host definitions — rejected: mixes SAP mask semantics with rule definitions, risks SAP branch.

#region Rls.RuleBuilder.Preview [C:5] [TYPE Function] [SEMANTICS rls,rule,preview]

@ingroup RlsServices

@BRIEF Live preview: SELECT DISTINCT unit_id_column FROM reference_table WHERE on DWH, limit 500 + count — the exact query the script CTE will run.

@PRE rule draft has reference_table, unit_id_column, ≥1 filter; dwh connection_id configured.

@POST RlsRulePreviewResponse (unit_count, sample) or validation error; empty result allowed (rule still saveable — zero roles at next run).

@SIDE_EFFECT Read-only SELECT on DWH; belief markers.

@DATA_CONTRACT Input: RlsRulePreviewRequest -> Output: RlsRulePreviewResponse

@RELATION DEPENDS_ON -> [DTO:RlsRulePreviewResponse]

@TEST_EDGE: zero_rows -> unit_count=0 with hint (rule saves, produces no roles until reference grows)

@TEST_EDGE: dwh_error -> error envelope with DWH message, retry

@TEST_EDGE: sql_injection_filter -> filters parameterized/escaped, never string-concatenated raw

@NOTE: db-executor-query-rows (R2: rows fetch capability)

#endregion Rls.RuleBuilder.Preview

#region Rls.RuleBuilder.SaveDefinition [C:5] [TYPE Function] [SEMANTICS rls,rule,save]

@ingroup RlsServices

@BRIEF Persist rule definition and membership: own DB + UPSERT rls_roles_filter + UPSERT rls_rule_users.

@PRE preview completed (expected slice known); rls WRITE permission; dwh reachable.

@POST Definition and unique (rule_id,user_id) memberships saved; duplicates skipped; effect_at surfaced.

@SIDE_EFFECT DWH write (UPSERT rls_roles_filter + UPSERT rls_rule_users); belief markers.

@DATA_CONTRACT Input: RlsRuleSaveRequest -> Output: RlsRuleSaveResponse {rule_id, roles_expected, rule_users_saved, duplicates_skipped, effect_at}

@RELATION DEPENDS_ON -> [DTO:RlsRuleSaveResponse]

@TEST_EDGE: duplicate_binding -> duplicate (rule_id,user_id) skipped

@TEST_EDGE: dwh_failure -> per-table error surfaced; own-DB row kept (retry-able)

@TEST_EDGE: rule_deactivated -> is_deleted=true upsert, roles vanish at next run (no confirm needed — definition-level soft toggle)

@TEST_EDGE: regexp_invalid -> validation error before save (FR-014)

#endregion Rls.RuleBuilder.SaveDefinition

#endregion Rls.RuleBuilder

#region Rls.BindingsService [C:4] [TYPE Module] [SEMANTICS rls,bindings,users] @defgroup RlsServices User binding to roles: manual AD-login, AD group/OU, org structure. @BRIEF Resolve bindings via IdmClient (autocomplete, group search, org tree), store RlsRuleBinding rows. @RELATION DEPENDS_ON -> [Rls.IdmClient] @RELATION DEPENDS_ON -> [Models.Rls] @RATIONALE Three clarified binding modes map to source_type enum; IDM search endpoints (mock extension) back group/org modes. @REJECTED Binding by department name free-text — rejected (requires IDM validation to avoid phantom users).

#region Rls.BindingsService.AddManual [C:4] [TYPE Function] [SEMANTICS rls,binding,manual]

@ingroup RlsServices

@BRIEF Add bindings by AD logins with IDM existence validation.

@POST Each login validated via IdmClient; unknown logins rejected with per-login error.

@TEST_EDGE: login_not_in_idm -> rejected with «нет в IDM» (distinct from «нет в RLS»)

@TEST_EDGE: fired_employee -> warning surfaced, binding still allowed (operator decision)

#endregion Rls.BindingsService.AddManual

#region Rls.BindingsService.AddGroup [C:4] [TYPE Function] [SEMANTICS rls,binding,group]

@ingroup RlsServices

@BRIEF Import all accounts of AD group/OU via SearchByOu.

@POST All found logins bound; 0 found → no binding + warning.

@TEST_EDGE: group_not_found -> no binding, warning (US4-4)

#endregion Rls.BindingsService.AddGroup

#region Rls.BindingsService.AddOrgUnit [C:4] [TYPE Function] [SEMANTICS rls,binding,org]

@ingroup RlsServices

@BRIEF Bind whole department/БЕ chunk via SearchByDepartment.

@POST Logins of department bound; count reported.

#endregion Rls.BindingsService.AddOrgUnit

#endregion Rls.BindingsService

#region Rls.UserAuditService [C:4] [TYPE Module] [SEMANTICS rls,audit,users,idm] @defgroup RlsServices bi_users × IDM enrichment and deactivation. @BRIEF Enrich RlsBiUserSnapshot rows with IDM statuses (employee status, account enabled, risk, VIP, adminRights); generate deactivation SQL; apply deactivation live. @RELATION DEPENDS_ON -> [Rls.IdmClient] @RELATION DEPENDS_ON -> [Models.Rls] @RELATION DEPENDS_ON -> [Rls.ScriptStore] @RATIONALE US3 core: fired/inactive detection = dateOut/leaveType/enabled from IDM. Deactivation = live UPDATE is_deleted + migration file (clarify). @REJECTED Deactivation as comment-only SQL — rejected (user: INSERT/UPDATE must be executable from UI).

#region Rls.UserAuditService.Enrich [C:4] [TYPE Function] [SEMANTICS rls,audit,enrich]

@ingroup RlsServices

@BRIEF Enrich bi_user rows with IDM statuses; per-row degradation on IDM failure.

@POST Rows carry idm_status or unknown; IDM-unavailable flagged globally (banner).

@TEST_EDGE: idm_down -> all statuses unknown + degraded flag, list still served

@TEST_EDGE: multi_account_user -> all accounts shown, matched by login without domain

#endregion Rls.UserAuditService.Enrich

#region Rls.UserAuditService.Deactivate [C:4] [TYPE Function] [SEMANTICS rls,audit,deactivate]

@ingroup RlsServices

@BRIEF Apply is_deleted=true for selected users on DWH + save migration SQL.

@PRE selected ad_users non-empty; rls WRITE permission.

@POST UPDATE executed per user; migration file saved; idempotent (already deleted → skipped).

@SIDE_EFFECT DWH write + git commit; belief markers.

@TEST_EDGE: already_deleted -> skipped, no error

@TEST_EDGE: dwh_failure -> error per user, applied ones persist

#endregion Rls.UserAuditService.Deactivate

#endregion Rls.UserAuditService

Backend: API routes (backend/src/api/routes/rls.py, prefix /api/rls)

#region Api.Rls [C:3] [TYPE Module] [SEMANTICS rls,api,routes] @defgroup RlsApi FastAPI routes for RLS workspace. @BRIEF Route group: overview, scripts, snapshot, audit, rules, bindings, datasets catalog. @RELATION DEPENDS_ON -> [Rls.ScriptStore] @RELATION DEPENDS_ON -> [Rls.SnapshotService] @RELATION DEPENDS_ON -> [Rls.AuditService] @RELATION DEPENDS_ON -> [Rls.RuleBuilder] @RELATION DEPENDS_ON -> [Rls.BindingsService] @RELATION DEPENDS_ON -> [Rls.UserAuditService] @RELATION DEPENDS_ON -> [Auth.Permissions]

#region Api.Rls.GetOverview [C:3] [TYPE Function] [SEMANTICS rls,api,overview]

@ingroup RlsApi

@BRIEF Dashboard cards: snapshot freshness, active version, audit counters.

@PRE has_permission("rls","READ").

@RELATION DEPENDS_ON -> [DTO:RlsOverviewResponse]

@DATA_CONTRACT GET /api/rls/overview -> RlsOverviewResponse

#endregion Api.Rls.GetOverview

#region Api.Rls.ListScripts [C:3] [TYPE Function] [SEMANTICS rls,api,scripts]

@ingroup RlsApi

@BRIEF List script versions.

@PRE has_permission("rls","READ").

@RELATION DEPENDS_ON -> [DTO:RlsScriptListItem]

@DATA_CONTRACT GET /api/rls/scripts -> RlsScriptListItem[]

#endregion Api.Rls.ListScripts

#region Api.Rls.CreateScriptVersion [C:3] [TYPE Function] [SEMANTICS rls,api,scripts]

@ingroup RlsApi

@BRIEF Create new version from active.

@PRE has_permission("rls_scripts","WRITE").

@DATA_CONTRACT POST /api/rls/scripts -> RlsScriptListItem

#endregion Api.Rls.CreateScriptVersion

#region Api.Rls.GetScriptDiff [C:3] [TYPE Function] [SEMANTICS rls,api,scripts]

@ingroup RlsApi

@BRIEF Diff between two versions.

@DATA_CONTRACT GET /api/rls/scripts/diff?a=v1&b=v2 -> RlsScriptDiffResponse

#endregion Api.Rls.GetScriptDiff

#region Api.Rls.PushScripts [C:4] [TYPE Function] [SEMANTICS rls,api,push]

@ingroup RlsApi

@BRIEF Push versions to corporate repo.

@PRE has_permission("rls_scripts","WRITE"); corp remote configured.

@POST Push attempted; envelope with ok/error (US1-7).

@DATA_CONTRACT POST /api/rls/scripts/push -> RlsPushResult

#endregion Api.Rls.PushScripts

#region Api.Rls.RefreshSnapshot [C:4] [TYPE Function] [SEMANTICS rls,api,snapshot]

@ingroup RlsApi

@BRIEF Manual snapshot refresh.

@PRE has_permission("rls","WRITE").

@POST Async or sync refresh with status envelope.

@DATA_CONTRACT POST /api/rls/snapshot/refresh -> RlsSnapshotRefreshResponse

#endregion Api.Rls.RefreshSnapshot

#region Api.Rls.AuditDatasetsByUser [C:3] [TYPE Function] [SEMANTICS rls,api,audit]

@ingroup RlsApi

@BRIEF User→datasets matrix.

@PRE has_permission("rls","READ").

@DATA_CONTRACT GET /api/rls/audit/datasets?user= -> RlsAuditMatrixResponse

#endregion Api.Rls.AuditDatasetsByUser

#region Api.Rls.AuditUsersList [C:3] [TYPE Function] [SEMANTICS rls,api,audit]

@ingroup RlsApi

@BRIEF Enriched bi_users list.

@PRE has_permission("rls","READ").

@DATA_CONTRACT GET /api/rls/audit/users -> RlsBiUserRow[]

#endregion Api.Rls.AuditUsersList

#region Api.Rls.AuditUserCard [C:3] [TYPE Function] [SEMANTICS rls,api,audit]

@ingroup RlsApi

@BRIEF IDM card for login.

@PRE has_permission("idm","READ_PERSONAL").

@DATA_CONTRACT GET /api/rls/audit/users/{login} -> RlsUserCardResponse

#endregion Api.Rls.AuditUserCard

#region Api.Rls.DeactivateUsers [C:4] [TYPE Function] [SEMANTICS rls,api,deactivate]

@ingroup RlsApi

@BRIEF Deactivate users (live UPDATE + migration).

@PRE has_permission("rls","WRITE").

@DATA_CONTRACT POST /api/rls/audit/users/deactivate -> RlsApplyResponse

#endregion Api.Rls.DeactivateUsers

#region Api.Rls.ListRules [C:3] [TYPE Function] [SEMANTICS rls,api,rules]

@ingroup RlsApi

@BRIEF Rule catalog with unit counts.

@PRE has_permission("rls","READ").

@DATA_CONTRACT GET /api/rls/rules -> RlsRuleDto[]

#endregion Api.Rls.ListRules

#region Api.Rls.CreateRule [C:3] [TYPE Function] [SEMANTICS rls,api,rules]

@ingroup RlsApi

@BRIEF Create rule draft.

@PRE has_permission("rls","WRITE").

@DATA_CONTRACT POST /api/rls/rules -> RlsRuleDto

#endregion Api.Rls.CreateRule

#region Api.Rls.PreviewRule [C:4] [TYPE Function] [SEMANTICS rls,api,preview]

@ingroup RlsApi

@BRIEF Live preview of rule filters on DWH.

@PRE has_permission("rls","WRITE"); rule has filters.

@POST Preview envelope or DWH error (US4-2/3).

@DATA_CONTRACT POST /api/rls/rules/{id}/preview -> RlsRulePreviewResponse

#endregion Api.Rls.PreviewRule

#region Api.Rls.AddBindings [C:4] [TYPE Function] [SEMANTICS rls,api,bindings]

@ingroup RlsApi

@BRIEF Add user bindings (manual/group/org).

@PRE has_permission("rls","WRITE").

@DATA_CONTRACT POST /api/rls/rules/{id}/bindings -> RlsApplyResponse

#endregion Api.Rls.AddBindings

#region Api.Rls.SaveRule [C:5] [TYPE Function] [SEMANTICS rls,api,rules]

@ingroup RlsApi

@BRIEF Save rule definition + rls_rule_users membership; effect at next script run.

@PRE has_permission("rls","WRITE"); preview completed; dwh reachable.

@DATA_CONTRACT POST /api/rls/rules/{id}/save -> RlsRuleSaveResponse

@TEST_EDGE: save_wo_preview -> 422 (preview required)

#endregion Api.Rls.SaveRule

#region Api.Rls.ListDatasetBindings [C:3] [TYPE Function] [SEMANTICS rls,api,catalog]

@ingroup RlsApi

@BRIEF Dataset catalog with auto-detected proposals.

@PRE has_permission("rls","READ").

@DATA_CONTRACT GET /api/rls/datasets/bindings -> RlsDatasetBindingDto[]

#endregion Api.Rls.ListDatasetBindings

#region Api.Rls.ConfirmDatasetBinding [C:4] [TYPE Function] [SEMANTICS rls,api,catalog]

@ingroup RlsApi

@BRIEF Confirm/override dataset binding (hybrid catalog).

@PRE has_permission("rls","WRITE").

@DATA_CONTRACT POST /api/rls/datasets/bindings -> RlsDatasetBindingDto

#endregion Api.Rls.ConfirmDatasetBinding

#endregion Api.Rls

Models & Schemas (backend)

#region Models.Rls [C:1] [TYPE Module] [SEMANTICS rls,models] @defgroup RlsModels SQLAlchemy entities: snapshot, script version, dataset binding, rule, binding. @BRIEF Pure ORM declarations per data-model.md §1 — no business logic (ADR-0001). @RELATION DEPENDS_ON -> [EXT:SQLAlchemy:declarative_base] #endregion Models.Rls

#region Schemas.Rls [C:1] [TYPE Module] [SEMANTICS rls,schemas,dto] @defgroup RlsSchemas Pydantic request/response schemas. @BRIEF DTOs per data-model.md §2; each has matching TS interface in frontend/src/types/rls.ts (cross-stack @RELATION). @RELATION DEPENDS_ON -> [DTO:RlsOverviewResponse] @RELATION DEPENDS_ON -> [DTO:RlsBiUserRow] @RELATION DEPENDS_ON -> [DTO:RlsRuleDto] @RELATION DEPENDS_ON -> [DTO:RlsApplyResponse] #endregion Schemas.Rls

Frontend

#region Rls.DashboardModel [C:4] [TYPE Model] [SEMANTICS rls,model,dashboard] @defgroup RlsModels Screen models for RLS workspace. @BRIEF Overview screen model: snapshot freshness, versions list, audit counters. @STATE idle/loading/loaded/stale/push-failed @STATE stale — age > threshold (banner + refresh action) @STATE push-failed — error envelope, retry @ACTION load(), createVersion(), activate(v), push() @INVARIANT Exactly one active version at any time. @INVARIANT push() never mutates local version data (failure is lossless). @RELATION CALLS -> [Lib.Api.Rls] #endregion Rls.DashboardModel

#region Rls.AuditMatrixModel [C:4] [TYPE Model] [SEMANTICS rls,model,audit] @defgroup RlsModels Screen models for RLS workspace. @BRIEF Audit matrix model: direction toggle, user/dataset search, matrix state. @STATE idle/searching/loaded/empty/degraded @ACTION search(), setDirection(), retry() @INVARIANT Changing direction resets matrix to idle. @INVARIANT search() with empty query is a no-op (idle stays). @INVARIANT degraded keeps last loaded matrix + age banner. @RELATION CALLS -> [Lib.Api.Rls] #endregion Rls.AuditMatrixModel

#region Rls.UserAuditModel [C:4] [TYPE Model] [SEMANTICS rls,model,useraudit] @defgroup RlsModels Screen models for RLS workspace. @BRIEF bi_users × IDM audit model: filters, selection, deactivation flow. @STATE loading/loaded/idm-degraded/empty @STATE idm-degraded — statuses unknown, rows still served (FR-017) @ACTION load(), setFilter(), toggleSelect(), deactivate() @INVARIANT Filtering applies to loaded rows only (no refetch). @INVARIANT deactivate() disabled while selection empty; double-submit guarded (DUP_01). @INVARIANT idm-degraded is sticky until successful retry. @RELATION CALLS -> [Lib.Api.Rls] #endregion Rls.UserAuditModel

#region Rls.RuleBuilderModel [C:5] [TYPE Model] [SEMANTICS rls,model,rulebuilder] @defgroup RlsModels Screen models for RLS workspace. @BRIEF 4-step rule builder: reference, filters, preview, bindings, save (definition + bindings). @STATE step-1/step-2/step-3/step-4, preview-loading/preview-empty/preview-loaded, saving/saved/error @ACTION next(), setReference(), addFilter(), validateRegexp(), preview(), bindManual(), bindGroup(), bindOrg(), save() @INVARIANT Changing reference table resets filters + preview + bindings (wizard restart guard). @INVARIANT save() blocked until preview completed (PREVIEW_REQUIRED); zero-unit preview allowed with hint. @INVARIANT save() runs once per press (saving state disables button, DUP_01). @INVARIANT Regexp validation happens live per filter row (FR-014). @INVARIANT saved result surfaces effect_at (latency until next script run). @RELATION CALLS -> [Lib.Api.Rls] #endregion Rls.RuleBuilderModel

#region Rls.RulesPage [C:3] [TYPE Component] [SEMANTICS rls,rules,page] @defgroup RlsComponents Components of RLS workspace. @BRIEF Rules catalog + builder page. @RELATION BINDS_TO -> [Rls.RuleBuilderModel] @RELATION DEPENDS_ON -> [Ui.Card] @RELATION DEPENDS_ON -> [Ui.Button] @RELATION DEPENDS_ON -> [Ui.Select] @RELATION DEPENDS_ON -> [Ui.Input] @RELATION DEPENDS_ON -> [Ui.Tabs] @RELATION DEPENDS_ON -> [Ui.ConfirmDialog] @RELATION DEPENDS_ON -> [Lib.Toasts] @UX_STATE step-1..4 -> wizard panels @UX_STATE preview-empty -> warning banner, apply disabled @UX_STATE saving -> button spinner + disabled @UX_STATE saved -> toast + summary modal with rule_id, users, roles_expected, effect_at @UX_RECOVERY preview error -> DWH error banner + retry @UX_RECOVERY CONFLICT -> ConfirmDialog reload/discard @RATIONALE Reuses Migration.WizardModel step pattern and Dashboards.SelectionModel submodel pattern (observed in lib/models/) instead of inventing new wizard/selection machinery. #endregion Rls.RulesPage

#region Rls.UserAuditPage [C:3] [TYPE Component] [SEMANTICS rls,audit,users,page] @defgroup RlsComponents Components of RLS workspace. @BRIEF bi_users table with IDM badges, filter chips, employee card, deactivation. @RELATION BINDS_TO -> [Rls.UserAuditModel] @RELATION DEPENDS_ON -> [Ui.Card] @RELATION DEPENDS_ON -> [Ui.Badge] @RELATION DEPENDS_ON -> [Ui.Button] @RELATION DEPENDS_ON -> [Ui.ConfirmDialog] @UX_STATE idm-degraded -> banner + «—» badges (FR-017) @UX_STATE empty -> EmptyState with CTA @UX_FEEDBACK deactivation toast + modal summary @UX_RECOVERY IDM retry button @RATIONALE Table pattern copied verbatim from components/backups/BackupList.svelte (th px-6 py-3 uppercase, td px-6 py-4) — pattern reuse, no new table component. #endregion Rls.UserAuditPage

#region Rls.AuditDatasetsPage [C:3] [TYPE Component] [SEMANTICS rls,audit,datasets,page] @defgroup RlsComponents Components of RLS workspace. @BRIEF Audit matrix page: direction tabs, search, matrix table, no-RLS group. @RELATION BINDS_TO -> [Rls.AuditMatrixModel] @RELATION DEPENDS_ON -> [Ui.Card] @RELATION DEPENDS_ON -> [Ui.Input] @RELATION DEPENDS_ON -> [Ui.Tabs] @RELATION DEPENDS_ON -> [Ui.EmptyState] @UX_STATE degraded -> snapshot-age banner (FR-019) @UX_STATE empty -> «не найден в RLS» + lowercase hint (US2-3) #endregion Rls.AuditDatasetsPage

#region Rls.DashboardPage [C:3] [TYPE Component] [SEMANTICS rls,dashboard,page] @defgroup RlsComponents Components of RLS workspace. @BRIEF Overview page: freshness cards, versions table, push action. @RELATION BINDS_TO -> [Rls.DashboardModel] @RELATION DEPENDS_ON -> [Ui.Card] @RELATION DEPENDS_ON -> [Ui.Badge] @RELATION DEPENDS_ON -> [Ui.Button] @RELATION DEPENDS_ON -> [Ui.PageHeader] @UX_STATE stale -> amber banner + refresh (US1-5) @UX_STATE push-failed -> destructive banner + retry (US1-7) #endregion Rls.DashboardPage

#region Lib.Api.Rls [C:3] [TYPE Module] [SEMANTICS rls,api,client] @defgroup RlsFrontend Frontend API client + DTO types. @BRIEF Typed requestApi wrappers for /api/rls/* endpoints; DTO interfaces in src/types/rls.ts mirror Schemas.Rls. @RELATION DEPENDS_ON -> [EXT:SvelteKit:requestApi] @RELATION DEPENDS_ON -> [DTO:RlsOverviewResponse] @RELATION DEPENDS_ON -> [DTO:RlsRuleDto] @RELATION DEPENDS_ON -> [DTO:RlsApplyResponse] @RATIONALE Native fetch forbidden (semantics-svelte §I) — requestApi enforces auth/trace_id/error normalization. #endregion Lib.Api.Rls

#region Rls.Routes [C:2] [TYPE Module] [SEMANTICS rls,routes,navigation] @defgroup RlsFrontend Route + navigation wiring. @BRIEF routes/rls/* pages + sidebarNavigation.ts category «RLS» with requiredPermission "rls". @RELATION DEPENDS_ON -> [Layout.SidebarNavigation] #endregion Rls.Routes

#endregion Rls.Modules


QUICKSTART — Dev Onboarding

Source: quickstart.md

#region Rls.Plan.Quickstart [C:2] [TYPE ADR] [SEMANTICS quickstart,verify,rls] @BRIEF Quickstart verification paths for 048-rls-management-workspace — real Makefile targets, timeout-protected.

Setup

# Backend deps
cd backend && source .venv/bin/activate && pip install -r requirements.txt

# Frontend deps
cd frontend && npm install

# IDM mock (dev)
cd research/rls/idm-mock && npm install && npm start   # http://localhost:3215

Конфиг: в backend/config.json блок rls — idm_url (dev: http://localhost:3215), idm_token (опционально), dwh_connection_id (существующий DatabaseConnection), corp_remote_url + corp_pat (шифрованный), snapshot_cron (дефолт 0 5 * * *), snapshot_retention (7).

Tier 1: Fast unit tests (<120s, no Docker)

make test            # backend + frontend unit tests (SQLite backend, vitest frontend)
make test-unit       # backend only
make test-frontend   # frontend only (L1 model invariants + L2 component UX)

Tier 2: Smart selection

make test-related F=backend/src/services/rls/rule_builder.py   # только тесты по @RELATION BINDS_TO

Tier 3: Integration (Docker required, <600s)

make test-integration    # testcontainers (PostgreSQL 16) — покрывает snapshot-загрузку и применение INSERT

Coverage & lint

make coverage     # pytest-cov + vitest v8
make lint         # ruff (backend) + eslint (frontend)

Docker

docker compose up --build

Локальный прогон IDM-мока

curl http://localhost:3215/health                      # статус + список пользователей
curl -X POST http://localhost:3215/api/Gather/GetPersonInfo -H 'Content-Type: application/json' \
  -d '{"FimSyncKey":"7272584452f9f8b2fd53291ee9f0bb2b"}'

#endregion Rls.Plan.Quickstart


TRACEABILITY — Requirements Matrix

Source: traceability.md

#region Rls.Traceability [C:3] [TYPE ADR] [SEMANTICS traceability,rtm,rls] @defgroup Trace Matrix Requirements → Screen+State → Model → API → Contract → Task → Test для 048-rls-management-workspace.

Applicability

  • Feature type: Fullstack
  • UI surface: Yes — раздел «RLS» (4 страницы) + state switcher в прототипе
  • API surface: Yes — GET|POST /api/rls/*

Traceability Matrix

Story / Req UX Screen + State Screen Model API operationId Contract Backend Task Frontend Task Test
US1: Dashboard cards /rls (loaded) Rls.DashboardModel getOverview Api.Rls.GetOverview T010 T012 Test.Rls.Overview
US1: Snapshot staleness /rls (stale) Rls.DashboardModel getOverview Api.Rls.GetOverview T010 T012 Test.Rls.Overview.Stale
US1: Push failure /rls (push-failed) Rls.DashboardModel pushScripts Api.Rls.PushScripts T00F T012 Test.Rls.Push.Failure
US1: Version list /rls (loaded) Rls.DashboardModel listScripts Api.Rls.ListScripts T00F T012 Test.Rls.Scripts.List
US1: Create version /rls (loaded) Rls.DashboardModel createScriptVersion Api.Rls.CreateScriptVersion T00F N/A — button reuse Test.Rls.Scripts.Create
US1: Diff /rls (loaded) Rls.DashboardModel getScriptDiff Api.Rls.GetScriptDiff T00F T012 Test.Rls.Scripts.Diff
US1: Freshness monitor N/A — backend cron N/A — no UI refreshSnapshot (cron) Rls.SnapshotService.Refresh T009 N/A — backend-only Test.Rls.Snapshot.Refresh
US1: Manual refresh /rls (stale) Rls.DashboardModel refreshSnapshot Api.Rls.RefreshSnapshot T009 T012 Test.Rls.Snapshot.Manual
US2: Matrix by user /rls/audit-datasets (loaded) Rls.AuditMatrixModel auditDatasetsByUser Api.Rls.AuditDatasetsByUser T018 T01A Test.Rls.Audit.MatrixByUser
US2: Search empty /rls/audit-datasets (empty) Rls.AuditMatrixModel auditDatasetsByUser Api.Rls.AuditDatasetsByUser T018 T01A Test.Rls.Audit.UserNotFound
US2: Degraded snapshot /rls/audit-datasets (degraded) Rls.AuditMatrixModel auditDatasetsByUser Api.Rls.AuditDatasetsByUser T018 T01A Test.Rls.Audit.StaleFlag
US2: Inverse matrix /rls/audit-datasets (loaded) Rls.AuditMatrixModel auditDatasetsByUser (direction=dataset) Rls.AuditService.MatrixByDataset T017 T01A Test.Rls.Audit.MatrixByDataset
US2: Dataset catalog /rls/audit-datasets (loaded) Rls.AuditMatrixModel listDatasetBindings Api.Rls.ListDatasetBindings T018 T01A Test.Rls.Catalog.List
US2: Confirm binding /rls/audit-datasets (loaded) Rls.AuditMatrixModel confirmDatasetBinding Api.Rls.ConfirmDatasetBinding T018 T01A Test.Rls.Catalog.Confirm
US3: Enriched users /rls/audit-users (loaded) Rls.UserAuditModel auditUsersList Api.Rls.AuditUsersList T021 T023 Test.Rls.UserAudit.Enrich
US3: IDM degraded /rls/audit-users (idm-degraded) Rls.UserAuditModel auditUsersList Rls.UserAuditService.Enrich T020 T023 Test.Rls.UserAudit.IdmDown
US3: User card /rls/audit-users (loaded) Rls.UserAuditModel auditUserCard Api.Rls.AuditUserCard T021 T023 Test.Rls.UserAudit.Card
US3: Deactivation /rls/audit-users (loaded) Rls.UserAuditModel deactivateUsers Api.Rls.DeactivateUsers T021 T023 Test.Rls.UserAudit.Deactivate
US3: Empty bi_users /rls/audit-users (empty) Rls.UserAuditModel auditUsersList Api.Rls.AuditUsersList T021 T023 Test.Rls.UserAudit.Empty
US4: Rule catalog /rls/rules (step-1) Rls.RuleBuilderModel listRules Api.Rls.ListRules T02D T02F Test.Rls.Rules.List
US4: Create rule /rls/rules (step-1) Rls.RuleBuilderModel createRule Api.Rls.CreateRule T02D T02F Test.Rls.Rules.Create
US4: Preview loaded /rls/rules (step-3-preview) Rls.RuleBuilderModel previewRule Api.Rls.PreviewRule T02D T02F Test.Rls.Rules.Preview
US4: Preview empty /rls/rules (step-3-empty) Rls.RuleBuilderModel previewRule Api.Rls.PreviewRule T02D T02F Test.Rls.Rules.PreviewEmpty
US4: Regexp validation /rls/rules (step-2) Rls.RuleBuilderModel N/A — client-side Rls.RuleBuilder.SaveDefinition T02B T02F Test.Rls.Rules.RegexpValidation
US4: Bind manual /rls/rules (step-4) Rls.RuleBuilderModel addBindings Api.Rls.AddBindings T02D T02F Test.Rls.Bindings.Manual
US4: Bind group /rls/rules (step-4) Rls.RuleBuilderModel addBindings Api.Rls.AddBindings T02D T02F Test.Rls.Bindings.Group
US4: Bind org /rls/rules (step-4) Rls.RuleBuilderModel addBindings Api.Rls.AddBindings T02D T02F Test.Rls.Bindings.Org
US4: Save rule /rls/rules (saving→saved) Rls.RuleBuilderModel saveRule Api.Rls.SaveRule T02D T02F Test.Rls.Rules.Save
US4: Save duplicate binding /rls/rules (saved) Rls.RuleBuilderModel saveRule Api.Rls.SaveRule T02D T02F Test.Rls.Rules.SaveIdempotent
US4: Preview after filter change /rls/rules (step-3-preview) Rls.RuleBuilderModel previewRule Api.Rls.PreviewRule T02D T02F Test.Rls.Rules.Preview
RLS-FR-004: corp push /rls (push-failed) Rls.DashboardModel pushScripts Rls.ScriptStore.PushToCorp T00E N/A — backend-only Test.Rls.Push.Unreachable
RLS-FR-020: snapshot hybrid N/A — infra N/A — infra refreshSnapshot Rls.SnapshotService.Refresh T009 N/A — infra Test.Rls.Snapshot.Refresh
RLS-FR-022: extensible rls_type N/A — data model N/A — data model N/A — no API Models.Rls T005 N/A — data model Test.Rls.Rules.NewRlsType
RLS-FR-023: idempotent bindings /rls/rules (saved) Rls.RuleBuilderModel saveRule Rls.RuleBuilder.SaveDefinition T02B N/A — backend-only Test.Rls.Rules.SaveIdempotent

N/A Rationale Key

  • N/A — backend-only: нет UI-поверхности (cron, push, idempotency)
  • N/A — frontend-only: клиентская валидация без API (regexp-проверка дублируется на бэке при SaveDefinition)
  • N/A — infra: фоновые/инфраструктурные требования без экрана и API-операции
  • N/A — data model: ограничение данных, не поведение
  • N/A — no API: внутренний сервис без HTTP-эндпоинта
  • N/A — button reuse: действие существующей кнопки без нового фронт-кода
  • N/A — from apply response: diff берётся из ответа apply, отдельного эндпоинта нет

Impact Analysis Quick Reference

If you change... These fixtures verify it These tests verify it These screens depend
Rls.RuleBuilder.Preview FX_Rls.RuleBuilder.Preview.* (5) Test.Rls.Rules.Preview* /rls/rules (step-3)
Rls.RuleBuilder.SaveDefinition FX_Rls.RuleBuilder.Save.* (7) Test.Rls.Rules.Save* /rls/rules (step-4)
Rls.UserAuditService.Deactivate FX_Rls.UserAudit.Deactivate.* (4) Test.Rls.UserAudit.Deactivate* /rls/audit-users
Rls.ScriptStore.PushToCorp FX_Rls.ScriptStore.Push.* (4) Test.Rls.Push.* /rls (push-failed)
Rls.SnapshotService.Refresh FX_Rls.Snapshot.Refresh.* (4) Test.Rls.Snapshot.* /rls, /rls/audit-*
Rls.BindingsService.AddManual FX_Rls.Bindings.AddManual.* (3) Test.Rls.Bindings.* /rls/rules (step-4)
Rls.IdmClient N/A — mock transport в тестах Test.Rls.IdmClient.* все экраны с IDM-данными

Coverage Gate

  • Every user story has at least one row (US1-4 — по 5-10 строк каждая)
  • Every functional requirement has a row OR explicit N/A rationale (FR-001..023 — репрезентативная выборка: US-строки покрывают FR-001..019, отдельные строки для FR-004/020/022/023; FR-021 покрыт строкой US2 Dataset catalog)
  • Every API endpoint has success AND error rows (Preview: loaded/empty/error; Save: saved/duplicate/failure; Push: valid/failure)
  • Every Screen Model has loaded AND error rows (DashboardModel: loaded/stale/push-failed; AuditMatrixModel: loaded/empty/degraded; UserAuditModel: loaded/idm-degraded/empty; RuleBuilderModel: preview/saving/saved/error)
  • Every N/A cell carries rationale from the key
  • Every contract referenced appears in contracts/modules.md
  • Every task ID (Txxx) appears in tasks.md (сгенерировано: 52 задачи, T001-T038 с hex-суффиксами)
  • Impact table covers every contract with downstream dependents

Gate status: ЗАКРЫТ — задачи привязаны (tasks.md, 52 задачи).

#endregion Rls.Traceability


TASKS — Implementation Tasks

Source: tasks.md

Tasks: RLS Management Workspace (048)

Input: Design documents from /specs/048-rls-management-workspace/ Prerequisites: plan.md, spec.md, research.md (R1-R13), data-model.md, contracts/modules.md

Tests: Включены для всех C3+ контрактов (фикстуры canonical в specs/048-rls-management-workspace/fixtures/, материализуются до написания тестов).

Organization: Tasks grouped by user story — каждая story независимо реализуема и тестируема.

Format: [ID] [P?] [Story] Description

Path Conventions

  • Backend: backend/src/**/*.py, тесты backend/tests/**/*.py
  • Frontend: frontend/src/routes/**/*.svelte, frontend/src/lib/**/*.svelte.ts|ts|svelte
  • Скрипты RLS: rls_scripts/** (корень репозитория)
  • Мок IDM: research/rls/idm-mock/**

Phase 1: Setup (Shared Infrastructure)

Purpose: Конфигурация домена RLS, скелет роутера, навигация, каталог скриптов.

  • T001 Расширить конфиг: блок rls в backend/src/core/config_models.py (GlobalSettings) + backend/config.json — поля: idm_url, idm_token, dwh_connection_id, corp_remote_url, corp_pat (шифруется через core/encryption.py), snapshot_cron (дефолт 0 5 * * *), snapshot_retention (7), rls_scripts_dir (default rls_scripts/) RATIONALE: один источник конфигурации для всех сервисов rls; секреты — существующий шифрованный механизм (research R3/R5)
  • T002 [P] Создать каталог rls_scripts/ в корне репозитория: seed-шаблон v001_rls_t_main_insert.sql (заготовка структуры из research/rls REPOSITORY_ANALYSIS §4.1) + version.json (version, author, comment, is_active) + git commit
  • T003 [P] Скелет роутера: backend/src/api/routes/rls.py (APIRouter prefix /api/rls, tags=["RLS"], зависимости get_current_user/get_db по образцу backend/src/api/routes/profile.py) + include_router(rls.router) в backend/src/app.py
  • T004 [P] Навигация: категория «RLS» в frontend/src/lib/components/layout/sidebarNavigation.ts (path /rls, requiredPermission: "rls", tone admin-категории #ffe4e6→#fecdd3 из tailwind category.admin) + пустые страницы frontend/src/routes/rls/+page.svelte, audit-datasets/+page.svelte, audit-users/+page.svelte, rules/+page.svelte

Phase 2: Foundational (Blocking Prerequisites)

Purpose: Модели, схемы, DTO, DbExecutor read-query, снапшот-сервис — без этого ни одна story не реализуема.

⚠️ CRITICAL: Работы по user stories не начинаются до завершения этой фазы.

  • T005 Модели Models.Rls в backend/src/models/rls.py (C1): RlsSnapshotMeta, RlsSnapshotRow, RlsBiUserSnapshot, RlsBiRoleSnapshot, RlsScriptVersion, RlsDatasetBinding, RlsRule, RlsRuleBinding — поля/индексы по data-model.md §1
  • T006 Схемы Schemas.Rls в backend/src/schemas/rls.py (C1): 18 Pydantic-моделей по data-model.md §2 (RlsOverviewResponse, RlsScriptListItem, RlsScriptDiffResponse, RlsPushResult, RlsSnapshotRefreshResponse, RlsBiUserRow, RlsUserCardResponse, RlsDeactivateRequest, RlsDatasetBindingDto, RlsAuditMatrixResponse, RlsRuleDto, RlsFilterDto, RlsRulePreviewResponse, RlsRulePreviewRequest, RlsRuleBindingRequest, RlsRuleSaveRequest, RlsRuleSaveResponse, RlsApplyResponse)
  • T007 [P] TS DTO frontend/src/types/rls.ts — интерфейсы, зеркалящие Schemas.Rls (cross-stack: каждый интерфейс несёт @RELATION на Pydantic-пару, общий @SEMANTICS rls)
  • T008 Проверить fetch-возможности Core.DbExecutor в backend/src/core/db_executor.py; при отсутствии чтения строк — добавить read-метод query(connection_id, sql, limit) -> list[dict] (read-only, без транзакции записи) RATIONALE: предпросмотр правил и снапшоты требуют SELECT-строк; расширение существующего контракта вместо нового слоя (research R2)
  • T009 Снапшот-сервис Rls.SnapshotService в backend/src/services/rls/snapshot_service.py (C4) @PRE: dwh connection_id из конфига; таблицы rls_t/bi_users/bi_roles доступны @POST: новая генерация RlsSnapshotMeta (success/failed) с row_count; retention 7 генераций; статус pending при активном прогоне @SIDE_EFFECT: SELECT на DWH через DbExecutor; bulk-insert в RlsSnapshot*; belief markers REASON/REFLECT/EXPLORE @TEST_EDGE: dwh_unreachable→failed+no partial rows, empty_source→success row_count=0, pending_conflict→409 Регистрация cron: RegisterCron через backend/src/core/scheduler.py (CronTrigger, паттерн scheduler.py:175) + POST /api/rls/snapshot/refresh в backend/src/api/routes/rls.py (has_permission("rls","WRITE"))
  • T00A [P] Материализовать фикстуры из specs/048-rls-management-workspace/fixtures/api/*.json (29 шт.) в backend/tests/fixtures/rls/ — копия as-is, без изменений содержимого
  • T00B Тесты снапшота Test.Rls.Snapshot в backend/tests/test_rls_snapshot.py (C2): мок DbExecutor (unittest.mock.patch), фикстуры snapshot_valid/dwh_unreachable/empty_source/pending_conflict; @RELATION BINDS_TO -> [Rls.SnapshotService]

Checkpoint: Foundation ready — user stories можно начинать параллельно.


Phase 3: User Story 1 - Версионирование RLS-скрипта (Priority: P1) 🎯 MVP

Goal: Версии SQL-скрипта rls_t в rls_scripts/ (git), diff, активация, push в corp remote, дашборд актуальности.

Independent Test: Раздел RLS → список версий, создание из активной, diff, активация, push с ошибкой и retry, карточка актуальности rls_t (age/строки/порог).

Tests for User Story 1 ⚠️ (пишутся ДО реализации, должны падать)

  • T00C [P] [US1] Контракт-тесты Test.Rls.Scripts в backend/tests/test_rls_scripts.py: create/diff/activate (временный git-репо через gitpython, tmp_path); Test.Rls.Push — фикстуры push_valid/corp_unreachable/no_corp_remote/no_credentials
  • T00D [P] [US1] L1-тест RlsDashboardModel в frontend/src/lib/models/__tests__/RlsDashboardModel.test.ts: инварианты (одна активная версия, push не мутирует локальные данные, stale = age>порог)

Implementation for User Story 1

  • T00E [US1] Реализовать Rls.ScriptStore в backend/src/services/rls/script_store.py (C4, gitpython Repo на репозитории) @PRE: rls_scripts_dir существует; git-репо доступно @POST: CreateVersion — новый файл vNNN + запись version.json, активная не меняется; Diff — unified-diff строк; Activate — ровно одна is_active @SIDE_EFFECT: файловая запись + git add/commit; belief markers @TEST_EDGE: no_active_version→seed из шаблона, duplicate_version→автоинкремент номера + PushToCorp (C4): remote corp с PAT из шифрованных credentials (паттерн Services.Git._clone_with_auth/push_changes) @PRE: corp remote настроен; PAT расшифрован @POST: push выполнен, pushed_at обновлён; ошибка — envelope без потери версии @TEST_EDGE: corp_unreachable→CORP_UNREACHABLE+local intact, no_corp_remote→CORP_REMOTE_MISSING, no_credentials→CREDENTIALS_MISSING
  • T00F [US1] Эндпоинты скриптов в backend/src/api/routes/rls.py: ListScripts (GET /api/rls/scripts, rls READ), CreateScriptVersion (POST, rls_scripts WRITE), GetScriptDiff (GET /api/rls/scripts/diff?a=&b=), ActivateScript (POST /api/rls/scripts/{id}/activate, rls_scripts WRITE), PushScripts (POST /api/rls/scripts/push, rls_scripts WRITE)
  • T010 [US1] GetOverview в backend/src/api/routes/rls.py (GET /api/rls/overview, rls READ): свежесть снапшота (taken_at/age/stale по порогу), активная версия, счётчики аудита (fired/inactive/no_rls) из снапшота
  • T011 [US1] Модель Rls.DashboardModel в frontend/src/lib/models/RlsDashboardModel.svelte.ts (C4) @STATE idle/loading/loaded/stale/push-failed @ACTION load(), createVersion(), activate(v), push() @INVARIANT: ровно одна активная версия; push() не мутирует локальные версии (lossless) @POST: load() → screenState=stale при age>threshold @SIDE_EFFECT: GET /api/rls/overview, POST /api/rls/scripts/... @TEST_EDGE: push_fail→push-failed state + retry, network_fail→error
  • T012 [US1] Страница Rls.DashboardPage в frontend/src/routes/rls/+page.svelte (C3): 3 карточки (актуальность/версии/аудит) через <Card> $lib/ui/Card.svelte, бейджи <Badge> $lib/ui/Badge.svelte (existing), таблица версий (паттерн BackupList: min-w-full divide-y divide-border, thead bg-surface-muted), кнопки <Button> $lib/ui/Button.svelte; stale-баннер (amber) + push-failed-баннер (destructive) + retry @UX_STATE stale → amber banner + «Обновить снапшот» (POST refresh) @UX_STATE push-failed → destructive banner + «Повторить push» @UX_FEEDBACK toast через notify() из $lib/toasts.svelte.ts (existing) @RELATION BINDS_TO -> [Rls.DashboardModel]
  • T013 [US1] Верификация US1: make test-unit (Test.Rls.Scripts/Test.Rls.Push), make test-frontend (L1 RlsDashboardModel), make test-related F=backend/src/services/rls/script_store.py

Checkpoint: US1 полностью функциональна и тестируема независимо.


Phase 4: User Story 2 - Аудит доступа к датасетам (Priority: P1)

Goal: Матрица «пользователь × датасет × unit_id» на снапшоте с учётом rls_type; каталог датасетов (автоопределение + подтверждение).

Independent Test: Поиск логина → матрица датасетов с unit_id/source_system/roles; direction=dataset; пустой результат с подсказкой про регистр; degraded-баннер; каталог с авто-предложениями и подтверждением.

Tests for User Story 2 ⚠️

  • T014 [P] [US2] Контракт-тесты Test.Rls.Audit в backend/tests/test_rls_audit.py: MatrixByUser (user_not_found→empty+hint, stale→STALE-флаг, rls_type-фильтрация по unit_type), MatrixByDataset (dataset_not_bound→no_rls-группа)
  • T015 [P] [US2] Контракт-тесты Test.Rls.Catalog в backend/tests/test_rls_catalog.py: автоопределение колонок unit_id/unit_balance_code/plant_code из метаданных Superset API, подтверждение/переопределение, multi-binding (оба типа единиц)
  • T016 [P] [US2] L1-тест RlsAuditMatrixModel в frontend/src/lib/models/__tests__/RlsAuditMatrixModel.test.ts: смена direction сбрасывает матрицу; пустой query — no-op; degraded сохраняет данные + баннер

Implementation for User Story 2

  • T017 [US2] Реализовать Rls.AuditService в backend/src/services/rls/audit_service.py (C4) @PRE: снапшот существует; user login нормализован lowercase @POST: MatrixByUser — строки (dataset, unit_type, unit_ids, source_systems, roles) + no_rls_datasets; join по rls_type×unit_type @SIDE_EFFECT: чтение снапшота; belief markers @TEST_EDGE: user_not_found→empty+hint lowercase, stale_snapshot→age флаг, dataset_not_bound→no_rls группа + автоопределение колонок: метаданные датасетов через SupersetClient (datasets API) → кандидаты detection=auto
  • T018 [US2] Эндпоинты в backend/src/api/routes/rls.py: AuditDatasetsByUser (GET /api/rls/audit/datasets?user=&direction=, rls READ), ListDatasetBindings (GET /api/rls/datasets/bindings), ConfirmDatasetBinding (POST /api/rls/datasets/bindings, rls WRITE)
  • T019 [US2] Модель Rls.AuditMatrixModel в frontend/src/lib/models/RlsAuditMatrixModel.svelte.ts (C4) @STATE idle/searching/loaded/empty/degraded @ACTION search(), setDirection(), retry() @INVARIANT: смена direction → idle; пустой query → no-op; degraded сохраняет матрицу @SIDE_EFFECT: GET /api/rls/audit/datasets @TEST_EDGE: not_found→empty+hint, network_fail→degraded
  • T01A [US2] Страница Rls.AuditDatasetsPage в frontend/src/routes/rls/audit-datasets/+page.svelte (C3): segmented-tabs direction через <Tabs> $lib/ui/Tabs.svelte (existing), поиск <Input> $lib/ui/Input.svelte, матрица-таблица (паттерн BackupList), группа «без RLS-фильтрации» <Card>, EmptyState <EmptyState> $lib/ui/EmptyState.svelte (existing), degraded-баннер @RELATION BINDS_TO -> [Rls.AuditMatrixModel] @UX_STATE empty → «не найден в RLS» + подсказка про регистр
  • T01B [US2] Верификация US2: make test-unit (Test.Rls.Audit/Test.Rls.Catalog), make test-frontend, make test-related F=backend/src/services/rls/audit_service.py

Checkpoint: US1 и US2 работают независимо.


Phase 5: User Story 3 - Аудит BI-пользователей через IDM (Priority: P1)

Goal: Обогащение rls_bi_users статусами IDM (уволен/отпуск/enabled/риск/VIP/adminRights), карточка сотрудника, деактивация через интерфейс (is_deleted=true + SQL-миграция).

Independent Test: Таблица bi_users со статусами IDM, чипы-фильтры, раскрытие карточки (аккаунты с OU/adminRights), деактивация выбранных с модалом и toast, idm-degraded-баннер.

Tests for User Story 3 ⚠️

  • T01C [P] [US3] Контракт-тесты Test.Rls.IdmClient в backend/tests/test_rls_idm_client.py: httpx MockTransport (мок-сервер research/rls/idm-mock shapes), degradation envelope (idm_unavailable→None+warning), malformed_id→opaque string
  • T01D [P] [US3] Контракт-тесты Test.Rls.UserAudit в backend/tests/test_rls_user_audit.py: Enrich (idm_down→unknown+degraded flag, multi_account→все аккаунты), Deactivate — фикстуры deactivate_valid/already_deleted/dwh_failure/empty_selection
  • T01E [P] [US3] L1-тест RlsUserAuditModel в frontend/src/lib/models/__tests__/RlsUserAuditModel.test.ts: фильтр по загруженным строкам без refetch; деактивация disabled при пустом selected; idm-degraded sticky до retry

Implementation for User Story 3

  • T01F [US3] Реализовать Rls.IdmClient (базовый) в backend/src/services/rls/idm_client.py (C4) @PRE: base_url из конфига rls.idm_url @POST: GetPersonInfo/GetPersonExtensionInfo/GetAccountInfo возвращают DTO или degraded None+warning @SIDE_EFFECT: внешние HTTP-вызовы httpx.AsyncClient (timeout 10s, retry ×2); belief markers @TEST_EDGE: idm_unavailable→degraded без raise, malformed_id→opaque string
  • T020 [US3] Реализовать Rls.UserAuditService в backend/src/services/rls/user_audit_service.py (C4) @PRE: снапшот bi_users существует @POST: Enrich — строки с idm_status (employee/account/risk/vip/admin) или unknown + degraded-флаг; Deactivate — UPDATE is_deleted=true per user + миграция через ScriptStore @SIDE_EFFECT: IDM-вызовы, DWH-UPDATE через DbExecutor, git-commit миграции; belief markers @TEST_EDGE: idm_down→все unknown+degraded, already_deleted→skipped, dwh_failure→ошибка per-user, empty_selection→422
  • T021 [US3] Эндпоинты в backend/src/api/routes/rls.py: AuditUsersList (GET /api/rls/audit/users, rls READ), AuditUserCard (GET /api/rls/audit/users/{login}, idm READ_PERSONAL), DeactivateUsers (POST /api/rls/audit/users/deactivate, rls WRITE)
  • T022 [US3] Модель Rls.UserAuditModel в frontend/src/lib/models/RlsUserAuditModel.svelte.ts (C4) @STATE loading/loaded/idm-degraded/empty @ACTION load(), setFilter(), toggleSelect(), deactivate() @INVARIANT: фильтр только по загруженным строкам; deactivate disabled при пустом selected; double-submit guard @SIDE_EFFECT: GET /api/rls/audit/users, POST deactivate @TEST_EDGE: idm_down→idm-degraded sticky, double_submit→один запрос
  • T023 [US3] Страница Rls.UserAuditPage в frontend/src/routes/rls/audit-users/+page.svelte (C3): таблица (паттерн BackupList) со статус-бейджами <Badge> (success/destructive/info/warning), чипы-фильтры <Tabs variant="pills">, карточка сотрудника (раскрытие строки), модал подтверждения деактивации <ConfirmDialog> $lib/ui/ConfirmDialog.svelte (existing), toasts notify() @RELATION BINDS_TO -> [Rls.UserAuditModel] @UX_STATE idm-degraded → баннер + «—» бейджи
  • T024 [US3] Верификация US3: make test-unit, make test-frontend, make test-related F=backend/src/services/rls/idm_client.py

Checkpoint: US1-US3 работают независимо.


Phase 6: User Story 4 - Конструктор кастомных RLS-правил (Priority: P2)

Goal: Динамические правила (R13): определение → preview → пользователи правила → сохранение в rls_roles_filter + rls_rule_users; скрипт v13 вычисляет user→unit при каждом запуске.

Independent Test: Мастер: справочник+колонка → фильтры (IN/LIKE/regexp с валидацией) → предпросмотр (счётчик unit_id) → привязки (3 вкладки) → сохранить (effect_at), деактивация правила.

Tests for User Story 4 ⚠️

  • T025 [P] [US4] Контракт-тесты Test.Rls.Rules в backend/tests/test_rls_rules.py: Preview (valid/zero_rows/dwh_error/sql_injection/missing_filter — фикстуры rule_preview_), SaveDefinition (valid/duplicate_binding/dwh_failure/deactivated/regexp_invalid/wo_preview/permission_denied — фикстуры rule_save_), в т.ч. rejected-path: sql_injection→FILTER_PARAMETERIZED, save_wo_preview→422
  • T026 [P] [US4] Контракт-тесты Test.Rls.Bindings в backend/tests/test_rls_bindings.py: AddManual (valid/login_not_in_idm/fired_employee — фикстуры binding_*), AddGroup (SearchByOu), AddOrgUnit (SearchByDepartment)
  • T027 [P] [US4] L1-тест RlsRuleBuilderModel в frontend/src/lib/models/__tests__/RlsRuleBuilderModel.test.ts: смена справочника сбрасывает фильтры+preview; save блокируется до preview; saving — один запрос; zero-unit preview с hint

Implementation for User Story 4

  • T028 [P] [US4] Расширить мок IDM: эндпоинты поиска по OU и подразделению в research/rls/idm-mock/server.js (POST /api/Gather/SearchPersons?ou=&department=, тестовые данные из data/users.json) + README.md
  • T029 [P] [US4] Rls.IdmClient.SearchByOu + SearchByDepartment в backend/src/services/rls/idm_client.py @PRE: ou/department non-empty @POST: список IdmAccountRef/IdmPersonRef или degraded [] @SIDE_EFFECT: HTTP-вызов; belief markers @TEST_EDGE: group_not_found→[]+warning, department_unknown→[]+warning
  • T02A [US4] Реализовать Rls.RuleBuilder.Preview в backend/src/services/rls/rule_builder.py (C5) @PRE: reference_table, unit_id_column, ≥1 фильтр; dwh connection_id; фильтры валидны (regexp проверен) @POST: RlsRulePreviewResponse (unit_count, sample ≤500); zero_rows — saveable+hint @SIDE_EFFECT: read-only SELECT на DWH (DbExecutor.query); параметризованные значения фильтров; belief markers @DATA_CONTRACT: RlsRulePreviewRequest → RlsRulePreviewResponse @TEST_EDGE: zero_rows→saveable, dwh_error→DWH_ERROR envelope, sql_injection→FILTER_PARAMETERIZED (никогда raw-конкатенация)
  • T02B [US4] Реализовать Rls.RuleBuilder.SaveDefinition в backend/src/services/rls/rule_builder.py (C5) @PRE: preview выполнен (иначе 422 PREVIEW_REQUIRED); rls WRITE; dwh доступна @POST: SSOT + UPSERT rls_roles_filter + UPSERT rls_rule_users (unique rule_id,user_id); effect_at — ближайший запуск скрипта @SIDE_EFFECT: DWH-UPSERT rls_roles_filter/rls_rule_users через DbExecutor; belief markers REASON/REFLECT/EXPLORE @DATA_CONTRACT: RlsRuleSaveRequest → RlsRuleSaveResponse @TEST_EDGE: duplicate_binding→skipped, dwh_failure→ошибка+локальная строка сохранена, deactivated→is_deleted=true toggle, regexp_invalid→422
  • T02C [US4] Реализовать Rls.BindingsService в backend/src/services/rls/bindings_service.py (C4) @PRE: rule существует; source_type ∈ manual|ad_group|org_unit @POST: AddManual — IDM-валидация логинов (not_found_in_idm→rejected per-login, fired→warning+allowed); AddGroup — все аккаунты OU; AddOrgUnit — все сотрудники отдела @SIDE_EFFECT: IDM-вызовы, запись RlsRuleBinding; belief markers @TEST_EDGE: login_not_in_idm→rejected, group_not_found→[]+warning
  • T02D [US4] Эндпоинты в backend/src/api/routes/rls.py: ListRules (GET /api/rls/rules, rls READ), CreateRule (POST, rls WRITE), PreviewRule (POST /api/rls/rules/{id}/preview, rls WRITE), SaveRule (POST /api/rls/rules/{id}/save, rls WRITE), AddBindings (POST /api/rls/rules/{id}/bindings, rls WRITE)
  • T02E [US4] Модель Rls.RuleBuilderModel в frontend/src/lib/models/RlsRuleBuilderModel.svelte.ts (C5) @STATE step-1..4, preview-loading/preview-empty/preview-loaded, saving/saved/error @ACTION next(), setReference(), addFilter(), validateRegexp(), preview(), bindManual(), bindGroup(), bindOrg(), save() @INVARIANT: смена справочника сбрасывает фильтры+preview+привязки; save блокируется до preview; saving — один запрос (DUP_01); regexp live-валидация @SIDE_EFFECT: POST preview/save; сохранённый результат показывает effect_at @TEST_EDGE: preview_empty→hint+saveable, save_wo_preview→422, network_fail→error
  • T02F [US4] Страница Rls.RulesPage в frontend/src/routes/rls/rules/+page.svelte (C3): каталог правил слева (список <Card>), мастер 4 шага (паттерн шагов из Migration.WizardModel.svelte.ts — reuse), фильтры-таблица (атрибут×тип×выражение), regexp-ошибка инлайн, предпросмотр (счётчик+таблица ≤500), 3 вкладки привязки (<Tabs>), SQL-модал <ConfirmDialog> с effect_at, toasts @RELATION BINDS_TO -> [Rls.RuleBuilderModel] @UX_STATE preview-empty → warning banner, save разрешён с hint @UX_STATE saving → спиннер+disabled
  • T030 [US4] Верификация US4: make test-unit (Test.Rls.Rules/Test.Rls.Bindings), make test-frontend, make test-related F=backend/src/services/rls/rule_builder.py

Checkpoint: Все 4 user stories функциональны независимо.


Phase 7: Polish & Cross-Cutting Concerns

Purpose: Скрипт v13 (динамические правила), rejected-path регрессия, семантические аудиты, полный прогон.

  • T031 Подготовить rls_scripts/v013_rls_t_main_insert.sql: CTE rule_units (active rls_roles_filter → rule_id,unit_id), JOIN active rls_rule_users → user_id,unit_id; статические rls_bi_users/rls_bi_roles и rls_roles_mask не меняются; diff v12→v13 через ScriptStore RATIONALE: динамические правила «прокручиваются» при каждом запуске; дрейф исключён по построению REJECTED: материализация ролей в bi_roles + sync (дрейф, confirm-флоу); изменение структуры rls_roles_mask (риск SAP-ветки)
  • T032 [P] Rejected-path регрессия: прогон фикстур rejected-сценариев (rule_preview_sql_injection, rule_save_permission_denied, rule_save_wo_preview, push_no_credentials, snapshot_pending_conflict) через make test-related F=backend/src/services/rls/ — каждый @TEST_EDGE: rejected_path обязан падать на сломанной реализации и проходить на рабочей
  • T033 [P] Attention compliance audit: ATTN_1 (однострочные anchors), ATTN_2 (иерархические Rls.*/Api.Rls.*), ATTN_3 (@SEMANTICS rls единообразие), ATTN_4 (контракты ≤150 строк, модули ≤400) — по semantics-core §VIII, проверка всех новых файлов backend/src/services/rls/ и frontend/src/lib/models/Rls*
  • T034 [P] Semantic index rebuild: axiom_search operation="rebuild" rebuild_mode="full" — 0 parse warnings (если Axiom MCP доступен; иначе grep-проверка пар #region/#endregion в новых файлах)
  • T035 [P] Belief runtime audit: axiom_audit operation="audit_belief_runtime" + audit_belief_protocol — все C4/C5 контракты имеют @RATIONALE/@REJECTED и REASON/REFLECT/EXPLORE маркеры (Rls.SnapshotService, Rls.ScriptStore, Rls.AuditService, Rls.RuleBuilder, Rls.BindingsService, Rls.UserAuditService, 4 модели)
  • T036 [P] Prototype validation: каждый @UX_STATE из contracts/modules.md достижим в specs/048-rls-management-workspace/prototype/index.html через state switcher; mobile-вьюпорт 375px без overflow
  • T037 Полный прогон: make test + make lint + make coverage; фронт: cd frontend && npm run build
  • T038 Pre-implementation validation gate: запустить /speckit.validate — PASS перед /speckit.implement

Dependencies & Execution Order

Phase Dependencies

  • Setup (Phase 1): без зависимостей
  • Foundational (Phase 2): после Setup — БЛОКИРУЕТ все user stories
  • US1/US2/US3 (Phase 3-5): после Foundational; независимы друг от друга (параллельны при наличии ресурсов)
  • US4 (Phase 6): после Foundational + зависим от IDM-клиента (T01F из US3) для bindings
  • Polish (Phase 7): после всех user stories

Within Each User Story

  • Тесты пишутся и падают ДО реализации
  • L1 (модели, без рендера) раньше L2 (компоненты)
  • Модели → сервисы → эндпоинты → страницы
  • Story завершена до перехода к следующей

Parallel Opportunities

  • Setup: T002/T003/T004 параллельно
  • Foundational: T007/T00A параллельно; T005→T006→T009 последовательно (модели→схемы→query)
  • US1: T00C/T00D параллельно
  • US2: T014/T015/T016 параллельно
  • US3: T01C/T01D/T01E параллельно
  • US4: T025/T026/T027 и T028/T029 параллельно (мок+IDM-поиск отдельными файлами)
  • Polish: T032-T036 параллельно

Implementation Strategy

MVP (US1 only)

  1. Phase 1+2 → 2. US1 → 3. STOP, validate independently → 4. demo (версии скрипта + дашборд актуальности)

Incremental Delivery

  1. Setup+Foundational → US1 (MVP: версии+push) → US2 (аудит датасетов) → US3 (аудит bi_users×IDM) → US4 (конструктор правил) → Polish (v13-скрипт, аудиты)

Notes

  • [P] = разные файлы, нет зависимостей
  • Верификация после каждой story: make test-unit + make test-frontend (+ make test-related для затронутых файлов)
  • Фикстуры canonical в specs/048-rls-management-workspace/fixtures/ — не редактируются в test-директориях
  • C4/C5 модули несут belief-маркеры (log/REASON/REFLECT/EXPLORE, belief_scope из ss_tools.lib.cot_logger)
  • Скрипт v13 — версия в rls_scripts/ (не исполняется из superset-tools; применяется внешним контуром)
  • rls_roles_mask/rls_bi_roles структуру НЕ меняем (SAP-совместимость, FR-025)

PROTOTYPE — State/Manifest

Source: prototype/manifest.md

#region Std.Opencode.PrototypeManifest [C:3] [TYPE ADR] [SEMANTICS prototype,manifest,rls,superset,idm] @defgroup Prototype Interactive HTML prototype manifest for RLS Management Workspace (048).

Prototype Metadata

  • Feature: RLS Management Workspace (048-rls-management-workspace)
  • Source contracts: ux_reference.md (contracts/ux/ не создавался — лёгкий прототип из reference)
  • Screens represented: 4
  • Total states: 24 (19 контрактных + 5 интерактивных переходов)
  • Accessibility validations: keyboard nav (Tab/Enter/Space/Escape), ARIA roles (tablist/tab/dialog/alert/status), touch targets, focus rings
  • Responsive breakpoints: 375px (mobile), 1280px (desktop)
  • Prototype path: specs/048-rls-management-workspace/prototype/index.html

State Coverage

Screen @UX_STATE Contract (ux_reference) Prototype State Reachable? Recovery Path
1 · RLS Dashboard loading loading (скелетоны) ✅ → loaded
1 · RLS Dashboard loaded loaded (карточки + таблица версий) ✅ —
1 · RLS Dashboard stale stale (amber-баннер «данные от 02.08.2026») ✅ «Обновить снапшот» → toast
1 · RLS Dashboard push-failed push-failed (destructive-баннер, версия сохранена) ✅ «Повторить push» → loaded + toast
2 · Аудит датасетов idle idle (поисковая форма, подсказка про регистр) ✅ —
2 · Аудит датасетов searching searching (скелетоны) ✅ авто → loaded/empty
2 · Аудит датасетов loaded loaded (матрица 5 датасетов, rls_type-нота, группа «без RLS») ✅ —
2 · Аудит датасетов empty empty (EmptyState «не найден в RLS» + подсказка) ✅ «Очистить поиск»
2 · Аудит датасетов degraded degraded (баннер + снапшот-данные) ✅ «Обновить снапшот» → toast
3 · Аудит пользователей loading users-loading (скелетоны) ✅ → loaded
3 · Аудит пользователей loaded users-loaded (таблица 6 строк + карточка IDM с OU/adminRights) ✅ —
3 · Аудит пользователей idm-degraded idm-degraded (баннер + статусы «—») ✅ «Повторить» / автоповтор
3 · Аудит пользователей empty users-empty (EmptyState «bi_users пуст») ✅ CTA → конструктор
4 · Конструктор step-1 step-1 (справочник + колонка unit_id) ✅ —
4 · Конструктор step-2 step-2 (фильтры IN/LIKE/regexp, ошибка regexp) ✅ исправить выражение
4 · Конструктор preview-loading step-3-loading (живой SELECT на DWH) ✅ → preview/empty
4 · Конструктор preview-empty step-3-empty (жёлтый баннер, «Применить» заблокирован) ✅ «Вернуться к фильтрам»
4 · Конструктор preview-loaded step-3-preview (34 unit_id, таблица первых 5) ✅ —
4 · Конструктор step-4 step-4 (3 вкладки привязки, статусы IDM) ✅ —
4 · Конструктор applying applying (спиннер на кнопке, disabled) ✅ → applied
4 · Конструктор applied/saved applied (toast + модал SQL/diff) ✅ модал: Отмена/Применить

Интерактивные переходы (recovery-пути): push-кнопка → push-failed; retry → loaded; search → searching → loaded/empty (по значению «kozlovakk»); preview → loading → preview; apply → модал → applying → applied; чекбоксы → активация «Сформировать SQL деактивации» → модал → toast; Escape/клик по backdrop закрывают модалы.

Screen ↔ Story Traceability

Prototype Screen User Story UX Reference Section Acceptance Criteria Verified
1 · RLS Dashboard US1: Версионирование скрипта Экран 1 US1-1 (список версий), US1-5 (актуальность/строки), US1-7 (push error)
2 · Аудит датасетов US2: Аудит доступа Экран 2 US2-1 (матрица с rls_type), US2-3 (empty + регистр), US2-5 (группа «без RLS»), US2-4 (degraded)
3 · Аудит пользователей US3: Аудит bi_users × IDM Экран 3 US3-1 (статусы IDM), US3-2 (чипы-фильтры), US3-3 (IDM degraded), US3-4 (карточка с OU/adminRights), US3-5 (SQL деактивации)
4 · Конструктор правил US4: Кастомные правила Экран 4 US4-1 (справочник+колонка), US4-2 (IN/LIKE/regexp), US4-3 (preview-empty блокировка), US4-4 (3 способа привязки), US4-5 (INSERT + SQL в репо), US4-6 (diff), US4-7 (regexp error)

Validation Results

  • All @UX_STATE contracts reachable via state switcher — 19/19 контрактных состояний, 5 интерактивных переходов
  • All @UX_RECOVERY paths traversable — retry push, refresh snapshot, retry IDM, очистка поиска, возврат к фильтрам, модальные отмена/применить
  • Keyboard navigation: Tab order verified; Enter/Space активируют кнопки; Escape закрывает модалы
  • Touch targets: ≥44×44px на мобильном вьюпорте (кнопки h-10/h-12)
  • ARIA: role=tablist/tab (aria-selected), role=dialog (aria-modal), role=status/aria-live (toasts), aria-label на иконках-кнопках
  • No broken links or dead-end states
  • Modals close: фикс [hidden] { display: none !important } в shim (авторский display: flex на .modal-backdrop/.alert переопределял UA-стиль hidden — модалы не закрывались); проверено скриптом: 21 скрываемый элемент, обработчики modal-cancel/deact-cancel/modal-confirm/Escape/backdrop на месте
  • Responsive layout: 375px — сайдбар скрыт, таблицы в overflow-x-auto
  • Design fidelity: 57/57 hex из tailwind.config.js; 0 классов без shim-определения (скриптовая проверка)

Примечание: браузерная валидация (DevTools accessibility tree, визуальное сравнение с dev-сервером) не выполнялась — окружение без браузера; статическая проверка классов/токенов выполнена скриптом. Рекомендуется ручной прогон в браузере на этапе /speckit.implement (browser-first практика).

Design System Reuse

Element Source Prototype Mapping
Button $lib/ui/Button.svelte .btn + .btn-primary/secondary/destructive/ghost + .btn-sm/md — те же токены: bg #2563eb, hover #1d4ed8, ring #3b82f6, h-10 px-4 py-2 text-sm
Card $lib/ui/Card.svelte .card = rounded-lg border border-border bg-surface-card text-text shadow-sm + card-title-row + card-pad-md
Badge $lib/ui/Badge.svelte .badge + .badge-pill + badge-{variant}: bg--light text-; dot-варианты
Skeleton $lib/ui/Skeleton.svelte .skeleton-line/.skeleton-card/.skeleton-row = animate-pulse + bg-surface-muted
EmptyState $lib/ui/EmptyState.svelte .empty-state = py-12 px-4 text-center + icon w-16 h-16 text-text-subtle
PageHeader $lib/ui/PageHeader.svelte flex items-center justify-between mb-8 + text-3xl font-bold tracking-tight text-text
Input $lib/ui/Input.svelte .input = h-10 w-full rounded-md border-border-strong px-3 py-2 text-sm + focus-visible ring + error border-destructive
Select $lib/ui/Select.svelte .select — те же токены
Tabs $lib/ui/Tabs.svelte .tab-segmented/.tab-pills/.tab-underline с active/inactive из Tabs.svelte
Switch $lib/ui/Switch.svelte .switch (h-6 w-11, knob translate) — включён в shim, используется для будущих настроек порога
Table components/backups/BackupList.svelte min-w-full divide-y divide-border, thead bg-surface-muted, th px-6 py-3 text-xs uppercase tracking-wider, td px-6 py-4
Modal $lib/ui/ConfirmDialog.svelte .modal-backdrop (bg-surface-overlay rgba(15,23,42,.5)) + .modal-card
Category chip tailwind category.admin #ffe4e6 → #fecdd3, text #be123c (nav-item active + chip)

Design Token Audit (MANDATORY)

Token (tailwind.config.js) Hex / Value Used in prototype (elements)
primary.DEFAULT #2563eb btn-primary, tab active text, switch-on, nav hover text, focus ring
primary.hover #1d4ed8 btn-primary hover
primary.ring #3b82f6 focus-visible rings, dot-primary
primary.light #eff6ff badge-primary, tab-pills active bg, tab-badge
secondary.DEFAULT #f3f4f6 btn-secondary, nav hover, tab-pills inactive
secondary.hover #e5e7eb btn-secondary hover
secondary.text #111827 btn-secondary text
secondary.ring #6b7280 btn-secondary/ghost focus ring
destructive.DEFAULT #dc2626 btn-destructive, badges destructive, input-error border
destructive.hover #b91c1c btn-destructive hover, alert-destructive text
destructive.ring #ef4444 btn-destructive focus ring
destructive.light #fef2f2 badge-destructive, alert-destructive bg
success.DEFAULT #22c55e badge-success, dot-success, step-num done
success.hover #16a34a btn-success hover
success.ring #4ade80 btn-success focus ring, toast-success border
success.light #f0fdf4 badge-success bg, toast-success bg
warning.DEFAULT #f59e0b badge-warning, dot-warning
warning.hover #d97706 btn-warning hover
warning.ring #fbbf24 btn-warning ring, alert-warning border
warning.light #fffbeb badge-warning bg, alert-warning bg
info.DEFAULT #0ea5e9 badge-info, dot-info
info.hover #0284c7 btn-info hover
info.ring #38bdf8 alert-info border, viewport-toggle active bg (chrome)
info.light #f0f9ff badge-info bg, alert-info bg
ghost.hover #f3f4f6 btn-ghost hover
ghost.text #374151 btn-ghost text
surface.page #f8fafc proto-app bg, thead bg, tab-segmented hover
surface.card #ffffff card bg, tbody bg, sidebar bg
surface.muted #f1f5f9 skeletons, badges muted, tab-pills inactive
surface.overlay rgba(15, 23, 42, 0.5) modal backdrop
border.DEFAULT #e2e8f0 card border, table borders, step-sep
border.strong #cbd5e1 input/select borders
text.DEFAULT #0f172a page titles, td-main, labels
text.muted #64748b secondary text, th, alerts
text.subtle #94a3b8 placeholders, empty-state icon, org-count
text.inverse #ffffff btn text (via .text-white)
terminal.bg #0f172a SQL preview .mono, tab-segmented active, chrome bg
terminal.border #334155 chrome border, mono borders
terminal.text.DEFAULT #cbd5e1 chrome text, mono text
terminal.text.bright #e2e8f0 chrome select text
terminal.accent #22d3ee chrome badge
terminal.surface #1e293b chrome select/button bg
category.admin.from/to/text #ffe4e6/#fecdd3/#be123c nav-item active, chip-category
category.dashboards.from/to #e0f2fe/#bae6fd sidebar nav-dot «Дашборды»
category.datasets.from/to #dcfce7/#bbf7d0 sidebar nav-dot «Датасеты»
category.migration.from/to #ffedd5/#fed7aa sidebar nav-dot «Миграция»
category.git.from/to #f3e8ff/#e9d5ff sidebar nav-dot «Git»
category.validation.from/to #fef3c7/#fde68a sidebar nav-dot «Валидация»
category.reports.from/to #ede9fe/#ddd6fe sidebar nav-dot «Отчёты»
brand.gradient-* #0ea5e9/#06b6d4/#4f46e5 sidebar logo

Аудит-скрипт: прогнан автоматически дважды — итог: 57 уникальных hex-кодов, неизвестных токенам: 0; классовых определений без правила в shim: 0.

#endregion Std.Opencode.PrototypeManifest


PROTOTYPE — Interactive HTML

Source: prototype/index.html

<html lang="ru"> <head> <style> /* ============================================================ PROTOTYPE CHROME (state switcher) — НЕ часть дизайн-системы ============================================================ */ #prototype-chrome { position: fixed; top: 0; left: 0; right: 0; z-index: 1000; display: flex; flex-wrap: wrap; align-items: center; gap: 10px; padding: 8px 14px; background: #0f172a; color: #cbd5e1; font: 12px/1.4 system-ui, sans-serif; border-bottom: 1px solid #334155; } #prototype-chrome label { display: flex; align-items: center; gap: 5px; } #prototype-chrome select, #prototype-chrome button { background: #1e293b; color: #e2e8f0; border: 1px solid #334155; border-radius: 4px; padding: 4px 8px; font: inherit; cursor: pointer; } #prototype-chrome .chrome-badge { background: #22d3ee; color: #0f172a; font-weight: 700; border-radius: 4px; padding: 2px 8px; } #prototype-chrome .viewport-toggle { display: inline-flex; gap: 4px; } #prototype-chrome .viewport-toggle button.active { background: #38bdf8; color: #0f172a; } body { margin: 0; } #app-frame { margin-top: 46px; transition: max-width 0.2s ease; } body.chrome-viewport-mobile #app-frame { max-width: 375px; margin-left: auto; margin-right: auto; } body.chrome-viewport-mobile .proto-sidebar { display: none; } body.chrome-viewport-mobile .proto-main { padding-left: 0; } body.chrome-viewport-mobile .proto-topbar { display: none; }

/* ============================================================ APP SHELL ============================================================ */ .proto-app { display: flex; min-height: 100vh; background: #f8fafc; } .proto-sidebar { width: 240px; flex-shrink: 0; background: #ffffff; border-right: 1px solid #e2e8f0; padding: 20px 12px; } .proto-sidebar .proto-brand { display: flex; align-items: center; gap: 8px; padding: 0 8px 18px; font-size: 14px; font-weight: 700; color: #0f172a; } .proto-sidebar .proto-brand .logo { width: 28px; height: 28px; border-radius: 8px; flex-shrink: 0; background: linear-gradient(135deg, #0ea5e9, #06b6d4, #4f46e5); } .proto-nav-item { display: flex; align-items: center; gap: 10px; width: 100%; padding: 8px 10px; margin-bottom: 2px; border-radius: 6px; font-size: 13px; font-weight: 500; color: #64748b; background: transparent; border: none; text-align: left; cursor: pointer; } .proto-nav-item:hover { background: #f3f4f6; } .proto-nav-item.active { background: #ffe4e6; color: #be123c; } .proto-nav-item .nav-dot { width: 8px; height: 8px; border-radius: 9999px; flex-shrink: 0; } .proto-main { flex: 1; min-width: 0; padding: 32px; } .proto-topbar { display: flex; align-items: center; gap: 12px; justify-content: flex-end; padding: 10px 32px; border-bottom: 1px solid #e2e8f0; background: #ffffff; }

/* ============================================================ DESIGN SYSTEM SHIM — классы скопированы из frontend/src/lib/ui/* Значения токенов ТОЛЬКО из frontend/tailwind.config.js ============================================================ / / --- Layout utilities --- */ .flex { display: flex; } .inline-flex { display: inline-flex; } .flex-col { flex-direction: column; } .flex-wrap { flex-wrap: wrap; } .flex-1 { flex: 1 1 0%; } .items-start { align-items: flex-start; } .min-w-0 { min-width: 0; } .w-full { width: 100%; } .items-center { align-items: center; } .justify-center { justify-content: center; } .justify-between { justify-content: space-between; } .justify-end { justify-content: flex-end; } .gap-1 { gap: 4px; } .gap-1.5 { gap: 6px; } .gap-2 { gap: 8px; } .gap-3 { gap: 12px; } .gap-4 { gap: 16px; } .gap-6 { gap: 24px; } .space-y-1 > * + * { margin-top: 4px; } .space-y-1.5 > * + * { margin-top: 6px; } .space-y-2 > * + * { margin-top: 8px; } .space-y-3 > * + * { margin-top: 12px; } .space-y-4 > * + * { margin-top: 16px; } .space-x-2 > * + * { margin-left: 8px; } .grid { display: grid; } .grid-cols-2 { grid-template-columns: repeat(2, minmax(0, 1fr)); } .grid-cols-3 { grid-template-columns: repeat(3, minmax(0, 1fr)); } .mb-1 { margin-bottom: 4px; } .mb-2 { margin-bottom: 8px; } .mb-3 { margin-bottom: 12px; } .mb-4 { margin-bottom: 16px; } .mb-6 { margin-bottom: 24px; } .mb-8 { margin-bottom: 32px; } .mt-1 { margin-top: 4px; } .mt-2 { margin-top: 8px; } .mt-3 { margin-top: 12px; } .mt-4 { margin-top: 16px; } .mt-6 { margin-top: 24px; } .mt-8 { margin-top: 32px; } .ml-1 { margin-left: 4px; } .ml-1.5 { margin-left: 6px; } .ml-2 { margin-left: 8px; } .mr-1 { margin-right: 4px; } .mr-2 { margin-right: 8px; } .p-3 { padding: 12px; } .p-6 { padding: 24px; } .p-8 { padding: 32px; } .px-2 { padding-left: 8px; padding-right: 8px; } .px-2.5 { padding-left: 10px; padding-right: 10px; } .px-3 { padding-left: 12px; padding-right: 12px; } .px-4 { padding-left: 16px; padding-right: 16px; } .px-6 { padding-left: 24px; padding-right: 24px; } .py-1 { padding-top: 4px; padding-bottom: 4px; } .py-1.5 { padding-top: 6px; padding-bottom: 6px; } .py-2 { padding-top: 8px; padding-bottom: 8px; } .py-2.5 { padding-top: 10px; padding-bottom: 10px; } .py-3 { padding-top: 12px; padding-bottom: 12px; } .py-4 { padding-top: 16px; padding-bottom: 16px; } .py-12 { padding-top: 48px; padding-bottom: 48px; } .pt-2 { padding-top: 8px; } .pb-4 { padding-bottom: 16px; } .px-1 { padding-left: 4px; padding-right: 4px; } .h-2 { height: 8px; } .h-2.5 { height: 10px; } .h-3 { height: 12px; } .h-4 { height: 16px; } .h-6 { height: 24px; } .h-8 { height: 32px; } .h-10 { height: 40px; } .h-12 { height: 48px; } .h-16 { height: 64px; } .h-24 { height: 96px; } .h-full { height: 100%; } .w-2 { width: 8px; } .w-2.5 { width: 10px; } .w-3 { width: 12px; } .w-4 { width: 16px; } .w-10 { width: 40px; } .w-11 { width: 44px; } .w-16 { width: 64px; } .relative { position: relative; } .absolute { position: absolute; } .inset-0 { inset: 0; } .translate-x-1 { transform: translateX(4px); } .translate-x-6 { transform: translateX(24px); } .rounded { border-radius: 4px; } .rounded-md { border-radius: 6px; } .rounded-lg { border-radius: 8px; } .rounded-full { border-radius: 9999px; } .rounded-t-lg { border-top-left-radius: 8px; border-top-right-radius: 8px; } .border { border-width: 1px; border-style: solid; } .border-b { border-bottom-width: 1px; border-bottom-style: solid; } .border-b-2 { border-bottom-width: 2px; border-bottom-style: solid; } .border-t { border-top-width: 1px; border-top-style: solid; } .border-border { border-color: #e2e8f0; } .border-border-strong { border-color: #cbd5e1; } .border-primary { border-color: #2563eb; } .border-destructive { border-color: #dc2626; } .border-transparent { border-color: transparent; } .-mb-px { margin-bottom: -1px; } .border-b-white { border-bottom-color: #ffffff; } .divide-y > * + * { border-top-width: 1px; border-top-style: solid; border-color: #e2e8f0; } .divide-border > * + * { border-top-width: 1px; border-top-style: solid; border-color: #e2e8f0; } .shadow-sm { box-shadow: 0 1px 2px 0 rgba(0, 0, 0, 0.05); } .overflow-hidden { overflow: hidden; } .overflow-x-auto { overflow-x: auto; } .whitespace-nowrap { white-space: nowrap; } .max-w-md { max-width: 28rem; } .max-w-xs { max-width: 20rem; } .min-w-full { min-width: 100%; } .text-left { text-align: left; } .text-right { text-align: right; } .text-center { text-align: center; } .uppercase { text-transform: uppercase; } .tracking-tight { letter-spacing: -0.025em; } .tracking-wider { letter-spacing: 0.05em; } .leading-none { line-height: 1; } .italic { font-style: italic; } .underline { text-decoration: underline; } .underline-offset-4 { text-underline-offset: 4px; } .cursor-pointer { cursor: pointer; } .cursor-not-allowed { cursor: not-allowed; } .fixed { position: fixed; } .z-50 { z-index: 50; } .transition-colors { transition-property: color, background-color, border-color; transition-duration: 150ms; } .transition-transform { transition-property: transform; transition-duration: 150ms; } .transition-all { transition-property: all; transition-duration: 150ms; } .transform { transform: none; } .opacity-25 { opacity: 0.25; }

/* --- Screen sections (prototype chrome) --- */ .proto-screen { display: block; }

/* --- Hidden attribute must win over author display rules (modals use display:flex, alerts use display:flex — without this the hidden attribute is overridden and dialogs never close) --- */ [hidden] { display: none !important; }

/* --- Typography --- */ .text-xs { font-size: 12px; line-height: 16px; } .text-sm { font-size: 14px; line-height: 20px; } .text-base { font-size: 16px; line-height: 24px; } .text-lg { font-size: 18px; line-height: 28px; } .text-3xl { font-size: 30px; line-height: 36px; } .font-medium { font-weight: 500; } .font-semibold { font-weight: 600; } .font-bold { font-weight: 700; } .text-text { color: #0f172a; } .text-text-muted { color: #64748b; } .text-text-subtle { color: #94a3b8; } .text-white { color: #ffffff; } .text-primary { color: #2563eb; } .text-destructive { color: #dc2626; } .text-success { color: #22c55e; } .text-warning { color: #f59e0b; } .text-info { color: #0ea5e9; } .text-current { color: currentColor; } .text-ghost { color: #374151; }

/* --- Backgrounds --- */ .bg-primary { background-color: #2563eb; } .bg-primary-hover { background-color: #1d4ed8; } .bg-primary-light { background-color: #eff6ff; } .bg-primary-ring { background-color: #3b82f6; } .bg-secondary { background-color: #f3f4f6; } .bg-secondary-hover { background-color: #e5e7eb; } .bg-destructive { background-color: #dc2626; } .bg-destructive-hover { background-color: #b91c1c; } .bg-destructive-light { background-color: #fef2f2; } .bg-success { background-color: #22c55e; } .bg-success-hover { background-color: #16a34a; } .bg-success-light { background-color: #f0fdf4; } .bg-warning { background-color: #f59e0b; } .bg-warning-hover { background-color: #d97706; } .bg-warning-light { background-color: #fffbeb; } .bg-info { background-color: #0ea5e9; } .bg-info-hover { background-color: #0284c7; } .bg-info-light { background-color: #f0f9ff; } .bg-surface-page { background-color: #f8fafc; } .bg-surface-card { background-color: #ffffff; } .bg-surface-muted { background-color: #f1f5f9; } .bg-terminal-bg { background-color: #0f172a; } .bg-transparent { background-color: transparent; } .bg-surface-overlay { background-color: rgba(15, 23, 42, 0.5); }

/* --- Buttons (verbatim from Button.svelte) --- */ .btn { display: inline-flex; align-items: center; justify-content: center; font-weight: 500; transition-property: color, background-color, border-color; transition-duration: 150ms; border-radius: 6px; } .btn:focus-visible { outline: none; box-shadow: 0 0 0 2px #ffffff, 0 0 0 4px var(--btn-ring, #3b82f6); } .btn:disabled { pointer-events: none; opacity: 0.5; } .btn-primary { background-color: #2563eb; color: #ffffff; } .btn-primary:hover:not(:disabled) { background-color: #1d4ed8; } .btn-primary:focus-visible { --btn-ring: #3b82f6; } .btn-secondary { background-color: #f3f4f6; color: #111827; } .btn-secondary:hover:not(:disabled) { background-color: #e5e7eb; } .btn-secondary:focus-visible { --btn-ring: #6b7280; } .btn-destructive { background-color: #dc2626; color: #ffffff; } .btn-destructive:hover:not(:disabled) { background-color: #b91c1c; } .btn-destructive:focus-visible { --btn-ring: #ef4444; } .btn-ghost { background-color: transparent; color: #374151; } .btn-ghost:hover:not(:disabled) { background-color: #f3f4f6; } .btn-ghost:focus-visible { --btn-ring: #6b7280; } .btn-success { background-color: #22c55e; color: #ffffff; } .btn-success:hover:not(:disabled) { background-color: #16a34a; } .btn-success:focus-visible { --btn-ring: #4ade80; } .btn-warning { background-color: #f59e0b; color: #ffffff; } .btn-warning:hover:not(:disabled) { background-color: #d97706; } .btn-warning:focus-visible { --btn-ring: #fbbf24; } .btn-info { background-color: #0ea5e9; color: #ffffff; } .btn-info:hover:not(:disabled) { background-color: #0284c7; } .btn-info:focus-visible { --btn-ring: #38bdf8; } .btn-link { background-color: transparent; color: #2563eb; } .btn-link:hover:not(:disabled) { color: #2563eb; text-decoration: underline; text-underline-offset: 4px; } .btn-sm { height: 32px; padding-left: 12px; padding-right: 12px; font-size: 12px; } .btn-md { height: 40px; padding: 8px 16px; font-size: 14px; } .btn-lg { height: 48px; padding-left: 24px; padding-right: 24px; font-size: 16px; } .spin { animation: spin 1s linear infinite; display: inline-block; width: 16px; height: 16px; margin-right: 8px; margin-left: -4px; vertical-align: middle; flex-shrink: 0; } @keyframes spin { from { transform: rotate(0deg); } to { transform: rotate(360deg); } }

/* --- Badge (verbatim from Badge.svelte) --- */ .badge { display: inline-flex; align-items: center; gap: 6px; } .badge-pill { border-radius: 9999px; font-size: 12px; font-weight: 500; } .badge-md { padding: 4px 10px; } .badge-sm { padding: 2px 8px; } .badge-success { background-color: #f0fdf4; color: #22c55e; } .badge-warning { background-color: #fffbeb; color: #f59e0b; } .badge-destructive { background-color: #fef2f2; color: #dc2626; } .badge-info { background-color: #f0f9ff; color: #0ea5e9; } .badge-primary { background-color: #eff6ff; color: #2563eb; } .badge-muted { background-color: #f1f5f9; color: #64748b; } .badge-dot { width: 10px; height: 10px; border-radius: 9999px; } .dot-success { background-color: #22c55e; } .dot-warning { background-color: #f59e0b; } .dot-destructive { background-color: #dc2626; } .dot-info { background-color: #0ea5e9; } .dot-primary { background-color: #3b82f6; } .dot-muted { background-color: #64748b; }

/* --- Card (verbatim from Card.svelte) --- */ .card { border-radius: 8px; border: 1px solid #e2e8f0; background-color: #ffffff; color: #0f172a; box-shadow: 0 1px 2px 0 rgba(0, 0, 0, 0.05); } .card-title-row { display: flex; flex-direction: column; gap: 6px; padding: 24px; border-bottom: 1px solid #e2e8f0; } .card-title { font-size: 18px; font-weight: 600; line-height: 1; letter-spacing: -0.025em; } .card-pad-md { padding: 24px; } .card-pad-sm { padding: 12px; }

/* --- Input (verbatim from Input.svelte) --- */ .field { display: flex; flex-direction: column; gap: 6px; width: 100%; } .field-label { display: flex; align-items: center; gap: 6px; } .field-label label { font-size: 14px; font-weight: 500; color: #0f172a; } .input { flex: 1; height: 40px; width: 100%; border-radius: 6px; border: 1px solid #cbd5e1; background-color: #ffffff; padding: 8px 12px; font-size: 14px; color: #0f172a; } .input::placeholder { color: #94a3b8; } .input:focus-visible { outline: none; box-shadow: 0 0 0 2px #ffffff, 0 0 0 4px #3b82f6; } .input:disabled { cursor: not-allowed; opacity: 0.5; } .input-error { border-color: #dc2626; } .select { height: 40px; width: 100%; border-radius: 6px; border: 1px solid #cbd5e1; background-color: #ffffff; padding: 8px 12px; font-size: 14px; color: #0f172a; } .select:focus-visible { outline: none; box-shadow: 0 0 0 2px #ffffff, 0 0 0 4px #3b82f6; } .select:disabled { cursor: not-allowed; opacity: 0.5; } .field-error { font-size: 12px; color: #dc2626; }

/* --- Tabs (verbatim from Tabs.svelte) --- */ .tablist { display: flex; } .tab-underline { padding: 8px 16px; font-size: 14px; font-weight: 500; transition-property: color, border-color; transition-duration: 150ms; border-bottom: 2px solid transparent; background: none; border-top: none; border-left: none; border-right: none; cursor: pointer; } .tab-underline:focus { outline: none; } .tab-underline.active { color: #2563eb; border-bottom-color: #2563eb; } .tab-underline.inactive { color: #64748b; } .tab-underline.inactive:hover { color: #0f172a; } .tab-segmented { border-radius: 6px; padding: 6px 12px; font-size: 14px; font-weight: 500; transition-property: color, background-color; transition-duration: 150ms; border: none; cursor: pointer; } .tab-segmented.active { background-color: #0f172a; color: #ffffff; } .tab-segmented.inactive { background-color: #f1f5f9; color: #0f172a; } .tab-segmented.inactive:hover { background-color: #f8fafc; } .tab-pills { position: relative; padding: 6px 12px; font-size: 12px; border-radius: 6px; transition-property: color, background-color; transition-duration: 150ms; border: none; cursor: pointer; } .tab-pills.active { background-color: #eff6ff; color: #2563eb; border: 1px solid #3b82f6; } .tab-pills.inactive { background-color: #f1f5f9; color: #64748b; border: 1px solid #e2e8f0; } .tab-pills.inactive:hover { background-color: #f8fafc; } .tab-badge { display: inline-flex; align-items: center; justify-content: center; height: 16px; width: 16px; border-radius: 9999px; font-size: 9px; font-weight: 700; background-color: #eff6ff; color: #2563eb; margin-left: 6px; }

/* --- Switch (verbatim from Switch.svelte) --- */ .switch { position: relative; display: inline-flex; height: 24px; width: 44px; align-items: center; border-radius: 9999px; transition-property: background-color; transition-duration: 150ms; border: none; cursor: pointer; } .switch:focus { outline: none; box-shadow: 0 0 0 2px #ffffff, 0 0 0 4px #3b82f6; } .switch:disabled { opacity: 0.5; cursor: not-allowed; } .switch-on { background-color: #2563eb; } .switch-off { background-color: #f1f5f9; } .switch-knob { display: inline-block; height: 16px; width: 16px; border-radius: 9999px; background-color: #ffffff; transition-property: transform; transition-duration: 150ms; }

/* --- Skeleton (verbatim from Skeleton.svelte) --- */ .skeleton-line { animation: pulse 2s cubic-bezier(0.4, 0, 0.6, 1) infinite; background-color: #f1f5f9; border-radius: 4px; height: 16px; } .skeleton-card { animation: pulse 2s cubic-bezier(0.4, 0, 0.6, 1) infinite; background-color: #f1f5f9; border-radius: 8px; height: 96px; } .skeleton-row { animation: pulse 2s cubic-bezier(0.4, 0, 0.6, 1) infinite; height: 16px; background-color: #f1f5f9; border-radius: 4px; flex: 1; } @keyframes pulse { 0%, 100% { opacity: 1; } 50% { opacity: 0.5; } }

/* --- EmptyState (verbatim from EmptyState.svelte) --- */ .empty-state { display: flex; flex-direction: column; align-items: center; justify-content: center; padding: 48px 16px; text-align: center; } .empty-state .es-icon { color: #94a3b8; margin-bottom: 16px; } .empty-state h3 { font-size: 18px; font-weight: 600; color: #0f172a; margin: 0 0 4px; } .empty-state p { font-size: 14px; color: #64748b; max-width: 28rem; margin: 0; } .empty-state .es-action { margin-top: 24px; }

/* --- Table (verbatim from BackupList.svelte) --- */ .table-wrap { border-radius: 8px; border: 1px solid #e2e8f0; background-color: #ffffff; box-shadow: 0 1px 2px 0 rgba(0, 0, 0, 0.05); overflow: hidden; } table.proto-table { min-width: 100%; } table.proto-table thead { background-color: #f1f5f9; } table.proto-table th { padding: 12px 24px; text-align: left; font-size: 12px; font-weight: 500; color: #64748b; text-transform: uppercase; letter-spacing: 0.05em; } table.proto-table tbody { background-color: #ffffff; } table.proto-table tbody tr { border-top: 1px solid #e2e8f0; } table.proto-table tbody tr:hover { background-color: #f1f5f9; } table.proto-table td { padding: 16px 24px; white-space: nowrap; font-size: 14px; } .td-main { font-weight: 500; color: #0f172a; } .td-muted { color: #64748b; } .td-right { text-align: right; }

/* --- Alert/banner --- */ .alert { display: flex; align-items: flex-start; gap: 12px; border-radius: 8px; border: 1px solid; padding: 12px 16px; font-size: 14px; margin-bottom: 16px; } .alert-warning { background-color: #fffbeb; border-color: #fbbf24; color: #a16207; } .alert-destructive { background-color: #fef2f2; border-color: #f87171; color: #b91c1c; } .alert-info { background-color: #f0f9ff; border-color: #38bdf8; color: #0c4a6e; } .alert .alert-actions { margin-left: auto; display: flex; gap: 8px; flex-shrink: 0; }

/* --- Modal (ConfirmDialog pattern) --- */ .modal-backdrop { position: fixed; inset: 0; z-index: 50; background-color: rgba(15, 23, 42, 0.5); display: flex; align-items: center; justify-content: center; padding: 16px; } .modal-card { border-radius: 8px; border: 1px solid #e2e8f0; background-color: #ffffff; color: #0f172a; box-shadow: 0 10px 15px -3px rgba(0,0,0,0.1); width: 100%; max-width: 560px; } .modal-body { padding: 24px; } .modal-footer { display: flex; justify-content: flex-end; gap: 12px; padding: 16px 24px; border-top: 1px solid #e2e8f0; }

/* --- Toast --- */ .toast-wrap { position: fixed; top: 56px; right: 16px; z-index: 2000; display: flex; flex-direction: column; gap: 8px; } .toast { display: flex; align-items: center; gap: 10px; border-radius: 8px; border: 1px solid; padding: 12px 16px; font-size: 14px; min-width: 260px; box-shadow: 0 4px 6px -1px rgba(0, 0, 0, 0.1); } .toast-success { background-color: #f0fdf4; border-color: #4ade80; color: #166534; } .toast-destructive { background-color: #fef2f2; border-color: #f87171; color: #b91c1c; }

/* --- Step indicator (wizard) --- */ .steps { display: flex; align-items: center; gap: 8px; margin-bottom: 24px; } .step-item { display: flex; align-items: center; gap: 8px; font-size: 13px; font-weight: 500; color: #64748b; } .step-item.active { color: #2563eb; } .step-item.done { color: #22c55e; } .step-num { display: inline-flex; align-items: center; justify-content: center; width: 24px; height: 24px; border-radius: 9999px; font-size: 12px; font-weight: 700; background-color: #f1f5f9; color: #64748b; } .step-item.active .step-num { background-color: #2563eb; color: #ffffff; } .step-item.done .step-num { background-color: #22c55e; color: #ffffff; } .step-sep { width: 24px; height: 1px; background-color: #e2e8f0; }

/* --- Org tree --- */ .org-node { display: flex; align-items: center; gap: 8px; padding: 6px 8px; border-radius: 6px; font-size: 14px; } .org-node:hover { background-color: #f1f5f9; } .org-node .org-caret { width: 16px; color: #94a3b8; text-align: center; flex-shrink: 0; } .org-node .org-label { color: #0f172a; font-weight: 500; } .org-node .org-count { color: #94a3b8; font-size: 12px; } .org-children { margin-left: 24px; }

/* --- Checkbox (native styled) --- */ input[type="checkbox"] { width: 16px; height: 16px; accent-color: #2563eb; cursor: pointer; }

/* --- Category gradient chips (category.admin token) --- */ .chip-category { background: linear-gradient(135deg, #ffe4e6, #fecdd3); color: #be123c; border-radius: 6px; padding: 4px 10px; font-size: 12px; font-weight: 600; }

/* --- Monospace (SQL preview) --- */ .mono { font-family: "JetBrains Mono", "Fira Code", monospace; font-size: 12px; line-height: 1.5; color: #cbd5e1; background-color: #0f172a; border-radius: 8px; padding: 16px; overflow-x: auto; white-space: pre; }

@media (prefers-reduced-motion: reduce) { .skeleton-line, .skeleton-card, .skeleton-row, .spin { animation: none !important; } .transition-colors, .transition-transform { transition-duration: 0.01ms !important; } } </style>

</head> PROTOTYPE Экран: 1 · RLS Dashboard 2 · Аудит датасетов 3 · Аудит пользователей 4 · Конструктор правил Состояние:
1280px 375px
<main class="proto-main" id="proto-main">
  <div class="proto-topbar" aria-hidden="true">
    <span class="chip-category">RLS</span>
    <span class="text-sm text-text-muted">Оператор: admin</span>
  </div>

  <!-- ================= SCREEN 1: RLS DASHBOARD ================= -->
  <section id="screen-dashboard" class="proto-screen">
    <header class="flex items-center justify-between mb-8">
      <div class="space-y-1">
        <h1 class="text-3xl font-bold tracking-tight text-text">RLS Management</h1>
        <p class="text-sm text-text-muted" data-role="subtitle">Версионирование скрипта, аудит доступа, ведение кастомных правил</p>
      </div>
      <div class="flex items-center gap-4">
        <button class="btn btn-secondary btn-md" type="button" data-role="new-version">Новая версия из активной</button>
        <button class="btn btn-primary btn-md" type="button" data-role="push">Push в корпоративный репозиторий</button>
      </div>
    </header>

    <!-- stale banner (state: stale) -->
    <div class="alert alert-warning" data-role="stale-banner" hidden>
      <span>⚠️ Снапшот rls_t от 02.08.2026 — данные устарели. Последнее обновление DWH: 03.08.2026.</span>
      <span class="alert-actions"><button class="btn btn-warning btn-sm" type="button">Обновить снапшот</button></span>
    </div>
    <!-- push-failed banner (state: push-failed) -->
    <div class="alert alert-destructive" data-role="push-failed-banner" hidden>
      <span>✕ Push в корпоративный репозиторий не выполнен: remote 'corp' недоступен (Connection refused). Версия v12 сохранена локально, данные не потеряны.</span>
      <span class="alert-actions"><button class="btn btn-destructive btn-sm" type="button">Повторить push</button></span>
    </div>

    <!-- Cards row -->
    <div class="grid grid-cols-3 gap-4 mb-6" data-role="cards">
      <div class="card">
        <div class="card-title-row"><h3 class="card-title">Актуальность rls_t</h3></div>
        <div class="card-pad-md">
          <div data-role="fresh-badge">
            <span class="badge badge-pill badge-md badge-success"><span class="badge-dot dot-success"></span> Актуальна</span>
          </div>
          <p class="text-sm text-text-muted mt-3 mb-1">Обновлено: <span class="text-sm font-medium text-text">03.08.2026 06:00</span></p>
          <p class="text-sm text-text-muted mb-1">Строк: <span class="text-sm font-medium text-text">214 503</span></p>
          <p class="text-sm text-text-muted">Порог устаревания: <span class="text-sm font-medium text-text">3 дня</span></p>
        </div>
      </div>
      <div class="card">
        <div class="card-title-row"><h3 class="card-title">Версии скрипта</h3></div>
        <div class="card-pad-md">
          <p class="text-sm text-text-muted mb-1">Активная версия: <span class="badge badge-pill badge-md badge-primary">v12</span></p>
          <p class="text-sm text-text-muted mb-1">Автор: <span class="text-sm font-medium text-text">a.ivanov</span></p>
          <p class="text-sm text-text-muted">Последний push: <span class="text-sm font-medium text-text">02.08.2026</span></p>
        </div>
      </div>
      <div class="card">
        <div class="card-title-row"><h3 class="card-title">Аудит</h3></div>
        <div class="card-pad-md">
          <p class="text-sm text-text-muted mb-1">Уволенные в bi_users: <span class="badge badge-pill badge-md badge-destructive">2</span></p>
          <p class="text-sm text-text-muted mb-1">Неактивные аккаунты: <span class="badge badge-pill badge-md badge-warning">1</span></p>
          <p class="text-sm text-text-muted">Датасетов без RLS: <span class="badge badge-pill badge-md badge-info">4</span></p>
        </div>
      </div>
    </div>

    <!-- Versions table -->
    <div class="table-wrap" data-role="versions-table">
      <div class="px-6 py-4 bg-surface-page border-b border-border flex items-center justify-between">
        <h3 class="text-lg font-semibold text-text">Версии скрипта rls_t main insert.sql</h3>
        <span class="badge badge-pill badge-md badge-muted">10 версий</span>
      </div>
      <div class="overflow-x-auto">
        <table class="proto-table">
          <thead>
            <tr>
              <th>Версия</th><th>Дата</th><th>Автор</th><th>Комментарий</th><th>Статус</th><th class="text-right">Действия</th>
            </tr>
          </thead>
          <tbody>
            <tr>
              <td class="td-main">v12</td><td class="td-muted">03.08.2026</td><td class="td-muted">a.ivanov</td><td class="td-muted">Поддержка wildcard для FM_COMMENTS_INPUT</td>
              <td><span class="badge badge-pill badge-md badge-success">активная</span></td>
              <td class="td-right"><button class="btn btn-ghost btn-sm" type="button">Diff</button></td>
            </tr>
            <tr>
              <td class="td-main">v11</td><td class="td-muted">28.07.2026</td><td class="td-muted">p.petrov</td><td class="td-muted">Маска SAP_ISO ^CS_…(WO|RO)$</td>
              <td><span class="badge badge-pill badge-md badge-muted">архив</span></td>
              <td class="td-right"><button class="btn btn-ghost btn-sm" type="button">Diff</button></td>
            </tr>
            <tr>
              <td class="td-main">v10</td><td class="td-muted">15.07.2026</td><td class="td-muted">a.ivanov</td><td class="td-muted">Оптимизация NOT EXISTS → LEFT JOIN</td>
              <td><span class="badge badge-pill badge-md badge-muted">архив</span></td>
              <td class="td-right"><button class="btn btn-ghost btn-sm" type="button">Diff</button></td>
            </tr>
          </tbody>
        </table>
      </div>
    </div>
  </section>

  <!-- ================= SCREEN 2: AUDIT DATASETS ================= -->
  <section id="screen-audit-datasets" class="proto-screen" hidden>
    <header class="flex items-center justify-between mb-8">
      <div class="space-y-1">
        <h1 class="text-3xl font-bold tracking-tight text-text">Аудит доступа к датасетам</h1>
        <p class="text-sm text-text-muted">Какой пользователь какие данные видит (rls_t × датасеты × unit_id)</p>
      </div>
      <div class="flex items-center gap-4">
        <span class="badge badge-pill badge-md badge-info">снапшот от 03.08.2026</span>
      </div>
    </header>

    <!-- degraded banner (state: degraded) -->
    <div class="alert alert-warning" data-role="degraded-banner" hidden>
      <span>⚠️ DWH недоступна — показан снапшот от 02.08.2026. Матрица может не отражать последние изменения rls_t.</span>
      <span class="alert-actions"><button class="btn btn-warning btn-sm" type="button">Обновить снапшот</button></span>
    </div>

    <!-- Direction toggle + search -->
    <div class="card card-pad-md mb-6">
      <div class="flex items-center gap-4 flex-wrap">
        <nav class="tablist" role="tablist" aria-label="Направление аудита">
          <button class="tab-segmented active" role="tab" aria-selected="true" type="button" data-dir="user">Пользователь → данные</button>
          <button class="tab-segmented inactive" role="tab" aria-selected="false" type="button" data-dir="dataset">Датасет → пользователи</button>
        </nav>
        <div class="field flex-1" style="max-width:360px">
          <div class="field-label"><label for="audit-search">AD-логин</label></div>
          <div class="flex items-center gap-2">
            <input id="audit-search" class="input" type="text" placeholder="например, ivanovii" aria-describedby="audit-hint">
            <button class="btn btn-primary btn-md" type="button" data-role="audit-search-btn">Найти</button>
          </div>
          <span id="audit-hint" class="field-error" style="color:#94a3b8">user_id хранится в нижнем регистре</span>
        </div>
      </div>
    </div>

    <!-- searching state -->
    <div data-role="searching" hidden>
      <div class="table-wrap">
        <div class="card-pad-md">
          <div class="skeleton-row mb-3" style="width:40%"></div>
          <div class="skeleton-row mb-3"></div>
          <div class="skeleton-row" style="width:70%"></div>
        </div>
      </div>
    </div>

    <!-- empty state -->
    <div data-role="empty" hidden>
      <div class="card">
        <div class="empty-state">
          <svg class="es-icon w-16 h-16" fill="none" stroke="currentColor" viewBox="0 0 24 24" aria-hidden="true"><path stroke-linecap="round" stroke-linejoin="round" stroke-width="1.5" d="M16 7a4 4 0 11-8 0 4 4 0 018 0zM12 14a7 7 0 00-7 7h14a7 7 0 00-7-7z"/></svg>
          <h3>Пользователь не найден в RLS</h3>
          <p>Записей для логина «KOZLOVAKK» нет в снапшоте rls_t. Проверьте регистр (user_id хранится в нижнем регистре) или наличие доступа в системе.</p>
          <div class="es-action"><button class="btn btn-secondary btn-md" type="button">Очистить поиск</button></div>
        </div>
      </div>
    </div>

    <!-- loaded matrix -->
    <div data-role="loaded" hidden>
      <div class="space-y-4">
        <div class="alert alert-info" data-role="rls-type-note">
          <span>ℹ️ Пользователь имеет rls_type: <b>user_to_unit_balance_code</b> (финансовые единицы) и <b>user_to_plant_code</b> (заводы). Сопоставление выполнено по типу единицы датасета.</span>
        </div>
        <div class="table-wrap">
          <div class="px-6 py-4 bg-surface-page border-b border-border flex items-center justify-between">
            <h3 class="text-lg font-semibold text-text">ivanovii — видимые данные</h3>
            <span class="badge badge-pill badge-md badge-muted">5 датасетов · 3 unit_id</span>
          </div>
          <div class="overflow-x-auto">
            <table class="proto-table">
              <thead>
                <tr><th>Датасет</th><th>Тип единицы</th><th>Видимые unit_id</th><th>Источник</th><th>Роли</th></tr>
              </thead>
              <tbody>
                <tr><td class="td-main">ФинРезультат_БЕ</td><td class="td-muted">unit_balance_code</td><td class="td-main">199, 210, 305</td><td class="td-muted">SAP</td><td class="td-muted">FI_BUKRS_199, FI_BUKRS_210</td></tr>
                <tr><td class="td-main">Заводы_Производство</td><td class="td-muted">plant_code</td><td class="td-main">1000, 1100</td><td class="td-muted">SAP</td><td class="td-muted">MM_WERKS_1000</td></tr>
                <tr><td class="td-main">Продажи_Регионы</td><td class="td-muted">unit_balance_code</td><td class="td-main">199</td><td class="td-muted">SAP_ISO</td><td class="td-muted">SD_BUKRS_199</td></tr>
                <tr><td class="td-main">Комментарии_ФМ</td><td class="td-muted">unit_balance_code</td><td class="td-main">210</td><td class="td-muted">FM_COMMENTS_INPUT</td><td class="td-muted">FM_…_PFM_0210_WO</td></tr>
                <tr><td class="td-main">Справочник_БЕ</td><td class="td-muted">unit_balance_code</td><td class="td-main">—</td><td class="td-muted">dict</td><td class="td-muted">справочник без RLS</td></tr>
              </tbody>
            </table>
          </div>
        </div>
        <!-- no-RLS group -->
        <div class="card card-pad-md">
          <div class="flex items-center justify-between">
            <div>
              <h3 class="text-lg font-semibold text-text">Датасеты без RLS-фильтрации</h3>
              <p class="text-sm text-text-muted">Не привязаны к unit_id — доступны без ограничений уровня строк</p>
            </div>
            <span class="badge badge-pill badge-md badge-info">4</span>
          </div>
        </div>
      </div>
    </div>
  </section>

  <!-- ================= SCREEN 3: AUDIT USERS (bi_users × IDM) ================= -->
  <section id="screen-audit-users" class="proto-screen" hidden>
    <header class="flex items-center justify-between mb-8">
      <div class="space-y-1">
        <h1 class="text-3xl font-bold tracking-tight text-text">Аудит пользователей</h1>
        <p class="text-sm text-text-muted">Сверка rls_bi_users с IDM: уволенные, неактивные, риск-скор</p>
      </div>
      <div class="flex items-center gap-4">
        <button class="btn btn-secondary btn-md" type="button" data-role="deactivate" disabled>Сформировать SQL деактивации</button>
      </div>
    </header>

    <!-- idm degraded banner -->
    <div class="alert alert-warning" data-role="idm-banner" hidden>
      <span>⚠️ Сервис IDM не отвечает. Статусы сотрудников недоступны, проверка по bi_users продолжается.</span>
      <span class="alert-actions"><button class="btn btn-warning btn-sm" type="button">Повторить</button></span>
    </div>

    <!-- Filter chips -->
    <nav class="flex items-center gap-2 mb-6 flex-wrap" role="tablist" aria-label="Фильтры аудита">
      <button class="tab-pills active" type="button" role="tab" aria-selected="true">Все <span class="tab-badge">12</span></button>
      <button class="tab-pills inactive" type="button" role="tab" aria-selected="false">Уволенные <span class="tab-badge" style="background-color:#fef2f2;color:#dc2626">2</span></button>
      <button class="tab-pills inactive" type="button" role="tab" aria-selected="false">Неактивные <span class="tab-badge" style="background-color:#fffbeb;color:#f59e0b">1</span></button>
      <button class="tab-pills inactive" type="button" role="tab" aria-selected="false">В отпуске <span class="tab-badge" style="background-color:#f0f9ff;color:#0ea5e9">1</span></button>
      <button class="tab-pills inactive" type="button" role="tab" aria-selected="false">adminRights <span class="tab-badge">1</span></button>
    </nav>

    <!-- loading state -->
    <div data-role="users-loading" hidden>
      <div class="table-wrap">
        <div class="card-pad-md">
          <div class="skeleton-row mb-3"></div>
          <div class="skeleton-row mb-3"></div>
          <div class="skeleton-row mb-3"></div>
          <div class="skeleton-row" style="width:60%"></div>
        </div>
      </div>
    </div>

    <!-- empty state -->
    <div data-role="users-empty" hidden>
      <div class="card">
        <div class="empty-state">
          <svg class="es-icon w-16 h-16" fill="none" stroke="currentColor" viewBox="0 0 24 24" aria-hidden="true"><path stroke-linecap="round" stroke-linejoin="round" stroke-width="1.5" d="M9 12h6m-6 4h6m2 5H7a2 2 0 01-2-2V5a2 2 0 012-2h5.586a1 1 0 01.707.293l5.414 5.414a1 1 0 01.293.707V19a2 2 0 01-2 2z"/></svg>
          <h3>Нет пользователей</h3>
          <p>rls_bi_users пуст — добавьте пользователей через конструктор правил или миграции.</p>
          <div class="es-action"><button class="btn btn-primary btn-md" type="button">Перейти к конструктору</button></div>
        </div>
      </div>
    </div>

    <!-- loaded table -->
    <div data-role="users-loaded" class="space-y-4">
      <div class="table-wrap">
        <div class="px-6 py-4 bg-surface-page border-b border-border flex items-center justify-between">
          <h3 class="text-lg font-semibold text-text">rls_bi_users × IDM</h3>
          <span class="badge badge-pill badge-md badge-muted">12 пользователей</span>
        </div>
        <div class="overflow-x-auto">
          <table class="proto-table">
            <thead>
              <tr><th style="width:32px"><input type="checkbox" aria-label="Выбрать всех"></th><th>AD-логин</th><th>Роль</th><th>Сотрудник</th><th>Аккаунт AD</th><th>Риск</th><th>VIP</th><th>admin</th><th>Обновлено</th></tr>
            </thead>
            <tbody>
              <tr>
                <td><input type="checkbox" aria-label="Выбрать novikovnn"></td>
                <td class="td-main">novikovnn</td><td class="td-muted">BI_ALL</td>
                <td><span class="badge badge-pill badge-md badge-destructive">уволен</span></td>
                <td><span class="badge badge-pill badge-md badge-warning">disabled</span></td>
                <td class="td-muted">70</td><td class="td-muted">да</td><td class="td-muted">да</td><td class="td-muted">03.08.2026</td>
              </tr>
              <tr>
                <td><input type="checkbox" aria-label="Выбрать sidorovss"></td>
                <td class="td-main">sidorovss</td><td class="td-muted">BI_bukrs_305</td>
                <td><span class="badge badge-pill badge-md badge-info">в отпуске</span></td>
                <td><span class="badge badge-pill badge-md badge-success">active</span></td>
                <td class="td-muted">40</td><td class="td-muted">—</td><td class="td-muted">—</td><td class="td-muted">03.08.2026</td>
              </tr>
              <tr>
                <td><input type="checkbox" aria-label="Выбрать ivanovii"></td>
                <td class="td-main">ivanovii</td><td class="td-muted">BI_bukrs_199</td>
                <td><span class="badge badge-pill badge-md badge-success">активен</span></td>
                <td><span class="badge badge-pill badge-md badge-success">active</span></td>
                <td class="td-muted">45</td><td class="td-muted">—</td><td class="td-muted">—</td><td class="td-muted">03.08.2026</td>
              </tr>
              <tr>
                <td><input type="checkbox" aria-label="Выбрать petrovpv"></td>
                <td class="td-main">petrovpv</td><td class="td-muted">BI_bukrs_199</td>
                <td><span class="badge badge-pill badge-md badge-success">активен</span></td>
                <td><span class="badge badge-pill badge-md badge-success">active</span></td>
                <td class="td-muted">30</td><td class="td-muted">—</td><td class="td-muted">—</td><td class="td-muted">03.08.2026</td>
              </tr>
              <tr>
                <td><input type="checkbox" aria-label="Выбрать kozlovakk"></td>
                <td class="td-main">kozlovakk</td><td class="td-muted">BI_bukrs_210</td>
                <td><span class="badge badge-pill badge-md badge-destructive">уволен</span></td>
                <td><span class="badge badge-pill badge-md badge-warning">disabled</span></td>
                <td class="td-muted">0</td><td class="td-muted">—</td><td class="td-muted">—</td><td class="td-muted">02.08.2026</td>
              </tr>
              <tr>
                <td><input type="checkbox" aria-label="Выбрать volkovaa"></td>
                <td class="td-main">volkovaa</td><td class="td-muted">BI_bukrs_210</td>
                <td><span class="badge badge-pill badge-md badge-muted">—</span></td>
                <td><span class="badge badge-pill badge-md badge-muted">—</span></td>
                <td class="td-muted">—</td><td class="td-muted">—</td><td class="td-muted">—</td><td class="td-muted">02.08.2026</td>
              </tr>
            </tbody>
          </table>
        </div>
      </div>

      <!-- Employee card (expanded row) -->
      <div class="card">
        <div class="card-title-row">
          <div class="flex items-center justify-between">
            <h3 class="card-title">Карточка сотрудника — Иванов Иван Иванович</h3>
            <span class="badge badge-pill badge-md badge-success">активен</span>
          </div>
        </div>
        <div class="card-pad-md">
          <div class="grid grid-cols-3 gap-4 mb-4">
            <div><p class="text-sm text-text-muted">Отдел</p><p class="text-sm font-medium text-text">Департамент информационных технологий</p></div>
            <div><p class="text-sm text-text-muted">Должность</p><p class="text-sm font-medium text-text">Разработчик</p></div>
            <div><p class="text-sm text-text-muted">Компания</p><p class="text-sm font-medium text-text">ООО «ТестСервис»</p></div>
            <div><p class="text-sm text-text-muted">Дата приёма</p><p class="text-sm font-medium text-text">17.04.2023</p></div>
            <div><p class="text-sm text-text-muted">Риск-скор</p><p class="text-sm font-medium text-text">45</p></div>
            <div><p class="text-sm text-text-muted">Учётная запись</p><p class="text-sm font-medium text-text">TEST.LOCAL\ivanovii</p></div>
          </div>
          <h4 class="text-sm font-semibold text-text mb-2">Аккаунты и приложения</h4>
          <div class="table-wrap">
            <table class="proto-table">
              <thead><tr><th>Приложение</th><th>Аккаунт</th><th>OU</th><th>adminRights</th></tr></thead>
              <tbody>
                <tr><td class="td-main">Active Directory (TEST.LOCAL)</td><td class="td-muted">IvanovII</td><td class="td-muted">OU=External,OU=Users,OU=Test,OU=Moscow,DC=TEST,DC=LOCAL</td><td><span class="badge badge-pill badge-md badge-muted">false</span></td></tr>
                <tr><td class="td-main">Active Directory (IE.CORP)</td><td class="td-muted">IvanovII</td><td class="td-muted">OU=Unassigned accounts,OU=Test,DC=ie,DC=corp</td><td><span class="badge badge-pill badge-md badge-muted">false</span></td></tr>
                <tr><td class="td-main">Exchange</td><td class="td-muted">IvanovII</td><td class="td-muted">TEST.LOCAL</td><td><span class="badge badge-pill badge-md badge-muted">false</span></td></tr>
                <tr><td class="td-main">RPAM / RPKI</td><td class="td-muted">Company PAM / PKI</td><td class="td-muted">—</td><td><span class="badge badge-pill badge-md badge-muted">false</span></td></tr>
              </tbody>
            </table>
          </div>
        </div>
      </div>
    </div>
  </section>

  <!-- ================= SCREEN 4: RULE BUILDER ================= -->
  <section id="screen-rule-builder" class="proto-screen" hidden>
    <header class="flex items-center justify-between mb-8">
      <div class="space-y-1">
        <h1 class="text-3xl font-bold tracking-tight text-text">Конструктор правил</h1>
        <p class="text-sm text-text-muted">Справочник → фильтр → предпросмотр → привязка → применение (INSERT на DWH + SQL в репозиторий)</p>
      </div>
    </header>

    <div class="flex gap-6 items-start">
      <!-- Rules catalog -->
      <div class="w-full" style="max-width:280px">
        <div class="card">
          <div class="card-title-row"><h3 class="card-title">Каталог правил</h3></div>
          <div class="card-pad-sm">
            <div class="space-y-1">
              <button class="btn btn-ghost btn-sm w-full text-left" type="button" style="justify-content:flex-start">БЕ по ЦО (ЦО-БЕ-ПФМ) <span class="badge badge-pill badge-sm badge-primary ml-2">34</span></button>
              <button class="btn btn-ghost btn-sm w-full text-left" type="button" style="justify-content:flex-start">Заводы по дивизиону <span class="badge badge-pill badge-sm badge-muted ml-2">12</span></button>
              <button class="btn btn-ghost btn-sm w-full text-left" type="button" style="justify-content:flex-start">БЕ по региону <span class="badge badge-pill badge-sm badge-muted ml-2">8</span></button>
            </div>
          </div>
          <div class="card-pad-md" style="padding-top:0">
            <button class="btn btn-primary btn-md w-full" type="button" data-role="new-rule">+ Новое правило</button>
          </div>
        </div>
      </div>

      <!-- Wizard -->
      <div class="flex-1 min-w-0">
        <!-- Step indicator -->
        <div class="steps" aria-label="Шаги мастера">
          <div class="step-item done"><span class="step-num">1</span> Справочник</div>
          <div class="step-sep"></div>
          <div class="step-item done"><span class="step-num">2</span> Фильтры</div>
          <div class="step-sep"></div>
          <div class="step-item active"><span class="step-num">3</span> Предпросмотр</div>
          <div class="step-sep"></div>
          <div class="step-item"><span class="step-num">4</span> Привязка</div>
        </div>

        <div class="card mb-6">
          <div class="card-title-row"><h3 class="card-title">Правило: БЕ по ЦО</h3></div>
          <div class="card-pad-md space-y-4">
            <div class="grid grid-cols-2 gap-4">
              <div class="field">
                <div class="field-label"><label for="dict-select">Справочник (витрина)</label></div>
                <select id="dict-select" class="select">
                  <option>dds.plant_and_subsidiary (ЦО-БЕ-ПФМ)</option>
                  <option>dds.unit_balance</option>
                </select>
              </div>
              <div class="field">
                <div class="field-label"><label for="unit-col">Колонка unit_id</label></div>
                <select id="unit-col" class="select">
                  <option>pfm (PFM-код)</option>
                  <option>unit_balance_code</option>
                </select>
              </div>
            </div>
            <div class="field">
              <div class="field-label"><label for="rule-name">Название правила</label></div>
              <input id="rule-name" class="input" type="text" value="БЕ по ЦО" placeholder="Например: БЕ по центру ответственности">
            </div>

            <div>
              <div class="flex items-center justify-between mb-2">
                <h4 class="text-sm font-semibold text-text">Фильтры (применяются на целевой БД)</h4>
                <button class="btn btn-secondary btn-sm" type="button">+ фильтр</button>
              </div>
              <div class="table-wrap">
                <table class="proto-table">
                  <thead><tr><th>Атрибут</th><th>Тип</th><th>Выражение</th><th style="width:40px"></th></tr></thead>
                  <tbody>
                    <tr>
                      <td><select class="select" style="min-width:180px"><option>centr_otv (центр ответственности)</option><option>be (бизнес-единица)</option><option>pfm</option></select></td>
                      <td><select class="select" style="min-width:120px"><option>WHERE IN</option><option>LIKE</option><option>regexp</option></select></td>
                      <td><input class="input" type="text" value="'ЦО-100', 'ЦО-110'"></td>
                      <td><button class="btn btn-ghost btn-sm" type="button" aria-label="Удалить фильтр">✕</button></td>
                    </tr>
                    <tr>
                      <td><select class="select" style="min-width:180px"><option>centr_otv (центр ответственности)</option><option>be (бизнес-единица)</option><option selected>be</option></select></td>
                      <td><select class="select" style="min-width:120px"><option>WHERE IN</option><option>LIKE</option><option selected>regexp</option></select></td>
                      <td>
                        <input class="input input-error" type="text" value="^(БЕ|BE)_[0-9]{4}(??" aria-invalid="true">
                        <span class="field-error" id="regexp-error">Ошибка regexp: unbalanced parenthesis at position 16</span>
                      </td>
                      <td><button class="btn btn-ghost btn-sm" type="button" aria-label="Удалить фильтр">✕</button></td>
                    </tr>
                  </tbody>
                </table>
              </div>
            </div>
          </div>
        </div>

        <!-- Step 3: preview -->
        <div class="card mb-6">
          <div class="card-title-row">
            <div class="flex items-center justify-between">
              <h3 class="card-title">Предпросмотр</h3>
              <button class="btn btn-secondary btn-md" type="button" data-role="preview-btn">Предпросмотр на целевой БД</button>
            </div>
          </div>
          <div class="card-pad-md">
            <div data-role="preview-loading" hidden>
              <div class="flex items-center gap-2 mb-3"><span class="text-sm text-text-muted">Выполняется SELECT DISTINCT pfm FROM dds.plant_and_subsidiary WHERE …</span></div>
              <div class="skeleton-row mb-3"></div><div class="skeleton-row mb-3" style="width:70%"></div>
            </div>
            <div data-role="preview-empty" hidden>
              <div class="alert alert-warning" style="margin-bottom:0">
                <span>⚠️ Фильтр не вернул ни одной записи. Проверьте атрибуты и выражение.</span>
                <span class="alert-actions"><button class="btn btn-warning btn-sm" type="button">Вернуться к фильтрам</button></span>
              </div>
            </div>
            <div data-role="preview-loaded" class="space-y-3">
              <p class="text-sm text-text-muted">В правило попадает <span class="badge badge-pill badge-md badge-primary">34 unit_id</span> из справочника (первые 5 из 500):</p>
              <div class="table-wrap">
                <table class="proto-table">
                  <thead><tr><th>pfm</th><th>be</th><th>centr_otv</th></tr></thead>
                  <tbody>
                    <tr><td class="td-main">0210</td><td class="td-muted">БЕ-001</td><td class="td-muted">ЦО-100</td></tr>
                    <tr><td class="td-main">0310</td><td class="td-muted">БЕ-001</td><td class="td-muted">ЦО-100</td></tr>
                    <tr><td class="td-main">0410</td><td class="td-muted">БЕ-002</td><td class="td-muted">ЦО-110</td></tr>
                    <tr><td class="td-main">0510</td><td class="td-muted">БЕ-002</td><td class="td-muted">ЦО-110</td></tr>
                    <tr><td class="td-main">0610</td><td class="td-muted">БЕ-003</td><td class="td-muted">ЦО-120</td></tr>
                  </tbody>
                </table>
              </div>
            </div>
          </div>
        </div>

        <!-- Step 4: bindings -->
        <div class="card mb-6">
          <div class="card-title-row"><h3 class="card-title">Привязка пользователей</h3></div>
          <div class="card-pad-md">
            <nav class="tablist gap-2 mb-4" role="tablist" aria-label="Способ привязки">
              <button class="tab-pills active" role="tab" aria-selected="true" type="button">По AD-логину</button>
              <button class="tab-pills inactive" role="tab" aria-selected="false" type="button">AD-группа / OU</button>
              <button class="tab-pills inactive" role="tab" aria-selected="false" type="button">Оргструктура</button>
            </nav>
            <div class="flex items-center gap-2 mb-4">
              <input class="input flex-1" type="text" placeholder="Введите AD-логин… (поиск через IDM)" aria-label="AD-логин">
              <button class="btn btn-secondary btn-md" type="button">Добавить</button>
            </div>
            <div class="space-y-2" data-role="bound-users">
              <div class="flex items-center justify-between border border-border rounded-md px-4 py-2">
                <span class="text-sm font-medium text-text">ivanovii</span>
                <span class="badge badge-pill badge-md badge-success">найден в IDM</span>
              </div>
              <div class="flex items-center justify-between border border-border rounded-md px-4 py-2">
                <span class="text-sm font-medium text-text">petrovpv</span>
                <span class="badge badge-pill badge-md badge-success">найден в IDM</span>
              </div>
              <div class="flex items-center justify-between border border-border rounded-md px-4 py-2">
                <span class="text-sm font-medium text-text">sidorovss</span>
                <span class="badge badge-pill badge-md badge-warning">в отпуске</span>
              </div>
              <div class="flex items-center justify-between border border-border rounded-md px-4 py-2">
                <span class="text-sm font-medium text-text">kozlovakk</span>
                <span class="badge badge-pill badge-md badge-destructive">уволен — будет исключена из роли</span>
              </div>
            </div>
          </div>
        </div>

        <!-- Footer actions -->
        <div class="flex items-center justify-end gap-4">
          <button class="btn btn-ghost btn-md" type="button">Сохранить черновик</button>
          <button class="btn btn-primary btn-md" type="button" data-role="apply-btn">Применить (INSERT на DWH)</button>
        </div>
      </div>
    </div>
  </section>
</main>

Сгенерированный SQL — v1

INSERT будет исполнен на целевой БД. Скрипт сохраняется в репозиторий как миграция (diff ниже).

-- rls_rule: БЕ по ЦО (v1) · rls_type: user_to_unit_balance_code
INSERT INTO rls.rls_bi_roles (role, unit_id, rls_type, dt_created, source_system)
SELECT DISTINCT 'BI_czo_' || pfm, pfm, 'user_to_unit_balance_code', NOW(), 'CUSTOM'
FROM dds.plant_and_subsidiary
WHERE centr_otv IN ('ЦО-100','ЦО-110');

INSERT INTO rls.rls_bi_users (ad_user, role, role_descr, timestamp, is_deleted) VALUES ('ivanovii','BI_czo_0210','048-rls','2026-08-04',false), ('petrovpv','BI_czo_0210','048-rls','2026-08-04',false);

Diff vs v0: +34 unit_id, −0

Отмена Применить

Деактивировать 2 пользователей?

novikovnn (уволен 15.01.2026), kozlovakk (уволен 10.03.2026). Будет выполнен UPDATE is_deleted=true на целевой БД и сохранена SQL-миграция.

Отмена Деактивировать
<script> /* ============================================================ STATE SWITCHER — прототип-хром ============================================================ */ const SCREENS = { 'screen-dashboard': { label: '1 · RLS Dashboard', states: { loaded: { label: 'loaded — данные, всё актуально' }, loading: { label: 'loading — скелетоны' }, stale: { label: 'stale — снапшот устарел (amber banner)' }, 'push-failed': { label: 'push-failed — ошибка push (destructive banner)' }, }, }, 'screen-audit-datasets': { label: '2 · Аудит датасетов', states: { idle: { label: 'idle — ожидание поиска' }, searching: { label: 'searching — выполняется поиск' }, loaded: { label: 'loaded — матрица с данными' }, empty: { label: 'empty — пользователь не найден' }, degraded: { label: 'degraded — снапшот устарел' }, }, }, 'screen-audit-users': { label: '3 · Аудит пользователей', states: { 'users-loaded': { label: 'loaded — таблица + карточка IDM' }, 'users-loading': { label: 'loading — скелетоны' }, 'idm-degraded': { label: 'idm-degraded — IDM недоступен' }, 'users-empty': { label: 'empty — bi_users пуст' }, }, }, 'screen-rule-builder': { label: '4 · Конструктор правил', states: { 'step-1': { label: 'step 1 — справочник' }, 'step-2': { label: 'step 2 — фильтры (regexp error)' }, 'step-3-preview': { label: 'step 3 — предпросмотр загружен' }, 'step-3-loading': { label: 'step 3 — предпросмотр выполняется' }, 'step-3-empty': { label: 'step 3 — предпросмотр пуст (0 unit_id)' }, 'step-4': { label: 'step 4 — привязка пользователей' }, applying: { label: 'applying — INSERT выполняется' }, applied: { label: 'applied — применено, SQL в репозитории' }, }, }, }; const screenSelect = document.getElementById('screen-select'); const stateSelect = document.getElementById('state-select'); const stateLabel = document.getElementById('state-label'); function applyScreenState(screenId, stateKey) { // hide all screens document.querySelectorAll('.proto-screen').forEach(s => { s.hidden = true; }); const screen = document.getElementById(screenId); screen.hidden = false; document.getElementById('proto-main').scrollTop = 0; // per-screen generic toggles const sc = SCREENS[screenId].states[stateKey]; // Dashboard const sb = document.querySelector('#screen-dashboard [data-role="stale-banner"]'); const pb = document.querySelector('#screen-dashboard [data-role="push-failed-banner"]'); if (sb) sb.hidden = true; if (pb) pb.hidden = true; // Audit datasets const dg = document.querySelector('#screen-audit-datasets [data-role="degraded-banner"]'); if (dg) dg.hidden = true; // Audit users const idmBanner = document.querySelector('#screen-audit-users [data-role="idm-banner"]'); if (idmBanner) idmBanner.hidden = true; // Rule builder preview ['preview-loading', 'preview-empty', 'preview-loaded'].forEach(k => { const el = document.querySelector('#screen-rule-builder [data-role="' + k + '"]'); if (el) el.hidden = true; }); if (screenId === 'screen-dashboard') { if (stateKey === 'stale') sb.hidden = false; if (stateKey === 'push-failed') pb.hidden = false; if (stateKey === 'loading') { showDashboardSkeletons(true); } else { showDashboardSkeletons(false); } } if (screenId === 'screen-audit-datasets') { const map = { idle: 'idle', searching: 'searching', loaded: 'loaded', empty: 'empty', degraded: 'degraded' }; ['searching', 'empty', 'loaded'].forEach(k => { const el = document.querySelector('#screen-audit-datasets [data-role="' + k + '"]'); if (el) el.hidden = k !== map[stateKey]; }); if (stateKey === 'degraded') { document.querySelector('#screen-audit-datasets [data-role="loaded"]').hidden = false; dg.hidden = false; } } if (screenId === 'screen-audit-users') { ['users-loading', 'users-empty', 'users-loaded'].forEach(k => { const el = document.querySelector('#screen-audit-users [data-role="' + k + '"]'); if (el) el.hidden = k !== stateKey; }); if (stateKey === 'idm-degraded') { document.querySelector('#screen-audit-users [data-role="users-loaded"]').hidden = false; idmBanner.hidden = false; } } if (screenId === 'screen-rule-builder') { const pre = document.querySelector('#screen-rule-builder [data-role="preview-loaded"]'); const preLoad = document.querySelector('#screen-rule-builder [data-role="preview-loading"]'); const preEmpty = document.querySelector('#screen-rule-builder [data-role="preview-empty"]'); if (stateKey === 'step-3-preview') { pre.hidden = false; } if (stateKey === 'step-3-loading') { preLoad.hidden = false; } if (stateKey === 'step-3-empty') { preEmpty.hidden = false; } // apply button spinner state const applyBtn = document.querySelector('#screen-rule-builder [data-role="apply-btn"]'); if (stateKey === 'applying') { applyBtn.innerHTML = 'Применяется…'; applyBtn.disabled = true; } else { applyBtn.innerHTML = 'Применить (INSERT на DWH)'; applyBtn.disabled = false; } } stateLabel.textContent = sc.label; } function showDashboardSkeletons(on) { const cards = document.querySelector('#screen-dashboard [data-role="cards"]'); const table = document.querySelector('#screen-dashboard [data-role="versions-table"]'); if (on) { cards.style.display = 'none'; table.style.display = 'none'; let sk = document.getElementById('dash-skeleton'); if (!sk) { sk = document.createElement('div'); sk.id = 'dash-skeleton'; sk.className = 'space-y-4'; sk.innerHTML = '
'; document.querySelector('#screen-dashboard').appendChild(sk); } } else { cards.style.display = ''; table.style.display = ''; const sk = document.getElementById('dash-skeleton'); if (sk) sk.remove(); } } function renderStateOptions(screenId) { const states = SCREENS[screenId].states; stateSelect.innerHTML = ''; Object.keys(states).forEach(k => { const opt = document.createElement('option'); opt.value = k; opt.textContent = states[k].label; stateSelect.appendChild(opt); }); const first = Object.keys(states)[0]; stateSelect.value = first; applyScreenState(screenId, first); } screenSelect.addEventListener('change', () => renderStateOptions(screenSelect.value)); stateSelect.addEventListener('change', () => applyScreenState(screenSelect.value, stateSelect.value)); /* Viewport toggles */ document.getElementById('vp-desktop').addEventListener('click', () => { document.body.classList.remove('chrome-viewport-mobile'); document.getElementById('vp-desktop').classList.add('active'); document.getElementById('vp-mobile').classList.remove('active'); }); document.getElementById('vp-mobile').addEventListener('click', () => { document.body.classList.add('chrome-viewport-mobile'); document.getElementById('vp-mobile').classList.add('active'); document.getElementById('vp-desktop').classList.remove('active'); }); /* ============================================================ TOAST HELPER ============================================================ */ function toast(text, variant) { const wrap = document.getElementById('toast-wrap'); const t = document.createElement('div'); t.className = 'toast toast-' + variant; t.setAttribute('role', 'status'); t.textContent = text; wrap.appendChild(t); setTimeout(() => t.remove(), 4000); } /* ============================================================ INTERACTIVE WIRING (minimal, для валидации recovery-путей) ============================================================ */ // Dashboard: push button → push-failed state const pushBtn = document.querySelector('#screen-dashboard [data-role="push"]'); if (pushBtn) { pushBtn.addEventListener('click', () => { screenSelect.value = 'screen-dashboard'; renderStateOptions('screen-dashboard'); stateSelect.value = 'push-failed'; applyScreenState('screen-dashboard', 'push-failed'); }); } // Push retry → loaded + toast document.addEventListener('click', (e) => { const retry = e.target.closest('#screen-dashboard [data-role="push-failed-banner"] button'); if (retry) { stateSelect.value = 'loaded'; applyScreenState('screen-dashboard', 'loaded'); toast('Push выполнен: активная версия v12 отправлена в корпоративный репозиторий (remote corp)', 'success'); } // Stale → refresh snapshot const refresh = e.target.closest('[data-role="stale-banner"] button, [data-role="degraded-banner"] button'); if (refresh) { toast('Снапшот обновлён: rls_t от 03.08.2026, 214 503 строки', 'success'); } // Audit users: search → loaded / empty const searchBtn = e.target.closest('[data-role="audit-search-btn"]'); if (searchBtn) { const q = document.getElementById('audit-search').value.toLowerCase(); stateSelect.value = 'searching'; applyScreenState('screen-audit-datasets', 'searching'); setTimeout(() => { stateSelect.value = (q === 'kozlovakk' || q === 'nobody') ? 'empty' : 'loaded'; applyScreenState('screen-audit-datasets', stateSelect.value); }, 900); } // Direction tabs const dirTab = e.target.closest('[data-dir]'); if (dirTab) { document.querySelectorAll('[data-dir]').forEach(t => { t.classList.remove('active'); t.classList.add('inactive'); t.setAttribute('aria-selected', 'false'); }); dirTab.classList.remove('inactive'); dirTab.classList.add('active'); dirTab.setAttribute('aria-selected', 'true'); } // Rule builder: preview button → loading → loaded const previewBtn = e.target.closest('[data-role="preview-btn"]'); if (previewBtn) { stateSelect.value = 'step-3-loading'; applyScreenState('screen-rule-builder', 'step-3-loading'); setTimeout(() => { stateSelect.value = 'step-3-preview'; applyScreenState('screen-rule-builder', 'step-3-preview'); }, 1200); } // Apply → modal const applyBtn = e.target.closest('[data-role="apply-btn"]'); if (applyBtn) { document.getElementById('sql-modal').hidden = false; } // Modal confirm → applying → applied + toast const modalConfirm = e.target.closest('[data-role="modal-confirm"]'); if (modalConfirm) { document.getElementById('sql-modal').hidden = true; stateSelect.value = 'applying'; applyScreenState('screen-rule-builder', 'applying'); setTimeout(() => { stateSelect.value = 'applied'; applyScreenState('screen-rule-builder', 'applied'); toast('Правило применено: 34 unit_id, 12 пользователей. SQL-миграция v1 сохранена в репозиторий', 'success'); }, 1400); } const modalCancel = e.target.closest('[data-role="modal-cancel"]'); if (modalCancel) document.getElementById('sql-modal').hidden = true; // Deactivation flow const deactBtn = e.target.closest('[data-role="deactivate"]'); if (deactBtn && !deactBtn.disabled) { document.getElementById('deactivate-modal').hidden = false; } const deactCancel = e.target.closest('[data-role="deact-cancel"]'); if (deactCancel) document.getElementById('deactivate-modal').hidden = true; const deactConfirm = e.target.closest('[data-role="deact-confirm"]'); if (deactConfirm) { document.getElementById('deactivate-modal').hidden = true; toast('Деактивировано: novikovnn, kozlovakk (is_deleted=true). SQL-миграция сохранена', 'success'); } // Modal backdrop click to dismiss if (e.target.classList.contains('modal-backdrop')) e.target.hidden = true; }); /* Checkbox → enable deactivate button */ document.querySelectorAll('#screen-audit-users input[type="checkbox"]').forEach(cb => { cb.addEventListener('change', () => { const checked = document.querySelectorAll('#screen-audit-users input[type="checkbox"]:checked').length; const btn = document.querySelector('#screen-audit-users [data-role="deactivate"]'); btn.disabled = checked === 0; btn.classList.toggle('btn-secondary', checked === 0); btn.classList.toggle('btn-primary', checked > 0); }); }); /* Pills filters (audit users) */ document.querySelectorAll('#screen-audit-users .tab-pills').forEach(p => { p.addEventListener('click', () => { document.querySelectorAll('#screen-audit-users .tab-pills').forEach(x => { x.classList.remove('active'); x.classList.add('inactive'); x.setAttribute('aria-selected', 'false'); }); p.classList.remove('inactive'); p.classList.add('active'); p.setAttribute('aria-selected', 'true'); }); }); /* Rule builder bind tabs */ document.querySelectorAll('#screen-rule-builder .tablist .tab-pills').forEach(p => { p.addEventListener('click', () => { document.querySelectorAll('#screen-rule-builder .tablist .tab-pills').forEach(x => { x.classList.remove('active'); x.classList.add('inactive'); x.setAttribute('aria-selected', 'false'); }); p.classList.remove('inactive'); p.classList.add('active'); p.setAttribute('aria-selected', 'true'); }); }); /* Escape closes modals */ document.addEventListener('keydown', (e) => { if (e.key === 'Escape') { document.getElementById('sql-modal').hidden = true; document.getElementById('deactivate-modal').hidden = true; } }); /* Initial render */ renderStateOptions('screen-dashboard'); </script> </html>

fixtures/manifest.md

Source: fixtures/manifest.md

#region Rls.Fixtures [C:3] [TYPE ADR] [SEMANTICS test,fixture,rls] @defgroup Fixtures Canonical test fixtures for 048-rls-management-workspace.

Fixture Index

@{ Fixture FX_Rls.RuleBuilder.Preview.Valid [C:2] [TYPE Block] [SEMANTICS test,rls,fixture]

@BRIEF Live preview returns unit_count + sample for IN-filter on DWH. @RELATION VERIFIES -> [Rls.RuleBuilder.Preview] @TEST_FIXTURE: valid_preview -> fixtures/api/rule_preview_valid.json @TEST_INVARIANT: PreviewCountsUnits -> VERIFIED_BY: [valid, zero_rows]

@} Fixture FX_Rls.RuleBuilder.Preview.Valid

@{ Fixture FX_Rls.RuleBuilder.Preview.ZeroRows [C:2] [TYPE Block] [SEMANTICS test,rls,fixture]

@BRIEF Filter with no matches → unit_count=0, saveable with hint (roles appear when reference grows). @RELATION VERIFIES -> [Rls.RuleBuilder.Preview] @TEST_EDGE: zero_rows -> unit_count=0 + saveable=true + hint @TEST_FIXTURE: preview_zero_rows -> fixtures/api/rule_preview_zero_rows.json

@} Fixture FX_Rls.RuleBuilder.Preview.ZeroRows

@{ Fixture FX_Rls.RuleBuilder.Preview.DwhError [C:2] [TYPE Block] [SEMANTICS test,rls,fixture]

@BRIEF DWH unreachable during preview → DWH_ERROR envelope, retryable. @RELATION VERIFIES -> [Rls.RuleBuilder.Preview] @TEST_EDGE: external_fail -> DWH_ERROR + retryable=true @TEST_FIXTURE: preview_dwh_error -> fixtures/api/rule_preview_dwh_error.json

@} Fixture FX_Rls.RuleBuilder.Preview.DwhError

@{ Fixture FX_Rls.RuleBuilder.Preview.SqlInjection [C:2] [TYPE Block] [SEMANTICS test,rls,fixture]

@BRIEF Injection attempt in filter value → parameterized binding enforced (rejected path). @RELATION VERIFIES -> [Rls.RuleBuilder.Preview] @TEST_EDGE: rejected_path -> FILTER_PARAMETERIZED, never raw concatenation @TEST_FIXTURE: preview_sql_injection -> fixtures/api/rule_preview_sql_injection.json

@} Fixture FX_Rls.RuleBuilder.Preview.SqlInjection

@{ Fixture FX_Rls.RuleBuilder.Preview.MissingFilter [C:2] [TYPE Block] [SEMANTICS test,rls,fixture]

@BRIEF Empty filter list → 422 VALIDATION_ERROR. @RELATION VERIFIES -> [Rls.RuleBuilder.Preview] @TEST_EDGE: missing_field -> 422, at least one filter required @TEST_FIXTURE: preview_missing_filter -> fixtures/api/rule_preview_missing_filter.json

@} Fixture FX_Rls.RuleBuilder.Preview.MissingFilter

@{ Fixture FX_Rls.RuleBuilder.Save.Valid [C:2] [TYPE Block] [SEMANTICS test,rls,fixture]

@BRIEF Save definition: own DB + UPSERT rls_roles_filter + bindings; effect at next script run. @RELATION VERIFIES -> [Rls.RuleBuilder.SaveDefinition] @TEST_FIXTURE: save_valid -> fixtures/api/rule_save_valid.json @TEST_INVARIANT: SaveIdempotent -> VERIFIED_BY: [valid, duplicate_binding]

@} Fixture FX_Rls.RuleBuilder.Save.Valid

@{ Fixture FX_Rls.RuleBuilder.Save.DuplicateBinding [C:2] [TYPE Block] [SEMANTICS test,rls,fixture]

@BRIEF Re-save → duplicate binding skipped (FR-023). @RELATION VERIFIES -> [Rls.RuleBuilder.SaveDefinition] @TEST_EDGE: duplicate_apply -> bindings_applied=0, duplicates_skipped=1 @TEST_FIXTURE: save_duplicate_binding -> fixtures/api/rule_save_duplicate_binding.json

@} Fixture FX_Rls.RuleBuilder.Save.DuplicateBinding

@{ Fixture FX_Rls.RuleBuilder.Save.DwhFailure [C:2] [TYPE Block] [SEMANTICS test,rls,fixture]

@BRIEF DWH write fails → local row kept, retryable. @RELATION VERIFIES -> [Rls.RuleBuilder.SaveDefinition] @TEST_EDGE: external_fail -> DWH_ERROR + local_row_kept + retryable @TEST_FIXTURE: save_dwh_failure -> fixtures/api/rule_save_dwh_failure.json

@} Fixture FX_Rls.RuleBuilder.Save.DwhFailure

@{ Fixture FX_Rls.RuleBuilder.Save.Deactivated [C:2] [TYPE Block] [SEMANTICS test,rls,fixture]

@BRIEF Rule deactivation = is_deleted=true definition toggle; roles vanish next run. @RELATION VERIFIES -> [Rls.RuleBuilder.SaveDefinition] @TEST_EDGE: deactivate -> is_deleted=true, no confirm needed @TEST_FIXTURE: save_deactivated -> fixtures/api/rule_save_deactivated.json

@} Fixture FX_Rls.RuleBuilder.Save.Deactivated

@{ Fixture FX_Rls.RuleBuilder.Save.RegexpInvalid [C:2] [TYPE Block] [SEMANTICS test,rls,fixture]

@BRIEF Invalid regexp blocks save (FR-014). @RELATION VERIFIES -> [Rls.RuleBuilder.SaveDefinition] @TEST_EDGE: rejected_path -> 422 REGEXP_INVALID @TEST_FIXTURE: save_regexp_invalid -> fixtures/api/rule_save_regexp_invalid.json

@} Fixture FX_Rls.RuleBuilder.Save.RegexpInvalid

@{ Fixture FX_Rls.RuleBuilder.Save.WoPreview [C:2] [TYPE Block] [SEMANTICS test,rls,fixture]

@BRIEF Save without preview → 422. @RELATION VERIFIES -> [Rls.RuleBuilder.SaveDefinition] @TEST_EDGE: missing_field -> 422 PREVIEW_REQUIRED @TEST_FIXTURE: save_wo_preview -> fixtures/api/rule_save_wo_preview.json

@} Fixture FX_Rls.RuleBuilder.Save.WoPreview

@{ Fixture FX_Rls.RuleBuilder.Save.PermissionDenied [C:2] [TYPE Block] [SEMANTICS test,rls,fixture]

@BRIEF rls_script_dev cannot save rules (RBAC separation). @RELATION VERIFIES -> [Rls.RuleBuilder.SaveDefinition] @TEST_EDGE: rejected_path -> 403 FORBIDDEN for rls_script_dev @TEST_FIXTURE: save_permission_denied -> fixtures/api/rule_save_permission_denied.json

@} Fixture FX_Rls.RuleBuilder.Save.PermissionDenied

@{ Fixture FX_Rls.UserAudit.Deactivate.Valid [C:2] [TYPE Block] [SEMANTICS test,rls,fixture]

@BRIEF Deactivation applies UPDATE + migration file. @RELATION VERIFIES -> [Rls.UserAuditService.Deactivate] @TEST_FIXTURE: deactivate_valid -> fixtures/api/deactivate_valid.json

@} Fixture FX_Rls.UserAudit.Deactivate.Valid

@{ Fixture FX_Rls.UserAudit.Deactivate.AlreadyDeleted [C:2] [TYPE Block] [SEMANTICS test,rls,fixture]

@BRIEF Re-deactivation is idempotent. @RELATION VERIFIES -> [Rls.UserAuditService.Deactivate] @TEST_EDGE: duplicate_apply -> deactivated=0, already_deleted=1 @TEST_FIXTURE: deactivate_already_deleted -> fixtures/api/deactivate_already_deleted.json

@} Fixture FX_Rls.UserAudit.Deactivate.AlreadyDeleted

@{ Fixture FX_Rls.UserAudit.Deactivate.DwhFailure [C:2] [TYPE Block] [SEMANTICS test,rls,fixture]

@BRIEF Partial DWH failure per user. @RELATION VERIFIES -> [Rls.UserAuditService.Deactivate] @TEST_EDGE: external_fail -> per-user error, applied persists @TEST_FIXTURE: deactivate_dwh_failure -> fixtures/api/deactivate_dwh_failure.json

@} Fixture FX_Rls.UserAudit.Deactivate.DwhFailure

@{ Fixture FX_Rls.UserAudit.Deactivate.EmptySelection [C:2] [TYPE Block] [SEMANTICS test,rls,fixture]

@BRIEF Empty selection → 422. @RELATION VERIFIES -> [Rls.UserAuditService.Deactivate] @TEST_EDGE: missing_field -> 422 VALIDATION_ERROR @TEST_FIXTURE: deactivate_empty_selection -> fixtures/api/deactivate_empty_selection.json

@} Fixture FX_Rls.UserAudit.Deactivate.EmptySelection

@{ Fixture FX_Rls.ScriptStore.Push.Valid [C:2] [TYPE Block] [SEMANTICS test,rls,fixture]

@BRIEF Push to corp remote succeeds, pushed_at recorded. @RELATION VERIFIES -> [Rls.ScriptStore.PushToCorp] @TEST_FIXTURE: push_valid -> fixtures/api/push_valid.json

@} Fixture FX_Rls.ScriptStore.Push.Valid

@{ Fixture FX_Rls.ScriptStore.Push.CorpUnreachable [C:2] [TYPE Block] [SEMANTICS test,rls,fixture]

@BRIEF Corp unreachable → error envelope, local version intact (US1-7). @RELATION VERIFIES -> [Rls.ScriptStore.PushToCorp] @TEST_EDGE: external_fail -> CORP_UNREACHABLE + local_version_intact @TEST_FIXTURE: push_corp_unreachable -> fixtures/api/push_corp_unreachable.json

@} Fixture FX_Rls.ScriptStore.Push.CorpUnreachable

@{ Fixture FX_Rls.ScriptStore.Push.NoCorpRemote [C:2] [TYPE Block] [SEMANTICS test,rls,fixture]

@BRIEF Missing corp remote config → config error with hint. @RELATION VERIFIES -> [Rls.ScriptStore.PushToCorp] @TEST_EDGE: missing_field -> CORP_REMOTE_MISSING @TEST_FIXTURE: push_no_corp_remote -> fixtures/api/push_no_corp_remote.json

@} Fixture FX_Rls.ScriptStore.Push.NoCorpRemote

@{ Fixture FX_Rls.ScriptStore.Push.NoCredentials [C:2] [TYPE Block] [SEMANTICS test,rls,fixture]

@BRIEF PAT missing/undecryptable → credentials error (rejected path). @RELATION VERIFIES -> [Rls.ScriptStore.PushToCorp] @TEST_EDGE: rejected_path -> CREDENTIALS_MISSING @TEST_FIXTURE: push_no_credentials -> fixtures/api/push_no_credentials.json

@} Fixture FX_Rls.ScriptStore.Push.NoCredentials

@{ Fixture FX_Rls.Snapshot.Refresh.Valid [C:2] [TYPE Block] [SEMANTICS test,rls,fixture]

@BRIEF Full snapshot refresh success. @RELATION VERIFIES -> [Rls.SnapshotService.Refresh] @TEST_FIXTURE: snapshot_valid -> fixtures/api/snapshot_valid.json

@} Fixture FX_Rls.Snapshot.Refresh.Valid

@{ Fixture FX_Rls.Snapshot.Refresh.DwhUnreachable [C:2] [TYPE Block] [SEMANTICS test,rls,fixture]

@BRIEF DWH down → meta failed, no partial rows. @RELATION VERIFIES -> [Rls.SnapshotService.Refresh] @TEST_EDGE: external_fail -> failed + partial_rows=false @TEST_FIXTURE: snapshot_dwh_unreachable -> fixtures/api/snapshot_dwh_unreachable.json

@} Fixture FX_Rls.Snapshot.Refresh.DwhUnreachable

@{ Fixture FX_Rls.Snapshot.Refresh.EmptySource [C:2] [TYPE Block] [SEMANTICS test,rls,fixture]

@BRIEF Empty rls_t source → success row_count=0, empty state. @RELATION VERIFIES -> [Rls.SnapshotService.Refresh] @TEST_EDGE: empty -> success + empty_state=true @TEST_FIXTURE: snapshot_empty_source -> fixtures/api/snapshot_empty_source.json

@} Fixture FX_Rls.Snapshot.Refresh.EmptySource

@{ Fixture FX_Rls.Snapshot.Refresh.PendingConflict [C:2] [TYPE Block] [SEMANTICS test,rls,fixture]

@BRIEF Concurrent refresh → 409 (rejected path). @RELATION VERIFIES -> [Rls.SnapshotService.Refresh] @TEST_EDGE: rejected_path -> 409 SNAPSHOT_IN_PROGRESS @TEST_FIXTURE: snapshot_pending_conflict -> fixtures/api/snapshot_pending_conflict.json

@} Fixture FX_Rls.Snapshot.Refresh.PendingConflict

@{ Fixture FX_Rls.Bindings.AddManual.Valid [C:2] [TYPE Block] [SEMANTICS test,rls,fixture]

@BRIEF Manual bindings validated via IDM. @RELATION VERIFIES -> [Rls.BindingsService.AddManual] @TEST_FIXTURE: binding_manual_valid -> fixtures/api/binding_manual_valid.json

@} Fixture FX_Rls.Bindings.AddManual.Valid

@{ Fixture FX_Rls.Bindings.AddManual.LoginNotInIdm [C:2] [TYPE Block] [SEMANTICS test,rls,fixture]

@BRIEF Unknown login rejected with «нет в IDM». @RELATION VERIFIES -> [Rls.BindingsService.AddManual] @TEST_EDGE: missing_field -> rejected not_found_in_idm (distinct from «нет в RLS») @TEST_FIXTURE: binding_login_not_in_idm -> fixtures/api/binding_login_not_in_idm.json

@} Fixture FX_Rls.Bindings.AddManual.LoginNotInIdm

@{ Fixture FX_Rls.Bindings.AddManual.FiredEmployee [C:2] [TYPE Block] [SEMANTICS test,rls,fixture]

@BRIEF Fired employee bindable with warning (operator decision). @RELATION VERIFIES -> [Rls.BindingsService.AddManual] @TEST_EDGE: edge_case -> bound + warning @TEST_FIXTURE: binding_fired_employee -> fixtures/api/binding_fired_employee.json

@} Fixture FX_Rls.Bindings.AddManual.FiredEmployee

#endregion Rls.Fixtures

================================================================================ FEATURE: 049-idm-account-integration Files: 5


SPEC — Feature Specification

Source: spec.md

#region Std.Specify.FeatureSpec [C:3] [TYPE ADR] [SEMANTICS spec,requirements,feature,idm,rls,account] @BRIEF Feature specification — WHAT the user needs and WHY. Implementation-free. Survives HCA 128× via @SEMANTICS grouping.

Navigation (DSA Indexer keywords)

@SEMANTICS: spec, requirements, feature, idm, rls, account, orgstructure

Feature Branch: (без ветки — пакет в specs/049-idm-account-integration/) Created: 2026-08-04 | Status: Draft Input: "Глубокая IDM-интеграция для RLS-инструмента: кэш профилей и статусов сотрудников (масштабируемое обогащение bi_users), канонизация основного AD-аккаунта при множественных аккаунтах, живые оргструктурные привязки (пересчёт состава отделов с авто-исключением уволенных/переведённых), поиск/автокомплит и расширение мока IDM"

Clarifications

Session 2026-08-04 (аудит пробелов 048, пред-спека)

  • Q: Что осталось непокрытым в базовой IDM-интеграции 048? → A: Пять пробелов: (1) Enrich делает N живых запросов на bi_users — нужен кэш; (2) оргструктурные привязки фиксируют срез — переводы не каскадятся; (3) правило «основного» аккаунта при нескольких AD-аккаунтах не определено; (4) автокомплит требует поиска по подстроке; (5) реальный API FIM (search/batch/группы) не подтверждён.
  • Q: Кэш или live при обогащении статусов? → A: Кэш: статусы меняются редко (уволен/отпуск/enabled), фоновая заливка по расписанию + live только для карточки.
  • Q: Может ли скрипт rls_t ходить в IDM? → A: Нет — IDM это сервис, не таблица DWH; пересчёт состава отделов — задача инструмента (cron), результат пишется в DWH-таблицу привязок.
  • Q: Первичный аккаунт при нескольких AD-аккаунтах? → A: Правило предлагается: активный (enabled) AD-аккаунт, иначе первый; требуется подтверждение владельцами IDM — [NEEDS CLARIFICATION: primary-account-rule].
  • Q: Реальные эндпоинты FIM для поиска/batch/групп? → A: Не подтверждены — [NEEDS CLARIFICATION: idm-api-surface].

User Scenarios

Каждая story — независимо тестируемая единица. Приоритеты P1 → P2. Доменные ключевые слова: idm, rls, account, orgstructure.

Story 1 — Кэш IDM-профилей и статусов (P1)

Why P1: Без кэша обогащение bi_users не масштабируется (N живых запросов), а это базис аудита 048.

Independent Test: Открыть аудит пользователей → статусы загружены из кэша (быстро, без N запросов); фоновая задача заливает кэш по расписанию; карточка сотрудника — live-запрос.

Acceptance:

  1. Given в системе есть bi_users любого объёма (до 50k), When открывается аудит пользователей, Then статусы IDM читаются из кэша, время загрузки не зависит от числа пользователей
  2. Given кэш пуст или устарел (TTL истёк), When открывается аудит, Then показывается «обогащение выполняется» и фоновая задача наполняет кэш
  3. Given фоновая задача обогащения, When она завершает проход, Then кэш содержит статусы (employee_status/account_status/risk/vip/admin) для 100% найденных в IDM пользователей, пропущенные зафиксированы
  4. Given пользователь открывает карточку сотрудника, When карточка загружена, Then данные получены live-запросом из IDM (актуальность), кэш при этом обновляется
  5. Given IDM недоступен, When аудит открыт, Then кэш продолжает обслуживать чтение, показывается возраст кэша и флаг деградации
  6. Given пользователь уволен в IDM (dateOut), When кэш обновляется, Then статус в кэше меняется на fired, и аудит 048 видит кандидата на деактивацию

Story 2 — Канонизация основного аккаунта (P1)

Why P1: Без единого правила сопоставления логинов (TEST.LOCAL vs IE.CORP) аудит и привязки могут работать с неверным аккаунтом.

Independent Test: У пользователя с двумя AD-аккаунтами карточка показывает primary-аккаунт, все аккаунты с доменами; сопоставление с rls_t идёт по каноническому логину.

Acceptance:

  1. Given пользователь имеет несколько AD-аккаунтов в разных доменах, When открывается карточка, Then показан primary-аккаунт (правило: активный enabled, иначе первый) и полный список аккаунтов с доменами и статусами
  2. Given у пользователя один аккаунт, When карточка загружена, Then он автоматически является primary
  3. Given primary-аккаунт выбран, When система сопоставляет с rls_t / привязками, Then используется канонический логин (без домена, lowercase) — единый для всех операций
  4. Given primary-аккаунт изменяется (старый отключён), When кэш обновляется, Then канонический логин пересчитывается, расхождения фиксируются в отчёте
  5. Given логин не найден ни в одном аккаунте, When идёт сопоставление, Then явное «нет в IDM» (не путать с «нет в RLS»)

Story 3 — Живые оргструктурные привязки (P1)

Why P1: Привязка «весь отдел → роль» замораживает срез; переводы и увольнения создают утечки/недодачи доступа.

Independent Test: Привязка по отделу → cron-пересчёт состава → diff: переведённые/уволенные деактивированы, новые добавлены; отчёт о пересчёте виден в UI.

Acceptance:

  1. Given существует оргструктурная привязка (отдел/БЕ → роль), When выполняется пересчёт по расписанию, Then состав отдела запрашивается из IDM (SearchByDepartment) и срез в bi_users синхронизируется: ушедшие (переведённые/уволенные) деактивируются (is_deleted=true), новые добавляются
  2. Given пересчёт завершён с изменениями, When открывается отчёт, Then показан diff: добавлено N, деактивировано M, с причинами по каждому (переведён/уволен/не найден)
  3. Given уволенный сотрудник (dateOut) состоит в орг-привязке, When пересчёт выполнен, Then он автоматически исключается (без ручной деактивации)
  4. Given IDM недоступен во время пересчёта, When задача завершается, Then пересчёт помечается failed, привязки не меняются, повтор при следующем расписании
  5. Given привязка создана вручную (по логинам), When пересчёт орг-привязок идёт, Then ручные привязки не затрагиваются (только org-тип)
  6. Given пересчёт изменяет привязки, When изменения применяются, Then генерируется SQL-миграция в версионируемый контур (история + согласование)

Story 4 — Поиск и расширенный API IDM (P2)

Why P2: Улучшает UX привязок (автокомплит) и требует подтверждения реального API; базовые кейсы 048 уже работают с точным логином.

Independent Test: Автокомплит по подстроке находит пользователей; batch-статусы заливают кэш пачками; поиск по AD-группе возвращает аккаунты; мок расширен.

Acceptance:

  1. Given пользователь вводит 3+ символа в поиске привязки, When запрос отправлен, Then автокомплит показывает до 20 кандидатов (логин, ФИО, отдел), выбранный кандидат валидируется
  2. Given batch-эндпоинт доступен, When фоновая заливка кэша запущена, Then статусы запрашиваются пачками (100/запрос), а не по одному
  3. Given поиск по AD-группе/OU, When группа найдена, Then возвращаются все аккаунты группы с OU и enabled
  4. Given мок IDM, When фича включена, Then мок расширен эндпоинтами: SearchPersons (подстрока), GetPersonsBatch, SearchByOu, SearchByDepartment
  5. Given реальный API FIM не поддерживает какой-то эндпоинт, When клиент вызывает его, Then деградация с явным сообщением «эндпоинт не поддерживается контуром» [NEEDS CLARIFICATION: idm-api-surface]

Edge Cases

  • bi_users > 50k → кэш заливается порциями (batch), аудит читает из кэша, обогащение не блокирует UI
  • Пользователь без AD-аккаунтов (только сервисные записи) → «аккаунтов нет», primary отсутствует, сопоставление невозможно — помечается в отчёте
  • Два сотрудника с одинаковым логином в разных доменах → канонический логин коллизия: фиксируется в отчёте, primary по правилу
  • Перевод между отделами в пределах одного пересчёта → привязка старого отдела деактивируется, нового — добавляется (если есть привязка)
  • Уволенный с датой увольнения в будущем (announced) → не исключается до dateOut
  • Отпуск (leaveType) → статус «в отпуске», не исключается (доступ сохраняется)
  • IDM возвращает мусор в accountId (пробелы, как в моке) → ID трактуются как непрозрачные строки
  • Кэш устарел более чем на N дней (TTL) и IDM недоступен → статусы «недоступно», возраст кэша показан
  • Пересчёт орг-привязок при 0 найденных в отделе → предупреждение, привязка не трогается (защита от массовой деактивации по ошибке поиска)
  • Спецсимволы в поисковой строке (% _ *) → экранирование, поиск безопасен

Requirements

Functional (IDs survive HCA 128× via hierarchical naming)

  • IDM-FR-001: Кэш-таблица idm_profile_cache (login, employee_status, account_status, risk_score, vip, admin_rights, primary_login, accounts JSON, cached_at, source) с TTL
  • IDM-FR-002: Фоновая задача обогащения кэша по расписанию (batch-запросы, порции), покрытие 100% найденных в IDM пользователей bi_users
  • IDM-FR-003: Аудит пользователей 048 читает статусы из кэша; время загрузки не зависит от числа пользователей
  • IDM-FR-004: Карточка сотрудника — live-запрос в IDM с обновлением кэша
  • IDM-FR-005: Деградация: кэш обслуживает чтение при недоступности IDM, возраст кэша отображается
  • IDM-FR-006: Правило primary-аккаунта: активный (enabled) AD-аккаунт, иначе первый; единственный аккаунт — primary автоматически
  • IDM-FR-007: Канонический логин (без домена, lowercase) — единый идентификатор для сопоставления с rls_t и привязками
  • IDM-FR-008: Коллизии канонических логинов фиксируются в отчёте
  • IDM-FR-009: Пересчёт оргструктурных привязок по расписанию: SearchByDepartment → синхронизация bi_users (деактивация ушедших is_deleted=true, INSERT новых)
  • IDM-FR-010: Авто-исключение уволенных (dateOut ≤ сегодня) из орг-привязок при пересчёте
  • IDM-FR-011: Отчёт пересчёта: diff (добавлено/деактивировано) с причинами (переведён/уволен/не найден)
  • IDM-FR-012: Защита от массовой деактивации: 0 найденных в отделе → предупреждение, привязка не изменяется
  • IDM-FR-013: Ручные привязки (по логинам) не затрагиваются пересчётом
  • IDM-FR-014: Пересчёт генерирует SQL-миграцию в версионируемый контур
  • IDM-FR-015: Автокомплит поиска по подстроке (3+ символа, до 20 кандидатов: логин/ФИО/отдел)
  • IDM-FR-016: Batch-запрос статусов (100/запрос) для фоновой заливки
  • IDM-FR-017: Расширение мока IDM: SearchPersons, GetPersonsBatch, SearchByOu, SearchByDepartment
  • IDM-FR-018: Неподдерживаемый реальным API эндпоинт → деградация «не поддерживается контуром» [NEEDS CLARIFICATION: idm-api-surface]
  • IDM-FR-019: Отпуск (leaveType) не исключает из привязок; будущая дата увольнения не исключает до dateOut
  • IDM-FR-020: RBAC: чтение персональных данных (риск/VIP) — idm READ_PERSONAL; управление кэшем/пересчётом — rls WRITE

Маркеры: [NEEDS CLARIFICATION: primary-account-rule], [NEEDS CLARIFICATION: idm-api-surface] — 2 маркера.

Key Entities

  • IdmProfileCache: Кэш статусов и профилей: login, employee_status (active/fired/vacation/unknown), account_status (enabled/disabled), risk_score, vip, admin_rights, primary_login, accounts (массив с доменами/OU/enabled), cached_at, ttl
  • IdmAccountCanon: Канонический логин пользователя: primary_login, домен, правило выбора, коллизии
  • OrgBindingRefresh: Прогон пересчёта орг-привязок: run_id, binding_id, department_ref, added[], deactivated[] с причинами, status, error, started/finished
  • IdmSearchQuery: Поиск по подстроке: query, limit, candidates (login, full_name, department, uid)
  • IdmBatchResult: Пачка статусов: logins[], statuses[], failures[]
  • IdmMockEndpoint: Расширение мока: SearchPersons, GetPersonsBatch, SearchByOu, SearchByDepartment

Success Criteria

  • SC-001: Аудит bi_users до 50k пользователей загружает статусы из кэша ≤ 3 секунд (без N живых запросов)
  • SC-002: Фоновая заливка кэша покрывает 100% найденных в IDM пользователей bi_users за один проход (≤ 15 минут на 50k)
  • SC-003: Пересчёт орг-привязок выявляет и деактивирует 100% уволенных/переведённых при доступном IDM
  • SC-004: Каждый пользователь имеет ровно один канонический логин; коллизии — 0 незафиксированных
  • SC-005: Автокомплит отвечает ≤ 300 мс на запрос (мок и реальный контур)

#endregion Std.Specify.FeatureSpec


UX REFERENCE — Interaction Narrative

Source: ux_reference.md

#region Std.Specify.UxReference [C:3] [TYPE ADR] [SEMANTICS ux,reference,idm,account,orgstructure] @BRIEF UX interaction reference — persona, flows, states, recovery paths for 049-idm-account-integration.

Feature: 049-idm-account-integration Created: 2026-08-04 | Status: Draft

1. User Persona & Context

  • Who is the user?: Администратор RLS (rls_operator) — ведёт привязки пользователей к ролям, аудирует bi_users; администратор IDM/DWH (rls_script_dev) — настраивает расписания кэша и пересчёта. Вторичная аудитория: специалист по безопасности (просмотр отчётов пересчёта и персональных данных).
  • What is their goal?: Масштабируемое обогащение статусами IDM без N запросов; уверенность, что оргструктурные привязки не устаревают (переводы/увольнения каскадятся); единый канонический логин у каждого пользователя.
  • Context: Браузер, корпоративный контур. Фича надстраивается над разделом RLS из 048 (аудит пользователей, конструктор правил). IDM: мок на dev, FIM в prod.

2. The "Happy Path" Narrative

Оператор открывает аудит пользователей: таблица 12k строк загружается мгновенно — статусы из кэша (бейджи «активен/уволен/в отпуске»), в шапке индикатор «кэш IDM: обновлён сегодня 06:00». Он открывает карточку сотрудника — данные подтянуты live, primary-аккаунт подсвечен, все три аккаунта с доменами видны. Затем оператор заходит в «Оргструктурные привязки»: вчерашний пересчёт показал «+3, −2 (переведён: sidorovss, уволен: kozlovakk)» — доступы отдела синхронизированы автоматически. При создании новой привязки автокомплит находит отдел за 200 мс.

3. Interface Mockups

UI Layout & Flow

Экран 1: Аудит пользователей (расширение экрана 048 US3)

  • Layout: Таблица bi_users (как в 048) + индикатор кэша в шапке: «Кэш IDM: обновлён 04.08 06:00 · 12 483 пользователя · 12 не найдено». Колонки статусов читаются из кэша. Карточка сотрудника раскрывается как раньше, но с блоком «Учётные записи»: primary-бейдж + список всех аккаунтов (домен, OU, enabled) + канонический логин.
  • Key Elements:
    • Индикатор кэша: бейдж (success — свежий, warning — устарел ≤TTL, destructive — обогащение упало) + кнопка «Обновить кэш» (ручной запуск фоновой заливки).
    • Блок «Учётные записи»: primary-аккаунт с бейджем primary, остальные с доменами; переключение primary недоступно в v1 (правило автоматическое).
    • Строка «не найдено в IDM»: бейдж muted + tooltip «проверьте домен или наличие аккаунтов».
  • Contract Mapping:
    • @UX_STATE: cache-fresh / cache-stale / cache-failed / enriching
    • @UX_STATE карточка: live-loaded / live-degraded
    • @UX_FEEDBACK: toast «Кэш обновлён: +12 483 статуса», бейдж возраста кэша
    • @UX_RECOVERY: cache-failed → «Повторить обогащение»; live-degraded → данные из кэша + бейдж «IDM недоступен»
    • @UX_REACTIVITY: модель IdmUserAuditModel (расширение RlsUserAuditModel из 048): $state cacheMeta/rows, $derived filteredRows
  • States:
    • Idle/Default: таблица + свежий кэш.
    • Loading: скелетон таблицы.
    • Enriching: баннер «Обогащение выполняется в фоне (порция 3/50)».
    • Error/Degraded: кэш устарел + IDM недоступен → статусы «—», возраст показан.

Экран 2: Оргструктурные привязки (новый раздел)

  • Layout: Список орг-привязок (отдел/БЕ → роль, N пользователей, дата последнего пересчёта, статус). Правая панель: отчёт последнего пересчёта — diff-таблица (пользователь, действие добавлено/деактивировано, причина: переведён/уволен/не найден/новый), кнопки «Пересчитать сейчас» (ручной запуск), «Настроить расписание» (cron).
  • Key Elements:
    • Карточка привязки: название отдела, роль, счётчик состава, бейдж статуса пересчёта (ok / pending / failed / warning-0-found).
    • Diff-таблица отчёта: строки с причинами; фильтр по причине.
    • Кнопка «Пересчитать сейчас»: запускает синхронизацию; при 0 найденных — подтверждающий модал «В отделе не найдено ни одного сотрудника. Привязка не будет изменена» (защита от массовой деактивации).
    • Настройки расписания: cron-выражение + «время последнего успешного прогона».
  • Contract Mapping:
    • @UX_STATE: list / refreshing / refresh-diff / refresh-failed / zero-found-warning
    • @UX_FEEDBACK: toast «Пересчёт: +3, −2», модал подтверждения при 0 найденных
    • @UX_RECOVERY: refresh-failed → «Повторить» (привязки не изменены)
    • @UX_REACTIVITY: модель OrgBindingModel — $state bindings/runs, $derived diffSummary
  • States:
    • Idle/Default: список привязок, последние пересчёты ok.
    • Refreshing: спиннер на карточке + прогресс порций.
    • Success: diff-таблица с причинами.
    • Error/Degraded: failed-бейдж, привязки не тронуты.

Экран 3: Поиск и привязка пользователей (расширение экрана 048 US4)

  • Layout: В конструкторе правил вкладка «По AD-логину» получает автокомплит: ввод 3+ символов → выпадающий список кандидатов (логин, ФИО, отдел, домен) → выбор валидируется. Вкладка «AD-группа / OU» — поиск группы с показом найденных аккаунтов и OU. Вкладка «Оргструктура» — дерево отделов с счётчиками состава (из кэша).
  • Key Elements:
    • Автокомплит: инлайн-список до 20 кандидатов, клавиатурная навигация (стрелки + Enter), debounce 200 мс.
    • Индикатор поиска: «ищем в IDM…» при запросе; «не найдено» — пустое состояние с подсказкой.
    • Вкладка Оргструктура: дерево отделов с количеством сотрудников (из кэша/IDM), чекбоксы мультивыбора.
  • Contract Mapping:
    • @UX_STATE: typing / searching / candidates / no-results / selected
    • @UX_FEEDBACK: подсветка primary-логина в кандидатах, toast при выборе уволенного (warning)
    • @UX_RECOVERY: no-results → «проверьте написание или домен»; поиск-ошибка → «Повторить»
    • @UX_REACTIVITY: модель IdmSearchModel — $state query/candidates/selected, $derived filtered
  • States:
    • Idle/Default: пустое поле, подсказка «минимум 3 символа».
    • Searching: спиннер в поле.
    • Candidates: список с primary-логином и ФИО.
    • No-results: пустое состояние с подсказкой.

4. The "Error" Experience

Philosophy: Кэш и пересчёт — фоновые процессы: UI никогда не блокируется; ошибки показываются с возрастом данных и путём восстановления.

Edge & Failure State Matrix Reference

State Class Trigger Applicable? Visual/Feedback Recovery
NET_01 — Offline Сеть недоступна Да Оффлайн-баннер Автоповтор при reconnect
NET_02 — Timeout IDM >10s Да (live-карточка) Toast + «из кэша» Retry (3 попытки)
NET_03 — Retry exhausted 3 неудачи live-карточки Да Данные из кэша + бейдж Ручной retry
STALE — Кэш устарел TTL истёк Да Бейдж «кэш от 02.08» Фоновая заливка / кнопка
PARTIAL — Часть не найдена 12 из 12k не в IDM Да Строки «не найдено» Перепроверка по домену
VAL_01 — Поиск <3 символов Слишком короткий запрос Да Подсказка «минимум 3» Продолжить ввод
429 — Rate limited Лимит IDM Да (заливка кэша) Портится: «пауза 60с» Автопродолжение
5XX — Server error Backend/IDM Да Баннер + retry Retry
CONF_01 — Два пересчёта Ручной + cron одновременно Да «Пересчёт уже идёт» Ожидание завершения
DUP_01 — Двойной клик «Пересчитать сейчас» Да Кнопка disabled Нормальное завершение
EMPTY — 0 найдено в отделе Ошибка поиска/пустой отдел Да Модал-предупреждение, привязка не меняется Подтвердить/отменить
LARGE — 50k пользователей Большой bi_users Да Порции + прогресс Без блокировки UI
MALFORMED — Мусор в ID accountId с пробелами Да ID как строки —
A11Y — Screen reader Смена состояния Да aria-live Встроено
RESP — Responsive Viewport <768px Да Stacked layout Встроено

Scenario A: Кэш устарел, IDM недоступен

  • System Response: Таблица показывает статусы из кэша (возраст «от 02.08»), бейдж destructive «IDM недоступен, кэш от 02.08». Фоновая задача помечена failed.
  • Recovery: Кнопка «Повторить обогащение» → задача в очереди; при восстановлении IDM заливка продолжается с порции, на которой упала.

Scenario B: Пересчёт нашёл 0 сотрудников в отделе

  • System Response: Модал: «В отделе «Департамент ИТ» не найдено сотрудников. Привязка НЕ будет изменена. Продолжить?» (защита от массовой деактивации).
  • Recovery: «Отмена» — привязка не тронута; «Продолжить» — пересчёт без изменений (0/0), отчёт фиксирует предупреждение.

5. Tone & Voice

  • Style: Технический, спокойный, на русском.
  • Terminology: «primary-аккаунт» (не «основной логин»), «канонический логин», «оргструктурная привязка», «пересчёт» (не «синхронизация»), «кэш IDM». Сохранять термины 048: bi_users, is_deleted, rls_operator.

#endregion Std.Specify.UxReference


CHECKLISTS — Requirements Quality — requirements.md

Source: checklists/requirements.md

[REQUIREMENTS] Checklist: 049-idm-account-integration

Purpose: Проверка полноты требований фичи IDM Account Integration перед планированием — каждое требование из spec.md покрыто проверяемым пунктом. Created: 2026-08-04 Feature: spec.md

Cache & Scaling (US1)

  • CHK001 Кэш idm_profile_cache: login, employee_status, account_status, risk_score, vip, admin_rights, primary_login, accounts JSON, cached_at, TTL ([IDM-FR-001])
  • CHK002 Фоновая заливка кэша по расписанию, порции/batch, покрытие 100% найденных в IDM пользователей bi_users ([IDM-FR-002])
  • CHK003 Аудит 048 читает статусы из кэша — время загрузки не зависит от числа пользователей ([IDM-FR-003, SC-001])
  • CHK004 Карточка сотрудника — live-запрос + обновление кэша ([IDM-FR-004])
  • CHK005 Деградация: кэш обслуживает чтение при недоступности IDM, возраст показан ([IDM-FR-005])
  • CHK006 Полный проход заливки ≤15 мин на 50k ([SC-002])
  • CHK007 Повтор неудачной заливки с порции сбоя (resume) (Scenario A)
  • CHK008 Rate-limit IDM: пауза и автопродолжение ([429])

Primary Account Canon (US2)

  • CHK009 Правило primary: активный enabled AD-аккаунт, иначе первый; единственный — primary автоматически ([IDM-FR-006])
  • CHK010 Канонический логин (без домена, lowercase) — единый идентификатор для rls_t и привязок ([IDM-FR-007])
  • CHK011 Карточка показывает primary-бейдж + все аккаунты с доменами/OU/enabled (acceptance US2-1)
  • CHK012 Пересчёт primary при отключении старого, расхождения в отчёте (acceptance US2-4)
  • CHK013 Коллизии канонических логинов фиксируются ([IDM-FR-008])
  • CHK014 «Нет в IDM» не путается с «нет в RLS» (acceptance US2-5)
  • CHK015 Ровно один канонический логин на пользователя, 0 незафиксированных коллизий ([SC-004])

Live Org Bindings (US3)

  • CHK016 Пересчёт орг-привязок по расписанию: SearchByDepartment → деактивация ушедших + INSERT новых ([IDM-FR-009])
  • CHK017 Авто-исключение уволенных (dateOut ≤ сегодня) ([IDM-FR-010, SC-003])
  • CHK018 Отчёт пересчёта: diff с причинами (переведён/уволен/не найден/новый) ([IDM-FR-011])
  • CHK019 Защита от массовой деактивации: 0 найденных → предупреждение, привязка не меняется ([IDM-FR-012, Scenario B])
  • CHK020 Ручные привязки не затрагиваются ([IDM-FR-013])
  • CHK021 SQL-миграция пересчёта в версионируемый контур ([IDM-FR-014])
  • CHK022 Неудачный пересчёт: привязки не меняются, повтор по расписанию ([IDM-FR-009 edge])
  • CHK023 Конфликт ручного и cron пересчёта: «уже идёт» (CONF_01)
  • CHK024 Отпуск не исключает; будущий dateOut не исключает до даты ([IDM-FR-019])

Search & Extended API (US4)

  • CHK025 Автокомплит 3+ символа, до 20 кандидатов (логин/ФИО/отдел/домен), debounce, ≤300 мс ([IDM-FR-015, SC-005])
  • CHK026 Batch-запрос статусов 100/запрос ([IDM-FR-016])
  • CHK027 Расширение мока: SearchPersons, GetPersonsBatch, SearchByOu, SearchByDepartment ([IDM-FR-017])
  • CHK028 Неподдерживаемый эндпоинт реального API → деградация «не поддерживается контуром» ([IDM-FR-018, NEEDS CLARIFICATION: idm-api-surface])
  • CHK029 Экранирование спецсимволов поиска (%, _, *) (edge case)

RBAC & Security

  • CHK030 Персональные данные (риск/VIP) — idm READ_PERSONAL ([IDM-FR-020])
  • CHK031 Управление кэшем/пересчётом — rls WRITE ([IDM-FR-020])

Edge Cases

  • CHK032 bi_users >50k: порции, обогащение не блокирует UI (LARGE)
  • CHK033 Пользователь без AD-аккаунтов: «аккаунтов нет», отчёт ([edge case])
  • CHK034 Одинаковый логин в разных доменах: коллизия в отчёте ([edge case])
  • CHK035 Мусор в accountId — непрозрачные строки (MALFORMED)

Notes

  • Check items off as completed: [x]
  • Два маркера: [NEEDS CLARIFICATION: primary-account-rule], [NEEDS CLARIFICATION: idm-api-surface] — закрываются в /speckit.clarify (требуют владельцев IDM).
  • Фича надстраивается над 048: кэш подменяет live-обогащение в Rls.UserAuditService.Enrich, орг-привязки используют Rls.BindingsService.AddOrgUnit из 048.
  • Скрипт rls_t в IDM не ходит — пересчёт состава отделов выполняет инструмент (cron), результат — в DWH-таблице привязок.

PROTOTYPE — State/Manifest

Source: prototype/manifest.md

#region Std.Opencode.PrototypeManifest [C:3] [TYPE ADR] [SEMANTICS prototype,manifest,idm,rls] @defgroup Prototype Interactive HTML prototype manifest for IDM Account Integration (049).

Prototype Metadata

  • Feature: IDM Account Integration (049-idm-account-integration)
  • Source contracts: ux_reference.md (contracts/ux/ не создавался — лёгкий прототип из reference)
  • Screens represented: 3
  • Total states: 15 (14 контрактных + интерактивные переходы)
  • Accessibility validations: keyboard nav (Tab/Enter/Escape), ARIA (tablist/tab/dialog/option/listbox/status), focus rings, touch targets
  • Responsive breakpoints: 375px (mobile), 1280px (desktop)
  • Prototype path: specs/049-idm-account-integration/prototype/index.html

State Coverage

Screen @UX_STATE Contract (ux_reference) Prototype State Reachable? Recovery Path
1 · Аудит + кэш cache-fresh cache-fresh (бейдж success, таблица из кэша) ✅ —
1 · Аудит + кэш cache-stale cache-stale (warning banner + «Повторить обогащение») ✅ retry → toast
1 · Аудит + кэш cache-failed cache-failed (destructive banner «IDM недоступен, кэш от 02.08») ✅ «Повторить» → cache-fresh
1 · Аудит + кэш enriching enriching (прогресс порции 3/50) ✅ авто → cache-fresh
1 · Аудит + кэш loading loading (скелетон) ✅ → loaded
1 · Аудит + кэш live-loaded (карточка) карточка с primary-блоком и 3 аккаунтами ✅ —
2 · Оргпривязки list list (карточки привязок + diff-отчёт с причинами) ✅ —
2 · Оргпривязки refreshing refreshing (прогресс порций) ✅ → list + toast diff
2 · Оргпривязки zero-found-warning zero-found (баннер + модал «0 сотрудников, привязка не меняется») ✅ Отмена/Продолжить 0/0
2 · Оргпривязки refresh-failed refresh-failed (destructive баннер, привязки не тронуты) ✅ «Повторить»
3 · Поиск typing typing (пустое поле + подсказка «минимум 3») ✅ —
3 · Поиск searching searching (спиннер в поле) ✅ → candidates
3 · Поиск candidates candidates (3 кандидата, primary-бейдж, клавиатура) ✅ выбор → selected
3 · Поиск no-results no-results («Ничего не найдено. Проверьте домен») ✅ изменить запрос
3 · Поиск selected selected (ivanovii добавлен + toast) ✅ —

Интерактивные переходы: «Обновить кэш» → enriching → cache-fresh + toast; retry-IDM → cache-fresh; «Пересчитать сейчас» → refreshing → list + toast с diff; zero-found → модал → подтверждение 0/0; ввод ≥3 символов в поиске → candidates; выбор кандидата → selected + toast; вкладки способов привязки; Escape закрывает модал.

Screen ↔ Story Traceability

Prototype Screen User Story UX Reference Section Acceptance Criteria Verified
1 · Аудит + кэш US1: Кэш профилей Экран 1 US1-1 (кэш, скорость), US1-2 (enriching), US1-4 (live-карточка), US1-5 (деградация), US1-6 (fired в кэше)
1 · Аудит + кэш (карточка) US2: Primary-аккаунт Экран 1 блок «Учётные записи» US2-1 (primary + все аккаунты с доменами), US2-5 (канонический логин)
2 · Оргпривязки US3: Живые привязки Экран 2 US3-1 (пересчёт → diff), US3-2 (отчёт с причинами), US3-3 (авто-исключение уволенных), US3-4 (IDM недоступен → failed), US3-5 (ручные не трогаются — в дизайне), US3-6 (SQL-миграция)
2 · Оргпривязки (модал) US3: Защита от массовой деактивации Scenario B US3 zero-found → предупреждение, привязка не меняется
3 · Поиск US4: Поиск и автокомплит Экран 3 US4-1 (3+ символа, 20 кандидатов), US4-3 (AD-группа с OU), US4-4 (оргструктура с составом)

Validation Results

  • All @UX_STATE contracts reachable via state switcher — 14/14 контрактных состояний
  • All @UX_RECOVERY paths traversable — retry-обогащение, retry-IDM, повтор пересчёта, модал 0/0, смена запроса
  • Keyboard navigation: Tab/Enter; Escape закрывает модал; стрелки в автокомплите (дизайн)
  • Touch targets ≥44×44px (кнопки h-10/h-12)
  • ARIA: role=tablist/tab (aria-selected), role=dialog (aria-modal), role=option/listbox в автокомплите, role=status/aria-live (toasts)
  • No dead-end states
  • Responsive: 375px — сайдбар скрыт, таблицы overflow-x-auto
  • Design fidelity: токен-аудит скриптом (см. ниже)

Примечание: браузерная валидация DevTools не выполнялась (окружение без браузера); статическая проверка выполнена скриптом — 0 неизвестных токенов, 0 классов без shim-определений.

Design System Reuse

Element Source Prototype Mapping
Button $lib/ui/Button.svelte .btn + variants/sizes — те же токены (bg #2563eb, hover #1d4ed8, ring #3b82f6)
Card $lib/ui/Card.svelte .card = rounded-lg border-border bg-surface-card shadow-sm + card-title-row/card-pad-md
Badge $lib/ui/Badge.svelte .badge + .badge-pill + badge-{variant} + dot-варианты
Table components/backups/BackupList.svelte min-w-full divide-y divide-border, thead bg-surface-muted, th px-6 py-3 uppercase
Input $lib/ui/Input.svelte .input h-10 rounded-md border-border-strong focus ring
Tabs $lib/ui/Tabs.svelte .tab-pills active/inactive (паттерн 048)
Modal $lib/ui/ConfirmDialog.svelte .modal-backdrop (bg-surface-overlay) + .modal-card
Skeleton $lib/ui/Skeleton.svelte animate-pulse + bg-surface-muted
Autocomplete Новый паттерн (нет в $lib) .autocomplete-list/.ac-item на токенах; кандидат на извлечение в $lib в plan-фазе при подтверждении
Nav sidebarNavigation.ts категория RLS с подразделами (раскрытие)

Design Token Audit (MANDATORY)

Все значения — из frontend/tailwind.config.js (тот же набор, что валидирован в 048):

Token Hex / Value Использование
primary.DEFAULT/hover/ring/light #2563eb/#1d4ed8/#3b82f6/#eff6ff кнопки, primary-бейджи, фокус-ринги, autocomplete hover
secondary.DEFAULT/hover/text #f3f4f6/#e5e7eb/#111827 secondary-кнопки
destructive.DEFAULT/hover/light #dc2626/#b91c1c/#fef2f2 баннеры, бейджи уволен, модал
success.DEFAULT/light #22c55e/#f0fdf4 бейджи активен/enabled
warning.DEFAULT/light #f59e0b/#fffbeb бейджи disabled/устарел, zero-found
info.DEFAULT/light #0ea5e9/#f0f9ff бейджи отпуск/live, enriching
surface.page/card/muted #f8fafc/#ffffff/#f1f5f9 фон, карточки, скелетоны
surface.overlay rgba(15,23,42,0.5) модальный backdrop
border.DEFAULT/strong #e2e8f0/#cbd5e1 границы, инпуты
text.DEFAULT/muted/subtle #0f172a/#64748b/#94a3b8 тексты
terminal.* #0f172a/#1e293b/#334155/#cbd5e1/#22d3ee chrome
category.admin #ffe4e6/#fecdd3/#be123c активный пункт навигации RLS
category.dashboards.to #bae6fd nav-dot «Дашборды»
category.datasets.to #bbf7d0 nav-dot «Датасеты»
category.migration.to #fed7aa nav-dot «Миграция»
category.reports.to #ddd6fe nav-dot «Отчёты»

Аудит-скрипт: прогнан — 49 уникальных hex, все из tailwind.config.js (49/49 = 100%); классов без shim-определений: 0.

#endregion Std.Opencode.PrototypeManifest


PROTOTYPE — Interactive HTML

Source: prototype/index.html

<html lang="ru"> <head> <style> /* ============ PROTOTYPE CHROME (не дизайн-система) ============ */ #prototype-chrome { position: fixed; top: 0; left: 0; right: 0; z-index: 1000; display: flex; flex-wrap: wrap; align-items: center; gap: 10px; padding: 8px 14px; background: #0f172a; color: #cbd5e1; font: 12px/1.4 system-ui, sans-serif; border-bottom: 1px solid #334155; } #prototype-chrome label { display: flex; align-items: center; gap: 5px; } #prototype-chrome select, #prototype-chrome button { background: #1e293b; color: #e2e8f0; border: 1px solid #334155; border-radius: 4px; padding: 4px 8px; font: inherit; cursor: pointer; } #prototype-chrome .chrome-badge { background: #22d3ee; color: #0f172a; font-weight: 700; border-radius: 4px; padding: 2px 8px; } #prototype-chrome .viewport-toggle { display: inline-flex; gap: 4px; } #prototype-chrome .viewport-toggle button.active { background: #38bdf8; color: #0f172a; } body { margin: 0; } #app-frame { margin-top: 46px; transition: max-width 0.2s ease; } body.chrome-viewport-mobile #app-frame { max-width: 375px; margin-left: auto; margin-right: auto; } body.chrome-viewport-mobile .proto-sidebar { display: none; } body.chrome-viewport-mobile .proto-main { padding-left: 0; } body.chrome-viewport-mobile .proto-topbar { display: none; }

/* ============ APP SHELL ============ */ .proto-app { display: flex; min-height: 100vh; background: #f8fafc; } .proto-sidebar { width: 240px; flex-shrink: 0; background: #ffffff; border-right: 1px solid #e2e8f0; padding: 20px 12px; } .proto-sidebar .proto-brand { display: flex; align-items: center; gap: 8px; padding: 0 8px 18px; font-size: 14px; font-weight: 700; color: #0f172a; } .proto-sidebar .proto-brand .logo { width: 28px; height: 28px; border-radius: 8px; flex-shrink: 0; background: linear-gradient(135deg, #0ea5e9, #06b6d4, #4f46e5); } .proto-nav-item { display: flex; align-items: center; gap: 10px; width: 100%; padding: 8px 10px; margin-bottom: 2px; border-radius: 6px; font-size: 13px; font-weight: 500; color: #64748b; background: transparent; border: none; text-align: left; cursor: pointer; } .proto-nav-item:hover { background: #f3f4f6; } .proto-nav-item.active { background: #ffe4e6; color: #be123c; } .proto-nav-item .nav-dot { width: 8px; height: 8px; border-radius: 9999px; flex-shrink: 0; } .proto-nav-item .nav-arrow { margin-left: auto; font-size: 10px; color: #94a3b8; } .proto-nav-sub { margin-left: 18px; } .proto-nav-sub .proto-nav-item { font-size: 12px; padding: 5px 10px; } .proto-main { flex: 1; min-width: 0; padding: 32px; } .proto-topbar { display: flex; align-items: center; gap: 12px; justify-content: flex-end; padding: 10px 32px; border-bottom: 1px solid #e2e8f0; background: #ffffff; }

/* ============ DESIGN SYSTEM SHIM (токены из tailwind.config.js) ============ */ .flex { display: flex; } .inline-flex { display: inline-flex; } .flex-col { flex-direction: column; } .flex-wrap { flex-wrap: wrap; } .flex-1 { flex: 1 1 0%; } .items-center { align-items: center; } .items-start { align-items: flex-start; } .justify-center { justify-content: center; } .justify-between { justify-content: space-between; } .gap-1 { gap: 4px; } .gap-1.5 { gap: 6px; } .gap-2 { gap: 8px; } .gap-3 { gap: 12px; } .gap-4 { gap: 16px; } .gap-6 { gap: 24px; } .space-y-1 > * + * { margin-top: 4px; } .space-y-2 > * + * { margin-top: 8px; } .space-y-3 > * + * { margin-top: 12px; } .space-y-4 > * + * { margin-top: 16px; } .grid { display: grid; } .grid-cols-2 { grid-template-columns: repeat(2, minmax(0, 1fr)); } .grid-cols-3 { grid-template-columns: repeat(3, minmax(0, 1fr)); } .mb-1 { margin-bottom: 4px; } .mb-2 { margin-bottom: 8px; } .mb-3 { margin-bottom: 12px; } .mb-4 { margin-bottom: 16px; } .mb-6 { margin-bottom: 24px; } .mb-8 { margin-bottom: 32px; } .mt-1 { margin-top: 4px; } .mt-2 { margin-top: 8px; } .mt-3 { margin-top: 12px; } .mt-4 { margin-top: 16px; } .mt-6 { margin-top: 24px; } .ml-1 { margin-left: 4px; } .ml-2 { margin-left: 8px; } .ml-auto { margin-left: auto; } .p-3 { padding: 12px; } .p-6 { padding: 24px; } .px-2 { padding-left: 8px; padding-right: 8px; } .px-2.5 { padding-left: 10px; padding-right: 10px; } .px-3 { padding-left: 12px; padding-right: 12px; } .px-4 { padding-left: 16px; padding-right: 16px; } .px-6 { padding-left: 24px; padding-right: 24px; } .py-1 { padding-top: 4px; padding-bottom: 4px; } .py-1.5 { padding-top: 6px; padding-bottom: 6px; } .py-2 { padding-top: 8px; padding-bottom: 8px; } .py-3 { padding-top: 12px; padding-bottom: 12px; } .py-4 { padding-top: 16px; padding-bottom: 16px; } .py-12 { padding-top: 48px; padding-bottom: 48px; } .h-2.5 { height: 10px; } .h-3 { height: 12px; } .h-4 { height: 16px; } .h-8 { height: 32px; } .h-10 { height: 40px; } .h-12 { height: 48px; } .w-2.5 { width: 10px; } .w-3 { width: 12px; } .w-4 { width: 16px; } .w-full { width: 100%; } .min-w-0 { min-width: 0; } .min-w-full { min-width: 100%; } .relative { position: relative; } .absolute { position: absolute; } .inset-0 { inset: 0; } .rounded { border-radius: 4px; } .rounded-md { border-radius: 6px; } .rounded-lg { border-radius: 8px; } .rounded-full { border-radius: 9999px; } .border { border-width: 1px; border-style: solid; } .border-b { border-bottom-width: 1px; border-bottom-style: solid; } .border-t { border-top-width: 1px; border-top-style: solid; } .border-border { border-color: #e2e8f0; } .border-border-strong { border-color: #cbd5e1; } .border-primary { border-color: #2563eb; } .border-destructive { border-color: #dc2626; } .border-success { border-color: #22c55e; } .border-warning { border-color: #f59e0b; } .divide-y > * + * { border-top-width: 1px; border-top-style: solid; border-color: #e2e8f0; } .shadow-sm { box-shadow: 0 1px 2px 0 rgba(0, 0, 0, 0.05); } .shadow-lg { box-shadow: 0 10px 15px -3px rgba(0, 0, 0, 0.1); } .overflow-hidden { overflow: hidden; } .overflow-x-auto { overflow-x: auto; } .whitespace-nowrap { white-space: nowrap; } .max-w-md { max-width: 28rem; } .text-left { text-align: left; } .text-right { text-align: right; } .text-center { text-align: center; } .uppercase { text-transform: uppercase; } .tracking-tight { letter-spacing: -0.025em; } .tracking-wider { letter-spacing: 0.05em; } .leading-none { line-height: 1; } .cursor-pointer { cursor: pointer; } .cursor-not-allowed { cursor: not-allowed; } .fixed { position: fixed; } .z-50 { z-index: 50; } .transition-colors { transition-property: color, background-color, border-color; transition-duration: 150ms; } [hidden] { display: none !important; }

.text-xs { font-size: 12px; line-height: 16px; } .text-sm { font-size: 14px; line-height: 20px; } .text-base { font-size: 16px; line-height: 24px; } .text-lg { font-size: 18px; line-height: 28px; } .text-3xl { font-size: 30px; line-height: 36px; } .font-medium { font-weight: 500; } .font-semibold { font-weight: 600; } .font-bold { font-weight: 700; } .text-text { color: #0f172a; } .text-text-muted { color: #64748b; } .text-text-subtle { color: #94a3b8; } .text-white { color: #ffffff; } .text-primary { color: #2563eb; } .text-destructive { color: #dc2626; } .text-success { color: #22c55e; } .text-warning { color: #f59e0b; } .text-info { color: #0ea5e9; } .text-current { color: currentColor; }

.bg-primary { background-color: #2563eb; } .bg-primary-light { background-color: #eff6ff; } .bg-primary-ring { background-color: #3b82f6; } .bg-secondary { background-color: #f3f4f6; } .bg-destructive { background-color: #dc2626; } .bg-destructive-light { background-color: #fef2f2; } .bg-success { background-color: #22c55e; } .bg-success-light { background-color: #f0fdf4; } .bg-warning { background-color: #f59e0b; } .bg-warning-light { background-color: #fffbeb; } .bg-info { background-color: #0ea5e9; } .bg-info-light { background-color: #f0f9ff; } .bg-surface-page { background-color: #f8fafc; } .bg-surface-card { background-color: #ffffff; } .bg-surface-muted { background-color: #f1f5f9; } .bg-terminal-bg { background-color: #0f172a; } .bg-transparent { background-color: transparent; } .bg-surface-overlay { background-color: rgba(15, 23, 42, 0.5); }

.btn { display: inline-flex; align-items: center; justify-content: center; font-weight: 500; transition-property: color, background-color; transition-duration: 150ms; border-radius: 6px; } .btn:focus-visible { outline: none; box-shadow: 0 0 0 2px #ffffff, 0 0 0 4px var(--btn-ring, #3b82f6); } .btn:disabled { pointer-events: none; opacity: 0.5; } .btn-primary { background-color: #2563eb; color: #ffffff; } .btn-primary:hover:not(:disabled) { background-color: #1d4ed8; } .btn-secondary { background-color: #f3f4f6; color: #111827; } .btn-secondary:hover:not(:disabled) { background-color: #e5e7eb; } .btn-destructive { background-color: #dc2626; color: #ffffff; } .btn-destructive:hover:not(:disabled) { background-color: #b91c1c; } .btn-ghost { background-color: transparent; color: #374151; } .btn-ghost:hover:not(:disabled) { background-color: #f3f4f6; } .btn-warning { background-color: #f59e0b; color: #ffffff; } .btn-warning:hover:not(:disabled) { background-color: #d97706; } .btn-sm { height: 32px; padding-left: 12px; padding-right: 12px; font-size: 12px; } .btn-md { height: 40px; padding: 8px 16px; font-size: 14px; } .spin { animation: spin 1s linear infinite; display: inline-block; width: 16px; height: 16px; margin-right: 8px; margin-left: -4px; vertical-align: middle; } @keyframes spin { from { transform: rotate(0deg); } to { transform: rotate(360deg); } }

.badge { display: inline-flex; align-items: center; gap: 6px; } .badge-pill { border-radius: 9999px; font-size: 12px; font-weight: 500; } .badge-md { padding: 4px 10px; } .badge-sm { padding: 2px 8px; } .badge-success { background-color: #f0fdf4; color: #22c55e; } .badge-warning { background-color: #fffbeb; color: #f59e0b; } .badge-destructive { background-color: #fef2f2; color: #dc2626; } .badge-info { background-color: #f0f9ff; color: #0ea5e9; } .badge-primary { background-color: #eff6ff; color: #2563eb; } .badge-muted { background-color: #f1f5f9; color: #64748b; } .badge-dot { width: 10px; height: 10px; border-radius: 9999px; } .dot-success { background-color: #22c55e; } .dot-warning { background-color: #f59e0b; } .dot-destructive { background-color: #dc2626; } .dot-info { background-color: #0ea5e9; } .dot-primary { background-color: #3b82f6; } .dot-muted { background-color: #64748b; }

.card { border-radius: 8px; border: 1px solid #e2e8f0; background-color: #ffffff; color: #0f172a; box-shadow: 0 1px 2px 0 rgba(0, 0, 0, 0.05); } .card-title-row { display: flex; flex-direction: column; gap: 6px; padding: 24px; border-bottom: 1px solid #e2e8f0; } .card-title { font-size: 18px; font-weight: 600; line-height: 1; letter-spacing: -0.025em; } .card-pad-md { padding: 24px; } .card-pad-sm { padding: 12px; }

.field { display: flex; flex-direction: column; gap: 6px; width: 100%; } .field-label label { font-size: 14px; font-weight: 500; color: #0f172a; } .input { flex: 1; height: 40px; width: 100%; border-radius: 6px; border: 1px solid #cbd5e1; background-color: #ffffff; padding: 8px 12px; font-size: 14px; color: #0f172a; } .input::placeholder { color: #94a3b8; } .input:focus-visible { outline: none; box-shadow: 0 0 0 2px #ffffff, 0 0 0 4px #3b82f6; } .field-error { font-size: 12px; color: #dc2626; }

.tablist { display: flex; } .tab-pills { position: relative; padding: 6px 12px; font-size: 12px; border-radius: 6px; transition-property: color, background-color; transition-duration: 150ms; border: none; cursor: pointer; } .tab-pills.active { background-color: #eff6ff; color: #2563eb; border: 1px solid #3b82f6; } .tab-pills.inactive { background-color: #f1f5f9; color: #64748b; border: 1px solid #e2e8f0; } .tab-pills.inactive:hover { background-color: #f8fafc; } .tab-badge { display: inline-flex; align-items: center; justify-content: center; height: 16px; width: 16px; border-radius: 9999px; font-size: 9px; font-weight: 700; background-color: #eff6ff; color: #2563eb; margin-left: 6px; }

.skeleton-line { animation: pulse 2s cubic-bezier(0.4, 0, 0.6, 1) infinite; background-color: #f1f5f9; border-radius: 4px; height: 16px; } .skeleton-card { animation: pulse 2s cubic-bezier(0.4, 0, 0.6, 1) infinite; background-color: #f1f5f9; border-radius: 8px; height: 96px; } @keyframes pulse { 0%, 100% { opacity: 1; } 50% { opacity: 0.5; } }

.empty-state { display: flex; flex-direction: column; align-items: center; justify-content: center; padding: 48px 16px; text-align: center; } .empty-state .es-icon { color: #94a3b8; margin-bottom: 16px; } .empty-state h3 { font-size: 18px; font-weight: 600; color: #0f172a; margin: 0 0 4px; } .empty-state p { font-size: 14px; color: #64748b; max-width: 28rem; margin: 0; } .empty-state .es-action { margin-top: 24px; }

.table-wrap { border-radius: 8px; border: 1px solid #e2e8f0; background-color: #ffffff; box-shadow: 0 1px 2px 0 rgba(0, 0, 0, 0.05); overflow: hidden; } table.proto-table { min-width: 100%; } table.proto-table thead { background-color: #f1f5f9; } table.proto-table th { padding: 12px 24px; text-align: left; font-size: 12px; font-weight: 500; color: #64748b; text-transform: uppercase; letter-spacing: 0.05em; } table.proto-table tbody { background-color: #ffffff; } table.proto-table tbody tr { border-top: 1px solid #e2e8f0; } table.proto-table tbody tr:hover { background-color: #f1f5f9; } table.proto-table td { padding: 16px 24px; white-space: nowrap; font-size: 14px; } .td-main { font-weight: 500; color: #0f172a; } .td-muted { color: #64748b; } .td-right { text-align: right; }

.alert { display: flex; align-items: flex-start; gap: 12px; border-radius: 8px; border: 1px solid; padding: 12px 16px; font-size: 14px; margin-bottom: 16px; } .alert-warning { background-color: #fffbeb; border-color: #fbbf24; color: #a16207; } .alert-destructive { background-color: #fef2f2; border-color: #f87171; color: #b91c1c; } .alert-info { background-color: #f0f9ff; border-color: #38bdf8; color: #0c4a6e; } .alert .alert-actions { margin-left: auto; display: flex; gap: 8px; flex-shrink: 0; }

.modal-backdrop { position: fixed; inset: 0; z-index: 50; background-color: rgba(15, 23, 42, 0.5); display: flex; align-items: center; justify-content: center; padding: 16px; } .modal-card { border-radius: 8px; border: 1px solid #e2e8f0; background-color: #ffffff; color: #0f172a; box-shadow: 0 10px 15px -3px rgba(0,0,0,0.1); width: 100%; max-width: 560px; } .modal-body { padding: 24px; } .modal-footer { display: flex; justify-content: flex-end; gap: 12px; padding: 16px 24px; border-top: 1px solid #e2e8f0; }

.toast-wrap { position: fixed; top: 56px; right: 16px; z-index: 2000; display: flex; flex-direction: column; gap: 8px; } .toast { display: flex; align-items: center; gap: 10px; border-radius: 8px; border: 1px solid; padding: 12px 16px; font-size: 14px; min-width: 260px; box-shadow: 0 4px 6px -1px rgba(0, 0, 0, 0.1); } .toast-success { background-color: #f0fdf4; border-color: #4ade80; color: #166534; } .toast-destructive { background-color: #fef2f2; border-color: #f87171; color: #b91c1c; }

/* autocomplete dropdown */ .autocomplete-wrap { position: relative; } .autocomplete-list { position: absolute; top: 44px; left: 0; right: 0; z-index: 30; background-color: #ffffff; border: 1px solid #cbd5e1; border-radius: 8px; box-shadow: 0 10px 15px -3px rgba(0, 0, 0, 0.1); overflow: hidden; } .ac-item { display: flex; align-items: center; gap: 10px; width: 100%; padding: 10px 12px; font-size: 13px; text-align: left; background: none; border: none; cursor: pointer; } .ac-item:hover, .ac-item.active { background-color: #eff6ff; } .ac-item .ac-login { font-weight: 600; color: #0f172a; } .ac-item .ac-name { color: #64748b; } .ac-item .ac-dept { color: #94a3b8; font-size: 12px; margin-left: auto; }

/* org tree */ .org-node { display: flex; align-items: center; gap: 8px; padding: 8px 10px; border-radius: 6px; font-size: 14px; } .org-node:hover { background-color: #f1f5f9; } .org-node .org-caret { width: 16px; color: #94a3b8; text-align: center; flex-shrink: 0; } .org-node .org-label { color: #0f172a; font-weight: 500; } .org-node .org-count { color: #94a3b8; font-size: 12px; } .org-children { margin-left: 24px; }

.proto-screen { display: block; } .proto-screen[hidden] { display: none !important; } input[type="checkbox"] { width: 16px; height: 16px; accent-color: #2563eb; cursor: pointer; }

@media (prefers-reduced-motion: reduce) { .skeleton-line, .skeleton-card, .spin { animation: none !important; } .transition-colors { transition-duration: 0.01ms !important; } } </style>

</head> PROTOTYPE Экран: 1 · Аудит + кэш IDM 2 · Оргструктурные привязки 3 · Поиск и автокомплит Состояние:
1280px 375px
<main class="proto-main" id="proto-main">
  <div class="proto-topbar" aria-hidden="true">
    <span class="text-sm text-text-muted">Оператор: admin (rls_operator)</span>
  </div>

  <!-- ============ SCREEN 1: AUDIT + CACHE ============ -->
  <section id="screen-audit" class="proto-screen">
    <header class="flex items-center justify-between mb-8">
      <div class="space-y-1">
        <h1 class="text-3xl font-bold tracking-tight text-text">Аудит пользователей</h1>
        <p class="text-sm text-text-muted">Статусы IDM из кэша · персональные данные — live по карточке</p>
      </div>
      <div class="flex items-center gap-4">
        <span class="badge badge-pill badge-md" data-role="cache-badge">
          <span class="badge-dot" data-role="cache-dot"></span>
          <span data-role="cache-text">Кэш IDM: обновлён 04.08 06:00</span>
        </span>
        <button class="btn btn-secondary btn-md" type="button" data-role="refresh-cache">Обновить кэш</button>
      </div>
    </header>

    <!-- cache banners -->
    <div class="alert alert-info" data-role="enriching-banner" hidden>
      <span>Обогащение выполняется в фоне: порция 3 из 50 · 6 214 / 12 483 обновлено</span>
      <span class="alert-actions"><span class="badge badge-pill badge-md badge-info">прогресс 50%</span></span>
    </div>
    <div class="alert alert-warning" data-role="stale-banner" hidden>
      <span>Кэш IDM устарел (от 02.08). Фоновая заливка не завершилась. Данные могут не отражать увольнения/переводы.</span>
      <span class="alert-actions"><button class="btn btn-warning btn-sm" type="button" data-role="retry-enrich">Повторить обогащение</button></span>
    </div>
    <div class="alert alert-destructive" data-role="degraded-banner" hidden>
      <span>IDM недоступен. Показан кэш от 02.08 (возраст 2 дня). Live-карточки недоступны.</span>
      <span class="alert-actions"><button class="btn btn-destructive btn-sm" type="button" data-role="retry-idm">Повторить</button></span>
    </div>

    <!-- loading -->
    <div data-role="loading" hidden>
      <div class="table-wrap"><div class="card-pad-md">
        <div class="skeleton-line mb-3"></div><div class="skeleton-line mb-3"></div><div class="skeleton-line" style="width:70%"></div>
      </div></div>
    </div>

    <!-- loaded table -->
    <div data-role="loaded" class="space-y-4">
      <div class="table-wrap">
        <div class="px-6 py-4 bg-surface-page border-b border-border flex items-center justify-between">
          <h3 class="text-lg font-semibold text-text">rls_bi_users × IDM (кэш)</h3>
          <span class="badge badge-pill badge-md badge-muted">12 483 · 12 не найдено</span>
        </div>
        <div class="overflow-x-auto">
          <table class="proto-table">
            <thead>
              <tr><th>AD-логин</th><th>Канон. логин</th><th>Роль</th><th>Сотрудник</th><th>Аккаунт AD</th><th>Риск</th><th>VIP</th><th>admin</th><th>Кэш</th></tr>
            </thead>
            <tbody>
              <tr>
                <td class="td-main">ivanovii</td>
                <td class="td-muted">ivanovii <span class="badge badge-pill badge-sm badge-primary ml-1">primary</span></td>
                <td class="td-muted">BI_bukrs_199</td>
                <td><span class="badge badge-pill badge-md badge-success">активен</span></td>
                <td><span class="badge badge-pill badge-md badge-success">active</span></td>
                <td class="td-muted">45</td><td class="td-muted">—</td><td class="td-muted">—</td>
                <td class="td-muted">04.08</td>
              </tr>
              <tr>
                <td class="td-main">novikovnn</td>
                <td class="td-muted">novikovnn <span class="badge badge-pill badge-sm badge-primary ml-1">primary</span></td>
                <td class="td-muted">BI_ALL</td>
                <td><span class="badge badge-pill badge-md badge-destructive">уволен</span></td>
                <td><span class="badge badge-pill badge-md badge-warning">disabled</span></td>
                <td class="td-muted">70</td><td class="td-muted">да</td><td class="td-muted">да</td>
                <td class="td-muted">04.08</td>
              </tr>
              <tr>
                <td class="td-main">sidorovss</td>
                <td class="td-muted">sidorovss <span class="badge badge-pill badge-sm badge-primary ml-1">primary</span></td>
                <td class="td-muted">BI_bukrs_305</td>
                <td><span class="badge badge-pill badge-md badge-info">в отпуске</span></td>
                <td><span class="badge badge-pill badge-md badge-success">active</span></td>
                <td class="td-muted">40</td><td class="td-muted">—</td><td class="td-muted">—</td>
                <td class="td-muted">04.08</td>
              </tr>
              <tr>
                <td class="td-main">kozlovakk</td>
                <td class="td-muted">kozlovakk <span class="badge badge-pill badge-sm badge-primary ml-1">primary</span></td>
                <td class="td-muted">BI_bukrs_210</td>
                <td><span class="badge badge-pill badge-md badge-destructive">уволен</span></td>
                <td><span class="badge badge-pill badge-md badge-warning">disabled</span></td>
                <td class="td-muted">0</td><td class="td-muted">—</td><td class="td-muted">—</td>
                <td class="td-muted">04.08</td>
              </tr>
              <tr>
                <td class="td-main">volkovaa</td>
                <td class="td-muted">—</td>
                <td class="td-muted">BI_bukrs_210</td>
                <td><span class="badge badge-pill badge-md badge-muted">не найдено</span></td>
                <td><span class="badge badge-pill badge-md badge-muted">—</span></td>
                <td class="td-muted">—</td><td class="td-muted">—</td><td class="td-muted">—</td>
                <td class="td-muted">—</td>
              </tr>
            </tbody>
          </table>
        </div>
      </div>

      <!-- employee card with primary account block -->
      <div class="card">
        <div class="card-title-row">
          <div class="flex items-center justify-between">
            <h3 class="card-title">Карточка — Иванов Иван Иванович</h3>
            <span class="badge badge-pill badge-md badge-info">live из IDM</span>
          </div>
        </div>
        <div class="card-pad-md">
          <div class="grid grid-cols-3 gap-4 mb-4">
            <div><p class="text-sm text-text-muted">Отдел</p><p class="text-sm font-medium text-text">Департамент информационных технологий</p></div>
            <div><p class="text-sm text-text-muted">Должность</p><p class="text-sm font-medium text-text">Разработчик</p></div>
            <div><p class="text-sm text-text-muted">Риск-скор / VIP</p><p class="text-sm font-medium text-text">45 / нет</p></div>
          </div>
          <h4 class="text-sm font-semibold text-text mb-2">Учётные записи (3)</h4>
          <div class="table-wrap">
            <table class="proto-table">
              <thead><tr><th>Приложение</th><th>Аккаунт</th><th>Домен</th><th>OU</th><th>Статус</th><th style="width:90px"></th></tr></thead>
              <tbody>
                <tr>
                  <td class="td-main">Active Directory</td><td class="td-muted">IvanovII</td><td class="td-muted">TEST.LOCAL</td>
                  <td class="td-muted">OU=External,OU=Users,DC=TEST,DC=LOCAL</td>
                  <td><span class="badge badge-pill badge-md badge-success">enabled</span></td>
                  <td><span class="badge badge-pill badge-sm badge-primary">primary</span></td>
                </tr>
                <tr>
                  <td class="td-main">Active Directory</td><td class="td-muted">IvanovII</td><td class="td-muted">IE.CORP</td>
                  <td class="td-muted">OU=Unassigned,DC=ie,DC=corp</td>
                  <td><span class="badge badge-pill badge-md badge-warning">disabled</span></td>
                  <td class="td-muted">—</td>
                </tr>
                <tr>
                  <td class="td-main">Exchange</td><td class="td-muted">IvanovII</td><td class="td-muted">TEST.LOCAL</td>
                  <td class="td-muted">—</td>
                  <td><span class="badge badge-pill badge-md badge-success">enabled</span></td>
                  <td class="td-muted">—</td>
                </tr>
              </tbody>
            </table>
          </div>
          <p class="text-sm text-text-muted mt-3">Канонический логин для RLS: <span class="text-sm font-medium text-text">ivanovii</span> (без домена, lowercase). Правило primary: активный enabled-аккаунт.</p>
        </div>
      </div>
    </div>
  </section>

  <!-- ============ SCREEN 2: ORG BINDINGS ============ -->
  <section id="screen-org" class="proto-screen" hidden>
    <header class="flex items-center justify-between mb-8">
      <div class="space-y-1">
        <h1 class="text-3xl font-bold tracking-tight text-text">Оргструктурные привязки</h1>
        <p class="text-sm text-text-muted">Живой состав отделов → роли · пересчёт по расписанию · авто-исключение уволенных/переведённых</p>
      </div>
      <div class="flex items-center gap-4">
        <span class="badge badge-pill badge-md" data-role="schedule-badge"><span class="badge-dot dot-success"></span> расписание: 0 3 * * *</span>
        <button class="btn btn-primary btn-md" type="button" data-role="refresh-now">Пересчитать сейчас</button>
      </div>
    </header>

    <!-- refreshing state -->
    <div data-role="refreshing" hidden>
      <div class="alert alert-info">
        <span>Пересчёт «Департамент ИТ»: запрос состава из IDM… порция 2/4</span>
        <span class="alert-actions"><span class="badge badge-pill badge-md badge-info">идёт</span></span>
      </div>
      <div class="card"><div class="card-pad-md">
        <div class="skeleton-line mb-3"></div><div class="skeleton-line mb-3"></div><div class="skeleton-line" style="width:60%"></div>
      </div></div>
    </div>

    <!-- zero-found warning modal trigger state (visible banner) -->
    <div data-role="zero-found" hidden>
      <div class="alert alert-warning">
        <span>В отделе «Департамент ИТ» не найдено сотрудников (поиск вернул 0). Привязка НЕ будет изменена — защита от массовой деактивации.</span>
        <span class="alert-actions"><button class="btn btn-warning btn-sm" type="button" data-role="open-zero-modal">Подробнее</button></span>
      </div>
    </div>

    <!-- list + last report -->
    <div data-role="list" class="space-y-4">
      <div class="grid grid-cols-2 gap-4">
        <div class="card">
          <div class="card-title-row"><div class="flex items-center justify-between">
            <h3 class="card-title">Департамент ИТ</h3>
            <span class="badge badge-pill badge-md badge-success">ok · 04.08 03:00</span>
          </div></div>
          <div class="card-pad-md">
            <p class="text-sm text-text-muted mb-1">Роль: <span class="text-sm font-medium text-text">BI_czo_0210</span></p>
            <p class="text-sm text-text-muted mb-1">Состав: <span class="text-sm font-medium text-text">12 сотрудников</span></p>
            <p class="text-sm text-text-muted">Последний пересчёт: <span class="text-sm font-medium text-text">+1 / −0</span></p>
          </div>
        </div>
        <div class="card">
          <div class="card-title-row"><div class="flex items-center justify-between">
            <h3 class="card-title">Блок финансов (БЕ-002)</h3>
            <span class="badge badge-pill badge-md badge-success">ok · 04.08 03:00</span>
          </div></div>
          <div class="card-pad-md">
            <p class="text-sm text-text-muted mb-1">Роль: <span class="text-sm font-medium text-text">BI_czo_0310</span></p>
            <p class="text-sm text-text-muted mb-1">Состав: <span class="text-sm font-medium text-text">47 сотрудников</span></p>
            <p class="text-sm text-text-muted">Последний пересчёт: <span class="text-sm font-medium text-text">+2 / −1</span></p>
          </div>
        </div>
      </div>

      <!-- diff report -->
      <div class="table-wrap">
        <div class="px-6 py-4 bg-surface-page border-b border-border flex items-center justify-between">
          <h3 class="text-lg font-semibold text-text">Отчёт пересчёта — 04.08 03:00</h3>
          <div class="flex items-center gap-3">
            <span class="badge badge-pill badge-md badge-success">+3 добавлено</span>
            <span class="badge badge-pill badge-md badge-destructive">−2 деактивировано</span>
          </div>
        </div>
        <div class="overflow-x-auto">
          <table class="proto-table">
            <thead><tr><th>Пользователь</th><th>Привязка</th><th>Действие</th><th>Причина</th><th>Время</th></tr></thead>
            <tbody>
              <tr><td class="td-main">petrovpv</td><td class="td-muted">Департамент ИТ</td><td><span class="badge badge-pill badge-md badge-success">добавлен</span></td><td class="td-muted">новый сотрудник отдела</td><td class="td-muted">03:00:12</td></tr>
              <tr><td class="td-main">sidorovss</td><td class="td-muted">Департамент ИТ</td><td><span class="badge badge-pill badge-md badge-destructive">деактивирован</span></td><td class="td-muted">переведён (отдел изменился)</td><td class="td-muted">03:00:14</td></tr>
              <tr><td class="td-main">kozlovakk</td><td class="td-muted">Департамент ИТ</td><td><span class="badge badge-pill badge-md badge-destructive">деактивирован</span></td><td class="td-muted">уволен (dateOut 10.03.2026)</td><td class="td-muted">03:00:14</td></tr>
              <tr><td class="td-main">volkovaa</td><td class="td-muted">Блок финансов</td><td><span class="badge badge-pill badge-md badge-success">добавлен</span></td><td class="td-muted">новый сотрудник отдела</td><td class="td-muted">03:00:20</td></tr>
            </tbody>
          </table>
        </div>
      </div>
      <p class="text-sm text-text-muted">SQL-миграция пересчёта: <span class="text-sm font-medium text-text">rls_scripts/migrations/org/20260804_0300.sql</span> (в версионируемом контуре)</p>
    </div>
  </section>

  <!-- ============ SCREEN 3: SEARCH / AUTOCOMPLETE ============ -->
  <section id="screen-search" class="proto-screen" hidden>
    <header class="flex items-center justify-between mb-8">
      <div class="space-y-1">
        <h1 class="text-3xl font-bold tracking-tight text-text">Привязка пользователей</h1>
        <p class="text-sm text-text-muted">Автокомплит по IDM · поиск групп · оргструктура с составом из кэша</p>
      </div>
    </header>

    <div class="card mb-6">
      <div class="card-title-row"><h3 class="card-title">Правило: БЕ по ЦО (шаг 4 — привязка)</h3></div>
      <div class="card-pad-md">
        <nav class="tablist gap-2 mb-4" role="tablist" aria-label="Способ привязки">
          <button class="tab-pills active" role="tab" aria-selected="true" type="button" data-tab="manual">По AD-логину</button>
          <button class="tab-pills inactive" role="tab" aria-selected="false" type="button" data-tab="group">AD-группа / OU</button>
          <button class="tab-pills inactive" role="tab" aria-selected="false" type="button" data-tab="org">Оргструктура</button>
        </nav>

        <!-- manual tab with autocomplete -->
        <div data-panel="manual">
          <div class="autocomplete-wrap">
            <div class="flex items-center gap-2">
              <input class="input flex-1" type="text" id="search-input" placeholder="Введите минимум 3 символа…" aria-label="Поиск по IDM" autocomplete="off">
              <button class="btn btn-secondary btn-md" type="button">Добавить</button>
            </div>
            <!-- autocomplete states -->
            <div class="autocomplete-list" data-role="ac-list" hidden role="listbox" aria-label="Кандидаты"></div>
            <div class="autocomplete-list" data-role="ac-loading" hidden>
              <div class="card-pad-sm"><div class="skeleton-line"></div></div>
            </div>
            <div class="autocomplete-list" data-role="ac-empty" hidden>
              <div class="empty-state" style="padding:16px">
                <p class="text-sm text-text-muted">Ничего не найдено. Проверьте написание или домен.</p>
              </div>
            </div>
          </div>
          <p class="text-sm text-text-muted mt-3">Подсказка: введите 3+ символа — поиск по логину, ФИО и отделу. Канонический логин (primary) подсвечивается.</p>
        </div>

        <!-- group tab -->
        <div data-panel="group" hidden>
          <div class="flex items-center gap-2 mb-3">
            <input class="input flex-1" type="text" placeholder="Поиск группы / OU (например, OU=External,OU=Users…)">
            <button class="btn btn-secondary btn-md" type="button">Найти группу</button>
          </div>
          <div class="table-wrap">
            <table class="proto-table">
              <thead><tr><th>Группа / OU</th><th>Аккаунтов</th><th>enabled</th><th style="width:90px"></th></tr></thead>
              <tbody>
                <tr><td class="td-main">OU=External,OU=Users,OU=Test,OU=Moscow,DC=TEST,DC=LOCAL</td><td class="td-muted">34</td><td class="td-muted">31</td><td><button class="btn btn-ghost btn-sm" type="button">Импорт</button></td></tr>
              </tbody>
            </table>
          </div>
        </div>

        <!-- org tab -->
        <div data-panel="org" hidden>
          <div class="mb-3">
            <div class="org-node"><span class="org-caret">▾</span><span class="org-label">ООО «ТестСервис»</span><span class="org-count">128</span></div>
            <div class="org-children">
              <div class="org-node"><span class="org-caret">▾</span><input type="checkbox" aria-label="Выбрать Департамент ИТ"><span class="org-label">Департамент информационных технологий</span><span class="org-count">12</span></div>
              <div class="org-children">
                <div class="org-node"><span class="org-caret"></span><input type="checkbox" aria-label="Выбрать Отдел разработки"><span class="org-label">Отдел разработки</span><span class="org-count">8</span></div>
                <div class="org-node"><span class="org-caret"></span><input type="checkbox" aria-label="Выбрать Отдел поддержки"><span class="org-label">Отдел поддержки</span><span class="org-count">4</span></div>
              </div>
              <div class="org-node"><span class="org-caret"></span><input type="checkbox" aria-label="Выбрать Блок финансов"><span class="org-label">Блок финансов (БЕ-002)</span><span class="org-count">47</span></div>
            </div>
          </div>
          <p class="text-sm text-text-muted">Выбрано: <span class="text-sm font-medium text-text">Департамент ИТ (12)</span> — привязка будет пересчитываться автоматически при переводах и увольнениях.</p>
        </div>
      </div>
    </div>

    <!-- bound users -->
    <div class="card">
      <div class="card-title-row"><h3 class="card-title">Привязанные пользователи</h3></div>
      <div class="card-pad-md space-y-2">
        <div class="flex items-center justify-between border border-border rounded-md px-4 py-2">
          <span class="text-sm font-medium text-text">ivanovii <span class="badge badge-pill badge-sm badge-primary ml-1">primary</span></span>
          <span class="badge badge-pill badge-md badge-success">найден в IDM</span>
        </div>
        <div class="flex items-center justify-between border border-border rounded-md px-4 py-2">
          <span class="text-sm font-medium text-text">petrovpv</span>
          <span class="badge badge-pill badge-md badge-success">найден в IDM</span>
        </div>
        <div class="flex items-center justify-between border border-border rounded-md px-4 py-2">
          <span class="text-sm font-medium text-text">kozlovakk</span>
          <span class="badge badge-pill badge-md badge-warning">уволен — будет исключена при пересчёте</span>
        </div>
      </div>
    </div>
  </section>
</main>

В отделе не найдено сотрудников

«Департамент ИТ»: поиск в IDM вернул 0 записей. Возможные причины: отдел переименован, поиск ошибочен, IDM недоступен. Привязка НЕ будет изменена (защита от массовой деактивации).

Отмена Всё равно пересчитать (0/0)
<script> /* ============ STATE SWITCHER ============ */ const SCREENS = { 'screen-audit': { label: '1 · Аудит + кэш IDM', states: { 'cache-fresh': { label: 'cache-fresh — статусы из кэша, всё актуально' }, 'cache-stale': { label: 'cache-stale — кэш устарел (warning banner)' }, 'cache-failed': { label: 'cache-failed — IDM недоступен, кэш от 02.08' }, 'enriching': { label: 'enriching — фоновая заливка (порция 3/50)' }, 'loading': { label: 'loading — скелетон таблицы' }, }, }, 'screen-org': { label: '2 · Оргструктурные привязки', states: { 'list': { label: 'list — привязки + отчёт пересчёта' }, 'refreshing': { label: 'refreshing — пересчёт идёт (прогресс)' }, 'zero-found': { label: 'zero-found — 0 сотрудников, привязка не меняется' }, 'refresh-failed': { label: 'refresh-failed — пересчёт не удался' }, }, }, 'screen-search': { label: '3 · Поиск и автокомплит', states: { 'typing': { label: 'typing — ввод, подсказка' }, 'searching': { label: 'searching — запрос в IDM' }, 'candidates': { label: 'candidates — список кандидатов' }, 'no-results': { label: 'no-results — ничего не найдено' }, 'selected': { label: 'selected — кандидат выбран' }, }, }, }; const screenSelect = document.getElementById('screen-select'); const stateSelect = document.getElementById('state-select'); const stateLabel = document.getElementById('state-label'); function applyScreenState(screenId, stateKey) { document.querySelectorAll('.proto-screen').forEach(s => { s.hidden = true; }); document.getElementById(screenId).hidden = false; // reset per-screen banners ['enriching-banner','stale-banner','degraded-banner'].forEach(k => { const el = document.querySelector('#screen-audit [data-role="' + k + '"]'); if (el) el.hidden = true; }); ['refreshing','zero-found','list'].forEach(k => { const el = document.querySelector('#screen-org [data-role="' + k + '"]'); if (el) el.hidden = true; }); ['ac-list','ac-loading','ac-empty'].forEach(k => { const el = document.querySelector('#screen-search [data-role="' + k + '"]'); if (el) el.hidden = true; }); const load = document.querySelector('#screen-audit [data-role="loading"]'); const loaded = document.querySelector('#screen-audit [data-role="loaded"]'); if (load) load.hidden = true; if (loaded) loaded.hidden = true; if (screenId === 'screen-audit') { const badge = document.querySelector('#screen-audit [data-role="cache-badge"]'); const dot = document.querySelector('#screen-audit [data-role="cache-dot"]'); const txt = document.querySelector('#screen-audit [data-role="cache-text"]'); if (stateKey === 'cache-fresh') { loaded.hidden = false; badge.className = 'badge badge-pill badge-md badge-success'; dot.className = 'badge-dot dot-success'; txt.textContent = 'Кэш IDM: обновлён 04.08 06:00'; } if (stateKey === 'cache-stale') { loaded.hidden = false; document.querySelector('#screen-audit [data-role="stale-banner"]').hidden = false; badge.className = 'badge badge-pill badge-md badge-warning'; dot.className = 'badge-dot dot-warning'; txt.textContent = 'Кэш IDM: от 02.08 (устарел)'; } if (stateKey === 'cache-failed') { loaded.hidden = false; document.querySelector('#screen-audit [data-role="degraded-banner"]').hidden = false; badge.className = 'badge badge-pill badge-md badge-destructive'; dot.className = 'badge-dot dot-destructive'; txt.textContent = 'IDM недоступен · кэш от 02.08'; } if (stateKey === 'enriching') { loaded.hidden = false; document.querySelector('#screen-audit [data-role="enriching-banner"]').hidden = false; badge.className = 'badge badge-pill badge-md badge-info'; dot.className = 'badge-dot dot-info'; txt.textContent = 'Обогащение: 6 214 / 12 483'; } if (stateKey === 'loading') load.hidden = false; } if (screenId === 'screen-org') { if (stateKey === 'list') document.querySelector('#screen-org [data-role="list"]').hidden = false; if (stateKey === 'refreshing') document.querySelector('#screen-org [data-role="refreshing"]').hidden = false; if (stateKey === 'zero-found') { document.querySelector('#screen-org [data-role="list"]').hidden = false; document.querySelector('#screen-org [data-role="zero-found"]').hidden = false; } if (stateKey === 'refresh-failed') { document.querySelector('#screen-org [data-role="list"]').hidden = false; const alert = document.createElement('div'); alert.className = 'alert alert-destructive'; alert.id = 'refresh-failed-alert'; alert.innerHTML = 'Пересчёт «Блок финансов» не выполнен: IDM недоступен. Привязки не изменены, повтор при следующем расписании.Повторить'; const list = document.querySelector('#screen-org [data-role="list"]'); if (!document.getElementById('refresh-failed-alert')) list.parentNode.insertBefore(alert, list); } else { const a = document.getElementById('refresh-failed-alert'); if (a) a.remove(); } } if (screenId === 'screen-search') { const input = document.getElementById('search-input'); if (stateKey === 'candidates') { input.value = 'iva'; const list = document.querySelector('#screen-search [data-role="ac-list"]'); list.hidden = false; list.innerHTML = ` ivanovii primary Иванов ИванДепартамент ИТ ivanovaia Иванова АлинаБлок финансов ivashkink Ивашкин КириллОтдел поддержки `; } if (stateKey === 'searching') { input.value = 'iva'; document.querySelector('#screen-search [data-role="ac-loading"]').hidden = false; } if (stateKey === 'no-results') { input.value = 'ghostuser'; document.querySelector('#screen-search [data-role="ac-empty"]').hidden = false; } if (stateKey === 'typing' || stateKey === 'selected') input.value = stateKey === 'selected' ? 'ivanovii' : ''; } stateLabel.textContent = SCREENS[screenId].states[stateKey].label; } function renderStateOptions(screenId) { stateSelect.innerHTML = ''; Object.keys(SCREENS[screenId].states).forEach(k => { const opt = document.createElement('option'); opt.value = k; opt.textContent = SCREENS[screenId].states[k].label; stateSelect.appendChild(opt); }); stateSelect.value = Object.keys(SCREENS[screenId].states)[0]; applyScreenState(screenId, stateSelect.value); } screenSelect.addEventListener('change', () => renderStateOptions(screenSelect.value)); stateSelect.addEventListener('change', () => applyScreenState(screenSelect.value, stateSelect.value)); document.getElementById('vp-desktop').addEventListener('click', () => { document.body.classList.remove('chrome-viewport-mobile'); document.getElementById('vp-desktop').classList.add('active'); document.getElementById('vp-mobile').classList.remove('active'); }); document.getElementById('vp-mobile').addEventListener('click', () => { document.body.classList.add('chrome-viewport-mobile'); document.getElementById('vp-mobile').classList.add('active'); document.getElementById('vp-desktop').classList.remove('active'); }); /* ============ TOASTS + WIRING ============ */ function toast(text, variant) { const t = document.createElement('div'); t.className = 'toast toast-' + variant; t.setAttribute('role', 'status'); t.textContent = text; document.getElementById('toast-wrap').appendChild(t); setTimeout(() => t.remove(), 4000); } document.addEventListener('click', (e) => { // refresh cache if (e.target.closest('[data-role="refresh-cache"]')) { stateSelect.value = 'enriching'; applyScreenState('screen-audit', 'enriching'); setTimeout(() => { stateSelect.value = 'cache-fresh'; applyScreenState('screen-audit', 'cache-fresh'); toast('Кэш обновлён: 12 483 статуса (12 не найдено)', 'success'); }, 1500); } // retry enrichment / idm if (e.target.closest('[data-role="retry-enrich"]')) { toast('Обогащение поставлено в очередь (с порции 41)', 'success'); } if (e.target.closest('[data-role="retry-idm"]')) { stateSelect.value = 'cache-fresh'; applyScreenState('screen-audit', 'cache-fresh'); toast('IDM доступен: кэш перезалит', 'success'); } // org refresh if (e.target.closest('[data-role="refresh-now"]')) { stateSelect.value = 'refreshing'; applyScreenState('screen-org', 'refreshing'); setTimeout(() => { stateSelect.value = 'list'; applyScreenState('screen-org', 'list'); toast('Пересчёт: +3 добавлено, −2 деактивировано (sidorovss: переведён, kozlovakk: уволен)', 'success'); }, 1600); } // zero-found modal if (e.target.closest('[data-role="open-zero-modal"]')) document.getElementById('zero-modal').hidden = false; if (e.target.closest('[data-role="zero-cancel"]')) document.getElementById('zero-modal').hidden = true; if (e.target.closest('[data-role="zero-confirm"]')) { document.getElementById('zero-modal').hidden = true; toast('Пересчёт без изменений (0/0). Предупреждение зафиксировано в отчёте', 'success'); } // tabs const tab = e.target.closest('[data-tab]'); if (tab) { document.querySelectorAll('#screen-search [data-tab]').forEach(t => { t.classList.remove('active'); t.classList.add('inactive'); t.setAttribute('aria-selected','false'); }); tab.classList.remove('inactive'); tab.classList.add('active'); tab.setAttribute('aria-selected','true'); document.querySelectorAll('#screen-search [data-panel]').forEach(p => { p.hidden = true; }); document.querySelector('#screen-search [data-panel="' + tab.dataset.tab + '"]').hidden = false; } // autocomplete: typing sim const searchInput = document.getElementById('search-input'); if (searchInput && e.target === searchInput && searchInput.value.length >= 3) { stateSelect.value = 'candidates'; applyScreenState('screen-search', 'candidates'); } if (e.target.classList.contains('ac-item')) { stateSelect.value = 'selected'; applyScreenState('screen-search', 'selected'); document.getElementById('search-input').value = 'ivanovii'; toast('ivanovii добавлен в привязку (primary TEST.LOCAL)', 'success'); } if (e.target.classList.contains('modal-backdrop')) e.target.hidden = true; }); document.addEventListener('keydown', (e) => { if (e.key === 'Escape') document.getElementById('zero-modal').hidden = true; }); renderStateOptions('screen-audit'); </script> </html>

================================================================================ FEATURE: 050-mcp-interface Files: 5


SPEC — Feature Specification

Source: spec.md

#region McpInterface.Spec [C:3] [TYPE ADR] [SEMANTICS mcp,interface,tools,agent,dashboard-testing,decommission] @BRIEF Simple MCP interface to ss-tools that replaces the Gradio chat agent as the primary agentic surface, including full dashboard-test creation. @RELATION DEPENDS_ON -> [Doc.Adr.ADR0001] @RELATION DEPENDS_ON -> [Doc.Adr.ADR0005] @RELATION DEPENDS_ON -> [Doc.Adr.ADR0006] @RELATION DEPENDS_ON -> [AgentTestStabilization.Spec] @RATIONALE ss-tools already concentrates truth and policy server-side: every agent tool is a thin forwarder over backend services (RBAC guard -> httpx -> truncated result), and 042-047 demand that any delegated actor works through the same server-owned contracts as an analyst. A conversational runtime coupled to a bespoke Gradio UI adds a parallel transport, a parallel HITL mechanism (LangGraph interrupt + confirmation module) and triple bookkeeping per capability (route + LangChain wrapper + allowlist). MCP externalizes reasoning to standard clients while the backend keeps deterministic validators, ActionApprovalGate, AgentAction provenance and InvestigationSignal semantics unchanged. @REJECTED Keeping the Gradio chat agent long-term — rejected 2026-08-24: the chat surface duplicates policy transport, blocks external clients, and every new spec (042-047) multiplies hand-written wrappers. @REJECTED Personal access tokens as the primary authentication mechanism — deferred 2026-08-24 in favor of full OAuth 2.1 (authorization code + PKCE) so that standard clients connect through discovery without manual token management; static tokens MAY return later as a convenience feature. @REJECTED Embedding roles or permissions into access tokens — rejected because rights must be revocable in real time; tokens carry identity only and permissions resolve live from the database. @REJECTED Adopting the vendored research/mcp-superset server as the tool surface — rejected because it talks to Superset directly, bypassing the ss-tools policy layer; direct SQL and raw Superset mutations violate 037 chart-truth boundaries and 038/044 executor contracts. It remains a reference implementation only. @REJECTED Replacing 37 tools with one generic api_call tool — rejected because it destroys curated schemas/discoverability, degrades small-model tool selection, and is not MCP. @REJECTED Auto-generating the tool catalog from the OpenAPI dump — rejected because curated descriptions, bounded response discipline and per-tool risk semantics are part of the safety contract; generation is allowed only as a scaffold for explicitly reviewed registrations.

Navigation (DSA Indexer keywords)

@SEMANTICS: spec, requirements, mcp, interface, tools, catalog, approval, provenance, decommission, gradio

Feature Branch: 050-mcp-interface Created: 2026-08-24 | Status: Ready for Implementation Input: "Убрать веб-интерфейс агента (Gradio, кнопка «Ассистент» и другие точки входа) и заменить его простым MCP-интерфейсом к ss-tools, через который внешние клиенты создают все тесты. Полноценный внешний доступ; первый срез — паритет текущих 37 инструментов; кнопка «Создать сценарий тестирования» выполняет handoff во внешний клиент."

Decisions (session 2026-08-24)

  • Direction: remove the Gradio agent (agent/ service, port 7860, /agent route, assistant drawer/API); its place is taken by ONE simple MCP server mounted in the FastAPI backend.
  • Hosting: /mcp inside the existing FastAPI app (FastMCP / official SDK, Streamable HTTP). No separate process, port or auth stack; tools call services in-process.
  • External access: full, via OAuth 2.1 from Phase 0 — the ss-tools backend acts as the Authorization Server (authorization code + PKCE on top of the existing session login incl. ADFS OIDC), the MCP endpoint acts as a Resource Server. Dynamic client registration for public clients; client_credentials/SERVICE_JWT fallback for machine clients.
  • Deployment boundary: "external MCP client" means a process external to the ss-tools backend, not an Internet-hosted client. Until a later ADR explicitly changes this, MCP clients and every LLM/VLM provider are enterprise-local deployments; PII exposure inside that perimeter is an accepted residual risk.
  • First slice: parity with the current 37 LangChain @tool wrappers before any new capability lands exclusively on MCP.
  • Dashboard entry: «Создать сценарий тестирования» becomes an explicit handoff to an external MCP client (instruction page / copyable prompt + connection parameters), not a chat launch.

Authorization Model

The backend already owns the canonical RBAC shape (Role(is_admin) → Permission(resource, action), pure predicate user_has_permission() in core/auth/permission_utils.py) and ships Authlib + cryptography. The drift reuses both; no parallel policy vocabulary is created.

Roles and endpoints

Part Role Endpoints / behavior
ss-tools backend OAuth 2.1 Authorization Server /oauth/authorize (requires existing web session: local password or ADFS OIDC), /oauth/token (authorization_code + PKCE S256, refresh_token, client_credentials), AS metadata (RFC 8414), JWKS
/mcp OAuth 2.1 Resource Server Protected-resource metadata (RFC 9728), 401 + WWW-Authenticate discovery hint, token validation per request
External client Public/confidential OAuth client Discovers AS via RFC 9728 → registers via DCR (RFC 7591) → authorization code + PKCE in the user's browser session → calls tools

Token and identity rules

  • Access tokens are short-lived signed JWTs (aud=mcp) carrying identity only (sub, scope, jti, session_id). Roles/permissions are NEVER embedded; they resolve live from the DB per request, so role changes take effect immediately.
  • Refresh tokens rotate on every use; reuse detection revokes the token family (extends the existing TokenBlacklist mechanism).
  • Machine clients authenticate with client_credentials or the existing SERVICE_JWT; both map to a service principal with its own scope set — never to a user's roles.
  • Authentication is a swappable adapter resolve_principal(credentials) -> McpSessionPrincipal; the tool catalog never sees transport credentials.

Tool catalog binding (hidden vs gated)

  1. Every tool registration declares required_permission=("resource","action") from the canonical vocabulary already used by specs (scenario:view/create/edit/archive, dataset:lineage:refresh, dashboard:loadtest:*, …). This is the single source of truth; the agent-side _TOOL_PERMISSIONS shadow is absorbed and deleted.
  2. tools/list filters the catalog through user_has_permission(principal.user, resource, action); is_admin=True sees everything. A tool the principal has no right for is absent from the listing (HIDDEN).
  3. Every invocation re-checks the same predicate plus service-level policies (defense in depth against stale client catalogs).
  4. Contextual risk does NOT hide a tool: a visible tool whose current invocation is risky (PROD environment, baseline publish) returns approval_required{gate_id,...} instead of failing silently (GATED).

Gates live entirely inside MCP

The MCP session MUST be self-sufficient for the whole approval loop: a principal operating only through MCP clients completes creation, gating, running and investigation without opening the ss-tools web UI. Rule set:

  • Gated invocations return approval_required{gate_id, targets, risk, reason_required, expires_at} synchronously — the client resolves it in-session.
  • list_pending_approvals(filter?) returns all gates visible to the principal (including gates raised by automation or other surfaces), each with target diff/provenance refs.
  • decide_approval(gate_id, confirm|deny, reason) consumes the gate atomically (CAS), records the reason, and unblocks the original workflow deterministically; the caller may then retry or consume per the workflow contract. Expired/already-decided gates return typed errors, never silent success.
  • Decision tools require an authenticated USER principal; service principals can list but never decide. Reasons are mandatory for confirm on high-risk classes (baseline publish, PROD dispatch, policy change) per AGSTAB-FR-005.
  • There is exactly ONE durable gate store (ActionApprovalGate). Web gate cards (042/043/045 product pages) and MCP decision tools are two renderers of the same rows — neither creates parallel state, and a decision through either path is immediately visible to the other.

Human checkpoints through MCP

HumanCheckpoint disposition (waiting_human runs, 044) is also completable inside MCP, amending SCEX-FR-013 for the post-chat world:

  • list_checkpoints(filter?) exposes waiting checkpoints visible to the principal with their evidence context; decide_checkpoint(run_id, disposition=confirm|false_positive|inconclusive, expected_version) consumes the checkpoint via the same CAS contract as the monitor.
  • Guards preserve the real invariant (no automated false PASS): decision tools accept USER principals only, CAS version mismatch is a typed error, every decision is audited with principal+reason, and no scheduled/deploy/API origin can ever reach these tools. A service principal calling them gets permission_denied.
  • The web monitor (045) remains an equal renderer of the same checkpoint rows; dispositions from either surface are immediately visible in both.

User Scenarios

Story 1 — Connect an External Client (P1)

Why P1: The MCP endpoint is the product surface; connecting must be boring and secure.

Independent Test: Point MCP Inspector or Claude Desktop at {backend}/mcp with a personal token and verify tools/list returns exactly the role-permitted catalog.

Acceptance:

  1. Given a valid user JWT When the client connects and lists tools Then the catalog contains only tools permitted for that user's role, with curated descriptions and JSON schemas.
  2. Given an invalid/expired token When the client connects Then it receives a typed authentication error and zero tool metadata beyond the public surface.
  3. Given a service JWT When a machine client lists tools Then machine-scoped tools resolve per service policy, indistinguishable in contract shape from user sessions.

Story 2 — Create a Dashboard Test Scenario End-to-End (P1)

Why P1: "Создавать все тесты через MCP" is the core value replacing the chat flow.

Independent Test: From an external client, drive inspect → compile → validate → resolve → draft-pack → save for a fixture dashboard and verify the saved immutable revision appears in the 042 registry.

Acceptance:

  1. Given a dashboard id and environment When the client calls the inspection and scenario-authoring tools Then the chain produces a validated graph with needs_context/needs_selector markers instead of invented data (038 semantics).
  2. Given delegated policy permits the actor When save is requested Then a server-stored WorkingDraft produces a new immutable revision with full provenance; the client never uploads a full graph for save (SCEDIT-FR-009 inheritance).
  3. Given policy does not delegate the save When save is requested Then the tool returns a typed approval_required envelope and nothing persists until the gate resolves.

Story 3 — Approval Round-Trip Fully Inside the Client (P1)

Why P1: HITL must survive the chat removal WITHOUT forcing a context switch to the web UI; the MCP session is self-sufficient.

Independent Test: Trigger a gated baseline-approval action via MCP, approve it via decide_approval in the same client session, and verify the subsequent consume succeeds; repeat with deny and with expiry.

Acceptance:

  1. Given a gated action When invoked via MCP Then the result is approval_required{gate_id, targets, risk, reason_required, expires_at} and no side effect occurs.
  2. Given the analyst calls list_pending_approvals in the same or a later session When gates exist (including gates raised by automation) Then they are listed with target diff/provenance refs.
  3. Given decide_approval(gate_id, confirm|deny, reason) When submitted by a user principal Then the gate is consumed atomically, the reason is recorded, and the original workflow continues deterministically (retry/consume).
  4. Given a denial, expiry, or a replayed decision When any path retries Then the gate refuses with the recorded outcome; no duplicate side effect.
  5. Given the same gate row When rendered as a web gate card OR decided via MCP from another session Then both surfaces show identical state — one durable store, two renderers.

Story 4 — Baseline and Verification Lifecycle via MCP (P2)

Why P2: 037 capture/approval/verification flows were the first agent tools and must reach parity first.

Independent Test: Execute capture_baseline_candidate → request_baseline_approval → decide → consume and create_verification_run via MCP against fixtures, matching legacy wrapper outputs.

Acceptance:

  1. Given fixture release/dashboard/metric inputs When capture runs Then the server computes hashes and creates the candidate exactly as the HTTP route does (no client hash logic).
  2. Given the full approval lifecycle When driven via MCP Then outcomes are byte-comparable with the legacy tool results on the same fixtures.

Story 5 — Decommission the Chat Surface (P2)

Why P2: The drift ends with the Gradio stack gone, not duplicated.

Independent Test: Flip the removal flag, rebuild frontend and backend, and verify no /agent references, no agent process, and functional handoff entry points remain.

Acceptance:

  1. Given parity is proven When the flag removes the chat Then run.sh/compose no longer start the agent service; port 7860 disappears from profiles.
  2. Given the dashboard page When the user clicks «Создать сценарий тестирования» Then a handoff surface opens (connection instructions + copyable prompt with dashboard context parameters), never a dead link.
  3. Given the frontend build When link-integrity tests run Then zero references to /agent, assistant API or Gradio proxy remain.

Edge & Failure Cases

# Scenario Expected Behavior Recovery
E1 JWT expires mid-session Typed auth error on next call; no privileged retry Client reconnects with fresh token
E2 Stale cached tools/list after upgrade Unknown tool call returns typed TOOL_NOT_FOUND + re-list hint Client refreshes catalog
E3 Oversized tool response Bounded payload; large artifacts returned as ref + digest, never raw dumps Client fetches artifact by ref
E4 Concurrent scenario saves 409 revision conflict with current hash (042 semantics) Reload / compare
E5 SQL-class tool requested in restricted context Server-side denial regardless of client claims Elevated permission required
E6 Rate/abuse from external client Standard throttling honored (Retry-After) Back off
E7 Client attempts raw graph upload to save Rejected; save only from server-stored draft Recompile server-side
E8 Rotated refresh token replayed Whole token family revoked; client re-authorizes Re-run consent flow
E9 ADFS unavailable at authorize time Local-password session login still completes the flow Retry SSO later
E10 Role changed after tools/list cached by client Invocation re-check denies with typed permission error + re-list hint Client refreshes catalog
E11 MCP request exceeds transport or domain limit Typed payload-limit error; no tool dispatch or mutable state Client reduces request / paginates
E12 Local LLM/VLM receives capture containing PII Accepted residual risk inside enterprise perimeter; no masking gate Optional display/export masking
E13 Local provider payload/log includes a credential, cookie, token, secret, or raw storage path Typed secret-exposure rejection; event is audit-safe Reconfigure producer; retry with sanitized metadata

Requirements

Functional

  • MCPX-FR-001: The system MUST expose exactly one MCP server at {backend}/mcp (Streamable HTTP) inside the FastAPI application; stdio transport MAY exist for local development only.
  • MCPX-FR-002: Tools MUST be registered explicitly with curated name, description and input schema; OpenAPI-derived bulk generation MUST NOT ship.
  • MCPX-FR-003: The /mcp endpoint MUST authenticate every request through the Authorization Model above; tools/list MUST filter by caller RBAC and every invocation MUST re-enforce permissions server-side (single server-side policy source; the agent-side _tool_filter/_guard_tool_permission logic is absorbed here).
  • MCPX-FR-004: Every tool invocation MUST persist AgentAction provenance (principal, tool, arguments digest, outcome, linked AgentRun where applicable) and MUST pass deterministic ACL/environment/capacity checks before execution, inheriting AGSTAB-FR-011.
  • MCPX-FR-005: Non-delegated risky actions MUST return the typed approval_required envelope bound to a durable ActionApprovalGate. The approval loop MUST be completable entirely within MCP (list_pending_approvals + decide_approval); web gate cards (042/043/045 product pages) remain an equal renderer of the same durable gates. Denial MUST record cancellation with zero side effects; decision tools MUST reject service principals and expired/decided gates with typed errors.
  • MCPX-FR-006: The scenario-authoring chain (inspect query model, compile, validate, resolve, draft-pack, initial bootstrap, request-save) MUST be fully available via MCP; all mutable state MUST live server-side (WorkingDraft, revisions, candidates).
  • MCPX-FR-007: Tool responses MUST be bounded; artifacts larger than the inline limit MUST be returned as typed refs with digests resolvable through existing artifact APIs.
  • MCPX-FR-007a: MCP transport and every tool schema MUST enforce server-owned request byte, JSON-depth, collection-count, string-length, graph-size, and per-session rate limits before dispatch. Rejection is typed and creates no mutable state; bounded responses alone are insufficient.
  • MCPX-FR-008: The initial catalog MUST provide 1:1 parity with the current 37 tools across domains: environments/health/tasks, git/deploy/migration/backup/maintenance, Superset operations (read + admin CRUD), baseline capture/approval/verification, scenario compile/validate/resolve/draft-pack/save. show_capabilities is retired in favor of standard tools/list.
  • MCPX-FR-009: Decommission MUST be flag-driven and ordered: parity proven → chat hidden → agent service dropped from run.sh/compose → code deletion. Frontend entry points MUST switch to the handoff surface in the same phase that hides /agent.
  • MCPX-FR-010: The catalog MUST carry a version; breaking changes (rename/schema change/removal) MUST bump the major version and remain listed with a deprecation marker for one minor cycle.
  • MCPX-FR-011: MCP tools MUST reuse existing services and contracts; a second Playwright/LLM/SQL/execution stack is forbidden (SCEX-FR-009 inheritance). No MCP path may bypass 038 validation, 037 baseline rules, 044 executors or a required gate.
  • MCPX-FR-012: InvestigationSignals and queue items MUST NOT trigger any MCP activity automatically; MCP is pull-only from the client side (SCAN-FR-001 inheritance).
  • MCPX-FR-013: The backend MUST implement the Authorization Server endpoints of the Authorization Model table; PKCE S256 MUST be mandatory for public clients, and /oauth/authorize MUST reuse the existing web session (local password or ADFS OIDC) without a second credential prompt inside an active session.
  • MCPX-FR-014: The MCP endpoint MUST publish RFC 9728 protected-resource metadata and answer unauthenticated calls with 401 + WWW-Authenticate so a compliant client completes discovery → registration → authorization → tool listing without manual token pasting.
  • MCPX-FR-015: Dynamic client registration (RFC 7591) MUST be supported for public clients with first-party-bounded scopes; registered clients MUST be visible and revocable in Admin.
  • MCPX-FR-016: Refresh tokens MUST rotate on use, and replay of a rotated refresh token MUST revoke the whole token family (extending TokenBlacklist).
  • MCPX-FR-017: Access tokens MUST carry identity only (sub, scope, jti, session_id, aud=mcp); permission resolution MUST hit live DB state per request so role changes apply to the next call without re-consent.
  • MCPX-FR-018: Catalog visibility MUST follow the hidden-vs-gated rule: missing permission hides the tool from tools/list; contextual risk (PROD environment, baseline publish, non-delegated mutation) keeps the tool listed but returns approval_required.
  • MCPX-FR-019: HumanCheckpoint disposition MUST be available via list_checkpoints + decide_checkpoint under the same CAS/audit contract as the 045 monitor, restricted to authenticated user principals; no automated origin may create, consume or bypass a checkpoint, and manual-run-only revisions remain ineligible for automation (SCEX-FR-004a stands).
  • MCPX-FR-020: Until an explicit architecture amendment states otherwise, MCP clients and all LLM/VLM providers are locally deployed inside the enterprise trust perimeter. PII exposure to those local providers is an accepted residual risk and is not blocked by masking. Credentials, cookies, access/refresh/service tokens, raw secrets, and raw storage paths MUST remain absent from tool payloads, telemetry, provenance, and artifact references.
  • MCPX-FR-021: Dynamic client registration MUST be rate-limited, fully audited, redirect-URI validated, constrained to reviewed first-party scopes, and immediately revocable. DCR must not grant database permissions or allow scope escalation beyond the principal's live RBAC.
  • MCPX-FR-022: The catalog MUST provide the persistent AgentAuthoringWorkspace operations create_authoring_session, propose_test_plan, start_exploration, get_exploration_result, propose_graph_revision, get_graph_diff, and promote_to_scenario (or an explicitly versioned equivalent) with server-owned state, provenance, idempotency and CAS.
  • MCPX-FR-023: Authoring exploration MUST use an isolated, allowlisted, time/size/network-bounded and cancellable sandbox with durable receipts and artifact ownership; shell, credential, filesystem/network escape and production side effects MUST be rejected. Sandbox readiness is a release gate, not an implementation claim.
  • MCPX-FR-024: MCP MUST pass authoring output only as typed 038 action candidates/graph proposals through deterministic compile/validate, user-reviewed diff, and 042 handle-based save. Raw code, browser URLs, cookies, secrets, filesystem paths and caller digests MUST never establish authority; code-backed production execution is a separate unimplemented contract.
  • MCPX-FR-025: MCP MUST provide bootstrap_authoring_scenario for a first scenario. It accepts only a bounded InitialScenarioIntent and server-issued valid compiled/draft-pack handles; atomically creates the registry entry, initial current revision, provenance/outbox and a workspace bound to that revision. It MUST NOT require or accept a client-created base revision, graph, content hash or revision ID.
  • MCPX-FR-026: MCP MUST expose the SCAUTO-FR-017 automation operations with identical REST policy/RBAC behavior. Schedule mutations require an explicit active revision, environment and idempotency key; they reject HumanStep revisions before any durable side effect and preserve the 046 PROD gate before dispatch.

Canonical Scenario MCP Names

The following snake_case names are the canonical MCP names for the complete authoring, persistence, activation, and run-target chain. The operationIds in the final column are transport-specific REST aliases only; they are not MCP names and MUST NOT be exposed by tools/list.

Canonical MCP name Responsibility Legacy REST operationId alias(es)
inspect_dashboard_context Inspect dashboard/query context inspectDashboardQueryModel
bootstrap_authoring_scenario Atomically create a first scenario, initial current revision and bound workspace from server-owned handles scenarioRegistry.create
create_authoring_session Create persistent authoring workspace none; MCP-only workspace operation
propose_test_plan Record typed test-plan intent none; MCP-only workspace operation
start_exploration Start bounded authoring exploration none; MCP-only workspace operation
get_exploration_result Read exploration result/artifact refs none; MCP-only workspace operation
propose_graph_revision Submit typed graph proposal scenarioEditor.agentPropose, scenarioEditor.apply
get_graph_diff Read server-computed proposal/revision diff scenarioRegistry.revisionDiff
promote_to_scenario Create/update validated server handles or a save request; never activate scenarioEditor.proposalSave, scenarioRegistry.create
scenario_compile Compile canonical ScenarioGraph compileDashboardScenario, compileScenario
scenario_validate Validate compiled graph validateDashboardScenario, validateScenario
scenario_resolve Resolve typed authoring inputs into a new compiled handle resolveDashboardScenario, resolveScenario
generate_draft_pack Generate server-owned draft pack compileScenarioDraftPack, buildScenarioDraftPack
request_save Create immutable candidate revision, or eligible initial current revision editor.save, scenarioRegistry.create
activate_revision CAS-activate an eligible materialized revision scenarioRegistry.activateCurrentRevision
start_scenario_run Start a run pinned to an explicit promoted revision scenarioRun.start
list_scenario_schedules / upsert_scenario_schedule / delete_scenario_schedule Manage schedule projections through 046 policy automation.listSchedules, automation.upsertSchedule, automation.deleteSchedule
list_scenario_trigger_rules / upsert_scenario_trigger_rule / delete_scenario_trigger_rule Manage event trigger rules through 046 policy automation.listTriggerRules, automation.upsertTriggerRule, automation.deleteTriggerRule
get_scenario_automation_policy / upsert_scenario_automation_policy / get_scenario_automation_metrics Read or manage automation policy and observability automation.getPolicy, automation.upsertPolicy, automation.metrics

promote_to_scenario does not activate and does not advance current_revision. It creates or updates server-owned validated handles or a server-stored save request. request_save creates a candidate; only the separate activate_revision operation can advance current_revision, after eligibility, materialization, policy, required approval, and CAS checks. Initial creation is the sole explicit exception when the initial eligible revision is created as current. start_scenario_run must receive an explicit promoted revision_id and verified content_hash, never a candidate or an implicit latest/current target.

Key Entities

  • McpToolCatalog: Versioned, explicitly registered set of tools grouped by domain; source of truth for names, schemas, risk class and required_permission(resource, action).
  • McpSessionPrincipal: Authenticated caller identity (user or service), role set resolved live from DB per invocation, and delegation scope.
  • OAuthClientRecord: Dynamically or statically registered client (public/confidential), first-party scope bound, owner visibility, revocation state.
  • ApprovalRequiredEnvelope: Typed result binding a refused-or-deferred action to its durable gate (gate_id, targets, risk, reason_required).
  • ToolInvocationRecord: AgentAction-provenance row per MCP tool call (principal, tool, argument digest, outcome, run linkage).
  • HandoffSurface: Frontend instruction/copy affordance replacing the «Создать сценарий тестирования» chat launch (dashboard context parameters + connection hint).
  • AgentAuthoringWorkspace: Persistent server-owned MCP co-authoring session with CAS state, proposals, exploration results, artifacts, review and promotion references.
  • AuthoringArtifact: Bounded server-owned source/patch/trace/screenshot/diagnostic/receipt reference used for review, never an execution program or production authority.

@{ McpInterface.ScenarioPipeline [C:5] [TYPE ADR]

@BRIEF Normative ownership and provenance contract for the MCP-to-scenario-run path. @RELATION DEPENDS_ON -> [ScenarioGraph.Compiler.Compile] @RELATION DEPENDS_ON -> [ScenarioRegistry.RevisionChain] @RELATION DEPENDS_ON -> [ScenarioExecution.RunnerPlan.Derive]

AgentAuthoringWorkspace MCP operations

The MCP catalog MUST expose bootstrap_authoring_scenario for initial creation and the following target contract for persistent co-authoring of an existing revision: create_authoring_session, propose_test_plan, start_exploration, get_exploration_result, propose_graph_revision, get_graph_diff, and promote_to_scenario. Names MAY be versioned during implementation, but the operation semantics and ownership boundary are normative.

All mutable workspace, proposal, exploration, artifact, review and promotion state is server-owned. Mutating operations require authenticated principal/delegation, idempotency key and expected workspace CAS version; replays return the original result and changed requests or stale versions return typed 409 without partial mutation. Read operations enforce object ACL and return bounded refs/digests for large traces/screenshots.

start_exploration uses only the isolated sandbox contract: allowlisted origins/APIs/actions, no shell/credential/filesystem escape, bounded time/size/network, cancellation, operation receipts, artifact ownership and no production side effects. The MCP server does not generate arbitrary Playwright code; it exposes reviewed templates/actions and consumes sandbox traces/candidates. Raw code, URLs, cookies, secrets, paths and caller digests are rejected as authority.

The initial chain is inspect -> compile/validate -> draft-pack -> bootstrap_authoring_scenario -> immutable current revision. The existing-scenario chain is session -> plan -> exploration -> typed proposal -> 038 compile/validate -> user-reviewed diff -> 042 save/promotion -> immutable revision. An unbound workspace is not a creation path and cannot propose or save a graph. promote_to_scenario cannot skip user review or required approval. 044 accepts only the resulting promoted revision. Code-backed production execution is a separate future, unimplemented contract.

The following table is the single cross-spec stage contract. A handle is an opaque server-issued reference; a digest is computed and verified by the owner named in the table. A client may pass a handle or request digest for correlation, but it cannot establish authority with a graph, caller-computed digest, path, or file.

Implementation status (audit 2026-09-06, Doc.Adr.ADR0023; Phase 2c T029d–T029h in tasks.md). The deterministic cores of compile/validate/resolve/draft-pack are IMPLEMENTED as pure 038 functions. The durable handle layer is IMPLEMENTED: CompiledScenarioHandle/ValidationResultHandle/DraftPackHandle entities (migrations 0019_scenario_handles, 0020_scenario_materialization, 0021_context_authority) are minted at the persisted boundaries — REST api_compile_scenario/api_validate_scenario/api_resolve_scenario/ api_draft_pack and the MCP register_draft_pack write tool — with content-addressed canonical-bytes storage, owner binding, single consumption (SELECT ... FOR UPDATE, PostgreSQL-proven) and binding/digest/staleness rejections (ScenarioGraph.Handles). bootstrap_authoring_scenario and 042 CreateScenario/CreateInitial accept ONLY stored handle ids; the transitional caller-composed compile:{run_id}:{digest} string path is removed for bootstrap. 042 create materializes graph_snapshot = the canonical DashboardTestScenario JSON (plus server-owned action-registry identity and context_authority) inside the create transaction, so a freshly bootstrapped current revision is directly runnable/schedulable; BOOTSTRAP_REVISION_NOT_RUNNABLE survives as defense-in-depth for provenance-only legacy rows. The OutboxEvent/RevisionMaterialization reference-artifact worker exists (materialize_pending_revisions); wiring it into the scheduler poll loop is the remaining operational step. Residuals: the 043 editor promote/save path and the REST-only api_draft_pack flow do not yet evaluate context_authority (persisted as NULL = legacy-allowed at the PROD gate). The inspect/context row is resolved as a hybrid (T029h, 2026-09-06): no persisted InspectionContextHandle is minted; the MCP read tool inspect_dashboard_context exposes the live authoritative DashboardQueryModel resolver, and the register_draft_pack boundary binds the client-carried context by recomputing its fingerprint against a live inspection (context_authority = verified/unverified; sentinel fingerprints never verify; unreachable environments fail open; PROD dispatch refuses an explicit non-verified marker via CONTEXT_AUTHORITY_REQUIRED_FOR_PROD; the validator recursively rejects query_context/SQL smuggling inside dashboard_context). The table remains the normative target; the InspectionContextHandle row records the declined persisted-handle alternative, superseded by the hybrid binding above.

Stage Input Authoritative owner Output handle Digest Persistence Failure / no-side-effect rule Test id
inspect/context dashboard/environment request 038 context/capability resolver, with server ACL InspectionContextHandle server context fingerprint server request record only unresolved facts become needs_context, needs_selector, or needs_baseline; no graph/revision/run Test.Scenario.Capability, Test.Scenario.Resolver
compile bounded intent + server context 038 ScenarioGraph.Compiler CompiledScenarioHandle server content_hash of canonical graph/program immutable server compile result deterministic validation first; errors persist no mutable artifact Test.Scenario.Compiler, Test.Scenario.Serializer
validate CompiledScenarioHandle or server graph 038 ScenarioGraph.Validator ValidationResultHandle server validation-result digest bound to compiled hash immutable server validation result invalid/blocking findings produce no draft, revision, gate, or run Test.Scenario.Validator, Test.Scenario.Validator.Edge
resolve server compiled handle + typed resolution + expected base 038 resolver; 042 owns revision CAS when persisted new CompiledScenarioHandle new server content_hash, parent hash link immutable server compile result; no in-place graph edit stale base is 409 STALE_REVISION; no partial graph or revision mutation Test.Scenario.Resolver, Test.Scenario.Resolver.Edge
draft-pack valid compiled/validation handles + registered templates 038 ScenarioGraph.PackCompiler DraftPackHandle server draft_pack_digest over rendered manifest/bytes server draft registry, save_eligible or preview_only unresolved/error/unsafe input is preview_only; no registry revision or launch side effect Test.Scenario.Pack, Test.Scenario.Pack.Security
initial bootstrap InitialScenarioIntent + valid compiled/draft-pack handles + idempotency key 042 ScenarioRegistry.CreateInitial ScenarioRegistryEntry + initial current ScenarioRevision + bound workspace server-verified content hash and draft digest PostgreSQL registry/revision/workspace + outbox invalid, stale, unauthorized or replay-conflicting input creates no partial row; no synthetic base revision test_mcp_initial_scenario_e2e
request-save compiled/draft handles + expected base + idempotency key 042 registry transaction ScenarioRevision (candidate, or initial current) server-verified graph/content hash and draft digest PostgreSQL revision + outbox; materialization async CAS/idempotency failure is atomic; no row/outbox on rejection; save never activates test_scenario_registry, test_revisions
run preflight selected ScenarioRevision + typed bindings + server target snapshot 044 RunPreflight RunPreflightHandle server preflight digest over revision/target/bindings/policy durable preflight record attached to launch request missing/invalid context, selector, baseline, policy, or revision mismatch blocks before dispatch/I/O test_scenario_runner, test_scenario_automation_api
runner-plan eligible preflight + immutable revision 044 RunnerPlan.Derive RunnerPlanHandle server descriptor/plan digest persisted on ScenarioRun; git runner.plan.json is reference only altered/unknown descriptor rejects before lease, execution, or provider I/O test_scenario_runner_plan, test_scenario_dispatch
launch request RunnerPlanHandle + idempotency key 044 server runner ScenarioRun server launch/request digest and pinned plan digest PostgreSQL run, queue/gate, audit same key+hash replays same run; changed request is 409 IDEMPOTENCY_KEY_REUSED; no duplicate dispatch test_scenario_queued_dispatch, test_scenario_scheduler_callbacks
dispatch/evidence pinned plan descriptor + capacity lease 044 executor/provider and evidence owner EvidenceRef / StepOutcome provider-verified receipt SHA-256 Artifact(owner_type=scenario_run) + immutable receipt + step outcome missing ownership/digest/ref is non-pass; late/unknown effect reconciles or terminalizes non-pass test_scenario_terminal_signals, test_provider_contract
signal terminal run + immutable evidence provenance 044 producer; 047 queue/episode projection idempotent InvestigationSignal / queue input server signal identity digest durable signal; 047 projection only failed/blocked/inconclusive emits once; passed emits none; never opens chat/action automatically test_scenario_terminal_signals

Handle and Digest Rules

  • CompiledScenarioHandle is minted only after 038 canonical serialization. Its content_hash is SHA-256 of the executable canonical Verification Program graph; it excludes timestamps, display fields, ParameterBindings, and runtime identity. 038 never mints scenario_id or revision_id.
  • ValidationResultHandle is valid only for the exact compiled-handle hash and validator/schema versions used to produce it. A validation digest does not become graph authority and cannot be supplied by a caller as proof of validity.
  • DraftPackHandle is minted by the server pack registry from registered, versioned templates. Its digest covers the server-rendered manifest and bytes; save_eligible requires valid validation and no unresolved required marker.
  • 042 alone mints scenario_id, revision_id, parent links, and server-owned revision records. Save verifies the compiled hash and draft digest against stored handles inside one transaction; client graph uploads are rejected.
  • RunPreflightHandle and RunnerPlanHandle are derived by 044 from persisted revision data, server target/policy context, and typed bindings. Their digests are server recomputed; a caller-supplied runner or digest is correlation data.
  • runner.plan.json and scenario.yaml are materialized reference artifacts. They can be regenerated and are never authority for validation, revision, preflight, dispatch, retry, recovery, or evidence.
  • Status (2026-09-06, updated after Phase 2c): CompiledScenarioHandle, ValidationResultHandle and DraftPackHandle ARE minted and stored (038 ScenarioGraph.Handles, migrations 0019–0021); the rules in this section are their implemented acceptance criteria. InspectionContextHandle is NOT persisted — T029h resolved the inspect stage as a live fingerprint binding (context_authority) instead. The authority that 042 materializes into ScenarioRevision.graph_snapshot is the persisted canonical DashboardTestScenario serialization behind CompiledScenarioHandle.canonical_bytes_ref — explicitly NOT the lossy pack templates above (038 ScenarioGraph.ServerOwnedPipeline, amendment 2026-09-06).

Resolution and Continuation Semantics

038 owns inspection and semantic resolution. needs_context means a required dashboard/query/change-request fact is absent; needs_selector means a browser target/action selector is unknown; needs_baseline means an approved baseline reference is absent or stale. None may be guessed or silently defaulted. 038 may compile a WorkingDraft/preview, but only 042 can save a revision and only 044 can reject launch bindings at RunPreflight.

Save is draft/working -> validated save request -> candidate, with the initial revision as the sole exception (candidate -> current at creation). Activation is a distinct CAS operation: candidate -> current, guarded by eligibility, materialization, policy and required approval. Approval continuation is pending_approval -> queued; deny/expiry is pending_approval -> blocked. Every continuation consumes its gate/checkpoint exactly once. Expected revision, gate, checkpoint and idempotency versions are CAS inputs; stale values return a typed 409 and create no side effect. Same idempotency key plus the same canonical request hash replays the existing result; the same key plus a different hash is 409 IDEMPOTENCY_KEY_REUSED.

The execution chain is strictly: ScenarioRevision -> RunPreflight -> RunnerPlan -> ScenarioRun -> queued|pending_approval -> dispatch -> EvidenceRef/StepOutcome -> 047 signal. ScenarioRevision owns immutable program input, RunPreflight owns launch eligibility, RunnerPlan owns descriptor order/policy, ScenarioRun owns the execution snapshot, 044 owns step/evidence receipts, and 047 owns queue/episode projection. No client graph, caller digest, or materialized plan can replace these owners.

@{ McpInterface.ScenarioPipeline.ReleaseGates [C:5] [TYPE Block]

@BRIEF Conjunctive release gates for the coordinated pipeline contract.

Release is GO only when every applicable gate passes: 038 fixture/schema and deterministic compiler/validator/serializer/pack checks; 042 registry create, revision CAS, idempotency, outbox and materialization checks; 044 preflight, plan, queued approval, exact dispatch, evidence ownership and terminal signal checks; 050 parity, bounded transport, hidden-vs-gated and end-to-end MCP checks; and structural anchor/Markdown audit. Partial green is NO-GO and does not mark implementation tasks complete.

Evidence commands:

cd backend && source .venv/bin/activate && python -m pytest -q tests/services/dashboard_testing/scenario/test_*.py tests/api/test_scenario_runs_api.py tests/api/test_scenario_automation_api.py tests/api/test_scenario_analytics_api.py
cd backend && source .venv/bin/activate && python -m ruff check src/services/dashboard_testing/execution src/api/routes/dashboard_testing/scenario_runs.py
python specs/044-dashboard-scenario-execution/prototype/validate_static.py
cd frontend && npm run test -- --run && npm run lint && npm run build

The required MCP parity and end-to-end evidence is test_mcp_* plus the 050 SC-001/SC-002 walkthrough; absent or partial evidence keeps the gate NO-GO.

Authoring E2E Traceability

These are required evidence rows for the authoring promotion path. They remain open until executable evidence is retained; unchecked rows keep this release gate NO-GO and do not mark T023 or T028 complete.

Evidence row Required trace Evidence required Status
E2E-AUTH-001 create_authoring_session -> propose_test_plan -> start_exploration -> get_exploration_result External MCP client trace proves persistent server-owned workspace, bounded exploration, durable receipt, and typed artifact refs [x] CLOSED 2026-09-03 — tests/test_mcp_authoring_promotion_e2e.py::test_sandbox_output_promotes_to_immutable_revision (plan → queued exploration → sandbox dispatch → exploration_passed, evidence draft:exploration-*, bounded projection)
E2E-AUTH-002 propose_graph_revision -> scenario_compile -> scenario_validate -> get_graph_diff External MCP client trace proves typed proposal, deterministic 038 validation, server-computed diff, and explicit user review boundary [x] CLOSED 2026-09-03 — tests/test_mcp_scenario_e2e.py (propose → promote validation → 038 inspect_scenario/scenario_resolve/validate_scenario over MCP → get_graph_diff; canonical tool names inspect_scenario/validate_scenario per FR-008 rename clause)
E2E-AUTH-003 promote_to_scenario -> request_save -> activate_revision -> start_scenario_run External MCP client trace proves handle-based save creates candidate, activation separately passes eligibility/materialization/policy/approval/CAS, and the run pins promoted revision_id + content_hash [x] CLOSED 2026-09-03 — tests/test_mcp_scenario_e2e.py extended: post-activation start_scenario_run creates a queued run pinned to scenario_revision_id + scenario_content_hash of the promoted revision

@} McpInterface.ScenarioPipeline.ReleaseGates

@} McpInterface.ScenarioPipeline

Success Criteria

  • SC-001: An external client completes the full scenario creation chain for a fixture dashboard and the revision appears in the registry with correct provenance.
  • SC-002: 100% of parity tests pass: MCP tool outcomes match legacy @tool wrapper outputs on shared fixtures.
  • SC-003: 100% of gated invocations produce zero side effects before approval; denial/expiry paths refuse cleanly.
  • SC-004: tools/list is RBAC-exact for admin/analyst/viewer roles in fixture tests.
  • SC-005: After decommission, builds and link-integrity suites pass with zero /agent, assistant-API or Gradio-proxy references; the stack starts without port 7860.
  • SC-006: No catalog path reaches arbitrary SQL or raw Superset mutation outside the governed tools; scenario contexts cannot invoke SQL-class tools.
  • SC-007: A compliant external client completes discovery (RFC 9728) → DCR → authorization code + PKCE → tools/list in one automated flow, with zero manual token management.
  • SC-008: Refresh-token replay revokes the token family in 100% of fault-injection cases; pre-revocation access tokens die at expiry, not silently extended.
  • SC-009: A role change is reflected in the next tools/list and the next invocation without new consent; a revoked permission hides the tool and denies cached-catalog calls.

Clarifications

Session 2026-08-24

  • Q: Separate MCP process or mounted in backend? → A: Mounted in the FastAPI app; simplest deployment, in-process service reuse, shared auth middleware.
  • Q: Does removing chat remove HITL? → A: No, and it does not force the web UI either (session 2026-08-24): gates are server-durable and fully resolvable inside MCP (list_pending_approvals + decide_approval); web gate cards on product pages remain an equal renderer of the same rows for browser-first users.
  • Q: Is the vendored mcp-superset adopted? → A: No — reference only; it bypasses the policy layer.
  • Q: What happens to AgentRun/DraftArtifact/InvestigationSignal contracts? → A: They persist unchanged; MCP invocations create the same provenance rows. Only the conversational transport and its UI retire.
  • Q: OAuth now or later? → A: Now (session 2026-08-24 decision). The backend becomes a full OAuth 2.1 Authorization Server in Phase 0 — Authlib==1.6.6 is already a dependency; personal access tokens are explicitly deferred, not rejected forever.
  • Q: Where do rights live? → A: In the existing DB RBAC (Role/Permission, user_has_permission). Tokens never embed roles; the catalog binds tools to canonical resource:action pairs and filters both listing and invocation through the same predicate.

Phases

Phase Scope Exit evidence
0 OAuth AS+RS skeleton: authorize/token/DCR/JWKS/metadata endpoints, /mcp validation middleware, 2–3 probe tools, Inspector + scripted-client connectivity SC-007 green in CI
1 Parity catalog for the 37 tools + contract tests mirroring wrapper tests; permission-bound hidden/gated matrix SC-002, SC-004, SC-009
2 Gates/provenance over MCP; end-to-end scenario creation walkthrough; refresh rotation/reuse tests SC-001, SC-003, SC-008
3 Frontend decommission behind flag; handoff surface; assistant API retirement SC-005 partial
4 Delete agent/ service; run.sh/compose updates; spec amortization closed SC-005 full

Spec Impact & Amortization Map

Spec Amendment
036 Transport rejection recorded; durable-runtime contracts carry over to MCP provenance
037 Baseline tools join the MCP catalog unchanged (thin forwarders)
038 Authoring entry becomes MCP-callable; validator remains sole gateway
039 AGUI-FR-001..013 chat workspace superseded; entry action becomes HandoffSurface
040 Delegated load experiments readable as MCP-driven, policy unchanged
041 Blast-radius explanation consumed by external clients; index contracts unchanged
042 Delegated actors explicitly include MCP principals; registry contracts unchanged
043 "Edit with agent" proposals callable via MCP; SCEDIT-FR-009 stands
044 Runner unaffected; authoring/investigation boundary inherited by MCP clients
045 "Investigate with agent" targets an external client session
046 Schedule management exercisable through MCP tools under same policy
047 Case threads continue in external clients; server keeps evidence/timeline

#endregion McpInterface.Spec


TASKS — Implementation Tasks

Source: tasks.md

050-mcp-interface — Tasks

Правило: [ ] не начата; [~] в работе; [x] только с доказательством (команда + вывод).

Phase 0 — OAuth skeleton + /mcp probe

  • T001 Зависимость MCP SDK/FastMCP в backend/requirements.txt; каркас backend/src/mcp_server/ с монтированием /mcp в FastAPI app.
  • T002 Authorization Server на Authlib: /oauth/authorize (поверх существующей web-сессии: local password + ADFS OIDC), JWKS, AS metadata (RFC 8414).
  • T003 /oauth/token: authorization_code + PKCE S256 (обязателен для публичных клиентов), refresh_token с ротацией, client_credentials для service-principal.
  • T004 DCR (RFC 7591): регистрация публичных клиентов с first-party scope bound; список/отзыв в Admin.
  • T005 Resource Server: RFC 9728 metadata для /mcp, 401 + WWW-Authenticate, валидация access-JWT (aud=mcp, jti/blacklist) на каждый запрос; resolve_principal() адаптер.
  • T005a Transport and domain input limits: server-owned request byte/JSON-depth/collection/string limits, ScenarioGraph and SQL/DSL size limits, and per-session rate limits; rejected payloads create no ToolInvocationRecord, gate, draft, run, or mutation.
  • T006 Refresh rotation + reuse-detection: повтор ротированного refresh токена отзывaет семейство (расширение TokenBlacklist); fault-injection тесты.
  • T007 2–3 пробных инструмента (read-only: list_environments, get_health_summary, search_dashboards) с явной регистрацией и курируемыми схемами.
  • T008 CI: скриптованный клиент проходит discovery → DCR → PKCE → token → tools/list без ручных шагов (SC-007); подключение MCP Inspector; документация в README/INSTALL. Доказательство (remediation round 1, 2026-09-03): pytest tests/test_mcp_client_flow_http.py -q → 2 passed — оба потока целиком поверх реального HTTP: (a) machine-поток: RFC 9728 discovery → AS metadata (grant client_credentials, client_secret_post) → confidential DCR с одноразовым client_secret → client_credentials → signed aud=mcp токен с principal_type=service → /mcp initialize → tools/list (gated/human-only тулы невидимы) → tools/call; (b) пользовательский поток: public DCR → PKCE S256 GET /oauth/authorize с Bearer веб-сессии (документированный SPA-mediated consent-путь) → 302 code → обмен → identity-only токен → /mcp tools/list с live-RBAC (RUN_PROD-тул скрыт для RUN-юзера); wrong secret → invalid_client. Реализовано: грант client_credentials (services/mcp_oauth.py), oauth_clients.secret_hash (миграция 0017, idempotent-guard), service-principal short-circuit в McpTokenVerifier, AS metadata расширена. Документация: INSTALL.md §«MCP клиент». MCP Inspector подключается по тому же discovery-контракту (Streamable HTTP /mcp). Ограничение браузера-без-Bearer (cookie-consent HTML): РЕШЕНО 2026-09-04 — не строим в 050. Документированный и протестированный путь — SPA-mediated authorize: аутентифицированный фронтенд вызывает GET /oauth/authorize с Bearer веб-сессии (SC-007 «zero manual token management» выполнен: DCR → PKCE → code → токены автоматизированы полностью). Отдельная интерактивная consent-страница (login+consent на cookie-сессии) — новая frontend-поверхность со своим UX/i18n/security-ревью, непропорциональная first-party perimeter scope (MCPX-FR-020: только локальные клиенты). Fallback для browser-only клиента описан в INSTALL.md §«MCP клиент». Если browser-only third-party клиенты станут требованием — это отдельное продуктовое решение через architecture amendment (MCPX-FR-020 stand).
  • T008a DCR abuse controls: rate limits, redirect-URI validation, reviewed first-party scope bounds, admin audit/revocation, and tests proving registration cannot grant permissions or escalate scopes.
  • T008b Local-perimeter deployment tests and docs: reject or disable non-local MCP/LLM/VLM endpoint configuration by default; assert PII is permitted only for configured enterprise-local providers while credentials/cookies/tokens/secrets remain rejected everywhere. (Closure review 2026-09-03: тестов на отклонение non-local эндпоинтов не найдено; секрет-гигиена provenance подтверждена — PARTIAL.) Доказательство (remediation round 5, 2026-09-04): deny-by-default guard Core.EndpointLocality (src/core/utils/endpoint_locality.py) в choke points LLMProviderService.create_provider/update_provider — типизированный отказ endpoint_not_local:<reason> ДО персистенции (create не добавляет row, update не мутирует), route-слой маппит в HTTP 400 (не 500). Локальность: loopback/local-литералы, RFC1918/ULA private IP, enterprise DNS-суффиксы (.local/.internal/.lan/.corp/.intranet, расширяемо), DNS-имена с полностью приватным резолвом; нерезолвимое = fail closed; substring-spoofing (https://api.openai.com/localhost) отклоняется по hostname; пустой base_url (публичное SDK-облако по умолчанию) отклоняется. Escape hatches env-документированы (LLM_ALLOW_NONLOCAL_ENDPOINTS, LLM_NONLOCAL_ENDPOINT_ALLOWED_HOSTS, LLM_LOCAL_HOST_SUFFIXES), по умолчанию закрыты. MCP-транспорт: server-owned JSON-depth limit (typed 400 pre-dispatch) + per-session rate limit с Retry-After (E6) + существующие body-limit/DNS-rebinding defaults. Docs: INSTALL.md §«Локальный периметр». Тесты: tests/test_endpoint_locality.py (22) + tests/test_mcp_transport_limits.py (5) + provider/route slices → pytest tests/test_endpoint_locality.py tests/services/test_llm_provider.py tests/api/test_llm.py tests/api/test_encryption_health.py -q → 107 passed. PII-часть (секрет-гигиена provenance) подтверждена closure review round 1 и E13-отклонениями authoring chain.

Phase 1 — Parity catalog (37 tools)

  • T010 Каталог-реестр инструментов по доменам: env/health/tasks, git/deploy/migration/backup/maintenance, superset ops, baseline, scenario. Каждый инструмент декларирует required_permission(resource, action) из канонического словаря; единая серверная политика поглощает _tool_filter.
  • T011 Инструменты env/health/tasks (list_environments, get_health_summary, get_task_status, maintenance CRUD, llm status).
  • T012 Инструменты git/deploy/migration/backup. Доказательство: pytest tests/test_mcp_ops_parity.py tests/test_mcp_server.py tests/test_mcp_approvals.py tests/test_mcp_maintenance.py tests/test_mcp_oauth.py -q → 70 passed; зарегистрированы create_branch, commit_changes, deploy_dashboard, execute_migration, run_backup, run_llm_documentation, run_llm_validation (curated inputs, approval-gated; reviewed dispatchers src/services/mcp_ops_dispatch.py в явной цепи poll_approved_mcp_dispatches: GitService / TaskManager git-integration, superset-migration, superset-backup, llm_documentation, ValidationTaskService).
  • T013 Superset-инструменты (databases/explore/sql/format/permissions/dashboard+dataset CRUD) — SQL-класс помечен risk-классом и отдельным правом. Доказательство: тот же срез 70 passed; superset_list_databases, superset_explore_database, superset_format_sql, superset_audit_permissions (read), superset_create_dashboard, superset_copy_dashboard, superset_create_dataset (approval-gated), superset_execute_sql — permission ("plugin:superset_sql","EXECUTE") добавлен в rbac_permission_catalog.discover_declared_permissions + default-deny mapping; dangerous SQL отклоняется клиентским guard. Доп. свидетельство (орт. аудит 2026-09-02): guard расширен (safety.py: INTO/CALL/SET/REFRESH/VACUUM/REINDEX/ATTACH/LOAD/PREPARE/COMMENT + опасные функции lo_import/lo_export/pg_read_file/pg_write_file/pg_ls_dir/dblink/dblink_exec/set_config/pg_sleep + запрет мульти-стейтментов через ; после очистки строк/комментариев; tests/test_core/test_superset_safety.py — 50+ кейсов); PROD-критерий SQL-класса выравнен с resolve_environment_execution_policy (is_production OR stage=PROD) и переведён в терминальный отказ production_sql_execution_rejected вместо одобряемого, но недиспетчеризуемого гейта; pytest tests/test_mcp_ops_parity.py tests/test_core/test_superset_safety.py -q → 85 passed; полный срез аудита 214 passed; полный бекенд 11202 passed.
  • T014 Baseline-домен: capture_baseline_candidate, request/decide/consume_baseline_approval, create_verification_run (паритет fixtures с tools_037). Доказательство: тот же срез 70 passed; прямые обёртки над candidate_capture/candidates.request|decide_approval/verification_service.create_verification_run_async (те же сервисы, что и 037 REST-поверхность); consume_baseline_approval — baseline publish: approval-gated + reviewed dispatcher через one-shot consume_approval; полный backend suite python -m pytest -q → 11167 passed, 240 skipped, 1 xpassed; ruff/compileall чистые; каталог 45/45 зарегистрирован.
  • T015 Scenario-домен: scenario_compile/scenario_validate/scenario_resolve/generate_draft_pack/request_save/activate_revision/start_scenario_run (паритет с tools_038 и 042/044; save только из server-stored draft, activation отдельная CAS-операция).
  • T016 Контрактные тесты паритета: выводы MCP-инструментов сопоставлены с legacy-обёртками на общих фикстурах (SC-002). Доказательство: pytest tests/api/test_mcp_parity_baseline_037.py -q → 3 passed: request/decide baseline-approval — MCP-gate валидируется моделью ApprovalGateResponse и совывает с REST по operation/risk_level/required_permission/status/reason_required/target_paths + actor parity на decide; create_verification_run — идентичные overall_status/category statuses/evidence_refs/created_by против REST /verification-runs; superset_format_sql — идентичная строка против legacy REST /api/agent/superset/sqllab/format (единственный дабл — [EXT:Superset] клиент). Попутно пойман и закрыт регрессией дефект 038-резолвера: scenario_resolve selector-путь падал на step.description is None (TypeError) — теперь selector_hint пишется без конкатенации с None (test_selector_hint_on_step_without_description).
  • T017 Bounded-response дисциплина: лимит инлайн-ответа, артефакты как ref+digest.
  • T018 Hidden/gated матрица: admin/analyst/viewer × каталог — отсутствие права скрывает инструмент из tools/list; role-change виден на следующем вызове без re-consent (SC-004, SC-009).

Phase 2 — Gates & provenance over MCP

  • T020 AgentAction-provenance на каждый вызов (ToolInvocationRecord): principal, tool, digest аргументов, исход, связь с AgentRun.
  • T021 approval_required конверт для негelegированных действий; durable ActionApprovalGate без side effects до решения (SC-003).
  • T022 Gate-инструменты как первичный путь полного цикла в клиенте: list_pending_approvals(filter) (+гейты от автоматизации) и decide_approval(gate_id, confirm|deny, reason) с CAS, обязательным reason для high-risk confirm, отказом service-principals и typed-ошибками на expired/replayed; web gate cards — равноправный рендер тех же строк (SC-003).
  • T023 E2E-walkthrough: внешний клиент создаёт сценарий фикстурного дашборда end-to-end → revision в registry (SC-001). Доказательство: pytest tests/test_mcp_scenario_e2e.py -q → 2 passed: один человеческий принципал с реальным scenario:EDIT/scenario:RUN (без RBAC-стабов) проходит tools/list-видимость → create_authoring_session → propose_graph_revision (server-derived op) → get_graph_diff → promote_to_scenario (awaiting_user_review) → request_save → ScenarioRevision(candidate) в registry → 038 inspect_scenario/scenario_resolve/validate_scenario поверх MCP → activate_revision → entry.current_revision_id продвинут отдельным CAS; второй тест — принципал без scenario:EDIT не видит и не может вызвать save/activate. Транспортный уровень (initialize → mcp-session-id → tools/list → tools/call) покрыт скриптовым клиентом T008.
  • T024 Gated-вызовы по контексту: PROD-окружение и baseline publish возвращают approval_required при видимом инструменте (hidden-vs-gated, MCPX-FR-018).
  • T024a Context revalidation: environment-policy and provider-binding fingerprints are rechecked at dispatcher admission and before provider I/O; target reclassification or binding drift invalidates a prior gate without I/O (SCEX-FR-026).
  • T025 Checkpoint-инструменты: list_checkpoints + decide_checkpoint(run_id, disposition, expected_version) через тот же CAS/аудит, что и монитор 045; user-principal only (service → permission_denied); тест что автоматизация не имеет пути к чекпоинтам (MCPX-FR-019). Доказательство (после closure-review داунгрейда, 2026-09-03): pytest tests/test_mcp_checkpoints.py -q → 3 passed: каталог ("scenario","RUN") + service_allowed=False на оба инструмента; сервис-принципал не видит и не может вызвать; list-проекция pending/decided чекпоинтов рана; decide повторяет REST /human/decision (pending-резолв на сервере, CAS decision_version, continue_after_human_decision, step→passed/HUMAN_*, run→queued/executing); stale expected_version → conflict без потребления; отсутствие pending/рана → checkpoint_not_found. (Closure review 2026-09-03 зафиксировал прежний фиктивный [x]; закрыт реальной реализацией.)
  • T026 AgentAuthoringWorkspace operations: create_authoring_session, propose_test_plan, start_exploration, get_exploration_result, propose_graph_revision, get_graph_diff, promote_to_scenario; persistent server-owned state, provenance, idempotency and workspace CAS.
  • T027 Authoring sandbox contract tests: isolated runtime, origin/API/action allowlists, no shell/credential/filesystem escape, limits, cancellation, receipts, artifact ownership and zero production side effects; retain exploratory traces/screenshots/diagnostics as bounded refs.
  • T028 Authoring promotion E2E: sandbox output -> typed proposal -> deterministic 038 compile/validate -> user diff review -> 042 handle-based save -> immutable revision; reject raw code/URLs/cookies/secrets/paths/caller digests and keep code-backed production execution unimplemented. Доказательство: pytest tests/test_mcp_authoring_promotion_e2e.py -q → 8 passed: полный прогон через tools/call — start_exploration теперь ставит queued при зарегистрированном раннере (wiring provider_available=get_registered_runner() is not None; регистрация default_runner+deployment context в bootstrap_live_execution_composition) → реальный execute_scheduled_exploration_dispatch из потокового контекста планировщика → exploration_passed с observations/proposed_graph/evidence_ref (draft:exploration-*) при [EXT:Browser]-дабле; get_exploration_result остаётся ограниченной проекцией без утечки наблюдений; значение observed_dashboard_title в сохранённой ревизии выводится только из наблюдения; propose→promote→diff→request_save→activate завершается current-ревизией без продвижения до отдельного CAS-активации. Негативные ветки: typed ops с SQL/..//\\/drop отклоняются до предложения и CAS; exploration-spec с code-токенами и незарегистрированными действиями не персистируется; GraphRevisionInput/PromoteScenarioInput/RequestSaveInput/ActivateRevisionInput/ExplorationInput структурно отвергают digest/content_hash. Дополнительно: диспетчерский soak (test_mcp_ops_parity.py) — 3 цикла поллера, каждое одобренное действие диспетчеризуется ровно один раз (14 passed). Полный срез: 83 passed (8 файлов MCP-вертикали); полный backend suite 11182 passed.

Phase 2b — Initial scenario and automation parity

  • T029 Implement ScenarioRegistry.CreateInitial and bootstrap_authoring_scenario; prove a fresh external MCP client creates a first registry scenario without a seeded base revision. Evidence: tests/test_mcp_initial_scenario_e2e.py creates real AgentRun/DraftArtifact rows and drives bootstrap_authoring_scenario without a seeded registry; replay and conflicting idempotency are covered by tests/test_mcp_t029_bootstrap_automation.py.
  • T029a Add typed InitialScenarioIntent/TestPackProfile validation and server-owned metric/filter/selector/baseline bindings; unresolved requirements must be preview_only, never guessed.
  • T029b Expose 046 schedule, trigger-rule, policy and automation-metrics management through curated MCP tools with REST-equivalent RBAC, idempotency and PROD gate behavior. Evidence: curated tools and catalog entries in src/mcp_server/tools_automation.py/rbac_server.py, idempotency migration alembic/versions/0018_automation_idempotency.py, and tests/test_mcp_t029_bootstrap_automation.py plus RBAC catalog tests.
  • T029c Add fresh-DB MCP E2E for bootstrap → visible registry entry → revision activation → manual run, plus scheduler eligibility/PROD-gate integration evidence. Evidence: tests/test_mcp_initial_scenario_e2e.py verifies fresh bootstrap, current revision, refusal to run an un-promoted bootstrap revision (BOOTSTRAP_REVISION_NOT_RUNNABLE), a pinned manual run against a genuinely materialized revision, human-step schedule rejection with zero side effects, and an idempotent PROD approval gate; focused MCP regression set: 145 passed (2026-09-06).

Phase 2c — Server-owned handle layer (root-cause closure of ADR-0023; contracts: 038 ScenarioGraph.ServerOwnedPipeline amendment 2026-09-06, 042 ScenarioRegistry.DataModel amendment, 050 stage table status note)

  • T029d Persist the 038 handle layer: CompiledScenarioHandle/ValidationResultHandle/DraftPackHandle tables (immutable, append-only, owner_principal+agent_run binding, digest/content_hash columns per 038 amendment) + Alembic migration + server content-store persistence of the canonical DashboardTestScenario bytes (ScenarioGraph.Models.CanonicalBytes) behind canonical_bytes_ref. Minting happens ONLY at the persisted REST boundaries (api_compile_scenario/api_validate_scenario/api_resolve_scenario/api_draft_pack); pure compiler functions keep @SIDE_EFFECT None. Single-consumption (consumed_by_revision_id), cross-binding rejection, GC via 036 draft-retention. Evidence: src/models/scenario_handles.py, src/services/dashboard_testing/scenario/handles.py (idempotent mint / verify / consume / purge), migration 0019_scenario_handles, REST minting in api/routes/dashboard_testing/scenario.py, authority tests tests/services/dashboard_testing/scenario/test_handles.py (3 passed).
  • T029e 042 create consumes handles: create_scenario/create_initial re-verify handle ownership/binding/save_eligible, materialize graph_snapshot = canonical DashboardTestScenario JSON from handle bytes inside the create transaction, and write OutboxEvent(type=materialize_revision) + RevisionMaterialization(pending); implement the idempotent reference-artifact worker (scenario.yaml/reference runner.plan.json) or record an explicit descope decision. BOOTSTRAP_REVISION_NOT_RUNNABLE demotes from primary guard to defense-in-depth; api_create_scenario stops hardcoding materialization_status="materialized". Evidence: handle-first branch in registry/create.py (materializes full canonical graph + server-owned action_registry_version/hash), src/models/scenario_materialization.py + migration 0020_scenario_materialization, worker registry/materialize.py::materialize_pending_revisions (idempotent, pending→materialized/failed), api_create_scenario returns the actual RevisionMaterialization.status.
  • T029f MCP minting surface: new register_draft_pack write tool returns bounded server-issued handle ids; bootstrap_authoring_scenario accepts ONLY stored handle ids (transitional compile:{run_id}:{digest} string check removed from create_initial); catalog bumped 1.0.0 → 2.0.0 (breaking bootstrap handle semantics) with pinned-major ritual PINNED_CATALOG_MAJOR=2. Legacy REST POST /scenarios/{id}/revisions raw-graph_snapshot path retired (returns 410 GONE REVISIONS_RAW_GRAPH_RETIRED; the server-side create_revision service remains for the 043 editor save path). Also reclassified generate_report out of _MUTATING_ACTIONS (local draft report, no external side effect), bumping ACTION_REGISTRY_VERSION 038.1.0 → 038.2.0 (fingerprint + scenario_execution/graph.json fixture updated).
  • T029g Full-chain fresh-DB MCP E2E without REST crutches and without hand-seeded revisions: register_draft_pack→bootstrap_authoring_scenario→direct start_scenario_run on a bootstrapped current revision (now materialized) returns queued. Evidence: tests/test_mcp_initial_scenario_e2e.py rewrote _pack() to drive the MCP register_draft_pack tool. PostgreSQL concurrency: verify_handle_chain uses SELECT ... FOR UPDATE + populate_existing() to serialize handle consumption, proven by tests/integration/test_scenario_handle_concurrency.py (one winner + one HANDLE_CONSUMED across two threads on a real PostgreSQL container; --run-integration green).
  • T029h Inspect-stage decision: resolved as hybrid option C (2026-09-06) — no persisted InspectionContextHandle; instead (a) new MCP read tool inspect_dashboard_context exposes the existing live resolver (BaselineEngine.QueryModel.Inspect via GET /query-model service) returning the full authoritative DashboardQueryModel + fingerprint for the agent to echo into compile; (b) the register boundary (register_draft_pack) evaluates context_authority by RECOMPUTING the fingerprint from the client-carried dashboard_context.query (claimed fingerprints are ignored — C2) against a live inspection, with sentinel/degraded fingerprints ("", sha256:error) never verifying (C1), unreachable/unconfigured environments failing OPEN to unverified (C3), and falsifiable model-shape claims on live environments rejecting typed (CONTEXT_FINGERPRINT_MISMATCH, CONTEXT_QUERY_MODEL_REQUIRED, CONTEXT_ENVIRONMENT_MISMATCH) — C5 inspect-first; (c) the marker persists on DraftPackHandle.context_authority (migration 0021_context_authority), materializes into graph_snapshot at 042 create, and ScenarioExecution.Runner.Start refuses PROD dispatch on an explicit non-verified marker (CONTEXT_AUTHORITY_REQUIRED_FOR_PROD, zero side effects; missing marker = legacy, allowed); (d) X1 smuggling closure: validator _check_dashboard_context recursively rejects query_context keys and SQL text inside dashboard_context. Catalog minor bump 2.0.0 → 2.1.0 (additive, pinned major untouched). Evidence: tests/services/dashboard_testing/scenario/test_context_authority.py (7), validator X1 tests (+2), test_prod_start_enforces_context_authority_marker, E2E marker-propagation + PROD-block assertions; broad regression 1338 passed; alembic heads → 0021_context_authority. Residual: editor promote/save path (request_save revisions) does not yet carry the marker — PROD gate treats it as legacy-allowed; tightening requires authority evaluation in the 043 save boundary (follow-up, not scoped by T029h).

(Orthogonal code review 2026-09-06, ADR-0023; specs amended the same day and the durable handle layer implemented — 038 contracts/modules.md, 042 data-model.md/contracts/modules.md, 050 spec.md): the bootstrap→run happy path is realizable through MCP alone (register_draft_pack → bootstrap_authoring_scenario → direct start_scenario_run on a materialized current revision). Phase 2c residuals now closed: api_create_scenario returns real materialization_status, legacy REST POST /scenarios/{id}/revisions retired (410 GONE), PostgreSQL single-consumption concurrency proven (SELECT ... FOR UPDATE + populate_existing(); tests/integration/test_scenario_handle_concurrency.py green on real PostgreSQL). T029h resolved as the hybrid (see above); follow-up residuals closed 2026-09-07: outbox worker wired into the scheduler poll loop (scenario_revision_materialization, 30s, singleton/coalesced, durable-tick test), REST api_draft_pack evaluates context_authority at parity with the MCP register boundary (async, typed 422 on falsifiable mismatch), and the 043 editor save path provably inherits the server-owned marker (test_save_proposal_inherits_server_context_authority_marker). Known-accepted limitation: legacy/NULL markers stay PROD-allowed by design; full-suite regression 3360 passed.)

Phase 3 — Frontend decommission (flag-driven)

  • T030 HandoffSurface: страница-инструкция + копируемый промпт с контекстом дашборда; кнопка «Создать сценарий тестирования» переключается на handoff.
  • T031 Скрыть /agent, AssistantChatPanel, кнопку «Ассистент» в TopNavbar за флагом; обновить link-integrity тесты.
  • T032 Retire /api/assistant/* и прокси /api/agent/gradio; retention-настройки assistant скрыть. (Closure review 2026-09-03: на HEAD роутеры /api/assistant и agent_* смонтированы безусловно (app.py), retention-UI в SystemSettings жив, frontend api/assistant.ts в поставке — фиктивный [x]. Флаг удалён в T041, поэтому retirement теперь безусловный.)
  • T033 vitest/build/link-integrity зелёные при включённом флаге демонтажа.

Phase 4 — Removal

  • T040 Удалить сервис agent/ из run.sh и compose-профилей (порт 7860); обновить AGENTS.md/INSTALL.md. Доказательство: rg "agent|7860" run.sh build.sh docker-compose.yml docker-compose.enterprise-clean.yml → пусто (кроме исторического секьюрити-комментария); bash -n/yaml.safe_load чистые; стенд ./run.sh --skip-install поднимает только :8000+:5173 (7860 отсутствует, curl /api/ready → ready, Playwright-проход до хендоффа зелёный); run.sh/build.sh/AGENTS.md/INSTALL.md обновлены; из build.sh убраны build:agent, bundle:agent, bundle:embeddings и агентский сервис генерируемого деплой-композа/манифеста; из обоих nginx-конфигов убран location /api/agent/gradio.
  • T041 Удалить код чата: agent/src (app, langgraph_setup, tools*.py, _confirmation, middleware...), frontend agent/assistant компоненты, i18n, типы. Доказательство: agent/ удалён целиком; удалены docker/Dockerfile.agent, docker/agent.entrypoint.sh, backend/tests/test_gradio_proxy_config.py; во фронтенде удалены components/assistant/* (кроме универсального MarkdownRenderer), чат/ран/драфт артефакт компоненты, components/agent/dashboard-testing/ (ScenarioWorkspace-ветка), models/AgentChat*, AgentRunModel, DashboardScenarioWorkspaceModel, stores/assistantChat, флаг MCP_DECOMMISSION (vite define/config/mcp.ts/global.d.ts), gradio-прокси из vite.config.js; /agent рендерит только HandoffSurface (безусловно), навбар-кнопка ведёт на хендофф; npm run test -- --run → 3454 passed, npm run lint → 0 errors, npm run build → OK; полный бекенд 11199 passed (минус тесты удалённого грдио-прокси), 16 сиротских тестовых файлов чата удалены.
  • T042 Финальные правки спек 036–047: перенести drift-amendments из статуса «planned» в «done» со ссылками на доказательства. Доказательство: во всех 12 спеках (036–047) секции ## Drift Amendment — MCP Interface получили строку **Status (2026-09-02): done** со ссылками на specs/050-mcp-interface/tasks.md (T012–T028), T030–T033 и чекпоинты specs/WORKSTATE-043-047.md.
  • T043 Полный прогон backend/frontend suites + стенд без 7860 (SC-005). Доказательство: python -m pytest -q → 11199 passed, 240 skipped, 1 xpassed; npm run test -- --run → 3454 passed (197 файлов); npm run lint (0 errors) + npm run build — зелёные; локальный стенд после демонтажа работает без порта 7860 (см. T040) и отдаёт рабочий хендофф на /agent. Browser E2E (изолированный стенд): docker compose -p ss-tools-e2e --env-file .env.e2e -f docker-compose.e2e.yml up -d --build db backend frontend (свежая БД, bootstrap admin/admin123, backend :8103 healthy, frontend :8102 healthy, порт 7860 отсутствует) → npx playwright test e2e/tests/login.e2e.js e2e/tests/agent.e2e.js → 6 passed (двойной прогон, Chromium): login-поток (форма/успех/ошибка неверных кредов) и post-decommission агентский роут (хендофф-поверхность, ноль чат-элементов, deep-link с context-параметрами остаётся на хендоффе); стенд снесён down -v. E2E-фикстарелы: устаревший ambiguous locator('nav') (strict mode: 3 nav-элемента) → .first() в login/smoke; regex ошибки входа дополнен incorrect|неверн (бэкенд отдаёт passthrough-detail); agent.e2e.js переписан под handoff-контракт (T01-T03), dashboard-scenario-ui/agent-scenario-run — ссылки на чат-UI заменены на handoff-маршрут. Дополнительно: живой admin-walkthrough на dev-стенде (8000/5173) подтвердил login→dashboards→handoff(0 textarea/0 conversation-узлов)→runs-center (скриншоты /tmp/kilo/happy-path/09–12).

plans/initial-scenario-test-pack-automation-handoff.md

Source: plans/initial-scenario-test-pack-automation-handoff.md

@{ McpInterface.InitialScenarioTestPackAutomation.Handoff [C:5] [TYPE ADR]

@BRIEF Implementation handoff: allow an external MCP client to create a first scenario, persist a complete test pack, and configure safe automation. @RELATION DEPENDS_ON -> [ScenarioGraph.AgentAuthoringWorkspace] @RELATION DEPENDS_ON -> [ScenarioRegistry.CreateInitial] @RELATION DEPENDS_ON -> [ScenarioExecution.RunPreflight] @RELATION DEPENDS_ON -> [Automation.Schedule] @RATIONALE Current MCP authoring can edit an existing revision, but an unbound workspace cannot propose a graph; tests mask this by pre-seeding ScenarioRegistryEntry and ScenarioRevision. @REJECTED Creating registry rows or synthetic base revisions in an MCP client, browser client, or test fixture as the production bootstrap path.

Goal

An authenticated MCP user can create a dashboard scenario from server-owned dashboard context; complete and save a test pack; activate the resulting immutable revision; and configure a schedule only for an automation-eligible revision.

Verified current gap

  • create_authoring_session permits scenario_id=null.
  • propose_graph_revision rejects that workspace because both scenario_id and base_revision_id are required.
  • Existing MCP E2E tests call _seed_registry(...) before their session.
  • MCP exposes no schedule, trigger-rule, policy or automation-metrics tools although specs 046 and 050 require schedule management through MCP.

Required design

1. Initial scenario bootstrap

Add a curated MCP tool named bootstrap_authoring_scenario.

Input: bounded InitialScenarioIntent (title, dashboard ID, allowed environment IDs, selected case IDs, objective) plus opaque server-issued compiled-scenario and draft-pack handles. It must not accept a graph, revision ID, content hash, SQL, arbitrary code, source path, URL, cookie, or secret.

Server transaction:

  1. Re-resolve dashboard/environment ACL and ownership.
  2. Verify compiled and draft-pack handles are valid, save-eligible and bound to the same server-computed content hash.
  3. Call ScenarioRegistry.CreateInitial to write one registry entry, one immutable current revision, audit provenance and materialization outbox atomically.
  4. Create a workspace bound to that scenario and base revision.
  5. Return opaque scenario_id, revision_id, workspace_id, content_hash, CAS version and activation state.

Invariants: idempotency replay returns the same IDs; failed validation/ACL/CAS creates no row; initial creation never relies on a synthetic base revision; only a server-owned graph handle is persisted.

2. Full test-pack profile

Introduce a typed server-owned TestPackProfile/InitialScenarioIntent mapping:

  • dashboard context and selected checklist case IDs;
  • chart/metric references, filter bindings and selector requirements;
  • parameter defaults and safe fixtures;
  • evidence requirements and metric/visual baseline references;
  • explicit unresolved state (needs_context, needs_selector, needs_baseline) rather than guessed values.

The compiler must produce preview_only when unresolved requirements exist. Only a valid, baseline-complete pack can be passed to initial creation or normal save.

3. MCP automation parity

Add curated tools, each reusing 046 services and the same REST RBAC/policy checks:

  • list_scenario_schedules, upsert_scenario_schedule, delete_scenario_schedule;
  • list_scenario_trigger_rules, upsert_scenario_trigger_rule, delete_scenario_trigger_rule;
  • get_scenario_automation_policy, upsert_scenario_automation_policy;
  • get_scenario_automation_metrics.

Schedule writes accept only an explicit active revision, environment, cron/timezone, misfire/concurrency/dedup settings and idempotency key. A revision containing human steps is rejected before any schedule/gate/run write. PROD schedule dispatch stays gate-controlled and requires a separate approval decision.

Implementation order

  1. Add/align contracts and schemas in specs 038, 042, 046 and 050.
  2. Implement ScenarioRegistry.CreateInitial with migration only if the existing registry model lacks fields for title/environment ownership; add transactional unit tests.
  3. Implement bootstrap_authoring_scenario MCP schema, tool registration, RBAC and invocation provenance.
  4. Implement TestPackProfile validation and draft-pack eligibility; add fixture dashboard tests with metric, visual, filter and unresolved branches.
  5. Add MCP automation wrappers around existing 046 routes/services; do not duplicate scheduling logic.
  6. Replace seeded-registry MCP E2E setup with a no-preseed initial-creation E2E; then cover save, activation, manual run and scheduled eligibility.
  7. Add real scheduler callback integration evidence for eligible PROD and verify pending_approval before any dispatch.

Acceptance checks

  • A fresh database, no registry seed: external MCP client creates a first scenario and it appears in GET /scenarios.
  • Repeating an identical bootstrap key returns the original IDs; changing the payload with the same key yields typed idempotency conflict.
  • Invalid handle, stale dashboard context, missing selector/baseline or revoked scope leaves registry/outbox/workspace unchanged.
  • Full test pack reports explicit coverage for selected checks and never invents metric expectations.
  • Only an active, automation-eligible revision can be scheduled; human-step revisions have zero schedule/run/gate/queue side effects.
  • A PROD scheduled due event creates at most one durable gate before dispatch; deny leaves no run dispatch.
  • MCP schedule CRUD matches 046 REST outcome and RBAC behavior.

Estimated scope

Workstream Production LOC Test/fixture LOC
Initial bootstrap + registry transaction + MCP registration 450-750 350-550
Test-pack profile and baseline binding 400-700 300-500
MCP schedule/trigger/policy parity 250-450 250-400
Real scheduling E2E and provider/scheduler proof 100-200 300-600
Total 1,200-2,100 1,200-2,050

Specification, OpenAPI, traceability and task updates are expected to add 300-550 Markdown/YAML lines. Total implementation is roughly 2,700-4,700 LOC including tests.

Non-goals

  • Do not run arbitrary browser code, SQL, shell commands or client-supplied graphs.
  • Do not execute a production deployment, migration, backup or scheduled run merely to prove the MCP path.
  • Do not make materialized scenario.yaml or runner.plan.json authoritative.

@} McpInterface.InitialScenarioTestPackAutomation.Handoff


plans/server-decomposition-gate.md

Source: plans/server-decomposition-gate.md

#region McpServer.DecompositionGatePlan [C:3] [TYPE ADR] [SEMANTICS mcp,decomposition,inv7,plan,gate] @BRIEF Binding decomposition-gate plan for backend/src/mcp_server/server.py (INV_7: 1571 LOC vs the 400-line cap). Referenced by the module's @INVARIANT DECOMPOSITION GATE. STATUS: EXECUTED 2026-09-03, phases A–D, all gates green — see the execution log below. @RELATION DEPENDS_ON -> [McpServer] @RATIONALE server.py grew 1135 -> 1545+ LOC across the 050 build-out (closure-gate audit 2026-09-03, worst INV_7 offender). Attention-decay risk (ATTN_4): tool bodies beyond ~150 lines leave the sliding window; nested per-tool regions inside one 660-line ProbeTools block make the seams invisible. Splitting on already-anchored contract boundaries keeps every region <= ~250 LOC without renaming a single contract ID, so all incoming @RELATION edges, test imports, and index entries keep resolving. @REJECTED Big-bang rewrite of the tool layer was rejected — behavior-neutral moves are gate-verifiable per phase; a rewrite is not. Splitting by tool-count instead of domain seam was rejected — it would cut through shared catalog/guard contracts. Leaving the file oversized with only a warning tag was rejected — INV_7 is a hard invariant, and the GATE demands a plan with line counts before decomposition starts.

Constraints (inviolable)

  1. Contract IDs are frozen. Every McpServer.* region moves verbatim; no renames, no re-tiering, no metadata rewrites beyond @RELATION DEPENDS_ON additions between the new modules.
  2. Import surface is frozen. server.py re-exports every moved public name explicitly (from .scenario_inputs import ...), so src.mcp_server.server keeps satisfying all existing importers (tests, app.py, ops modules) with zero edits.
  3. Behavior diff is zero. No logic, ordering, decorator, or registration-timing changes ride along with a move. Registration order of tools must be byte-preserved (catalog + FastMCP registration order are observable via tools/list tests).
  4. INV_6 during moves: a moved contract is never deleted-then-recreated inside one edit — the region block is cut and pasted atomically per file, anchors verified balanced immediately after.

Phases (one phase = one reviewed commit-sized unit)

Phase New module Moves (regions, approx LOC) server.py LOC after
A mcp_server/scenario_inputs.py McpServer.ScenarioModels, AuthoringSessionInput, TestPlanContentInput, TestPlanInput, ExplorationInput, ExplorationResultInput, GraphRevisionInput, GraphDiffInput, PromoteScenarioInput, RequestSaveInput, ActivateRevisionInput, ScenarioResolveInput, DraftPackInput (~245) ~1330
B mcp_server/auth.py McpServer.Configuration, Principal, TokenVerifier, TransportAuth (~200) ~1130
C mcp_server/rbac_server.py McpServer.GateArguments, ScenarioGate, Catalog, RbacServer (~305) ~830
D1 mcp_server/tools_authoring.py authoring-chain tool bodies from McpServer.ProbeTools (session/plan/exploration/result/proposal/diff/promote/save/activate, ~330) ~500
D2 mcp_server/tools_scenario.py scenario-chain + probe/checkpoint/approval tool bodies (~330) ~170

Phase D splits McpServer.ProbeTools itself: the parent region becomes a thin registration seam that stays in server.py (imports the two tool modules, calls their register_* functions in the current order); nested per-tool regions move under matching McpServer.ProbeTools.* IDs in the new files.

Per-phase gate (all must pass before the next phase starts):

  • anchor stack balance in every touched file (#region/#endregion pairing script);
  • MCP slice green: pytest tests/test_mcp_*.py tests/api/test_mcp_parity_baseline_037.py -q (95 tests at 2026-09-03 baseline — never fewer);
  • ruff check src + compileall -q src clean;
  • tools/list catalog snapshot identical (covered by test_mcp_server.py catalog assertions);
  • any failure -> git checkout the phase start; fold the attempt note into this plan, do not retry in a poisoned state.

Risk register

  • Circular imports: tools_* modules need RbacServer/catalog types (Phase C output) and auth helpers (Phase B). Dependency direction is strictly server.py -> tools -> rbac_server -> auth -> scenario_inputs; enforce by importing modules (not names) where a cycle threatens, verified by the per-phase test gate.
  • Registration timing: FastMCP decorators execute at import; the seam must import tool modules before the guarded app is constructed (current CreateApp order preserved verbatim).
  • ContextVar locality: _access_token_context stays in server.py or moves to auth.py with a single definition site — never duplicated (a second ContextVar instance silently forks the context).
  • Test patches by path: tests patching src.mcp_server.server.<name> keep working through the re-exports; any patch of a moved CLASS's internals follows the re-exported binding — grep for mcp_server.server in tests during each phase gate.

Execution log — 2026-09-03 (phases A–D, all gates green)

Phase Module Result LOC Gate
A scenario_inputs.py (13 typed input models, verbatim) 268 anchors balanced; MCP slice 63 then 95 green; ruff/compileall clean
B auth.py (Configuration/Principal/TokenVerifier/TransportAuth + single _access_token_context site) 238 first run 42 failed — frozen-surface catch: tests build fake tokens via mcp_server.AccessToken; import retained → 95 green
C rbac_server.py (GateArguments/ScenarioGate/Catalog/RbacServer, verbatim) 393 first run 1 failed — monkeypatch seam server.request_mcp_approval relocated to owning module → 95 green
D tools_authoring.py + tools_scenario.py (ProbeTools closure split at the register-seam precedent register_ops_tools) 369 + 373 ruff F821 caught the one moved-body import (timedelta); seams relocated → 95 green
— server.py residual (composition seam only) 177 (was 1571) INV_7 satisfied package-wide
E* tools_review.py — checkpoint + baseline blocks carved out of register_ops_tools (pre-existing 420-LOC INV_7 offender the closure audit missed; split under the same gate rules) 253 (+ ops_tools 215) anchors balanced; ruff F821 silent; MCP slice 95 green; shared helpers (_SRC/_current_user/_resolve_environment) stay single-sited in ops_tools, seam import is function-local so the import graph stays acyclic

Deviations from the table above (recorded, none behavior-affecting):

  • Seam relocations (constraint-2 refinement): a module-attribute monkeypatch seam is NOT healed by re-export — the moved code resolves its own module's binding. Relocated to owning modules: request_mcp_approval → rbac_server (test_mcp_server); SQL-gate get_config_manager → rbac_server (ops_parity ×3); start-gate get_config_manager → rbac_server (scenario_e2e); tool-body get_config_manager (list_environments, start_scenario_run) and get_task_manager (maintenance sentinel) → tools_scenario (ops_parity, scenario_e2e, test_mcp_server). mcp_server.AccessToken and all import-from-server surfaces unchanged.
  • Registration order: preserved by calling the four register seams in the original in-function sequence (probe-read → authoring → scenario → maintenance/approval → ops); the @server.tool decorators execute per seam call, exactly as before.
  • Imports trimmed in server.py to the composition surface + frozen re-export blocks (auth names, 13 input models, catalog/RBAC names, AccessToken); every moved name still resolves via src.mcp_server.server for external importers. #endregion McpServer.DecompositionGatePlan

tests/test-report-2026-09-01.md

Source: tests/test-report-2026-09-01.md

Orthogonal Test Report — MCP Interface (050) + Authoring Workspace slice

Date: 2026-09-01 Scope: uncommitted working tree (backend/src/mcp_server/server.py, backend/src/services/agent_authoring_workspace/*, backend/src/services/dashboard_testing/editor/agent.py, backend/src/services/dashboard_testing/registry/revisions.py, backend/src/models/agent_authoring_workspace.py, alembic 0010–0013 renames; tests backend/tests/test_mcp_server.py, backend/tests/services/test_agent_authoring_workspace.py).

1. Mocking Audit Report

Summary

Total tests scanned Total mocks found Valid mocks Violations Logic Mirrors Uncertain
2 files (58 tests) 4 distinct patch targets 4 0 0 0

Files scanned: backend/tests/test_mcp_server.py, backend/tests/services/test_agent_authoring_workspace.py, backend/tests/conftest.py (setup), backend/tests/services/dashboard_testing/registry/conftest.py (setup).

Violations

None. test_agent_authoring_workspace.py contains zero mocks — it exercises the real service against the SQLite test DB.

Logic Mirrors

None. Assertion values are hardcoded fixtures (typed statuses, counts, IDs), no algorithmic re-computation of production logic detected.

Mock classification (per target)

Target (file:line) Classification Rationale
monkeypatch.setenv/delenv("SERVICE_JWT") (test_mcp_server.py:100,114,135,1033) ✅ VALID Environment configuration — infrastructure, not logic
monkeypatch.setattr(server, "_can_use_tool", ...) (multiple) ✅ VALID Auth/RBAC policy boundary stub; the SUT (approval envelope, catalog visibility invariant, retry idempotency) is never mocked; RBAC itself covered by test_rbac_server_hides_and_denies_the_same_tool and test_service_principal_cannot_see_or_use_gated_tools
monkeypatch.setattr(server, "_record", ...) + request_mcp_approval (test_mcp_server.py:170–176) ✅ VALID [EXT:Database] persistence boundaries (invocation record + durable gate) isolated for envelope-shape test; real gate/record path proven by test_direct_gated_retry_reuses_invocation_and_gate against real DB
monkeypatch.setattr(mcp_server, "get_task_manager", ...) (test_mcp_server.py:249) ✅ VALID External subsystem boundary (TaskManager/scheduler); used to prove absence of side effects for read-only listing

Integration Test Boundaries

File Type Real deps Mocked deps Verdict
tests/integration/test_mcp_postgres_concurrency.py integration Testcontainers PostgreSQL none ✅ CLEAN — 4/4 passed --run-integration (evidence 2026-09-01 work-state checkpoint, code unchanged since)

Clean tests (no violations)

  • backend/tests/services/test_agent_authoring_workspace.py
  • backend/tests/test_smoke_migration_chain.py
  • backend/tests/services/dashboard_testing/registry/test_revisions.py

Global setup (not violations)

  • backend/tests/conftest.py — SQLite global DB fixture, --run-integration registration
  • backend/tests/services/dashboard_testing/registry/conftest.py — seeded registry fixtures

Uncertain (requires human review)

None.

Size flags

File Lines Limit Verdict
backend/tests/test_mcp_server.py 1076 600 (unit) ⚠️ SIZE — split recommended by domain: transport/auth, gated maintenance, scenario tools, authoring chain, streamable-HTTP. Deferred: split is a mechanical refactor; no behavioral risk introduced by deferring.

2. Coverage Summary

Commands executed (scoped; make coverage deferred — not required for a working-tree audit slice):

  • pytest tests/test_mcp_server.py tests/services/test_agent_authoring_workspace.py → 51 passed (5.85s)
  • pytest tests/test_smoke_migration_chain.py tests/services/dashboard_testing/registry/test_revisions.py → 7 passed (1.34s)
  • ruff check on all touched src paths → All checks passed
  • compileall src → passed
  • Frontend: no working-tree changes → vitest not applicable
  • Integration (PostgreSQL): alembic upgrade head full chain + alembic check clean + concurrency 4/4 — captured in specs/WORKSTATE-043-047.md (2026-09-01), no code changed since

3. Semantic Audit Verdict

  • Anchors: server.py 32/32, service.py 12/12, revisions.py 6/6 — stack-nested, IDs pair correctly
  • Contract density: C4/C5 contracts in service.py carry @INVARIANT/@REJECTED/@RATIONALE matching effective complexity (CAS identity, ownership disclosure, server-side graph derivation)
  • Belief runtime: editor/agent.py (agent flow) has REASON/REFLECT/EXPLORE; service.py is a persistence/authorization service — belief markers not applicable
  • Rejected-path regression: client graph injection rejected (ScenarioStartInput closed union; tested), MCP dynamic dispatcher removal fail-closed (approved start_scenario_run stays queued), promotion never saves/activates — all guarded by @REJECTED + tests

AUDIT: PASS — no [AUDIT_FAIL:*] conditions.

4. Issues Found and Resolutions

  • False-positive region mismatch from linear open/close comparison during audit; corrected methodology (stack-based nesting check) confirms proper pairing. No code fix needed.
  • No mock violations, no logic mirrors, no rejected-path regressions found; zero fixes required.

5. Remaining Risk or Debt

  • test_mcp_server.py exceeds the 600-line unit limit → split by domain in a follow-up slice
  • PostgreSQL concurrency re-run deferred (Docker evidence from same day, code unchanged)
  • Product status remains NO-GO: isolated sandbox runtime (start_exploration real provider), browser E2E, full external MCP authoring E2E, 044 live providers + shared ExecutionCapacityManager

================================================================================ Features merged: 15 | Files merged: 213