Files
ss-tools/specs/044-dashboard-scenario-execution/quickstart.md

7.2 KiB
Raw Blame History

Quickstart: Scenario Execution Engine (044)

Refresh 2026-09-08 (production contract): production baseline-backed evaluation (SCEX-FR-028) is normative in contracts/production-chain.md, contracts/artifact-content.openapi.yaml, contracts/result-evidence.schema.json — implemented=false / acceptance OPEN. Canonical chain: browser→capture→durable artifact→deterministic baseline comparison→optional declared AgentEvaluation→DecisionPolicy→StepOutcome. Offline executable check (schemas + truth tables + evaluation-evidence binding):

backend/.venv/bin/python specs/044-dashboard-scenario-execution/prototype/validate_contract_refresh.py \
  specs/044-dashboard-scenario-execution/fixtures/production-contract-refresh.json

Historical result: PASS: 11 schemas / 5 positive / 49 negative. Schema/fixture PASS is not production GO.

Target vs observed

Layer Target (refresh) Observed (do not treat as GO)
Runner Deterministic DAG walker, pinned ActionRegistry, human checkpoint, cancel/retry Local SQLite/unit profile exists; fail-closed adapters exist
Provider loop FastAPI lifespan owns exactly one ProviderEventLoop Runtime wiring landed and committed 2026-09-10 (83727aa7); live stand: loop ready, browser-preflight startup deadlock fixed (ba2f1f45)
AgentEvaluation Immutable record + DecisionPolicy mapper; production adapter builds input_manifest from prior-step artifacts Landed, committed and live-proven 2026-09-10 (canary v2 run 4eebfab3: real qwen3.8-flash evaluation, AgentEvaluation row persisted, DecisionPolicy row 11 LOW_CONFIDENCE fired; trace: docs/reports/agentic-runtime-live-canary-v2-2026-09-10.md). OPEN: multimodal image attachment — the prompt is text-only, so visual verdicts cannot be trusted yet
Artifact bytes Authenticated GET/HEAD with MIME/digest/length (Slice G) Runtime GET/HEAD landed and committed 2026-09-10. Live storage canary PASSED 2026-09-10 (run 597274d3: 9 scenario_artifacts with digests, bytes on disk)
Visual chain Compiler emits browser→capture→comparison→evaluation→policy (Slice E) One-action templates; live Playwright canary PASSED 2026-09-10 (browser open_dashboard + screenshot capture vs ss-prod dashboard 11; comparison/evaluation steps not in the canary graph)
Live providers Registered browser/screenshot + multimodal LLM Binding ss-prod-d11-live-001 resolved server-side at REST start (no binding in the request body); /api/ready: browser/screenshot ready, bindings 1; canary v2 live trace: docs/reports/agentic-runtime-live-canary-v2-2026-09-10.md

Open live gaps (2026-09-10, for the next agent): multimodal image attachment (prompts are text-only → visual verdicts stay inconclusive); baseline_pin in the persisted AgentEvaluation record is {} for baseline-backed plans; screenshot/browser evidence MIME is hardcoded (image/jpeg/image/png) and would mismatch a WebP archive → all three CLOSED 2026-09-10/11 (commit 792bb125): multimodal evidence attachment (evaluation_images.py + _evaluation_messages), server-authoritative baseline_pin stamping (baseline_resolver.stamp_baseline_pin, live canary v3 9e6f59c0: verdict=pass, visual explanation), per-ref MIME sniffing (mime_sniff.py). Still open: live baseline pin is blocked on a valid Gitea PAT (publisher ready: backend/src/scripts/publish_catalog.py).

The former agent-driven product-UI flow (dashboard → /agent workspace → agent-generated scenario) is SUPERSEDED. 044 is a headless execution engine; agent interaction is external MCP (050). Product UI (045) is read-only evidence plus human approvals.

Prerequisites

  • PostgreSQL reachable through a real DATABASE_URL for migration/integration checks.
  • backend/.venv activated.
  • SERVICE_JWT set when using Docker Compose.
  • 042 registry and 044 migrations applied.
  • Enabled live bindings configured only through server-owned startup settings.

Available local verification

cd backend
source .venv/bin/activate
python -m pytest -q \
  tests/services/dashboard_testing/registry/test_scenario_*.py \
  tests/api/test_scenario_runs_api.py \
  tests/api/test_scenario_automation_api.py \
  tests/api/test_scenario_analytics_api.py
python -m pytest -q \
  tests/services/dashboard_testing/registry/test_scenario_executors.py \
  tests/services/dashboard_testing/registry/test_live_execution_binding.py \
  tests/services/dashboard_testing/registry/test_scenario_cancel_timeout.py \
  tests/services/dashboard_testing/registry/test_scenario_crash_recovery.py \
  tests/services/dashboard_testing/registry/test_scenario_worker.py \
  tests/services/dashboard_testing/registry/test_scenario_queued_dispatch.py
python -m pytest -q tests/services/dashboard_testing/execution \
  tests/services/dashboard_testing/registry/test_agent_evaluation_store.py \
  tests/services/dashboard_testing/scenario/test_agent_evaluation_models.py
python -m ruff check src/services/dashboard_testing/execution src/api/routes/dashboard_testing/scenario_runs.py
python -m compileall -q src/services/dashboard_testing/execution src/api/routes/dashboard_testing/scenario_runs.py
cd ..
python3 specs/044-dashboard-scenario-execution/prototype/validate_static.py

2026-09-10 handoff recorded 292 passed on the first 044-slice command and 21+18 on the evaluation/policy suites. Those counts are historical runtime evidence for slices A+B/C/F, not this documentation pass and not production GO.

Production verification (still OPEN)

cd backend
source .venv/bin/activate
alembic check
alembic upgrade head
python -m pytest -q --run-integration tests/integration/

Provider contract profile (T028–T042; live Browser/Screenshot still required):

python -m pytest -q tests/services/dashboard_testing/registry/test_provider_contract.py
python -m pytest -q tests/services/dashboard_testing/registry/test_provider_*.py

Measurable exit gates (acceptance OPEN)

  1. SC-001..011 each has a passing named test or deployment evidence record.
  2. 100/100 cancellation trials terminate by cancel_drain_deadline_at + 5 seconds.
  3. 100% of provider I/O attempts have a valid CapacityLease and operation receipt.
  4. 100% of accepted evidence refs have an ownership receipt and verified SHA-256.
  5. 100% of unknown external effects are reconciled or terminalized non-pass before retry.
  6. Every enabled provider has passing liveness, readiness and dependency-health checks.
  7. Browser and Screenshot perform one real authorized deployment run with durable evidence.
  8. No unresolved P0/P1 traceability row remains in 044 or execution-critical dependencies 036, 037, 038, 041, 042, 046 and 047.

Current boundary

The local profile proves fail-closed adapters, exact Superset binding behavior, lifecycle closure, prototype state coverage, DecisionPolicy unit coverage, AgentEvaluation store/parser/executor, production evaluation adapter (mock provider), and walker EVALUATION_UNAVAILABLE / BASELINE_AND_SEMANTIC_PASS. It does not prove Browser/Screenshot live composition, shared provider capacity, live LLM evaluation, authenticated artifact GET/HEAD, real PostgreSQL migration validity, scheduler deployment behavior, or 047 case ingestion on a live stand.