7.2 KiB
Quickstart: Scenario Execution Engine (044)
Refresh 2026-09-08 (production contract): production baseline-backed evaluation (SCEX-FR-028) is normative in
contracts/production-chain.md,contracts/artifact-content.openapi.yaml,contracts/result-evidence.schema.json— implemented=false / acceptance OPEN. Canonical chain: browser→capture→durable artifact→deterministic baseline comparison→optional declared AgentEvaluation→DecisionPolicy→StepOutcome. Offline executable check (schemas + truth tables + evaluation-evidence binding):backend/.venv/bin/python specs/044-dashboard-scenario-execution/prototype/validate_contract_refresh.py \ specs/044-dashboard-scenario-execution/fixtures/production-contract-refresh.jsonHistorical result:
PASS: 11 schemas / 5 positive / 49 negative. Schema/fixture PASS is not production GO.
Target vs observed
| Layer | Target (refresh) | Observed (do not treat as GO) |
|---|---|---|
| Runner | Deterministic DAG walker, pinned ActionRegistry, human checkpoint, cancel/retry | Local SQLite/unit profile exists; fail-closed adapters exist |
| Provider loop | FastAPI lifespan owns exactly one ProviderEventLoop |
Runtime wiring landed and committed 2026-09-10 (83727aa7); live stand: loop ready, browser-preflight startup deadlock fixed (ba2f1f45) |
| AgentEvaluation | Immutable record + DecisionPolicy mapper; production adapter builds input_manifest from prior-step artifacts |
Landed, committed and live-proven 2026-09-10 (canary v2 run 4eebfab3: real qwen3.8-flash evaluation, AgentEvaluation row persisted, DecisionPolicy row 11 LOW_CONFIDENCE fired; trace: docs/reports/agentic-runtime-live-canary-v2-2026-09-10.md). OPEN: multimodal image attachment — the prompt is text-only, so visual verdicts cannot be trusted yet |
| Artifact bytes | Authenticated GET/HEAD with MIME/digest/length (Slice G) | Runtime GET/HEAD landed and committed 2026-09-10. Live storage canary PASSED 2026-09-10 (run 597274d3: 9 scenario_artifacts with digests, bytes on disk) |
| Visual chain | Compiler emits browser→capture→comparison→evaluation→policy (Slice E) | One-action templates; live Playwright canary PASSED 2026-09-10 (browser open_dashboard + screenshot capture vs ss-prod dashboard 11; comparison/evaluation steps not in the canary graph) |
| Live providers | Registered browser/screenshot + multimodal LLM | Binding ss-prod-d11-live-001 resolved server-side at REST start (no binding in the request body); /api/ready: browser/screenshot ready, bindings 1; canary v2 live trace: docs/reports/agentic-runtime-live-canary-v2-2026-09-10.md |
Open live gaps (2026-09-10, for the next agent): multimodal image attachment (prompts are text-only → visual verdicts stay inconclusive); → all three CLOSED 2026-09-10/11 (commit baseline_pin in the persisted AgentEvaluation record is {} for baseline-backed plans; screenshot/browser evidence MIME is hardcoded (image/jpeg/image/png) and would mismatch a WebP archive792bb125): multimodal evidence attachment (evaluation_images.py + _evaluation_messages), server-authoritative baseline_pin stamping (baseline_resolver.stamp_baseline_pin, live canary v3 9e6f59c0: verdict=pass, visual explanation), per-ref MIME sniffing (mime_sniff.py). Still open: live baseline pin is blocked on a valid Gitea PAT (publisher ready: backend/src/scripts/publish_catalog.py).
The former agent-driven product-UI flow (dashboard → /agent workspace → agent-generated scenario) is SUPERSEDED. 044 is a headless execution engine; agent interaction is external MCP (050). Product UI (045) is read-only evidence plus human approvals.
Prerequisites
- PostgreSQL reachable through a real
DATABASE_URLfor migration/integration checks. backend/.venvactivated.SERVICE_JWTset when using Docker Compose.- 042 registry and 044 migrations applied.
- Enabled live bindings configured only through server-owned startup settings.
Available local verification
cd backend
source .venv/bin/activate
python -m pytest -q \
tests/services/dashboard_testing/registry/test_scenario_*.py \
tests/api/test_scenario_runs_api.py \
tests/api/test_scenario_automation_api.py \
tests/api/test_scenario_analytics_api.py
python -m pytest -q \
tests/services/dashboard_testing/registry/test_scenario_executors.py \
tests/services/dashboard_testing/registry/test_live_execution_binding.py \
tests/services/dashboard_testing/registry/test_scenario_cancel_timeout.py \
tests/services/dashboard_testing/registry/test_scenario_crash_recovery.py \
tests/services/dashboard_testing/registry/test_scenario_worker.py \
tests/services/dashboard_testing/registry/test_scenario_queued_dispatch.py
python -m pytest -q tests/services/dashboard_testing/execution \
tests/services/dashboard_testing/registry/test_agent_evaluation_store.py \
tests/services/dashboard_testing/scenario/test_agent_evaluation_models.py
python -m ruff check src/services/dashboard_testing/execution src/api/routes/dashboard_testing/scenario_runs.py
python -m compileall -q src/services/dashboard_testing/execution src/api/routes/dashboard_testing/scenario_runs.py
cd ..
python3 specs/044-dashboard-scenario-execution/prototype/validate_static.py
2026-09-10 handoff recorded 292 passed on the first 044-slice command and 21+18 on the evaluation/policy suites. Those counts are historical runtime evidence for slices A+B/C/F, not this documentation pass and not production GO.
Production verification (still OPEN)
cd backend
source .venv/bin/activate
alembic check
alembic upgrade head
python -m pytest -q --run-integration tests/integration/
Provider contract profile (T028–T042; live Browser/Screenshot still required):
python -m pytest -q tests/services/dashboard_testing/registry/test_provider_contract.py
python -m pytest -q tests/services/dashboard_testing/registry/test_provider_*.py
Measurable exit gates (acceptance OPEN)
- SC-001..011 each has a passing named test or deployment evidence record.
- 100/100 cancellation trials terminate by
cancel_drain_deadline_at + 5 seconds. - 100% of provider I/O attempts have a valid CapacityLease and operation receipt.
- 100% of accepted evidence refs have an ownership receipt and verified SHA-256.
- 100% of unknown external effects are reconciled or terminalized non-pass before retry.
- Every enabled provider has passing liveness, readiness and dependency-health checks.
- Browser and Screenshot perform one real authorized deployment run with durable evidence.
- No unresolved P0/P1 traceability row remains in 044 or execution-critical dependencies 036, 037, 038, 041, 042, 046 and 047.
Current boundary
The local profile proves fail-closed adapters, exact Superset binding behavior, lifecycle closure, prototype state coverage, DecisionPolicy unit coverage, AgentEvaluation store/parser/executor, production evaluation adapter (mock provider), and walker EVALUATION_UNAVAILABLE / BASELINE_AND_SEMANTIC_PASS. It does not prove Browser/Screenshot live composition, shared provider capacity, live LLM evaluation, authenticated artifact GET/HEAD, real PostgreSQL migration validity, scheduler deployment behavior, or 047 case ingestion on a live stand.