Files
ss-tools/specs/044-dashboard-scenario-execution/data-model.md

26 KiB

#region ScenarioExecution.DataModel [C:5] [TYPE ADR] [SEMANTICS data-model,scenario,execution,run,step,state] @BRIEF Canonical ScenarioRun, ScenarioStepRun, RunnerPlan, executor registry, and lifecycle state model. @RELATION DEPENDS_ON -> [ScenarioExecution.Research] @RATIONALE A typed run/step state model is required for reload during a test, recovery, resume, and reproducibility. Without it, human-checkpoint resume and page reload are impossible. @REJECTED In-memory-only run state — rejected because runs must survive disconnect and be recoverable by run id.

ScenarioRun and PROD approval lifecycle

ScenarioRun is created before dispatch. Fields: id, scenario_id, scenario_revision_id (revision_id UUID), scenario_content_hash, verification_program_hash, action_registry_version, dashboard_id, environment_id, status (pending_approval|queued|running|waiting_human|blocked|cancel_requested|cancelled|passed|failed|inconclusive), phase (preflight|setup|executing|waiting_human|draining|terminal), parameter_bindings (immutable JSON), baseline_pin (exact 037 BaselineSelectionPin; null only for a revision declaring no baseline-backed checks), target_snapshot, execution_principal_fingerprint, execution_toggles (optional evidence only), trigger_source (server-owned), agent_run_id? (provenance), verification_run_id? (aggregation), idempotency_key (unique), started_at, finished_at, resume_token, error_code, runner_version.

live_execution_binding_ref? plus live_execution_binding_snapshot? are the only persisted live-I/O coordinates. The snapshot is an exact allowlist: binding ref; environment/release/query-model, execution-principal and RLS/security fingerprints; browser-safe checkpoint/action refs; and evidence owner/ref policy. It contains no credential, client, cookie, raw browser context, callable principal, or capture bytes. LiveExecutionCompositionRoot resolves this immutable identity server-side at startup; a missing provider is typed unavailable, and any exact-snapshot mismatch is typed non-pass before I/O.

The deployment-owned settings.scenario_live_execution_bindings[] record stores enabled, binding_snapshot, and query_model_snapshot only. Startup resolves its Environment credentials server-side to build the existing SupersetClient, then registers the exact model/storage tuple. Browser and screenshot providers are process-local registrations and are never serialized in this record; absent registrations yield stable configured-unavailable outcomes.

If the selected revision contains a human step, it is derived as manual_run_only=true: it may start only from the authenticated analyst manual-run route. Scheduler, deploy/release/ETL/API trigger and any background runner or recovery worker are ineligible. The eligibility check occurs before idempotency, ScenarioRun, ActionApprovalGate, notification, queue insertion and dispatcher CAS; no HumanCheckpoint may be created, skipped, defaulted or converted into an automated approval path.

For a PROD request, the service atomically creates ScenarioRun(status=pending_approval) and ActionApprovalGate(owner_type=scenario_run, owner_id=run_id, operation=scenario_execution). Approval transitions only pending_approval → queued; denial/expiry transitions to blocked. Schedulers and external triggers therefore receive a durable run/intent, never an unusable 403. trigger_source is set only by the trusted entry point (manual route, scheduler, deploy connector, or API key), never by a bearer client field. Idempotency uses canonical execution-request hash: same key + same hash returns the existing run; same key + different hash returns 409 IDEMPOTENCY_KEY_REUSED.

EnvironmentExecutionPolicy is resolved server-side from the configured Environment record (stage=PROD or is_production=true) before RBAC, idempotency, run, gate, queue, notification, or adapter work. Any compatibility is_prod request field is non-authoritative and cannot upgrade or downgrade that class; an unknown environment fails closed without creating a run or gate.

ParameterBinding, ExecutionPrincipal, and TargetSnapshot

ParameterBinding: parameter_name, resolved_value, source, resolved_at, validation_fingerprint. It is derived at start from the 038 ParameterDefinition and launch input; it is never embedded in ScenarioRevision or its content_hash.

ExecutionPrincipal: auth_mode, actor_id?, service_identity?, impersonated_user?, effective_roles_hash, rls_context_hash. Its immutable fingerprint is stored on the run and used for every Superset request.

TargetSnapshot: environment_id, dashboard_release_id, dashboard_fingerprint, dataset_lineage_fingerprint, captured_at. It is mandatory even when no release was selected, so environment_id=preprod is never treated as an immutable target.

AnalyticsContextKey is a server-derived SHA-256 over environment_class + compatibility_family + baseline_family + dashboard_release_id + dashboard_fingerprint + dataset_lineage_fingerprint + execution_principal_fingerprint. It is captured on the run and every step result; analytics never groups runs merely by environment or revision id.

Statuses: queued, running, waiting_human, blocked, cancel_requested, cancelled, passed, failed, inconclusive.

Terminal failed, inconclusive and blocked outcomes emit a 036 InvestigationSignal with this immutable execution snapshot. 047 deterministically creates/updates the Queue/Episode from that signal. Signal delivery never changes run status and never auto-starts an agent run; an analyst opens the case explicitly from 045/047.

ScenarioStepRun

Fields: id, run_id FK, logical_step_id (immutable UUID, from #8), step_position (mutable), step_content_hash (mutable), attempt, status (queued|running|waiting_human|passed|failed|inconclusive|blocked|skipped), started_at, finished_at, inputs_snapshot (JSON, no secrets), outputs (JSON), artifact_refs (JSON), error_code, progress, timeout_ms, side_effect_key (nullable), step_outcome (StepOutcome).

StepOutcome { status, reason_codes[], deterministic_evidence_refs[], agent_evaluation_ids[], decision_policy_id?, decision_policy_version?, decided_at } is the authoritative result of one logical step. It is distinct from executor output and from a model verdict.

AgentEvaluation and DecisionPolicy

AgentEvaluation { evaluation_id, scenario_run_id, logical_step_id, attempt, provider_id, model_id, model_version, prompt_template_id, prompt_template_version, input_manifest_hash, evidence_refs, verdict, confidence, findings, reason_codes, raw_response_artifact_ref, started_at, finished_at } is immutable evidence generated only for a declared 038 AgentEvaluationSpec. Its tool/evidence access is bounded by that spec; it cannot mutate program content, invoke mutation actions, change run state or choose downstream scheduling.

DecisionPolicy { policy_id, version, deterministic_hard_failure, high_confidence_failure, low_confidence, disagreement, missing_evidence } is a versioned deterministic mapper. Defaults: deterministic hard failure→failed; high-confidence policy-qualified agent failure→failed; low confidence→inconclusive; evaluation/evidence disagreement→inconclusive; missing required evidence→blocked; provider error for required evaluation→inconclusive. Exact first-match baseline-semantic/1.0.0 cases and strict schemas are 038 DecisionPolicy; no dynamic checkpoint creation. Scenario aggregation consumes StepOutcome, not AgentEvaluation verdicts directly.

RunnerPlan — deterministic derivation from revision (#2)

RunnerPlan is derived deterministically at run start from the selected immutable ScenarioRevision/Verification Program, NOT read from a stored runner.plan.json. Fields: scenario_revision_id, scenario_content_hash, verification_program_hash, action_registry_version, action_registry_hash, env targets, resolved params, exact baseline_pin plus canonical pin digest, topological order, and one immutable ActionExecutionDescriptor per step. A descriptor contains the exact {tool, action}, typed input/output contracts, idempotency, retry safety, side-effect-key policy, timeout, mutation contract/risk. The runner persists the descriptor snapshot in both steps and executor_mapping; it derives leases/recovery/retry policy only from that snapshot. A missing, altered, unknown, version/hash-mismatched descriptor, invalid input/output shape, or mutating action without its required contract rejects before run/lease/I/O. Run refuses if its revision/program/action-registry hashes differ from the selected revision.

@INVARIANT RunnerPlan descriptor resolution is exact and version/hash pinned; tool is never a dispatch or retry-policy fallback. @REJECTED A universal idempotent=true, retry_safe=true claim based on a non-human tool was rejected because it can repeat unsafe effects.

The materialized runner.plan.json in git is a reference artifact, never the runtime source of truth; it may be regenerated from any revision.

ScenarioExecutionContext (#11)

Fields: run_id, browser_session (ref, not raw cookies), page/context ref, auth_context ref, current_dashboard, current_filters, artifact_namespace (owner_type=scenario_run), environment client. Browser secrets/cookies live in a separate secure context, NEVER in inputs_snapshot JSON (keeps reproducibility snapshot secret-free). Browser workers resume only by deterministic replay from the last browser-safe checkpoint (dashboard_open, filters_applied, etc.); replay records reconstruction_replay=true and never reclassifies already completed logical steps as rerun. API/XLSX/pure assertion steps may resume directly only when their executor declares retry-safe/idempotent. Browser session/context refs are valid only on the single application-owned provider event loop (loop_binding); they are never shared across runs and never cross loops (see 044 ProviderRuntime contract).

Artifact ownership (#4)

Artifacts use a generic owner: Artifact { id, owner_type: agent_run|scenario_run|verification_run|load_run, owner_id, kind, sha256, content_ref, retention_class, ... }. ScenarioRun evidence/screenshots/report/xlsx use owner_type=scenario_run; no artificial AgentRun is created. Retention_class ties to 046 tiers. Playwright trace/video dumps are not Artifact kinds: an operator debug flag may keep them locally, and they are never registered as Artifact or EvidenceReceipt.

Decision gates (#3)

  • ActionApprovalGate — authorization approval for PROD execution, baseline approval, repository mutation. Generalized 036 gate with owner_type + owner_id.
  • HumanCheckpoint — checkpoint_id, run_id, logical_step_id, checkpoint_type, decision_policy, status, created_at, expires_at, eligible_role?, eligible_actor_ids?, assigned_to?, evidence_refs, decision_version, decided_by?, decided_at?, disposition?, comment?. Status is pending|decided|expired|cancelled; decision is CAS on decision_version, stale/concurrent decision returns 409. v1 disposition maps confirm→passed, false_positive→inconclusive, inconclusive→inconclusive; manual_assertion maps pass→passed, fail→failed, inconclusive→inconclusive. It is NOT a 036 ApprovalGate decision.

A HumanCheckpoint is never delegated to the agent: it is a manual-run-only analyst decision inside a currently executing run. External MCP clients may display evidence; only authenticated human USER decision tools may consume the checkpoint under 050. Product frontend offers human review, never agent workspace or invocation.

Worker semantics — at-least-once execution (#6)

Runtime primitives: worker lease, heartbeat, run claim, step claim, lease expiration, idempotency key, recovery scheduler. Each executor declares: idempotent? | retry-safe? | side_effect_key? | external_request_id?. POST /scenario-runs requires Idempotency-Key (unique) to prevent double-run on double-click. A crashed step with an external side effect is only re-run if idempotent/retry-safe or keyed.

Retry semantics (#14)

Retry of a failed step invalidates its downstream closure (descendants depending on its output) and re-runs them; retry after the run advanced beyond the step is rejected unless the whole closure re-runs. Bounded attempts per policy.

Result aggregation truth table (#16)

  • any step failed → scenario failed
  • blocked descendants counted as blocked (NOT failed)
  • skipped does not count against pass
  • inconclusive → scenario inconclusive unless a failed also present (then failed)
  • warning is an evidence-level qualifier, not an execution status (045 maps from evidence)

Lifecycle

pending_approval → queued → running → waiting_human | blocked → passed | failed | inconclusive; cancel_requested → cancelled. Human decision atomically consumes the checkpoint and resumes internally. Public /resume is reserved for a recoverable infrastructure pause and requires a typed resume token/reason; it cannot consume a HumanCheckpoint. Cancel drains in-flight within a bounded window.

ScenarioExecutorRegistry and BrowserExecutor

The registry resolves ActionExecutionDescriptor -> executor; the following list states each descriptor's tool family, not a tool-only fallback:

  • browser -> BrowserExecutor → version-pinned 038 ActionRegistry → Playwright/session infrastructure
  • superset_api -> 037 metric_executor_async / SupersetClient.ChartData.Execute
  • sql_evidence -> SqlEvidenceExecutor → Superset SQL Lab backend/API → configured database connection (no credentials exposed to agent)
  • transform -> version-pinned bounded 038 TransformSpec DSL executor
  • xlsx -> xlsx parser + 037 normalization
  • assertion -> 037 comparison.py + 038 ComparisonSpec/AssertionSpec executor
  • agent_evaluation -> bounded provider adapter executing a declared 038 AgentEvaluationSpec and emitting AgentEvaluation; DecisionPolicy owns StepOutcome
  • screenshot -> 038 capture.py + ScreenshotService (owner_type=scenario_run)
  • report -> report-template render + artifact (owner_type=scenario_run)
  • artifact -> generic artifact register (owner_type=scenario_run)
  • human -> EXCLUDED (runner-lifecycle HumanCheckpoint control, not an executor)

BrowserExecutor implements registered actions: open_dashboard, navigate_tab, navigate_tabs, apply_native_filter, inspect_filter_state, inspect_filter_options, apply_table_filter, extract_table, scroll_to, inspect_columns, click, select_rows, assert_dom, wait_for_selector, edit_row, bulk_edit, download, refresh, wait_for_state. pagination and navigate_dashboard remain registry-declared but disabled until their drivers and live canary gates close. ScreenshotService is evidence infrastructure only, not the browser action executor. It validates the registry-declared inputs/outputs/risk/timeout/idempotency before dispatch. Composite navigate_tabs is bounded to the visible tab nav and returns the visited labels; assert_dom evaluates only declared structural DOM predicates and cannot embed business metric truth; inspect_filter_options is observe-only; wait_for_selector is bounded and fail-closed.

Provider runtime protocol

Every provider implements the following logical records. They are persisted or emitted as immutable operation/evidence receipts; a Python callback alone is not a sufficient provider implementation.

ProviderExecutionContext: run_id, logical_step_id, attempt, descriptor_fingerprint, live_binding_ref, execution_principal_fingerprint, deadline_at, idempotency_key, capacity_lease_id, trace_id, cancellation_requested, cancellation_deadline_at.

ProviderOperationReceipt: operation_id, context identity, provider_id, provider_version, started_at, finished_at, effect_state, external_request_id?, resource_refs, cleanup_state, reconciliation_state, result_digest?. Receipts are append-only; late responses are historical records and cannot mutate the winning attempt.

ProviderEvidenceReceipt: evidence_ref, owner_type=scenario_run, owner_id, run_id, logical_step_id, attempt, operation_id, provider_id/version, descriptor_fingerprint, binding_ref, execution_principal_fingerprint, content_type, byte_length, sha256, retention_class, created_at. The durable store must atomically commit the content and receipt or return no usable ref.

Provider state is registered|health_checked|capacity_claimed|admitted|running|completed|failed| cancel_requested|finalized|reconciliation_required. A provider exposes separate liveness, readiness and dependency-health checks and a redacted capability/limit document. Readiness gates new admissions only; it does not remove active leases or evidence.

Cancellation calls cancel(operation_id) and records stopped|completed|unknown. unknown means the provider cannot prove whether an external effect happened: retry is forbidden until reconcile(operation_id, idempotency_key) returns a terminal receipt. If reconciliation is unavailable before the bounded recovery deadline, the runner terminalizes inconclusive or blocked according to effect risk. Capacity is released only after receipt and cleanup state are durable.

Provider-specific schemas and obligations

Provider Input boundary Required result/evidence Resource and recovery rules
Browser 038 action descriptor, server-owned auth/session binding, typed action input operation receipt, safe checkpoint, page/action provenance and durable evidence where declared isolated context, bounded pages/downloads, cleanup on finalization, safe-checkpoint replay only
Superset API exact 037 query envelope and pinned model/principal/RLS request/response fingerprints, raw response SHA-256 and durable ref request/response limits; 429/5xx bounded retry; unknown timeout reconciles external request
SQL evidence immutable 038 SqlEvidenceSpec plus typed ParameterBindings DB identity, model/principal/RLS fingerprints, raw response digest/ref read-only; caller SQL rejected; reconcile before retry
XLSX verified server-owned artifact ref source ownership/digest and canonical normalized output archive/cell/formula limits; local parse is idempotent
Assertion canonical values plus pinned 037 ComparisonPolicy deterministic input manifest, policy/version and StepOutcome pure, no external resources, retry-safe
Transform bounded 038 TransformSpec DSL source manifest, transform fingerprint and canonical output digest hard operation/row/column/output limits; no code/SQL/network
Screenshot bound capture session/profile and durable evidence policy atomic evidence receipt with media type, length and SHA-256 temp cleanup; unknown capture reconciles before retry
Report pinned template and immutable evidence manifest template/manifest/output digest and durable ref bounded deterministic renderer; partial writes reconcile or discard
Artifact authorized producer ref or verified source artifact producer receipt, owner tuple, digest and retention never accept caller digest alone; registration idempotent
AgentEvaluation immutable 038 spec, evidence allowlist, provider/model/prompt versions immutable AgentEvaluation and raw response artifact receipt token/cost/time limits; verdict cannot directly set StepOutcome

BrowserProvider action contract

Each browser step persists the exact 038 action descriptor plus input_schema_version, output_schema_version, risk, checkpoint_policy, idempotency, retry_safe, limits and mutation contract fingerprint. Defaults are context/authentication timeout 120000ms, action timeout 30000ms, maximum 3 pages, maximum download 26214400 bytes and maximum screenshot 10485760 bytes. A descriptor may lower a limit but cannot raise it from scenario payload.

open_dashboard, navigate_tab, filter application, inspection, scroll, extraction, refresh and wait are read-only/evidence actions; click is rejected unless its effect class is explicit; edit_row and bulk_edit are mutation actions; select_rows is selection-only; download produces a server-owned artifact ref, never a local path. extract_table is bounded to 10,000 rows, 100 columns and 10 MiB output.

Browser resource state is lease_claimed|context_created|authenticated|action_running|checkpointed| effect_recorded|cleanup_started|receipt_finalized|context_closed. One active context is allowed per (run_id, capacity_lease_id) and cleanup failure prevents PASS. Concurrency defaults are 2 concurrent browser contexts for DEV/PREPROD and 1 for PROD per environment, admitted by the shared CapacityManager; screenshot capture shares the same lease accounting, and capacity exhaustion keeps the run queued with CAPACITY_BLOCKED. Recovery never attaches to a dead context; it reconstructs from a safe checkpoint and records reconstruction_replay=true. Unknown mutation effects are held for reconciliation and cannot retry.

Readiness requires browser executable/version, auth binding, allowed-origin, evidence-storage, cleanup-worker and cancellation-capability checks. Deployment requires a read-only PREPROD canary and forced timeout/ cleanup canary before enabling; mutation requires a separate fixture-lease canary.

AgentEvaluation and DecisionPolicy integration

The AgentEvaluation provider claims capacity as workload_class=agent_evaluation, creates an immutable input manifest from deterministic evidence refs, executes within the declared budget, and stores the raw response before publishing a verdict. Provider/model/prompt versions and manifest hash are mandatory. DecisionPolicy(policy_id, version) consumes only the immutable evaluation record plus deterministic evidence and emits StepOutcome. Low confidence, missing evidence or disagreement follow the pinned policy table; the model cannot select a policy, mutate the graph, invoke a provider action, consume a HumanCheckpoint or schedule downstream work.

ExecutionCapacityManager integration

Before any provider I/O, the dispatcher calls the shared allocator with environment, workload_class, provider_id, priority, requested_units, quota, reserved_capacity, run/step and deadline. The atomic result is a durable CapacityLease; no lease means queued or blocked according to policy and no provider invocation. Heartbeat, expiry, release and forced reconciliation are idempotent. Retries claim new leases, while active unknown operations retain their lease until reconciliation or terminal closure.

Verification thresholds

  • scenario_content_hash, verification_program_hash, descriptor_fingerprint, query_model_fingerprint, execution_principal_fingerprint, rls_security_fingerprint and evidence sha256 are canonical lowercase hexadecimal SHA-256 values with exactly 64 characters.
  • attempt starts at 1 and increases by exactly 1 for each new logical-step attempt. Historical attempts are immutable; only one attempt may be the active projection for a logical step.
  • byte_length is a positive integer equal to the durable content length. A zero-byte evidence object is never eligible for PASS.
  • deadline_at is later than started_at; a result received after the deadline cannot become the winning outcome. The runner records the late result as historical operation data.
  • CapacityLease validity is [claimed_at, expires_at]; provider I/O requires now < expires_at and a matching run/step/attempt/provider context. Lease release is idempotent: repeated release changes no terminal state or accounting total.
  • Cancellation finalization must occur no later than cancel_drain_deadline_at + 5 seconds under the default scheduler interval. Any exception is a release failure and keeps the run non-GO.

Mutation authorization and execution policy

Authorization is evaluated per run and step as environment_class + scenario risk profile + action mutation profile. Read-only PROD steps require the ScenarioExecution approval policy. Every mutation requires the immutable 038 mutation contract, scoped target keys, precondition evidence, side-effect identity, cleanup/reconciliation outcome and retry_safe=false unless the registry proves an idempotent compensating action. Mutating browser steps in PROD are prohibited. Non-PROD test-data mutation may be delegated only inside an authorized fixture lease; this mutation policy is separate from PROD execution approval.

Immutable Execution Snapshot

A ScenarioRun pins scenario_revision_id + scenario_content_hash at start. Results carry provenance: scenario_revision_id, runner_version, template_version, baseline_revision, target_snapshot, parameter_bindings, execution_principal_fingerprint, query fingerprints. Later edits never alter completed runs.

ExecutionCapacityManager

All AgentRun, VerificationRun, LoadRun, and ScenarioRun claims pass through one environment-scoped capacity manager: environment, workload_class, priority, quota, reserved_capacity. A 046 scenario policy is a consumer of this global allocator, not an independent PROD/PREPROD concurrency limit.

Production records and compatibility — 2026-09-08

Result schema defines the unified run projection. AgentEvaluation schema replaces the abbreviated field list above for new records: operation/spec identity, input artifact MIME/length/SHA manifest, provider/model/prompt/schema versions, criterion-bound findings, raw response provenance, usage/pricing and full baseline pin. Evaluations/comparisons/outcomes are append-only; unique active attempt is selected by CAS. Historical late attempts remain visible but never replace a winner.

Run request hash includes the server-resolved baseline set/version, catalog/release/publication commit and entry IDs/digests before run/gate creation. The exact pin is persisted in RunnerPlan, ScenarioRun, result and AnalyticsContextKey provenance; comparison never uses first-map-entry or moving latest fallback. Missing legacy pins are ineligible, not fabricated.

Artifact content uses protected GET/HEAD, not public URLs. Metadata carries owner/run/step/attempt/operation, SHA-256, actual MIME, byte length, retention/expiry and availability. Required baseline bytes have retention holds. Production chain owns provider startup/shutdown/cancel/reconcile. New records and lifecycle acceptance are live-verified for the provider/artifact/evaluation chain (2026-09-17: T042b/T044/T045 canaries + T022 audit); the graph-level terminal PASS remains pending (T046 binding). Frontend only renders read-only evaluation evidence; no provider/prompt/retry-agent controls. Approved performance baseline is outside scope.

#endregion ScenarioExecution.DataModel