26 KiB
#region ScenarioExecution.DataModel [C:5] [TYPE ADR] [SEMANTICS data-model,scenario,execution,run,step,state] @BRIEF Canonical ScenarioRun, ScenarioStepRun, RunnerPlan, executor registry, and lifecycle state model. @RELATION DEPENDS_ON -> [ScenarioExecution.Research] @RATIONALE A typed run/step state model is required for reload during a test, recovery, resume, and reproducibility. Without it, human-checkpoint resume and page reload are impossible. @REJECTED In-memory-only run state — rejected because runs must survive disconnect and be recoverable by run id.
ScenarioRun and PROD approval lifecycle
ScenarioRun is created before dispatch. Fields: id, scenario_id, scenario_revision_id (revision_id UUID), scenario_content_hash, verification_program_hash, action_registry_version, dashboard_id, environment_id, status (pending_approval|queued|running|waiting_human|blocked|cancel_requested|cancelled|passed|failed|inconclusive), phase (preflight|setup|executing|waiting_human|draining|terminal), parameter_bindings (immutable JSON), baseline_pin (exact 037 BaselineSelectionPin; null only for a revision declaring no baseline-backed checks), target_snapshot, execution_principal_fingerprint, execution_toggles (optional evidence only), trigger_source (server-owned), agent_run_id? (provenance), verification_run_id? (aggregation), idempotency_key (unique), started_at, finished_at, resume_token, error_code, runner_version.
live_execution_binding_ref? plus live_execution_binding_snapshot? are the only persisted live-I/O
coordinates. The snapshot is an exact allowlist: binding ref; environment/release/query-model,
execution-principal and RLS/security fingerprints; browser-safe checkpoint/action refs; and evidence
owner/ref policy. It contains no credential, client, cookie, raw browser context, callable principal,
or capture bytes. LiveExecutionCompositionRoot resolves this immutable identity server-side at
startup; a missing provider is typed unavailable, and any exact-snapshot mismatch is typed non-pass
before I/O.
The deployment-owned settings.scenario_live_execution_bindings[] record stores enabled,
binding_snapshot, and query_model_snapshot only. Startup resolves its Environment credentials
server-side to build the existing SupersetClient, then registers the exact model/storage tuple.
Browser and screenshot providers are process-local registrations and are never serialized in this
record; absent registrations yield stable configured-unavailable outcomes.
If the selected revision contains a human step, it is derived as manual_run_only=true: it may start
only from the authenticated analyst manual-run route. Scheduler, deploy/release/ETL/API trigger and any
background runner or recovery worker are ineligible. The eligibility check occurs before idempotency,
ScenarioRun, ActionApprovalGate, notification, queue insertion and dispatcher CAS; no HumanCheckpoint may
be created, skipped, defaulted or converted into an automated approval path.
For a PROD request, the service atomically creates ScenarioRun(status=pending_approval) and ActionApprovalGate(owner_type=scenario_run, owner_id=run_id, operation=scenario_execution). Approval transitions only pending_approval → queued; denial/expiry transitions to blocked. Schedulers and external triggers therefore receive a durable run/intent, never an unusable 403. trigger_source is set only by the trusted entry point (manual route, scheduler, deploy connector, or API key), never by a bearer client field. Idempotency uses canonical execution-request hash: same key + same hash returns the existing run; same key + different hash returns 409 IDEMPOTENCY_KEY_REUSED.
EnvironmentExecutionPolicy is resolved server-side from the configured Environment record
(stage=PROD or is_production=true) before RBAC, idempotency, run, gate, queue, notification, or
adapter work. Any compatibility is_prod request field is non-authoritative and cannot upgrade or
downgrade that class; an unknown environment fails closed without creating a run or gate.
ParameterBinding, ExecutionPrincipal, and TargetSnapshot
ParameterBinding: parameter_name, resolved_value, source, resolved_at, validation_fingerprint. It is derived at start from the 038 ParameterDefinition and launch input; it is never embedded in ScenarioRevision or its content_hash.
ExecutionPrincipal: auth_mode, actor_id?, service_identity?, impersonated_user?, effective_roles_hash, rls_context_hash. Its immutable fingerprint is stored on the run and used for every Superset request.
TargetSnapshot: environment_id, dashboard_release_id, dashboard_fingerprint, dataset_lineage_fingerprint, captured_at. It is mandatory even when no release was selected, so environment_id=preprod is never treated as an immutable target.
AnalyticsContextKey is a server-derived SHA-256 over environment_class + compatibility_family + baseline_family + dashboard_release_id + dashboard_fingerprint + dataset_lineage_fingerprint + execution_principal_fingerprint. It is captured on the run and every step result; analytics never groups runs merely by environment or revision id.
Statuses: queued, running, waiting_human, blocked, cancel_requested, cancelled, passed, failed, inconclusive.
Terminal failed, inconclusive and blocked outcomes emit a 036 InvestigationSignal with this immutable execution snapshot. 047 deterministically creates/updates the Queue/Episode from that signal. Signal delivery never changes run status and never auto-starts an agent run; an analyst opens the case explicitly from 045/047.
ScenarioStepRun
Fields: id, run_id FK, logical_step_id (immutable UUID, from #8), step_position (mutable), step_content_hash (mutable), attempt, status (queued|running|waiting_human|passed|failed|inconclusive|blocked|skipped), started_at, finished_at, inputs_snapshot (JSON, no secrets), outputs (JSON), artifact_refs (JSON), error_code, progress, timeout_ms, side_effect_key (nullable), step_outcome (StepOutcome).
StepOutcome { status, reason_codes[], deterministic_evidence_refs[], agent_evaluation_ids[], decision_policy_id?, decision_policy_version?, decided_at } is the authoritative result of one logical step. It is distinct from executor output and from a model verdict.
AgentEvaluation and DecisionPolicy
AgentEvaluation { evaluation_id, scenario_run_id, logical_step_id, attempt, provider_id, model_id, model_version, prompt_template_id, prompt_template_version, input_manifest_hash, evidence_refs, verdict, confidence, findings, reason_codes, raw_response_artifact_ref, started_at, finished_at } is immutable evidence generated only for a declared 038 AgentEvaluationSpec. Its tool/evidence access is bounded by that spec; it cannot mutate program content, invoke mutation actions, change run state or choose downstream scheduling.
DecisionPolicy { policy_id, version, deterministic_hard_failure, high_confidence_failure, low_confidence, disagreement, missing_evidence } is a versioned deterministic mapper. Defaults: deterministic hard failure→failed; high-confidence policy-qualified agent failure→failed; low confidence→inconclusive; evaluation/evidence disagreement→inconclusive; missing required evidence→blocked; provider error for required evaluation→inconclusive. Exact first-match baseline-semantic/1.0.0 cases and strict schemas are 038 DecisionPolicy; no dynamic checkpoint creation. Scenario aggregation consumes StepOutcome, not AgentEvaluation verdicts directly.
RunnerPlan — deterministic derivation from revision (#2)
RunnerPlan is derived deterministically at run start from the selected immutable ScenarioRevision/Verification Program, NOT read from a stored runner.plan.json. Fields: scenario_revision_id, scenario_content_hash, verification_program_hash, action_registry_version, action_registry_hash, env targets, resolved params, exact baseline_pin plus canonical pin digest, topological order, and one immutable ActionExecutionDescriptor per step. A descriptor contains the exact {tool, action}, typed input/output contracts, idempotency, retry safety, side-effect-key policy, timeout, mutation contract/risk. The runner persists the descriptor snapshot in both steps and executor_mapping; it derives leases/recovery/retry policy only from that snapshot. A missing, altered, unknown, version/hash-mismatched descriptor, invalid input/output shape, or mutating action without its required contract rejects before run/lease/I/O. Run refuses if its revision/program/action-registry hashes differ from the selected revision.
@INVARIANT RunnerPlan descriptor resolution is exact and version/hash pinned; tool is never a dispatch or retry-policy fallback.
@REJECTED A universal idempotent=true, retry_safe=true claim based on a non-human tool was rejected because it can repeat unsafe effects.
The materialized runner.plan.json in git is a reference artifact, never the runtime source of truth; it may be regenerated from any revision.
ScenarioExecutionContext (#11)
Fields: run_id, browser_session (ref, not raw cookies), page/context ref, auth_context ref, current_dashboard, current_filters, artifact_namespace (owner_type=scenario_run), environment client. Browser secrets/cookies live in a separate secure context, NEVER in inputs_snapshot JSON (keeps reproducibility snapshot secret-free). Browser workers resume only by deterministic replay from the last browser-safe checkpoint (dashboard_open, filters_applied, etc.); replay records reconstruction_replay=true and never reclassifies already completed logical steps as rerun. API/XLSX/pure assertion steps may resume directly only when their executor declares retry-safe/idempotent. Browser session/context refs are valid only on the single application-owned provider event loop (loop_binding); they are never shared across runs and never cross loops (see 044 ProviderRuntime contract).
Artifact ownership (#4)
Artifacts use a generic owner: Artifact { id, owner_type: agent_run|scenario_run|verification_run|load_run, owner_id, kind, sha256, content_ref, retention_class, ... }. ScenarioRun evidence/screenshots/report/xlsx use owner_type=scenario_run; no artificial AgentRun is created. Retention_class ties to 046 tiers. Playwright trace/video dumps are not Artifact kinds: an operator debug flag may keep them locally, and they are never registered as Artifact or EvidenceReceipt.
Decision gates (#3)
ActionApprovalGate— authorization approval for PROD execution, baseline approval, repository mutation. Generalized 036 gate withowner_type+owner_id.HumanCheckpoint—checkpoint_id, run_id, logical_step_id, checkpoint_type, decision_policy, status, created_at, expires_at, eligible_role?, eligible_actor_ids?, assigned_to?, evidence_refs, decision_version, decided_by?, decided_at?, disposition?, comment?. Status ispending|decided|expired|cancelled; decision is CAS ondecision_version, stale/concurrent decision returns 409. v1 disposition maps confirm→passed, false_positive→inconclusive, inconclusive→inconclusive;manual_assertionmaps pass→passed, fail→failed, inconclusive→inconclusive. It is NOT a 036 ApprovalGate decision.
A HumanCheckpoint is never delegated to the agent: it is a manual-run-only analyst decision inside a currently executing run. External MCP clients may display evidence; only authenticated human USER decision tools may consume the checkpoint under 050. Product frontend offers human review, never agent workspace or invocation.
Worker semantics — at-least-once execution (#6)
Runtime primitives: worker lease, heartbeat, run claim, step claim, lease expiration, idempotency key, recovery scheduler. Each executor declares: idempotent? | retry-safe? | side_effect_key? | external_request_id?. POST /scenario-runs requires Idempotency-Key (unique) to prevent double-run on double-click. A crashed step with an external side effect is only re-run if idempotent/retry-safe or keyed.
Retry semantics (#14)
Retry of a failed step invalidates its downstream closure (descendants depending on its output) and re-runs them; retry after the run advanced beyond the step is rejected unless the whole closure re-runs. Bounded attempts per policy.
Result aggregation truth table (#16)
- any step
failed→ scenariofailed blockeddescendants counted asblocked(NOT failed)skippeddoes not count against passinconclusive→ scenarioinconclusiveunless afailedalso present (then failed)- warning is an evidence-level qualifier, not an execution status (045 maps from evidence)
Lifecycle
pending_approval → queued → running → waiting_human | blocked → passed | failed | inconclusive; cancel_requested → cancelled. Human decision atomically consumes the checkpoint and resumes internally. Public /resume is reserved for a recoverable infrastructure pause and requires a typed resume token/reason; it cannot consume a HumanCheckpoint. Cancel drains in-flight within a bounded window.
ScenarioExecutorRegistry and BrowserExecutor
The registry resolves ActionExecutionDescriptor -> executor; the following list states each descriptor's tool family, not a tool-only fallback:
- browser ->
BrowserExecutor→ version-pinned 038ActionRegistry→ Playwright/session infrastructure - superset_api -> 037 metric_executor_async / SupersetClient.ChartData.Execute
- sql_evidence ->
SqlEvidenceExecutor→ Superset SQL Lab backend/API → configured database connection (no credentials exposed to agent) - transform -> version-pinned bounded 038 TransformSpec DSL executor
- xlsx -> xlsx parser + 037 normalization
- assertion -> 037 comparison.py + 038 ComparisonSpec/AssertionSpec executor
- agent_evaluation -> bounded provider adapter executing a declared 038 AgentEvaluationSpec and emitting
AgentEvaluation; DecisionPolicy owns StepOutcome - screenshot -> 038 capture.py + ScreenshotService (owner_type=scenario_run)
- report -> report-template render + artifact (owner_type=scenario_run)
- artifact -> generic artifact register (owner_type=scenario_run)
- human -> EXCLUDED (runner-lifecycle HumanCheckpoint control, not an executor)
BrowserExecutor implements registered actions: open_dashboard, navigate_tab, navigate_tabs, apply_native_filter, inspect_filter_state, inspect_filter_options, apply_table_filter, extract_table, scroll_to, inspect_columns, click, select_rows, assert_dom, wait_for_selector, edit_row, bulk_edit, download, refresh, wait_for_state. pagination and navigate_dashboard remain registry-declared but disabled until their drivers and live canary gates close. ScreenshotService is evidence infrastructure only, not the browser action executor. It validates the registry-declared inputs/outputs/risk/timeout/idempotency before dispatch. Composite navigate_tabs is bounded to the visible tab nav and returns the visited labels; assert_dom evaluates only declared structural DOM predicates and cannot embed business metric truth; inspect_filter_options is observe-only; wait_for_selector is bounded and fail-closed.
Provider runtime protocol
Every provider implements the following logical records. They are persisted or emitted as immutable operation/evidence receipts; a Python callback alone is not a sufficient provider implementation.
ProviderExecutionContext: run_id, logical_step_id, attempt, descriptor_fingerprint,
live_binding_ref, execution_principal_fingerprint, deadline_at, idempotency_key,
capacity_lease_id, trace_id, cancellation_requested, cancellation_deadline_at.
ProviderOperationReceipt: operation_id, context identity, provider_id, provider_version,
started_at, finished_at, effect_state, external_request_id?, resource_refs,
cleanup_state, reconciliation_state, result_digest?. Receipts are append-only; late responses are
historical records and cannot mutate the winning attempt.
ProviderEvidenceReceipt: evidence_ref, owner_type=scenario_run, owner_id, run_id,
logical_step_id, attempt, operation_id, provider_id/version, descriptor_fingerprint,
binding_ref, execution_principal_fingerprint, content_type, byte_length, sha256,
retention_class, created_at. The durable store must atomically commit the content and receipt or
return no usable ref.
Provider state is registered|health_checked|capacity_claimed|admitted|running|completed|failed| cancel_requested|finalized|reconciliation_required. A provider exposes separate liveness, readiness and
dependency-health checks and a redacted capability/limit document. Readiness gates new admissions only;
it does not remove active leases or evidence.
Cancellation calls cancel(operation_id) and records stopped|completed|unknown. unknown means the
provider cannot prove whether an external effect happened: retry is forbidden until
reconcile(operation_id, idempotency_key) returns a terminal receipt. If reconciliation is unavailable
before the bounded recovery deadline, the runner terminalizes inconclusive or blocked according to
effect risk. Capacity is released only after receipt and cleanup state are durable.
Provider-specific schemas and obligations
| Provider | Input boundary | Required result/evidence | Resource and recovery rules |
|---|---|---|---|
| Browser | 038 action descriptor, server-owned auth/session binding, typed action input | operation receipt, safe checkpoint, page/action provenance and durable evidence where declared | isolated context, bounded pages/downloads, cleanup on finalization, safe-checkpoint replay only |
| Superset API | exact 037 query envelope and pinned model/principal/RLS | request/response fingerprints, raw response SHA-256 and durable ref | request/response limits; 429/5xx bounded retry; unknown timeout reconciles external request |
| SQL evidence | immutable 038 SqlEvidenceSpec plus typed ParameterBindings | DB identity, model/principal/RLS fingerprints, raw response digest/ref | read-only; caller SQL rejected; reconcile before retry |
| XLSX | verified server-owned artifact ref | source ownership/digest and canonical normalized output | archive/cell/formula limits; local parse is idempotent |
| Assertion | canonical values plus pinned 037 ComparisonPolicy | deterministic input manifest, policy/version and StepOutcome | pure, no external resources, retry-safe |
| Transform | bounded 038 TransformSpec DSL | source manifest, transform fingerprint and canonical output digest | hard operation/row/column/output limits; no code/SQL/network |
| Screenshot | bound capture session/profile and durable evidence policy | atomic evidence receipt with media type, length and SHA-256 | temp cleanup; unknown capture reconciles before retry |
| Report | pinned template and immutable evidence manifest | template/manifest/output digest and durable ref | bounded deterministic renderer; partial writes reconcile or discard |
| Artifact | authorized producer ref or verified source artifact | producer receipt, owner tuple, digest and retention | never accept caller digest alone; registration idempotent |
| AgentEvaluation | immutable 038 spec, evidence allowlist, provider/model/prompt versions | immutable AgentEvaluation and raw response artifact receipt | token/cost/time limits; verdict cannot directly set StepOutcome |
BrowserProvider action contract
Each browser step persists the exact 038 action descriptor plus input_schema_version,
output_schema_version, risk, checkpoint_policy, idempotency, retry_safe, limits and mutation
contract fingerprint. Defaults are context/authentication timeout 120000ms, action timeout 30000ms,
maximum 3 pages, maximum download 26214400 bytes and maximum screenshot 10485760 bytes. A descriptor
may lower a limit but cannot raise it from scenario payload.
open_dashboard, navigate_tab, filter application, inspection, scroll, extraction, refresh and wait are
read-only/evidence actions; click is rejected unless its effect class is explicit; edit_row and
bulk_edit are mutation actions; select_rows is selection-only; download produces a server-owned
artifact ref, never a local path. extract_table is bounded to 10,000 rows, 100 columns and 10 MiB output.
Browser resource state is lease_claimed|context_created|authenticated|action_running|checkpointed| effect_recorded|cleanup_started|receipt_finalized|context_closed. One active context is allowed per
(run_id, capacity_lease_id) and cleanup failure prevents PASS. Concurrency defaults are 2 concurrent
browser contexts for DEV/PREPROD and 1 for PROD per environment, admitted by the shared CapacityManager;
screenshot capture shares the same lease accounting, and capacity exhaustion keeps the run queued with
CAPACITY_BLOCKED. Recovery never attaches to a dead context;
it reconstructs from a safe checkpoint and records reconstruction_replay=true. Unknown mutation effects
are held for reconciliation and cannot retry.
Readiness requires browser executable/version, auth binding, allowed-origin, evidence-storage, cleanup-worker and cancellation-capability checks. Deployment requires a read-only PREPROD canary and forced timeout/ cleanup canary before enabling; mutation requires a separate fixture-lease canary.
AgentEvaluation and DecisionPolicy integration
The AgentEvaluation provider claims capacity as workload_class=agent_evaluation, creates an immutable
input manifest from deterministic evidence refs, executes within the declared budget, and stores the raw
response before publishing a verdict. Provider/model/prompt versions and manifest hash are mandatory.
DecisionPolicy(policy_id, version) consumes only the immutable evaluation record plus deterministic
evidence and emits StepOutcome. Low confidence, missing evidence or disagreement follow the pinned
policy table; the model cannot select a policy, mutate the graph, invoke a provider action, consume a
HumanCheckpoint or schedule downstream work.
ExecutionCapacityManager integration
Before any provider I/O, the dispatcher calls the shared allocator with environment, workload_class,
provider_id, priority, requested_units, quota, reserved_capacity, run/step and deadline. The
atomic result is a durable CapacityLease; no lease means queued or blocked according to policy and no
provider invocation. Heartbeat, expiry, release and forced reconciliation are idempotent. Retries claim
new leases, while active unknown operations retain their lease until reconciliation or terminal closure.
Verification thresholds
scenario_content_hash,verification_program_hash,descriptor_fingerprint,query_model_fingerprint,execution_principal_fingerprint,rls_security_fingerprintand evidencesha256are canonical lowercase hexadecimal SHA-256 values with exactly 64 characters.attemptstarts at 1 and increases by exactly 1 for each new logical-step attempt. Historical attempts are immutable; only one attempt may be the active projection for a logical step.byte_lengthis a positive integer equal to the durable content length. A zero-byte evidence object is never eligible for PASS.deadline_atis later thanstarted_at; a result received after the deadline cannot become the winning outcome. The runner records the late result as historical operation data.- CapacityLease validity is
[claimed_at, expires_at]; provider I/O requiresnow < expires_atand a matching run/step/attempt/provider context. Lease release is idempotent: repeated release changes no terminal state or accounting total. - Cancellation finalization must occur no later than
cancel_drain_deadline_at + 5 secondsunder the default scheduler interval. Any exception is a release failure and keeps the run non-GO.
Mutation authorization and execution policy
Authorization is evaluated per run and step as environment_class + scenario risk profile + action mutation profile. Read-only PROD steps require the ScenarioExecution approval policy. Every mutation requires the immutable 038 mutation contract, scoped target keys, precondition evidence, side-effect identity, cleanup/reconciliation outcome and retry_safe=false unless the registry proves an idempotent compensating action. Mutating browser steps in PROD are prohibited. Non-PROD test-data mutation may be delegated only inside an authorized fixture lease; this mutation policy is separate from PROD execution approval.
Immutable Execution Snapshot
A ScenarioRun pins scenario_revision_id + scenario_content_hash at start. Results carry provenance: scenario_revision_id, runner_version, template_version, baseline_revision, target_snapshot, parameter_bindings, execution_principal_fingerprint, query fingerprints. Later edits never alter completed runs.
ExecutionCapacityManager
All AgentRun, VerificationRun, LoadRun, and ScenarioRun claims pass through one environment-scoped capacity manager: environment, workload_class, priority, quota, reserved_capacity. A 046 scenario policy is a consumer of this global allocator, not an independent PROD/PREPROD concurrency limit.
Production records and compatibility — 2026-09-08
Result schema defines the unified run projection. AgentEvaluation schema replaces the abbreviated field list above for new records: operation/spec identity, input artifact MIME/length/SHA manifest, provider/model/prompt/schema versions, criterion-bound findings, raw response provenance, usage/pricing and full baseline pin. Evaluations/comparisons/outcomes are append-only; unique active attempt is selected by CAS. Historical late attempts remain visible but never replace a winner.
Run request hash includes the server-resolved baseline set/version, catalog/release/publication commit and entry IDs/digests before run/gate creation. The exact pin is persisted in RunnerPlan, ScenarioRun, result and AnalyticsContextKey provenance; comparison never uses first-map-entry or moving latest fallback. Missing legacy pins are ineligible, not fabricated.
Artifact content uses protected GET/HEAD, not public URLs. Metadata carries owner/run/step/attempt/operation, SHA-256, actual MIME, byte length, retention/expiry and availability. Required baseline bytes have retention holds. Production chain owns provider startup/shutdown/cancel/reconcile. New records and lifecycle acceptance are live-verified for the provider/artifact/evaluation chain (2026-09-17: T042b/T044/T045 canaries + T022 audit); the graph-level terminal PASS remains pending (T046 binding). Frontend only renders read-only evaluation evidence; no provider/prompt/retry-agent controls. Approved performance baseline is outside scope.
#endregion ScenarioExecution.DataModel