feat(scenario): 044 provider INV_7 decomposition + T044/T045 live canaries; evidence-backed spec status
The 044 provider lifecycle is decomposed to satisfy INV_7 and closed with live canaries; the remaining acceptance row (T046 graph-level terminal PASS) is diagnosed with executable proof. - split browser_readonly_actions (534->160), browser_session (501->58) and browser.py (407->398) into sibling modules with frozen contract IDs and frozen import surface (facades re-export every moved public name); monkeypatch seams follow the owning modules - gates per phase: scoped 044 suite 1590 passed and full backend 11269 passed / 0 failed, no delta - T042b: six PREPROD browser canaries GREEN under the Wave-C run-scoped session architecture; harness gains SS_CANARY_STORAGE_ROOT and fixes the dead-loop scheduler drain hang - T044: new artifact-content live HTTP canary 11/11 (real app, PostgreSQL, JWTs, live-captured bytes) - T045: new fault-injection canary 16/16 (bounded loop refusals, deadline cancellation, shutdown drain, capacity quarantine + TTL reconcile, receipt CAS, late response, cancel drain) - T046: walker-level evaluation->comparison binding diagnosed (compare-only + mandatory -> EVALUATION_UNAVAILABLE; bound evaluation -> BASELINE_AND_SEMANTIC_PASS); T043/T046 stay open - refresh 044 spec/traceability/checklist/quickstart/SESSION_STATE/contracts/plan/data-model and the 037/050 rows to evidence-backed states; split three malformed multi-target @RELATION lines - fix skill drift in semantics-python/semantics-svelte (logger facade, notify(), i18n proxy)
@@ -20,42 +20,30 @@ Load this skill when implementing Python backend code under the GRACE-Poly proto
|
||||
|
||||
superset-tools uses the canonical **Molecular CoT Logging** protocol for belief markers. For the full wire-format specification, see the `molecular-cot-logging` skill.
|
||||
|
||||
**ALWAYS import from the shared module — never copy-paste inline:**
|
||||
**SSOT layout (ADR-0022 absorbed the former `shared/` package into backend — `ss_tools.shared.cot_logger` no longer exists):**
|
||||
|
||||
- primitives: `backend/src/core/cot_logger.py` (`src.core.cot_logger`) — ContextVars (`trace_id`/`span_id`/`task_id`/`contract_id`), `build_cot_event`, `resolve_contract_id`, module-internal `log(src, marker, intent, ...)`;
|
||||
- facade — the ONLY call-site convention in `backend/src`: `src.core.logger` — `belief_scope`, `logger.reason/reflect/explore`, `level=` kwarg;
|
||||
- formatter: `src.core.cot_formatter`.
|
||||
|
||||
```python
|
||||
from ss_tools.shared.cot_logger import log, push_span, pop_span
|
||||
from src.core.logger import belief_scope, logger
|
||||
|
||||
# Usage:
|
||||
# log("src_id", "REASON", "intent", payload_dict)
|
||||
# log("src_id", "EXPLORE", "message", payload_dict, error="assumption violated")
|
||||
# log("src_id", "REFLECT", "outcome", payload_dict)
|
||||
_SRC = "Migration.RunTask" # contract id — mirrors into contract_id
|
||||
|
||||
logger.reason("Starting migration task", src=_SRC, payload={"task_id": task_id})
|
||||
logger.reflect("Migration completed", src=_SRC, payload={"dashboards": len(result)})
|
||||
logger.explore("Migration failed, rolling back", src=_SRC, claim="POST: status terminal",
|
||||
error_code="MIGRATION_ROLLBACK", payload={"task_id": task_id}, error=str(exc))
|
||||
with belief_scope("Core.Auth.Login", claim="POST: token issued"):
|
||||
...
|
||||
```
|
||||
|
||||
Thin context-manager wrappers (backward-compatible aliases for `push_span`/`pop_span`):
|
||||
|
||||
```python
|
||||
from contextlib import contextmanager
|
||||
|
||||
@contextmanager
|
||||
def belief_scope(contract_id: str):
|
||||
prev_span = push_span(contract_id)
|
||||
log(contract_id, "REASON", "enter")
|
||||
try:
|
||||
yield
|
||||
except Exception as e:
|
||||
log(contract_id, "EXPLORE", "error", error=str(e))
|
||||
raise
|
||||
else:
|
||||
log(contract_id, "REFLECT", "exit")
|
||||
finally:
|
||||
pop_span(prev_span)
|
||||
```
|
||||
|
||||
**CRITICAL:** Import CoT helpers from `ss_tools.shared.cot_logger` (shared package SSOT). Backend call sites may use the facade `from src.core.logger import log, belief_scope, logger`. Never define `reason()`, `explore()`, `reflect()` inline — use the canonical `log()` function. Do NOT manually type `[REASON]` in message strings; `log()` emits the marker field automatically in the molecular-cot JSON wire format. Do not invent `ss_tools.lib.cot_logger` — that module does not exist.
|
||||
**CRITICAL:** the FIRST positional binds as `intent`; `src=`/`payload=`/`error=`/`level=`/`contract_id=`/`claim=`/`error_code=` are keywords. The src-first primitive `log(src, marker, intent, ...)` is module-internal — direct production use in `backend/src` is forbidden and pinned executable by `backend/tests/test_core/test_logger_wire_format.py` (repo-wide AST sweeps over the two-positional facade misuse and direct `log` imports). Never define `reason()`/`explore()`/`reflect()` inline; never type `[REASON]` into message strings — the facade emits the marker field in the molecular-cot JSON wire format. Do not import `ss_tools.shared.cot_logger` or invent `ss_tools.lib.cot_logger` — neither exists.
|
||||
|
||||
## II. PYTHON COMPLEXITY EXAMPLES
|
||||
|
||||
Live exemplars in this repo (prefer these over the sketches): `shared/src/ss_tools/shared/cot_logger.py`, `backend/src/core/task_manager/manager.py`. Sketches below show shape only — do not copy their `@`-tags into unrelated files.
|
||||
Live exemplars in this repo (prefer these over the sketches): `backend/src/core/cot_logger.py` (SSOT primitive), `backend/src/core/logger.py` (facade), `backend/src/core/task_manager/manager.py`. Sketches below show shape only — do not copy their `@`-tags into unrelated files. All logging calls in the sketches use the intent-first facade (`logger.reason/reflect/explore`), never the src-first primitive.
|
||||
|
||||
### C1 (Atomic) — DTOs, Pydantic schemas, simple constants
|
||||
```python
|
||||
@@ -114,28 +102,31 @@ def migrate_dashboard(source_client, target_client, dashboard_id: str, db_mappin
|
||||
# @RELATION DEPENDS_ON -> [MigrationService]
|
||||
# @RELATION DEPENDS_ON -> [WebSocketNotifier]
|
||||
async def run_migration_task(task_id: str, db_session) -> dict:
|
||||
log("Migration.RunTask", "REASON", "Starting migration task", {"task_id": task_id})
|
||||
logger.reason("Starting migration task", src=_SRC, payload={"task_id": task_id})
|
||||
task = await db_session.get(Task, task_id)
|
||||
if not task:
|
||||
log("Migration.RunTask", "EXPLORE", "Task not found", error="TaskNotFound")
|
||||
logger.explore("Task not found", src=_SRC, claim="PRE: task row exists",
|
||||
error_code="MIGRATION_TASK_NOT_FOUND", payload={"task_id": task_id})
|
||||
raise TaskNotFoundError(task_id)
|
||||
try:
|
||||
task.status = "RUNNING"
|
||||
await db_session.commit()
|
||||
log("Migration.RunTask", "REASON", "Task status set to RUNNING", {"task_id": task_id})
|
||||
logger.reason("Task status set to RUNNING", src=_SRC, payload={"task_id": task_id})
|
||||
result = await execute_migration_plan(task.migration_plan)
|
||||
task.status = "COMPLETED"
|
||||
task.result = result
|
||||
await db_session.commit()
|
||||
await notify_frontend(task_id, "completed", result)
|
||||
log("Migration.RunTask", "REFLECT", "Migration completed", {"task_id": task_id, "dashboards": len(result)})
|
||||
logger.reflect("Migration completed", src=_SRC, claim="POST: status terminal",
|
||||
payload={"task_id": task_id, "dashboards": len(result)})
|
||||
return result
|
||||
except Exception as e:
|
||||
log("Migration.RunTask", "EXPLORE", "Migration failed, rolling back", {"task_id": task_id}, error=str(e))
|
||||
logger.explore("Migration failed, rolling back", src=_SRC,
|
||||
error_code="MIGRATION_FAILED", payload={"task_id": task_id}, error=str(e))
|
||||
task.status = "FAILED"
|
||||
task.error = str(e)
|
||||
await db_session.commit()
|
||||
await notify_frontend(task_id, "failed", {"error": str(e)})
|
||||
await notify_frontend(task_id, "failed", error=str(e))
|
||||
raise
|
||||
# #endregion Migration.RunTask
|
||||
```
|
||||
@@ -157,17 +148,19 @@ async def run_migration_task(task_id: str, db_session) -> dict:
|
||||
# @REJECTED Incremental-only update was rejected — it leaves stale edges when contracts
|
||||
# are deleted; only full scan guarantees consistency.
|
||||
def rebuild_index(root_path: str) -> dict:
|
||||
log("Index.Rebuild", "REASON", "Scanning source files", {"root": root_path})
|
||||
logger.reason("Scanning source files", src=_SRC, payload={"root": root_path})
|
||||
contracts = []
|
||||
for filepath in scan_files(root_path):
|
||||
try:
|
||||
parsed = parse_contract(filepath)
|
||||
contracts.append(parsed)
|
||||
except Exception as e:
|
||||
log("Index.Rebuild", "EXPLORE", "Parse failure, skipping file", {"file": filepath}, error=str(e))
|
||||
logger.explore("Parse failure, skipping file", src=_SRC,
|
||||
error_code="INDEX_PARSE_FAILURE", payload={"file": filepath}, error=str(e))
|
||||
snapshot = {"contracts": contracts, "timestamp": datetime.utcnow().isoformat()}
|
||||
write_checkpoint(root_path, snapshot)
|
||||
log("Index.Rebuild", "REFLECT", "Rebuild complete", {"contracts": len(contracts)})
|
||||
logger.reflect("Rebuild complete", src=_SRC, claim="INVARIANT: ids map to nodes",
|
||||
payload={"contracts": len(contracts)})
|
||||
return snapshot
|
||||
# #endregion Index.Rebuild
|
||||
```
|
||||
@@ -260,27 +253,20 @@ python -m mypy src/
|
||||
```
|
||||
|
||||
## V. FASTAPI / ASYNC PATTERNS
|
||||
|
||||
### Async belief scope
|
||||
```python
|
||||
from contextlib import asynccontextmanager
|
||||
from ss_tools.shared.cot_logger import log, push_span, pop_span
|
||||
|
||||
@asynccontextmanager
|
||||
async def async_belief_scope(contract_id: str):
|
||||
prev_span = push_span(contract_id)
|
||||
log(contract_id, "REASON", "enter")
|
||||
try:
|
||||
yield
|
||||
except Exception as e:
|
||||
log(contract_id, "EXPLORE", "error", error=str(e))
|
||||
raise
|
||||
else:
|
||||
log(contract_id, "REFLECT", "exit")
|
||||
finally:
|
||||
pop_span(prev_span)
|
||||
`belief_scope` is a plain (sync) context manager — use it directly around `await` blocks; there is no async wrapper and none is needed. Bind the contract once, emit markers per step:
|
||||
|
||||
```python
|
||||
from src.core.logger import belief_scope, logger
|
||||
|
||||
with belief_scope("Migration.RunTask"):
|
||||
result = await execute_migration_plan(task.migration_plan)
|
||||
logger.reflect("Step persisted", src=_SRC, payload={"rows": len(result)})
|
||||
```
|
||||
|
||||
`seed_trace_id()` / `push_span()` / `pop_span()` are imported from the SSOT module `src.core.cot_logger` where a component seeds or scopes traces. Do not hand-roll span plumbing in call sites.
|
||||
|
||||
### Dependency injection convention
|
||||
- Use FastAPI `Depends()` for injecting services
|
||||
- Services are singletons or request-scoped
|
||||
|
||||
@@ -40,7 +40,7 @@ You are bound by strict repository-level design rules:
|
||||
6. **Component Reuse:** Before creating any new component, scan the existing library:
|
||||
- **Atoms:** `$lib/ui/Button.svelte`, `$lib/ui/Select.svelte`, `$lib/ui/Input.svelte`, `$lib/ui/Card.svelte`
|
||||
- **Widgets:** `$lib/components/ui/SearchableMultiSelect.svelte`, `$lib/components/ui/MultiSelect.svelte`
|
||||
- **Infrastructure:** `addToast()` from `$lib/toasts.js` (Toast already mounted in root layout)
|
||||
- **Infrastructure:** `notify()` from `$lib/toasts.svelte.ts` (Toast already mounted in root layout)
|
||||
- **Patterns (no component needed):** badges (`rounded-full px-2.5 py-0.5 text-xs font-medium`), tooltips (native `title`), skeletons (`animate-pulse bg-gray-200`), collapsibles (`<details><summary>`), empty states (`border-dashed bg-gray-50`), confirmations (`confirm()`)
|
||||
Refer to `.agents/commands/speckit.plan.md` §"Frontend Component Reuse Scan" for the mandatory scan workflow.
|
||||
|
||||
@@ -64,7 +64,7 @@ Key stores in `frontend/src/lib/stores/` (bind with `@RELATION BINDS_TO` only wh
|
||||
- `translationRunStore` — Active translation run
|
||||
- `activityStore` — Activity feed
|
||||
- `environmentContext` — Selected environment
|
||||
- Toasts: `addToast()` / `notifications` from `$lib/toasts` — not a domain store named `notificationStore`
|
||||
- Toasts: `notify({ type, message })` / `notifications` from `$lib/toasts.svelte.ts` — not a domain store named `notificationStore`
|
||||
- Screen-level dashboards/migration/git state lives in `[TYPE Model]` (`DashboardHubModel`, `MigrationModel`, `GitManagerModel`, `AgentChatModel`), not in a global `dashboardStore` / `migrationStore`
|
||||
|
||||
**Store subscription rules:**
|
||||
@@ -316,8 +316,8 @@ Region format for HTML/Svelte comments:
|
||||
<!-- @LAYER UI -->
|
||||
<!-- @RELATION DEPENDS_ON -> [StatusBadge] -->
|
||||
<!-- @RELATION DEPENDS_ON -> [ProgressBar] -->
|
||||
<!-- @RELATION DEPENDS_ON -> [$lib/toasts] -->
|
||||
<!-- @RELATION BINDS_TO -> [taskDrawerStore] -->
|
||||
<!-- @RELATION BINDS_TO -> [notificationStore] -->
|
||||
<!-- @UX_STATE Idle -> Default card view with task summary. -->
|
||||
<!-- @UX_STATE Loading -> Action button disabled, spinner active, progress bar animated. -->
|
||||
<!-- @UX_STATE Error -> Card border + bg use destructive tokens, retry button visible. -->
|
||||
@@ -331,10 +331,10 @@ Region format for HTML/Svelte comments:
|
||||
<script lang="ts">
|
||||
import { fetchApi } from "$lib/api";
|
||||
import { log } from "$lib/cot-logger";
|
||||
import { t } from "$lib/i18n";
|
||||
import { t } from "$lib/i18n/index.svelte.js";
|
||||
import { Button } from "$lib/ui";
|
||||
import { taskDrawerStore } from "$lib/stores";
|
||||
import { notificationStore } from "$lib/stores";
|
||||
import { notify } from "$lib/toasts.svelte.ts";
|
||||
import StatusBadge from "./StatusBadge.svelte";
|
||||
import ProgressBar from "./ProgressBar.svelte";
|
||||
|
||||
@@ -343,6 +343,10 @@ Region format for HTML/Svelte comments:
|
||||
let isLoading = $state(false);
|
||||
let error: string | null = $state(null);
|
||||
let status: "idle" | "loading" | "success" | "error" = $state("idle");
|
||||
// i18n: `t` is a reactive dictionary proxy — subscribe via $t in templates,
|
||||
// never call it as a function. Slice the namespace once with $derived.
|
||||
const m = $derived($t.migration ?? {});
|
||||
const actions = $derived($t.actions ?? {});
|
||||
|
||||
async function handleRunMigration() {
|
||||
isLoading = true;
|
||||
@@ -355,12 +359,12 @@ Region format for HTML/Svelte comments:
|
||||
const result = await fetchApi(`/api/tasks/${taskId}/run`, { method: "POST" });
|
||||
status = "success";
|
||||
log("MigrationTaskCard", "REFLECT", "Migration completed", { taskId, result });
|
||||
notificationStore.add({ type: "success", message: $t("migration.completed", { name: dashboardName }) });
|
||||
notify({ type: "success", message: m.completed });
|
||||
} catch (e) {
|
||||
status = "error";
|
||||
error = e instanceof Error ? e.message : "Migration failed";
|
||||
log("MigrationTaskCard", "EXPLORE", "Migration failed", { taskId }, error);
|
||||
notificationStore.add({ type: "error", message: $t("migration.failed", { name: dashboardName }) });
|
||||
notify({ type: "error", message: m.failed });
|
||||
} finally {
|
||||
isLoading = false;
|
||||
}
|
||||
@@ -379,7 +383,7 @@ Region format for HTML/Svelte comments:
|
||||
{status === 'success' ? 'border border-success-DEFAULT bg-success-light' : ''}
|
||||
{status !== 'error' && status !== 'success' ? 'border border-border bg-surface-card' : ''}"
|
||||
role="region"
|
||||
aria-label={$t("migration.task_card", { name: dashboardName })}
|
||||
aria-label={m.task_card}
|
||||
>
|
||||
<div class="flex items-center justify-between mb-2">
|
||||
<h3 class="font-semibold text-text">{dashboardName}</h3>
|
||||
@@ -387,7 +391,7 @@ Region format for HTML/Svelte comments:
|
||||
</div>
|
||||
|
||||
<div class="text-sm text-text-muted mb-3">
|
||||
{$t("migration.from")}: {sourceEnv} → {$t("migration.to")}: {targetEnv}
|
||||
{m.from}: {sourceEnv} → {m.to}: {targetEnv}
|
||||
</div>
|
||||
|
||||
{#if status === "loading"}
|
||||
@@ -405,10 +409,10 @@ Region format for HTML/Svelte comments:
|
||||
onclick={handleRunMigration}
|
||||
isLoading={isLoading}
|
||||
>
|
||||
{status === "error" ? $t("actions.retry") : $t("actions.run")}
|
||||
{status === "error" ? actions.retry : actions.run}
|
||||
</Button>
|
||||
<Button variant="secondary" size="sm" onclick={handleViewLogs}>
|
||||
{$t("actions.view_logs")}
|
||||
{actions.view_logs}
|
||||
</Button>
|
||||
</div>
|
||||
</div>
|
||||
|
||||
@@ -20,42 +20,30 @@ Load this skill when implementing Python backend code under the GRACE-Poly proto
|
||||
|
||||
superset-tools uses the canonical **Molecular CoT Logging** protocol for belief markers. For the full wire-format specification, see the `molecular-cot-logging` skill.
|
||||
|
||||
**ALWAYS import from the shared module — never copy-paste inline:**
|
||||
**SSOT layout (ADR-0022 absorbed the former `shared/` package into backend — `ss_tools.shared.cot_logger` no longer exists):**
|
||||
|
||||
- primitives: `backend/src/core/cot_logger.py` (`src.core.cot_logger`) — ContextVars (`trace_id`/`span_id`/`task_id`/`contract_id`), `build_cot_event`, `resolve_contract_id`, module-internal `log(src, marker, intent, ...)`;
|
||||
- facade — the ONLY call-site convention in `backend/src`: `src.core.logger` — `belief_scope`, `logger.reason/reflect/explore`, `level=` kwarg;
|
||||
- formatter: `src.core.cot_formatter`.
|
||||
|
||||
```python
|
||||
from ss_tools.shared.cot_logger import log, push_span, pop_span
|
||||
from src.core.logger import belief_scope, logger
|
||||
|
||||
# Usage:
|
||||
# log("src_id", "REASON", "intent", payload_dict)
|
||||
# log("src_id", "EXPLORE", "message", payload_dict, error="assumption violated")
|
||||
# log("src_id", "REFLECT", "outcome", payload_dict)
|
||||
_SRC = "Migration.RunTask" # contract id — mirrors into contract_id
|
||||
|
||||
logger.reason("Starting migration task", src=_SRC, payload={"task_id": task_id})
|
||||
logger.reflect("Migration completed", src=_SRC, payload={"dashboards": len(result)})
|
||||
logger.explore("Migration failed, rolling back", src=_SRC, claim="POST: status terminal",
|
||||
error_code="MIGRATION_ROLLBACK", payload={"task_id": task_id}, error=str(exc))
|
||||
with belief_scope("Core.Auth.Login", claim="POST: token issued"):
|
||||
...
|
||||
```
|
||||
|
||||
Thin context-manager wrappers (backward-compatible aliases for `push_span`/`pop_span`):
|
||||
|
||||
```python
|
||||
from contextlib import contextmanager
|
||||
|
||||
@contextmanager
|
||||
def belief_scope(contract_id: str):
|
||||
prev_span = push_span(contract_id)
|
||||
log(contract_id, "REASON", "enter")
|
||||
try:
|
||||
yield
|
||||
except Exception as e:
|
||||
log(contract_id, "EXPLORE", "error", error=str(e))
|
||||
raise
|
||||
else:
|
||||
log(contract_id, "REFLECT", "exit")
|
||||
finally:
|
||||
pop_span(prev_span)
|
||||
```
|
||||
|
||||
**CRITICAL:** Import CoT helpers from `ss_tools.shared.cot_logger` (shared package SSOT). Backend call sites may use the facade `from src.core.logger import log, belief_scope, logger`. Never define `reason()`, `explore()`, `reflect()` inline — use the canonical `log()` function. Do NOT manually type `[REASON]` in message strings; `log()` emits the marker field automatically in the molecular-cot JSON wire format. Do not invent `ss_tools.lib.cot_logger` — that module does not exist.
|
||||
**CRITICAL:** the FIRST positional binds as `intent`; `src=`/`payload=`/`error=`/`level=`/`contract_id=`/`claim=`/`error_code=` are keywords. The src-first primitive `log(src, marker, intent, ...)` is module-internal — direct production use in `backend/src` is forbidden and pinned executable by `backend/tests/test_core/test_logger_wire_format.py` (repo-wide AST sweeps over the two-positional facade misuse and direct `log` imports). Never define `reason()`/`explore()`/`reflect()` inline; never type `[REASON]` into message strings — the facade emits the marker field in the molecular-cot JSON wire format. Do not import `ss_tools.shared.cot_logger` or invent `ss_tools.lib.cot_logger` — neither exists.
|
||||
|
||||
## II. PYTHON COMPLEXITY EXAMPLES
|
||||
|
||||
Live exemplars in this repo (prefer these over the sketches): `shared/src/ss_tools/shared/cot_logger.py`, `backend/src/core/task_manager/manager.py`. Sketches below show shape only — do not copy their `@`-tags into unrelated files.
|
||||
Live exemplars in this repo (prefer these over the sketches): `backend/src/core/cot_logger.py` (SSOT primitive), `backend/src/core/logger.py` (facade), `backend/src/core/task_manager/manager.py`. Sketches below show shape only — do not copy their `@`-tags into unrelated files. All logging calls in the sketches use the intent-first facade (`logger.reason/reflect/explore`), never the src-first primitive.
|
||||
|
||||
### C1 (Atomic) — DTOs, Pydantic schemas, simple constants
|
||||
```python
|
||||
@@ -114,28 +102,31 @@ def migrate_dashboard(source_client, target_client, dashboard_id: str, db_mappin
|
||||
# @RELATION DEPENDS_ON -> [MigrationService]
|
||||
# @RELATION DEPENDS_ON -> [WebSocketNotifier]
|
||||
async def run_migration_task(task_id: str, db_session) -> dict:
|
||||
log("Migration.RunTask", "REASON", "Starting migration task", {"task_id": task_id})
|
||||
logger.reason("Starting migration task", src=_SRC, payload={"task_id": task_id})
|
||||
task = await db_session.get(Task, task_id)
|
||||
if not task:
|
||||
log("Migration.RunTask", "EXPLORE", "Task not found", error="TaskNotFound")
|
||||
logger.explore("Task not found", src=_SRC, claim="PRE: task row exists",
|
||||
error_code="MIGRATION_TASK_NOT_FOUND", payload={"task_id": task_id})
|
||||
raise TaskNotFoundError(task_id)
|
||||
try:
|
||||
task.status = "RUNNING"
|
||||
await db_session.commit()
|
||||
log("Migration.RunTask", "REASON", "Task status set to RUNNING", {"task_id": task_id})
|
||||
logger.reason("Task status set to RUNNING", src=_SRC, payload={"task_id": task_id})
|
||||
result = await execute_migration_plan(task.migration_plan)
|
||||
task.status = "COMPLETED"
|
||||
task.result = result
|
||||
await db_session.commit()
|
||||
await notify_frontend(task_id, "completed", result)
|
||||
log("Migration.RunTask", "REFLECT", "Migration completed", {"task_id": task_id, "dashboards": len(result)})
|
||||
logger.reflect("Migration completed", src=_SRC, claim="POST: status terminal",
|
||||
payload={"task_id": task_id, "dashboards": len(result)})
|
||||
return result
|
||||
except Exception as e:
|
||||
log("Migration.RunTask", "EXPLORE", "Migration failed, rolling back", {"task_id": task_id}, error=str(e))
|
||||
logger.explore("Migration failed, rolling back", src=_SRC,
|
||||
error_code="MIGRATION_FAILED", payload={"task_id": task_id}, error=str(e))
|
||||
task.status = "FAILED"
|
||||
task.error = str(e)
|
||||
await db_session.commit()
|
||||
await notify_frontend(task_id, "failed", {"error": str(e)})
|
||||
await notify_frontend(task_id, "failed", error=str(e))
|
||||
raise
|
||||
# #endregion Migration.RunTask
|
||||
```
|
||||
@@ -157,17 +148,19 @@ async def run_migration_task(task_id: str, db_session) -> dict:
|
||||
# @REJECTED Incremental-only update was rejected — it leaves stale edges when contracts
|
||||
# are deleted; only full scan guarantees consistency.
|
||||
def rebuild_index(root_path: str) -> dict:
|
||||
log("Index.Rebuild", "REASON", "Scanning source files", {"root": root_path})
|
||||
logger.reason("Scanning source files", src=_SRC, payload={"root": root_path})
|
||||
contracts = []
|
||||
for filepath in scan_files(root_path):
|
||||
try:
|
||||
parsed = parse_contract(filepath)
|
||||
contracts.append(parsed)
|
||||
except Exception as e:
|
||||
log("Index.Rebuild", "EXPLORE", "Parse failure, skipping file", {"file": filepath}, error=str(e))
|
||||
logger.explore("Parse failure, skipping file", src=_SRC,
|
||||
error_code="INDEX_PARSE_FAILURE", payload={"file": filepath}, error=str(e))
|
||||
snapshot = {"contracts": contracts, "timestamp": datetime.utcnow().isoformat()}
|
||||
write_checkpoint(root_path, snapshot)
|
||||
log("Index.Rebuild", "REFLECT", "Rebuild complete", {"contracts": len(contracts)})
|
||||
logger.reflect("Rebuild complete", src=_SRC, claim="INVARIANT: ids map to nodes",
|
||||
payload={"contracts": len(contracts)})
|
||||
return snapshot
|
||||
# #endregion Index.Rebuild
|
||||
```
|
||||
@@ -260,27 +253,20 @@ python -m mypy src/
|
||||
```
|
||||
|
||||
## V. FASTAPI / ASYNC PATTERNS
|
||||
|
||||
### Async belief scope
|
||||
```python
|
||||
from contextlib import asynccontextmanager
|
||||
from ss_tools.shared.cot_logger import log, push_span, pop_span
|
||||
|
||||
@asynccontextmanager
|
||||
async def async_belief_scope(contract_id: str):
|
||||
prev_span = push_span(contract_id)
|
||||
log(contract_id, "REASON", "enter")
|
||||
try:
|
||||
yield
|
||||
except Exception as e:
|
||||
log(contract_id, "EXPLORE", "error", error=str(e))
|
||||
raise
|
||||
else:
|
||||
log(contract_id, "REFLECT", "exit")
|
||||
finally:
|
||||
pop_span(prev_span)
|
||||
`belief_scope` is a plain (sync) context manager — use it directly around `await` blocks; there is no async wrapper and none is needed. Bind the contract once, emit markers per step:
|
||||
|
||||
```python
|
||||
from src.core.logger import belief_scope, logger
|
||||
|
||||
with belief_scope("Migration.RunTask"):
|
||||
result = await execute_migration_plan(task.migration_plan)
|
||||
logger.reflect("Step persisted", src=_SRC, payload={"rows": len(result)})
|
||||
```
|
||||
|
||||
`seed_trace_id()` / `push_span()` / `pop_span()` are imported from the SSOT module `src.core.cot_logger` where a component seeds or scopes traces. Do not hand-roll span plumbing in call sites.
|
||||
|
||||
### Dependency injection convention
|
||||
- Use FastAPI `Depends()` for injecting services
|
||||
- Services are singletons or request-scoped
|
||||
|
||||
@@ -40,7 +40,7 @@ You are bound by strict repository-level design rules:
|
||||
6. **Component Reuse:** Before creating any new component, scan the existing library:
|
||||
- **Atoms:** `$lib/ui/Button.svelte`, `$lib/ui/Select.svelte`, `$lib/ui/Input.svelte`, `$lib/ui/Card.svelte`
|
||||
- **Widgets:** `$lib/components/ui/SearchableMultiSelect.svelte`, `$lib/components/ui/MultiSelect.svelte`
|
||||
- **Infrastructure:** `addToast()` from `$lib/toasts.js` (Toast already mounted in root layout)
|
||||
- **Infrastructure:** `notify()` from `$lib/toasts.svelte.ts` (Toast already mounted in root layout)
|
||||
- **Patterns (no component needed):** badges (`rounded-full px-2.5 py-0.5 text-xs font-medium`), tooltips (native `title`), skeletons (`animate-pulse bg-gray-200`), collapsibles (`<details><summary>`), empty states (`border-dashed bg-gray-50`), confirmations (`confirm()`)
|
||||
Refer to `.agents/commands/speckit.plan.md` §"Frontend Component Reuse Scan" for the mandatory scan workflow.
|
||||
|
||||
@@ -64,7 +64,7 @@ Key stores in `frontend/src/lib/stores/` (bind with `@RELATION BINDS_TO` only wh
|
||||
- `translationRunStore` — Active translation run
|
||||
- `activityStore` — Activity feed
|
||||
- `environmentContext` — Selected environment
|
||||
- Toasts: `addToast()` / `notifications` from `$lib/toasts` — not a domain store named `notificationStore`
|
||||
- Toasts: `notify({ type, message })` / `notifications` from `$lib/toasts.svelte.ts` — not a domain store named `notificationStore`
|
||||
- Screen-level dashboards/migration/git state lives in `[TYPE Model]` (`DashboardHubModel`, `MigrationModel`, `GitManagerModel`, `AgentChatModel`), not in a global `dashboardStore` / `migrationStore`
|
||||
|
||||
**Store subscription rules:**
|
||||
@@ -316,8 +316,8 @@ Region format for HTML/Svelte comments:
|
||||
<!-- @LAYER UI -->
|
||||
<!-- @RELATION DEPENDS_ON -> [StatusBadge] -->
|
||||
<!-- @RELATION DEPENDS_ON -> [ProgressBar] -->
|
||||
<!-- @RELATION DEPENDS_ON -> [$lib/toasts] -->
|
||||
<!-- @RELATION BINDS_TO -> [taskDrawerStore] -->
|
||||
<!-- @RELATION BINDS_TO -> [notificationStore] -->
|
||||
<!-- @UX_STATE Idle -> Default card view with task summary. -->
|
||||
<!-- @UX_STATE Loading -> Action button disabled, spinner active, progress bar animated. -->
|
||||
<!-- @UX_STATE Error -> Card border + bg use destructive tokens, retry button visible. -->
|
||||
@@ -331,10 +331,10 @@ Region format for HTML/Svelte comments:
|
||||
<script lang="ts">
|
||||
import { fetchApi } from "$lib/api";
|
||||
import { log } from "$lib/cot-logger";
|
||||
import { t } from "$lib/i18n";
|
||||
import { t } from "$lib/i18n/index.svelte.js";
|
||||
import { Button } from "$lib/ui";
|
||||
import { taskDrawerStore } from "$lib/stores";
|
||||
import { notificationStore } from "$lib/stores";
|
||||
import { notify } from "$lib/toasts.svelte.ts";
|
||||
import StatusBadge from "./StatusBadge.svelte";
|
||||
import ProgressBar from "./ProgressBar.svelte";
|
||||
|
||||
@@ -343,6 +343,10 @@ Region format for HTML/Svelte comments:
|
||||
let isLoading = $state(false);
|
||||
let error: string | null = $state(null);
|
||||
let status: "idle" | "loading" | "success" | "error" = $state("idle");
|
||||
// i18n: `t` is a reactive dictionary proxy — subscribe via $t in templates,
|
||||
// never call it as a function. Slice the namespace once with $derived.
|
||||
const m = $derived($t.migration ?? {});
|
||||
const actions = $derived($t.actions ?? {});
|
||||
|
||||
async function handleRunMigration() {
|
||||
isLoading = true;
|
||||
@@ -355,12 +359,12 @@ Region format for HTML/Svelte comments:
|
||||
const result = await fetchApi(`/api/tasks/${taskId}/run`, { method: "POST" });
|
||||
status = "success";
|
||||
log("MigrationTaskCard", "REFLECT", "Migration completed", { taskId, result });
|
||||
notificationStore.add({ type: "success", message: $t("migration.completed", { name: dashboardName }) });
|
||||
notify({ type: "success", message: m.completed });
|
||||
} catch (e) {
|
||||
status = "error";
|
||||
error = e instanceof Error ? e.message : "Migration failed";
|
||||
log("MigrationTaskCard", "EXPLORE", "Migration failed", { taskId }, error);
|
||||
notificationStore.add({ type: "error", message: $t("migration.failed", { name: dashboardName }) });
|
||||
notify({ type: "error", message: m.failed });
|
||||
} finally {
|
||||
isLoading = false;
|
||||
}
|
||||
@@ -379,7 +383,7 @@ Region format for HTML/Svelte comments:
|
||||
{status === 'success' ? 'border border-success-DEFAULT bg-success-light' : ''}
|
||||
{status !== 'error' && status !== 'success' ? 'border border-border bg-surface-card' : ''}"
|
||||
role="region"
|
||||
aria-label={$t("migration.task_card", { name: dashboardName })}
|
||||
aria-label={m.task_card}
|
||||
>
|
||||
<div class="flex items-center justify-between mb-2">
|
||||
<h3 class="font-semibold text-text">{dashboardName}</h3>
|
||||
@@ -387,7 +391,7 @@ Region format for HTML/Svelte comments:
|
||||
</div>
|
||||
|
||||
<div class="text-sm text-text-muted mb-3">
|
||||
{$t("migration.from")}: {sourceEnv} → {$t("migration.to")}: {targetEnv}
|
||||
{m.from}: {sourceEnv} → {m.to}: {targetEnv}
|
||||
</div>
|
||||
|
||||
{#if status === "loading"}
|
||||
@@ -405,10 +409,10 @@ Region format for HTML/Svelte comments:
|
||||
onclick={handleRunMigration}
|
||||
isLoading={isLoading}
|
||||
>
|
||||
{status === "error" ? $t("actions.retry") : $t("actions.run")}
|
||||
{status === "error" ? actions.retry : actions.run}
|
||||
</Button>
|
||||
<Button variant="secondary" size="sm" onclick={handleViewLogs}>
|
||||
{$t("actions.view_logs")}
|
||||
{actions.view_logs}
|
||||
</Button>
|
||||
</div>
|
||||
</div>
|
||||
|
||||
@@ -4,7 +4,8 @@
|
||||
# SSOT for trace_id/span_id/_contract_id ContextVars, derive_src, log(), cot_span.
|
||||
# Single consumer: backend (absorbed from the former shared/ package per ADR-0022).
|
||||
# @LAYER Core
|
||||
# @RELATION CALLED_BY -> [Core.Logger.LoggerModule, Shared.CotJsonFormatter]
|
||||
# @RELATION CALLED_BY -> [Core.Logger.LoggerModule]
|
||||
# @RELATION CALLED_BY -> [Shared.CotJsonFormatter]
|
||||
# @PRE Python 3.7+ (ContextVar available).
|
||||
# @POST JSON log records written to the 'cot' Python logger.
|
||||
# @SIDE_EFFECT Writes structured JSON to the 'cot' Python logger.
|
||||
@@ -174,7 +175,9 @@ _DEFAULT_SKIP: tuple[str, ...] = (
|
||||
# @BRIEF Derive qualified src (and file:line) for logs to ensure agent-readable traces (stdlib only).
|
||||
# @INVARIANT Never emits generic "superset_tools_app" or "root" when caller can be determined.
|
||||
# @POST derive_src() returns the src string; _derive_src_loc() returns (src, loc) from ONE frame walk.
|
||||
# @RELATION CALLED_BY -> [Shared.log, Core.Logger.LoggerModule, Shared.CotJsonFormatter]
|
||||
# @RELATION CALLED_BY -> [Shared.log]
|
||||
# @RELATION CALLED_BY -> [Core.Logger.LoggerModule]
|
||||
# @RELATION CALLED_BY -> [Shared.CotJsonFormatter]
|
||||
_REPO_ROOT = Path(__file__).resolve().parents[3] # backend/src/core/ -> repo root
|
||||
|
||||
|
||||
@@ -549,7 +552,9 @@ def log(
|
||||
# @BRIEF Decorator for C4/C5 functions that auto-emits REASON on entry and
|
||||
# REFLECT/EXPLORE on exit with elapsed_ms timing.
|
||||
# @INVARIANT Always emits exactly one entry marker and one exit/explore marker.
|
||||
# @RELATION DEPENDS_ON -> [push_span, pop_span, log]
|
||||
# @RELATION DEPENDS_ON -> [Shared.push_span]
|
||||
# @RELATION DEPENDS_ON -> [Shared.pop_span]
|
||||
# @RELATION DEPENDS_ON -> [Shared.log]
|
||||
# @NOTE
|
||||
# @cot_span("REASON", "High level description")
|
||||
# async def my_important_operation(...):
|
||||
|
||||
@@ -62,6 +62,9 @@ from src.services.dashboard_testing.execution.providers.browser_admission import
|
||||
store_browser_evidence,
|
||||
store_download_artifact,
|
||||
)
|
||||
from src.services.dashboard_testing.execution.providers.browser_factory_helpers import (
|
||||
mutation_receipt_summary,
|
||||
)
|
||||
from src.services.dashboard_testing.execution.providers.browser_session import (
|
||||
BrowserCheckpointMissing,
|
||||
BrowserSessionManager,
|
||||
@@ -157,19 +160,7 @@ def build_browser_provider(
|
||||
idempotency_key=f"{run_id}:{metadata.get('logical_step_id')}:{int(metadata.get('attempt') or 1)}",
|
||||
capacity_lease_id=lease["lease_id"],
|
||||
effect_state="unknown",
|
||||
summary={
|
||||
"mutating": True,
|
||||
"precondition_hash": metadata.get("mutation_contract", {}).get("precondition_hash"),
|
||||
"mutation_context": {
|
||||
"dashboard_id": binding.dashboard_id,
|
||||
"table": descriptor.get("inputs", {}).get("table"),
|
||||
"key_columns": descriptor.get("inputs", {}).get("key_columns"),
|
||||
"assignments": descriptor.get("inputs", {}).get("assignments"),
|
||||
"target_keys": metadata.get("mutation_contract", {}).get("target_keys"),
|
||||
"field_allowlist": metadata.get("mutation_contract", {}).get("field_allowlist"),
|
||||
"database_id": descriptor.get("inputs", {}).get("database_id"),
|
||||
},
|
||||
},
|
||||
summary=mutation_receipt_summary(binding=binding, descriptor=descriptor, metadata=metadata),
|
||||
)
|
||||
operation_id = receipt["operation_id"]
|
||||
db.commit()
|
||||
|
||||
@@ -0,0 +1,35 @@
|
||||
# #region ScenarioExecution.BrowserProvider.FactoryHelpers [C:2] [TYPE Module] [SEMANTICS provider,browser,factory,helpers,receipt,summary]
|
||||
# @ingroup ScenarioExecution
|
||||
# @BRIEF Pure summary/payload builders for the browser provider factory closure; extracted per
|
||||
# INV_7 (browser.py 407 -> <400) with no behavior change.
|
||||
# @POST Builders return plain JSON-safe dicts; they perform no I/O and mutate nothing.
|
||||
# @RATIONALE The receipt summary is pure data shaping of (binding, descriptor, metadata);
|
||||
# keeping it module-level lets browser.py stay under the 400-line INV_7 limit without
|
||||
# touching the closure's typed control flow.
|
||||
# @REJECTED Extracting the typed exception->result ladder into helpers was rejected — the typed
|
||||
# reason/effect mapping per branch is the provider's review surface and must stay inline.
|
||||
from __future__ import annotations
|
||||
|
||||
from typing import Any
|
||||
|
||||
|
||||
# #region ScenarioExecution.BrowserProvider.FactoryHelpers.MutationSummary [C:2] [TYPE Function] [SEMANTICS provider,browser,receipt,summary]
|
||||
# @ingroup ScenarioExecution
|
||||
# @BRIEF Build the mutation receipt summary (mutating flag, precondition hash, mutation context).
|
||||
# @POST Returns a JSON-safe dict; fields come verbatim from binding/descriptor/metadata.
|
||||
def mutation_receipt_summary(*, binding: Any, descriptor: dict[str, Any], metadata: dict[str, Any]) -> dict[str, Any]:
|
||||
return {
|
||||
"mutating": True,
|
||||
"precondition_hash": metadata.get("mutation_contract", {}).get("precondition_hash"),
|
||||
"mutation_context": {
|
||||
"dashboard_id": binding.dashboard_id,
|
||||
"table": descriptor.get("inputs", {}).get("table"),
|
||||
"key_columns": descriptor.get("inputs", {}).get("key_columns"),
|
||||
"assignments": descriptor.get("inputs", {}).get("assignments"),
|
||||
"target_keys": metadata.get("mutation_contract", {}).get("target_keys"),
|
||||
"field_allowlist": metadata.get("mutation_contract", {}).get("field_allowlist"),
|
||||
"database_id": descriptor.get("inputs", {}).get("database_id"),
|
||||
},
|
||||
}
|
||||
# #endregion ScenarioExecution.BrowserProvider.FactoryHelpers.MutationSummary
|
||||
# #endregion ScenarioExecution.BrowserProvider.FactoryHelpers
|
||||
@@ -1,101 +1,72 @@
|
||||
# #region ScenarioExecution.BrowserProvider.ReadOnlyActions [C:5] [TYPE Module] [SEMANTICS scenario,execution,provider,browser,readonly,transport,navigate,extract,download,table-filter]
|
||||
# @ingroup ScenarioExecution
|
||||
# @BRIEF Read-only browser action flows for the T034 round-2 catalog: navigate_tab,
|
||||
# inspect_filter_state, apply_table_filter, extract_table, scroll_to, inspect_columns,
|
||||
# click, select_rows, download. Each action carries typed input validation (BROWSER_*_INVALID
|
||||
# before any I/O), fail-closed selector miss (BROWSER_SELECTOR_NOT_FOUND, no retry),
|
||||
# bounded output (extract: 10 MiB / 10 000 rows / 100 columns; download: 25 MiB), and
|
||||
# evidence (PNG for UI actions, structured rows for extract/inspect, artifact ref for
|
||||
# download) following the browser_native_filter.py pattern.
|
||||
# @BRIEF Facade for the T034 round-2 read-only browser catalog: typed input validation
|
||||
# (BROWSER_*_INVALID before any I/O) and the dispatch table onto the nine flows.
|
||||
# The flows live in browser_readonly_flows_nav.py (navigate_tab, inspect_filter_state,
|
||||
# apply_table_filter, extract_table) and browser_readonly_flows_interact.py (scroll_to,
|
||||
# inspect_columns, click, select_rows, download); limits/allowlists/selector candidates
|
||||
# and the shared resolver live in browser_readonly_limits.py.
|
||||
# @RELATION IMPLEMENTS -> [ScenarioExecution.BrowserProvider.Transport]
|
||||
# @RELATION DEPENDS_ON -> [Plugin.Service.ScreenshotService]
|
||||
# @RELATION DEPENDS_ON -> [ScenarioExecution.BrowserProvider.ReadOnlyActions.Limits]
|
||||
# @PRE The page holds an authenticated Superset dashboard session opened by _launch_and_login; all
|
||||
# inputs are admission-validated before reaching these flows.
|
||||
# @POST Every action returns typed details dict for the transport outcome; invalid input raises
|
||||
# ValueError("BROWSER_*_INVALID") before any locator work; a locator miss raises
|
||||
# BrowserTransportSelectorNotFound (fail-closed, no retry).
|
||||
# @INVARIANT Read-only: no SQL, no API writes; download is bounded to 25 MiB and stored as a
|
||||
# server-owned artifact ref with sha256; extract_table is bounded to 10 000 rows, 100
|
||||
# columns and 10 MiB output.
|
||||
# @SIDE_EFFECT Clicks, scrolls and DOM reads inside the isolated browser context; no server-side
|
||||
# data mutation; download captures bytes for provider-side storage.
|
||||
# @RATIONALE Factoring these flows into a dedicated module keeps browser_transport.py below INV_7
|
||||
# and mirrors the browser_native_filter.py isolation pattern — each action's selectors,
|
||||
# validation and UI flow are co-located for review and test.
|
||||
# @REJECTED Inlining all nine actions into browser_transport.py was rejected — the module would
|
||||
# exceed INV_7 and lose the per-action selector/timeout constants that make each flow
|
||||
# independently testable.
|
||||
# @POST validate_readonly_action_input returns a stable BROWSER_*_INVALID code or None;
|
||||
# run_readonly_action_flow returns (details, bytes|None) and raises
|
||||
# ValueError("BROWSER_ACTION_NOT_SUPPORTED") for unknown actions.
|
||||
# @INVARIANT Frozen import surface: external consumers (browser_transport, browser_admission)
|
||||
# import run_readonly_action_flow and validate_readonly_action_input from THIS module;
|
||||
# the flow-function names stay re-exported here for test/debug access. INV_7 split
|
||||
# per specs/044-dashboard-scenario-execution/plans/provider-decomposition-gate.md.
|
||||
# @SIDE_EFFECT None in this module itself; the flows perform read-only browser I/O (clicks,
|
||||
# scrolls, DOM reads, bounded downloads) inside the isolated browser context.
|
||||
# @RATIONALE Original 534-line module split to satisfy INV_7 while keeping every flow's region
|
||||
# ID, metadata and code verbatim; only the module header reflects the facade role.
|
||||
# @REJECTED Merging the flows back into browser_transport.py was rejected — see the original
|
||||
# module rationale and the decomposition-gate plan.
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
from typing import Any
|
||||
|
||||
from src.core.logger import logger
|
||||
from src.services.dashboard_testing.execution.providers.browser_native_filter import (
|
||||
BrowserTransportSelectorNotFound,
|
||||
from src.services.dashboard_testing.execution.providers.browser_readonly_flows_interact import (
|
||||
click_flow,
|
||||
download_flow,
|
||||
inspect_columns_flow,
|
||||
scroll_to_flow,
|
||||
select_rows_flow,
|
||||
)
|
||||
from src.services.dashboard_testing.execution.providers.browser_readonly_flows_nav import (
|
||||
apply_table_filter_flow,
|
||||
extract_table_flow,
|
||||
inspect_filter_state_flow,
|
||||
navigate_tab_flow,
|
||||
)
|
||||
from src.services.dashboard_testing.execution.providers.browser_readonly_limits import (
|
||||
_ALLOWED_ARTIFACT_TYPES,
|
||||
_ALLOWED_SCROLL_DIRECTIONS,
|
||||
_MAX_CLICK_SELECTOR_LENGTH,
|
||||
_MAX_COLUMN_LENGTH,
|
||||
_MAX_EXTRACT_COLUMNS,
|
||||
_MAX_EXTRACT_ROWS,
|
||||
_MAX_SCROLL_SELECTOR_LENGTH,
|
||||
_MAX_SELECT_ROWS,
|
||||
_MAX_TABLE_FILTER_VALUE_LENGTH,
|
||||
_MAX_TAB_LENGTH,
|
||||
)
|
||||
|
||||
_SRC = "ScenarioExecution.BrowserProvider.ReadOnlyActions"
|
||||
|
||||
# -- Limits (T034 spec) ---------------------------------------------------
|
||||
_MAX_EXTRACT_ROWS = 10_000
|
||||
_MAX_EXTRACT_COLUMNS = 100
|
||||
_MAX_EXTRACT_OUTPUT_BYTES = 10 * 1024 * 1024 # 10 MiB
|
||||
_MAX_DOWNLOAD_BYTES = 25 * 1024 * 1024 # 25 MiB
|
||||
_MAX_CLICK_SELECTOR_LENGTH = 512
|
||||
_MAX_SCROLL_SELECTOR_LENGTH = 512
|
||||
_MAX_SELECT_ROWS = 100
|
||||
_MAX_TAB_LENGTH = 128
|
||||
_MAX_COLUMN_LENGTH = 256
|
||||
_MAX_TABLE_FILTER_VALUE_LENGTH = 256
|
||||
|
||||
# -- Allowlists -----------------------------------------------------------
|
||||
_ALLOWED_TABS = frozenset({
|
||||
"overview", "charts", "data", "filters", "settings", "annotations",
|
||||
"css", "properties", "json", "code", "sql", "table", "dashboard",
|
||||
})
|
||||
_ALLOWED_SCROLL_DIRECTIONS = frozenset({"up", "down", "left", "right", "top", "bottom"})
|
||||
_ALLOWED_ARTIFACT_TYPES = frozenset({"xlsx", "csv", "pdf", "png", "json"})
|
||||
|
||||
# -- Multi-strategy locator candidates (Superset version drift expected) ---
|
||||
_TAB_SELECTORS = (
|
||||
'[data-test="tab-{tab}"]',
|
||||
'.nav-item a[data-tab="{tab}"]',
|
||||
'button[role="tab"][aria-controls*="{tab}"]',
|
||||
'a[href*="{tab}"]',
|
||||
)
|
||||
_TABLE_FILTER_SELECTORS = (
|
||||
'[data-test="table-filter"]',
|
||||
".table-filter-container",
|
||||
"th .ant-table-filter-column",
|
||||
)
|
||||
_TABLE_CONTAINER_SELECTORS = (
|
||||
'[data-test="table-container"]',
|
||||
".ant-table-wrapper",
|
||||
".table-container",
|
||||
".grid-content table",
|
||||
)
|
||||
_COLUMN_HEADER_SELECTORS = (
|
||||
"th .ant-table-column-title",
|
||||
"th span",
|
||||
"th",
|
||||
)
|
||||
_ROW_CHECKBOX_SELECTORS = (
|
||||
'td .ant-checkbox-input',
|
||||
'td input[type="checkbox"]',
|
||||
"td .row-select",
|
||||
)
|
||||
_DOWNLOAD_TRIGGER_SELECTORS = (
|
||||
'[data-test="download-button"]',
|
||||
'button:has-text("Download")',
|
||||
'button:has-text("Export")',
|
||||
'a[download]',
|
||||
)
|
||||
|
||||
|
||||
# =========================================================================
|
||||
# Input validation (BROWSER_*_INVALID before any I/O)
|
||||
# =========================================================================
|
||||
__all__ = [
|
||||
"apply_table_filter_flow",
|
||||
"click_flow",
|
||||
"download_flow",
|
||||
"extract_table_flow",
|
||||
"inspect_columns_flow",
|
||||
"inspect_filter_state_flow",
|
||||
"navigate_tab_flow",
|
||||
"run_readonly_action_flow",
|
||||
"scroll_to_flow",
|
||||
"select_rows_flow",
|
||||
"validate_readonly_action_input",
|
||||
]
|
||||
|
||||
# #region ScenarioExecution.BrowserProvider.ReadOnlyActions.Validate [C:3] [TYPE Function] [SEMANTICS provider,browser,readonly,validate]
|
||||
# @ingroup ScenarioExecution
|
||||
@@ -153,350 +124,6 @@ def validate_readonly_action_input(action: str, action_input: dict[str, Any]) ->
|
||||
return None
|
||||
# #endregion ScenarioExecution.BrowserProvider.ReadOnlyActions.Validate
|
||||
|
||||
|
||||
# =========================================================================
|
||||
# Selector resolution helpers
|
||||
# =========================================================================
|
||||
|
||||
# #region ScenarioExecution.BrowserProvider.ReadOnlyActions.ResolveSelector [C:2] [TYPE Function] [SEMANTICS provider,browser,readonly,selector]
|
||||
# @ingroup ScenarioExecution
|
||||
# @BRIEF Resolve the first visible locator from a list of selector templates.
|
||||
# @POST Returns a visible locator or None (caller maps to BrowserTransportSelectorNotFound).
|
||||
async def _resolve_first_visible(service: Any, page: Any, selectors: tuple[str, ...], **fmt: str) -> Any:
|
||||
candidates = [page.locator(selector.format(**fmt)) for selector in selectors]
|
||||
return await service._find_first_visible_locator(candidates)
|
||||
# #endregion ScenarioExecution.BrowserProvider.ReadOnlyActions.ResolveSelector
|
||||
|
||||
|
||||
# =========================================================================
|
||||
# Per-action UI flows
|
||||
# =========================================================================
|
||||
|
||||
# #region ScenarioExecution.BrowserProvider.ReadOnlyActions.NavigateTab [C:3] [TYPE Function] [SEMANTICS provider,browser,navigate-tab]
|
||||
# @ingroup ScenarioExecution
|
||||
# @BRIEF Click the dashboard tab identified by typed input or selector_hint.
|
||||
# @POST Returns typed details {tab, tab_navigated}; raises BrowserTransportSelectorNotFound on miss.
|
||||
async def navigate_tab_flow(
|
||||
service: Any, page: Any, action_input: dict[str, Any], *, timeout_seconds: float,
|
||||
) -> dict[str, Any]:
|
||||
tab = str(action_input.get("tab") or "").strip()
|
||||
hint = action_input.get("selector_hint")
|
||||
timeout_ms = int(timeout_seconds * 1000)
|
||||
locator = None
|
||||
if isinstance(hint, str) and hint.strip():
|
||||
locator = await service._find_first_visible_locator([page.locator(str(hint))])
|
||||
if locator is None:
|
||||
locator = await _resolve_first_visible(service, page, _TAB_SELECTORS, tab=tab)
|
||||
if locator is None:
|
||||
logger.explore("Tab locator not found", src=_SRC, payload={"tab": tab}, error_code="BROWSER_SELECTOR_NOT_FOUND")
|
||||
raise BrowserTransportSelectorNotFound("BROWSER_SELECTOR_NOT_FOUND")
|
||||
await locator.click(timeout=timeout_ms)
|
||||
await page.wait_for_load_state("domcontentloaded", timeout=timeout_ms)
|
||||
logger.reflect("Dashboard tab navigated", src=_SRC, payload={"tab": tab})
|
||||
return {"tab": tab, "tab_navigated": True}
|
||||
# #endregion ScenarioExecution.BrowserProvider.ReadOnlyActions.NavigateTab
|
||||
|
||||
|
||||
# #region ScenarioExecution.BrowserProvider.ReadOnlyActions.InspectFilterState [C:3] [TYPE Function] [SEMANTICS provider,browser,inspect-filter]
|
||||
# @ingroup ScenarioExecution
|
||||
# @BRIEF Read the current filter-bar state without modifying it (observe-only).
|
||||
# @POST Returns typed details {filter_controls, filter_count}; no filter state is modified.
|
||||
async def inspect_filter_state_flow(
|
||||
service: Any, page: Any, *, timeout_seconds: float,
|
||||
) -> dict[str, Any]:
|
||||
from src.services.dashboard_testing.execution.providers.browser_native_filter import (
|
||||
_FILTER_BAR_SELECTORS,
|
||||
_FILTER_CONTROL_SELECTORS,
|
||||
_SELECTED_VALUE_SELECTOR,
|
||||
)
|
||||
bar = await service._find_first_visible_locator([page.locator(s) for s in _FILTER_BAR_SELECTORS])
|
||||
if bar is None:
|
||||
logger.explore("Filter bar not found for inspect_filter_state", src=_SRC, error_code="BROWSER_SELECTOR_NOT_FOUND")
|
||||
raise BrowserTransportSelectorNotFound("BROWSER_SELECTOR_NOT_FOUND")
|
||||
controls: list[dict[str, Any]] = []
|
||||
for selector in _FILTER_CONTROL_SELECTORS:
|
||||
loc = bar.locator(selector)
|
||||
count = min(await loc.count(), 50)
|
||||
for i in range(count):
|
||||
item = loc.nth(i)
|
||||
if not await item.is_visible():
|
||||
continue
|
||||
text = str(await item.text_content() or "").strip()
|
||||
selected = item.locator(_SELECTED_VALUE_SELECTOR)
|
||||
sel_count = min(await selected.count(), 20)
|
||||
values = []
|
||||
for j in range(sel_count):
|
||||
val = str(await selected.nth(j).text_content() or "").strip()
|
||||
if val:
|
||||
values.append(val)
|
||||
controls.append({"text": text, "selected_values": values})
|
||||
if controls:
|
||||
break
|
||||
logger.reflect("Filter state inspected", src=_SRC, payload={"filter_count": len(controls)})
|
||||
return {"filter_controls": controls, "filter_count": len(controls)}
|
||||
# #endregion ScenarioExecution.BrowserProvider.ReadOnlyActions.InspectFilterState
|
||||
|
||||
|
||||
# #region ScenarioExecution.BrowserProvider.ReadOnlyActions.ApplyTableFilter [C:3] [TYPE Function] [SEMANTICS provider,browser,table-filter]
|
||||
# @ingroup ScenarioExecution
|
||||
# @BRIEF Apply a typed column/value filter to the dashboard table UI.
|
||||
# @POST Returns typed details {column, value, table_filter_applied}; selector miss is typed.
|
||||
async def apply_table_filter_flow(
|
||||
service: Any, page: Any, action_input: dict[str, Any], *, timeout_seconds: float,
|
||||
) -> dict[str, Any]:
|
||||
column = str(action_input.get("column") or "").strip()
|
||||
value = action_input.get("value")
|
||||
timeout_ms = int(timeout_seconds * 1000)
|
||||
hint = action_input.get("selector_hint")
|
||||
locator = None
|
||||
if isinstance(hint, str) and hint.strip():
|
||||
locator = await service._find_first_visible_locator([page.locator(str(hint))])
|
||||
if locator is None:
|
||||
locator = await _resolve_first_visible(service, page, _TABLE_FILTER_SELECTORS)
|
||||
if locator is None:
|
||||
logger.explore("Table filter control not found", src=_SRC, payload={"column": column}, error_code="BROWSER_SELECTOR_NOT_FOUND")
|
||||
raise BrowserTransportSelectorNotFound("BROWSER_SELECTOR_NOT_FOUND")
|
||||
await locator.click(timeout=timeout_ms)
|
||||
if isinstance(value, str) and value.strip():
|
||||
input_loc = await service._find_first_visible_locator([
|
||||
page.locator(".ant-table-filter-dropdown input"),
|
||||
page.locator(".table-filter-input input"),
|
||||
page.locator('input[type="search"]'),
|
||||
])
|
||||
if input_loc is not None:
|
||||
await input_loc.fill(value, timeout=timeout_ms)
|
||||
confirm = await service._find_first_visible_locator([
|
||||
page.locator(".ant-table-filter-dropdown .ant-btn-primary"),
|
||||
page.locator('button:has-text("OK")'),
|
||||
page.locator('button:has-text("Apply")'),
|
||||
])
|
||||
if confirm is not None:
|
||||
await confirm.click(timeout=timeout_ms)
|
||||
await page.wait_for_load_state("domcontentloaded", timeout=timeout_ms)
|
||||
logger.reflect("Table filter applied", src=_SRC, payload={"column": column, "value": value})
|
||||
return {"column": column, "value": value, "table_filter_applied": True}
|
||||
# #endregion ScenarioExecution.BrowserProvider.ReadOnlyActions.ApplyTableFilter
|
||||
|
||||
|
||||
# #region ScenarioExecution.BrowserProvider.ReadOnlyActions.ExtractTable [C:4] [TYPE Function] [SEMANTICS provider,browser,extract,table,bounded]
|
||||
# @ingroup ScenarioExecution
|
||||
# @BRIEF Extract bounded table data from the dashboard DOM (10 000 rows, 100 columns, 10 MiB).
|
||||
# @POST Returns typed details {columns, rows, row_count, column_count}; oversized output raises
|
||||
# ValueError("BROWSER_EXTRACT_TOO_LARGE") before evidence is produced.
|
||||
async def extract_table_flow(
|
||||
service: Any, page: Any, action_input: dict[str, Any], *, timeout_seconds: float,
|
||||
) -> dict[str, Any]:
|
||||
max_rows = int(action_input.get("max_rows") or _MAX_EXTRACT_ROWS)
|
||||
max_cols = int(action_input.get("max_columns") or _MAX_EXTRACT_COLUMNS)
|
||||
hint = action_input.get("selector_hint")
|
||||
table_loc = None
|
||||
if isinstance(hint, str) and hint.strip():
|
||||
table_loc = await service._find_first_visible_locator([page.locator(str(hint))])
|
||||
if table_loc is None:
|
||||
table_loc = await _resolve_first_visible(service, page, _TABLE_CONTAINER_SELECTORS)
|
||||
if table_loc is None:
|
||||
logger.explore("Table container not found for extract_table", src=_SRC, error_code="BROWSER_SELECTOR_NOT_FOUND")
|
||||
raise BrowserTransportSelectorNotFound("BROWSER_SELECTOR_NOT_FOUND")
|
||||
raw = await table_loc.evaluate(
|
||||
"""(el, opts) => {
|
||||
const maxRows = opts.maxRows; const maxCols = opts.maxCols;
|
||||
const headers = Array.from(el.querySelectorAll('thead th, tr:first-child th, th'))
|
||||
.slice(0, maxCols)
|
||||
.map(th => (th.innerText || '').trim());
|
||||
const bodyRows = Array.from(el.querySelectorAll('tbody tr, tr'))
|
||||
.slice(0, maxRows);
|
||||
const rows = bodyRows.map(tr =>
|
||||
Array.from(tr.querySelectorAll('td')).slice(0, maxCols).map(td => (td.innerText || '').trim())
|
||||
);
|
||||
return { columns: headers, rows: rows };
|
||||
}""",
|
||||
{"maxRows": max_rows, "maxCols": max_cols},
|
||||
)
|
||||
columns = list(raw.get("columns") or [])[:max_cols]
|
||||
rows = list(raw.get("rows") or [])[:max_rows]
|
||||
payload = {"columns": columns, "rows": rows}
|
||||
serialized = json.dumps(payload, separators=(",", ":"), ensure_ascii=False)
|
||||
byte_size = len(serialized.encode("utf-8"))
|
||||
if byte_size > _MAX_EXTRACT_OUTPUT_BYTES:
|
||||
logger.explore(
|
||||
"Extract table output exceeds the 10 MiB limit", src=_SRC,
|
||||
payload={"bytes": byte_size, "rows": len(rows), "columns": len(columns)},
|
||||
error_code="BROWSER_EXTRACT_TOO_LARGE",
|
||||
)
|
||||
raise ValueError("BROWSER_EXTRACT_TOO_LARGE")
|
||||
logger.reflect(
|
||||
"Table extracted", src=_SRC,
|
||||
payload={"row_count": len(rows), "column_count": len(columns), "bytes": byte_size},
|
||||
)
|
||||
return {
|
||||
"columns": columns,
|
||||
"rows": rows,
|
||||
"row_count": len(rows),
|
||||
"column_count": len(columns),
|
||||
"byte_size": byte_size,
|
||||
}
|
||||
# #endregion ScenarioExecution.BrowserProvider.ReadOnlyActions.ExtractTable
|
||||
|
||||
|
||||
# #region ScenarioExecution.BrowserProvider.ReadOnlyActions.ScrollTo [C:2] [TYPE Function] [SEMANTICS provider,browser,scroll]
|
||||
# @ingroup ScenarioExecution
|
||||
# @BRIEF Scroll the page or a specific element into view.
|
||||
# @POST Returns typed details {selector, direction, scrolled}; selector miss is typed.
|
||||
async def scroll_to_flow(
|
||||
service: Any, page: Any, action_input: dict[str, Any], *, timeout_seconds: float,
|
||||
) -> dict[str, Any]:
|
||||
selector = str(action_input.get("selector") or "").strip()
|
||||
direction = str(action_input.get("direction") or "down")
|
||||
timeout_ms = int(timeout_seconds * 1000)
|
||||
target = await service._find_first_visible_locator([page.locator(selector)])
|
||||
if target is None:
|
||||
logger.explore("Scroll target not found", src=_SRC, payload={"selector": selector}, error_code="BROWSER_SELECTOR_NOT_FOUND")
|
||||
raise BrowserTransportSelectorNotFound("BROWSER_SELECTOR_NOT_FOUND")
|
||||
await target.scroll_into_view_if_needed(timeout=timeout_ms)
|
||||
logger.reflect("Page scrolled to target", src=_SRC, payload={"selector": selector, "direction": direction})
|
||||
return {"selector": selector, "direction": direction, "scrolled": True}
|
||||
# #endregion ScenarioExecution.BrowserProvider.ReadOnlyActions.ScrollTo
|
||||
|
||||
|
||||
# #region ScenarioExecution.BrowserProvider.ReadOnlyActions.InspectColumns [C:3] [TYPE Function] [SEMANTICS provider,browser,inspect-columns]
|
||||
# @ingroup ScenarioExecution
|
||||
# @BRIEF Read visible table column headers (observe-only, bounded).
|
||||
# @POST Returns typed details {columns, column_count}; selector miss is typed.
|
||||
async def inspect_columns_flow(
|
||||
service: Any, page: Any, action_input: dict[str, Any], *, timeout_seconds: float,
|
||||
) -> dict[str, Any]:
|
||||
hint = action_input.get("selector_hint")
|
||||
table_loc = None
|
||||
if isinstance(hint, str) and hint.strip():
|
||||
table_loc = await service._find_first_visible_locator([page.locator(str(hint))])
|
||||
if table_loc is None:
|
||||
table_loc = await _resolve_first_visible(service, page, _TABLE_CONTAINER_SELECTORS)
|
||||
if table_loc is None:
|
||||
logger.explore("Table container not found for inspect_columns", src=_SRC, error_code="BROWSER_SELECTOR_NOT_FOUND")
|
||||
raise BrowserTransportSelectorNotFound("BROWSER_SELECTOR_NOT_FOUND")
|
||||
columns: list[str] = []
|
||||
for selector in _COLUMN_HEADER_SELECTORS:
|
||||
loc = table_loc.locator(selector)
|
||||
count = min(await loc.count(), _MAX_EXTRACT_COLUMNS)
|
||||
for i in range(count):
|
||||
text = str(await loc.nth(i).text_content() or "").strip()
|
||||
if text:
|
||||
columns.append(text)
|
||||
if columns:
|
||||
break
|
||||
logger.reflect("Columns inspected", src=_SRC, payload={"column_count": len(columns)})
|
||||
return {"columns": columns, "column_count": len(columns)}
|
||||
# #endregion ScenarioExecution.BrowserProvider.ReadOnlyActions.InspectColumns
|
||||
|
||||
|
||||
# #region ScenarioExecution.BrowserProvider.ReadOnlyActions.Click [C:2] [TYPE Function] [SEMANTICS provider,browser,click]
|
||||
# @ingroup ScenarioExecution
|
||||
# @BRIEF Click a typed selector target (read-only interaction).
|
||||
# @POST Returns typed details {selector, clicked}; selector miss is typed.
|
||||
async def click_flow(
|
||||
service: Any, page: Any, action_input: dict[str, Any], *, timeout_seconds: float,
|
||||
) -> dict[str, Any]:
|
||||
selector = str(action_input.get("selector") or "").strip()
|
||||
hint = action_input.get("selector_hint")
|
||||
timeout_ms = int(timeout_seconds * 1000)
|
||||
locator = None
|
||||
if isinstance(hint, str) and hint.strip():
|
||||
locator = await service._find_first_visible_locator([page.locator(str(hint))])
|
||||
if locator is None:
|
||||
locator = await service._find_first_visible_locator([page.locator(selector)])
|
||||
if locator is None:
|
||||
logger.explore("Click target not found", src=_SRC, payload={"selector": selector}, error_code="BROWSER_SELECTOR_NOT_FOUND")
|
||||
raise BrowserTransportSelectorNotFound("BROWSER_SELECTOR_NOT_FOUND")
|
||||
await locator.click(timeout=timeout_ms)
|
||||
logger.reflect("Click target activated", src=_SRC, payload={"selector": selector})
|
||||
return {"selector": selector, "clicked": True}
|
||||
# #endregion ScenarioExecution.BrowserProvider.ReadOnlyActions.Click
|
||||
|
||||
|
||||
# #region ScenarioExecution.BrowserProvider.ReadOnlyActions.SelectRows [C:3] [TYPE Function] [SEMANTICS provider,browser,select-rows]
|
||||
# @ingroup ScenarioExecution
|
||||
# @BRIEF Select table rows by typed row_keys (selection-only, never proves mutation).
|
||||
# @POST Returns typed details {row_keys, selected_count, rows_selected}; selector miss is typed.
|
||||
async def select_rows_flow(
|
||||
service: Any, page: Any, action_input: dict[str, Any], *, timeout_seconds: float,
|
||||
) -> dict[str, Any]:
|
||||
row_keys = [str(k).strip() for k in action_input.get("row_keys") or []]
|
||||
timeout_ms = int(timeout_seconds * 1000)
|
||||
selected = 0
|
||||
for key in row_keys:
|
||||
candidates = [
|
||||
page.get_by_text(key, exact=True),
|
||||
page.locator(f'tr:has-text("{key}")'),
|
||||
]
|
||||
row_loc = await service._find_first_visible_locator(candidates)
|
||||
if row_loc is None:
|
||||
logger.explore("Row key not found for select_rows", src=_SRC, payload={"key": key}, error_code="BROWSER_SELECTOR_NOT_FOUND")
|
||||
raise BrowserTransportSelectorNotFound("BROWSER_SELECTOR_NOT_FOUND")
|
||||
checkbox = await service._find_first_visible_locator([
|
||||
row_loc.locator(s) for s in _ROW_CHECKBOX_SELECTORS
|
||||
])
|
||||
if checkbox is None:
|
||||
await row_loc.click(timeout=timeout_ms)
|
||||
else:
|
||||
await checkbox.click(timeout=timeout_ms)
|
||||
selected += 1
|
||||
logger.reflect("Rows selected", src=_SRC, payload={"selected_count": selected})
|
||||
return {"row_keys": row_keys, "selected_count": selected, "rows_selected": True}
|
||||
# #endregion ScenarioExecution.BrowserProvider.ReadOnlyActions.SelectRows
|
||||
|
||||
|
||||
# #region ScenarioExecution.BrowserProvider.ReadOnlyActions.Download [C:4] [TYPE Function] [SEMANTICS provider,browser,download,artifact,bounded]
|
||||
# @ingroup ScenarioExecution
|
||||
# @BRIEF Bounded download (25 MiB limit): trigger and return bytes for provider-side storage.
|
||||
# @POST Returns (details, download_bytes); oversized downloads raise
|
||||
# ValueError("BROWSER_DOWNLOAD_TOO_LARGE") before returning; the transport stays
|
||||
# storage-agnostic — the provider stores the artifact ref + sha256 separately.
|
||||
async def download_flow(
|
||||
service: Any,
|
||||
page: Any,
|
||||
action_input: dict[str, Any],
|
||||
*,
|
||||
timeout_seconds: float,
|
||||
) -> tuple[dict[str, Any], bytes]:
|
||||
artifact_type = str(action_input.get("artifact_type") or "xlsx")
|
||||
hint = action_input.get("selector_hint")
|
||||
timeout_ms = int(timeout_seconds * 1000)
|
||||
trigger = None
|
||||
if isinstance(hint, str) and hint.strip():
|
||||
trigger = await service._find_first_visible_locator([page.locator(str(hint))])
|
||||
if trigger is None:
|
||||
trigger = await _resolve_first_visible(service, page, _DOWNLOAD_TRIGGER_SELECTORS)
|
||||
if trigger is None:
|
||||
logger.explore("Download trigger not found", src=_SRC, error_code="BROWSER_SELECTOR_NOT_FOUND")
|
||||
raise BrowserTransportSelectorNotFound("BROWSER_SELECTOR_NOT_FOUND")
|
||||
async with page.expect_download(timeout=timeout_ms) as download_info:
|
||||
await trigger.click(timeout=timeout_ms)
|
||||
download = await download_info.value
|
||||
data = b""
|
||||
path = await download.path()
|
||||
if path is not None:
|
||||
with open(str(path), "rb") as fh:
|
||||
data = fh.read()
|
||||
if len(data) > _MAX_DOWNLOAD_BYTES:
|
||||
logger.explore(
|
||||
"Download exceeds the 25 MiB limit", src=_SRC,
|
||||
payload={"bytes": len(data), "artifact_type": artifact_type},
|
||||
error_code="BROWSER_DOWNLOAD_TOO_LARGE",
|
||||
)
|
||||
raise ValueError("BROWSER_DOWNLOAD_TOO_LARGE")
|
||||
logger.reflect(
|
||||
"Download captured (storage is the provider's)", src=_SRC,
|
||||
payload={"artifact_type": artifact_type, "bytes": len(data)},
|
||||
)
|
||||
details = {
|
||||
"artifact_type": artifact_type,
|
||||
"byte_size": len(data),
|
||||
"downloaded": True,
|
||||
}
|
||||
return details, data
|
||||
# #endregion ScenarioExecution.BrowserProvider.ReadOnlyActions.Download
|
||||
|
||||
|
||||
# #region ScenarioExecution.BrowserProvider.ReadOnlyActions.Dispatch [C:3] [TYPE Function] [SEMANTICS provider,browser,readonly,dispatch]
|
||||
# @ingroup ScenarioExecution
|
||||
# @BRIEF Dispatch a round-2 read-only action to its typed flow; returns (details, download_bytes).
|
||||
@@ -530,5 +157,4 @@ async def run_readonly_action_flow(
|
||||
return await download_flow(service, page, action_input, timeout_seconds=timeout_seconds)
|
||||
raise ValueError("BROWSER_ACTION_NOT_SUPPORTED")
|
||||
# #endregion ScenarioExecution.BrowserProvider.ReadOnlyActions.Dispatch
|
||||
|
||||
# #endregion ScenarioExecution.BrowserProvider.ReadOnlyActions
|
||||
|
||||
@@ -0,0 +1,192 @@
|
||||
# #region ScenarioExecution.BrowserProvider.ReadOnlyActions.FlowsInteract [C:4] [TYPE Module] [SEMANTICS provider,browser,readonly,scroll,select,click,download]
|
||||
# @ingroup ScenarioExecution
|
||||
# @BRIEF Interaction and bounded-download flows of the T034 round-2 read-only catalog:
|
||||
# scroll_to, inspect_columns, click, select_rows, download.
|
||||
# @POST Every flow returns a typed details dict (download returns (details, bytes)); invalid input
|
||||
# raises ValueError("BROWSER_*_INVALID") before I/O and a locator miss raises
|
||||
# BrowserTransportSelectorNotFound (fail-closed, no retry).
|
||||
# @INVARIANT download is bounded to 25 MiB and stays storage-agnostic (the provider stores the
|
||||
# artifact ref + sha256); select_rows proves selection only, never a mutation.
|
||||
# @RELATION DEPENDS_ON -> [ScenarioExecution.BrowserProvider.ReadOnlyActions.Limits]
|
||||
# @RATIONALE Split from browser_readonly_actions.py (INV_7) keeping each region verbatim; the facade
|
||||
# re-exports these names so the transport import surface stays frozen.
|
||||
# @REJECTED Leaving all nine flows in one 534-line module was rejected — see
|
||||
# specs/044-dashboard-scenario-execution/plans/provider-decomposition-gate.md.
|
||||
from __future__ import annotations
|
||||
|
||||
from typing import Any
|
||||
|
||||
from src.core.logger import logger
|
||||
from src.services.dashboard_testing.execution.providers.browser_native_filter import (
|
||||
BrowserTransportSelectorNotFound,
|
||||
)
|
||||
from src.services.dashboard_testing.execution.providers.browser_readonly_limits import (
|
||||
_COLUMN_HEADER_SELECTORS,
|
||||
_DOWNLOAD_TRIGGER_SELECTORS,
|
||||
_MAX_DOWNLOAD_BYTES,
|
||||
_MAX_EXTRACT_COLUMNS,
|
||||
_ROW_CHECKBOX_SELECTORS,
|
||||
_TABLE_CONTAINER_SELECTORS,
|
||||
_resolve_first_visible,
|
||||
)
|
||||
|
||||
_SRC = "ScenarioExecution.BrowserProvider.ReadOnlyActions"
|
||||
|
||||
# #region ScenarioExecution.BrowserProvider.ReadOnlyActions.ScrollTo [C:2] [TYPE Function] [SEMANTICS provider,browser,scroll]
|
||||
# @ingroup ScenarioExecution
|
||||
# @BRIEF Scroll the page or a specific element into view.
|
||||
# @POST Returns typed details {selector, direction, scrolled}; selector miss is typed.
|
||||
async def scroll_to_flow(
|
||||
service: Any, page: Any, action_input: dict[str, Any], *, timeout_seconds: float,
|
||||
) -> dict[str, Any]:
|
||||
selector = str(action_input.get("selector") or "").strip()
|
||||
direction = str(action_input.get("direction") or "down")
|
||||
timeout_ms = int(timeout_seconds * 1000)
|
||||
target = await service._find_first_visible_locator([page.locator(selector)])
|
||||
if target is None:
|
||||
logger.explore("Scroll target not found", src=_SRC, payload={"selector": selector}, error_code="BROWSER_SELECTOR_NOT_FOUND")
|
||||
raise BrowserTransportSelectorNotFound("BROWSER_SELECTOR_NOT_FOUND")
|
||||
await target.scroll_into_view_if_needed(timeout=timeout_ms)
|
||||
logger.reflect("Page scrolled to target", src=_SRC, payload={"selector": selector, "direction": direction})
|
||||
return {"selector": selector, "direction": direction, "scrolled": True}
|
||||
# #endregion ScenarioExecution.BrowserProvider.ReadOnlyActions.ScrollTo
|
||||
|
||||
|
||||
# #region ScenarioExecution.BrowserProvider.ReadOnlyActions.InspectColumns [C:3] [TYPE Function] [SEMANTICS provider,browser,inspect-columns]
|
||||
# @ingroup ScenarioExecution
|
||||
# @BRIEF Read visible table column headers (observe-only, bounded).
|
||||
# @POST Returns typed details {columns, column_count}; selector miss is typed.
|
||||
async def inspect_columns_flow(
|
||||
service: Any, page: Any, action_input: dict[str, Any], *, timeout_seconds: float,
|
||||
) -> dict[str, Any]:
|
||||
hint = action_input.get("selector_hint")
|
||||
table_loc = None
|
||||
if isinstance(hint, str) and hint.strip():
|
||||
table_loc = await service._find_first_visible_locator([page.locator(str(hint))])
|
||||
if table_loc is None:
|
||||
table_loc = await _resolve_first_visible(service, page, _TABLE_CONTAINER_SELECTORS)
|
||||
if table_loc is None:
|
||||
logger.explore("Table container not found for inspect_columns", src=_SRC, error_code="BROWSER_SELECTOR_NOT_FOUND")
|
||||
raise BrowserTransportSelectorNotFound("BROWSER_SELECTOR_NOT_FOUND")
|
||||
columns: list[str] = []
|
||||
for selector in _COLUMN_HEADER_SELECTORS:
|
||||
loc = table_loc.locator(selector)
|
||||
count = min(await loc.count(), _MAX_EXTRACT_COLUMNS)
|
||||
for i in range(count):
|
||||
text = str(await loc.nth(i).text_content() or "").strip()
|
||||
if text:
|
||||
columns.append(text)
|
||||
if columns:
|
||||
break
|
||||
logger.reflect("Columns inspected", src=_SRC, payload={"column_count": len(columns)})
|
||||
return {"columns": columns, "column_count": len(columns)}
|
||||
# #endregion ScenarioExecution.BrowserProvider.ReadOnlyActions.InspectColumns
|
||||
|
||||
|
||||
# #region ScenarioExecution.BrowserProvider.ReadOnlyActions.Click [C:2] [TYPE Function] [SEMANTICS provider,browser,click]
|
||||
# @ingroup ScenarioExecution
|
||||
# @BRIEF Click a typed selector target (read-only interaction).
|
||||
# @POST Returns typed details {selector, clicked}; selector miss is typed.
|
||||
async def click_flow(
|
||||
service: Any, page: Any, action_input: dict[str, Any], *, timeout_seconds: float,
|
||||
) -> dict[str, Any]:
|
||||
selector = str(action_input.get("selector") or "").strip()
|
||||
hint = action_input.get("selector_hint")
|
||||
timeout_ms = int(timeout_seconds * 1000)
|
||||
locator = None
|
||||
if isinstance(hint, str) and hint.strip():
|
||||
locator = await service._find_first_visible_locator([page.locator(str(hint))])
|
||||
if locator is None:
|
||||
locator = await service._find_first_visible_locator([page.locator(selector)])
|
||||
if locator is None:
|
||||
logger.explore("Click target not found", src=_SRC, payload={"selector": selector}, error_code="BROWSER_SELECTOR_NOT_FOUND")
|
||||
raise BrowserTransportSelectorNotFound("BROWSER_SELECTOR_NOT_FOUND")
|
||||
await locator.click(timeout=timeout_ms)
|
||||
logger.reflect("Click target activated", src=_SRC, payload={"selector": selector})
|
||||
return {"selector": selector, "clicked": True}
|
||||
# #endregion ScenarioExecution.BrowserProvider.ReadOnlyActions.Click
|
||||
|
||||
|
||||
# #region ScenarioExecution.BrowserProvider.ReadOnlyActions.SelectRows [C:3] [TYPE Function] [SEMANTICS provider,browser,select-rows]
|
||||
# @ingroup ScenarioExecution
|
||||
# @BRIEF Select table rows by typed row_keys (selection-only, never proves mutation).
|
||||
# @POST Returns typed details {row_keys, selected_count, rows_selected}; selector miss is typed.
|
||||
async def select_rows_flow(
|
||||
service: Any, page: Any, action_input: dict[str, Any], *, timeout_seconds: float,
|
||||
) -> dict[str, Any]:
|
||||
row_keys = [str(k).strip() for k in action_input.get("row_keys") or []]
|
||||
timeout_ms = int(timeout_seconds * 1000)
|
||||
selected = 0
|
||||
for key in row_keys:
|
||||
candidates = [
|
||||
page.get_by_text(key, exact=True),
|
||||
page.locator(f'tr:has-text("{key}")'),
|
||||
]
|
||||
row_loc = await service._find_first_visible_locator(candidates)
|
||||
if row_loc is None:
|
||||
logger.explore("Row key not found for select_rows", src=_SRC, payload={"key": key}, error_code="BROWSER_SELECTOR_NOT_FOUND")
|
||||
raise BrowserTransportSelectorNotFound("BROWSER_SELECTOR_NOT_FOUND")
|
||||
checkbox = await service._find_first_visible_locator([
|
||||
row_loc.locator(s) for s in _ROW_CHECKBOX_SELECTORS
|
||||
])
|
||||
if checkbox is None:
|
||||
await row_loc.click(timeout=timeout_ms)
|
||||
else:
|
||||
await checkbox.click(timeout=timeout_ms)
|
||||
selected += 1
|
||||
logger.reflect("Rows selected", src=_SRC, payload={"selected_count": selected})
|
||||
return {"row_keys": row_keys, "selected_count": selected, "rows_selected": True}
|
||||
# #endregion ScenarioExecution.BrowserProvider.ReadOnlyActions.SelectRows
|
||||
|
||||
|
||||
# #region ScenarioExecution.BrowserProvider.ReadOnlyActions.Download [C:4] [TYPE Function] [SEMANTICS provider,browser,download,artifact,bounded]
|
||||
# @ingroup ScenarioExecution
|
||||
# @BRIEF Bounded download (25 MiB limit): trigger and return bytes for provider-side storage.
|
||||
# @POST Returns (details, download_bytes); oversized downloads raise
|
||||
# ValueError("BROWSER_DOWNLOAD_TOO_LARGE") before returning; the transport stays
|
||||
# storage-agnostic — the provider stores the artifact ref + sha256 separately.
|
||||
async def download_flow(
|
||||
service: Any,
|
||||
page: Any,
|
||||
action_input: dict[str, Any],
|
||||
*,
|
||||
timeout_seconds: float,
|
||||
) -> tuple[dict[str, Any], bytes]:
|
||||
artifact_type = str(action_input.get("artifact_type") or "xlsx")
|
||||
hint = action_input.get("selector_hint")
|
||||
timeout_ms = int(timeout_seconds * 1000)
|
||||
trigger = None
|
||||
if isinstance(hint, str) and hint.strip():
|
||||
trigger = await service._find_first_visible_locator([page.locator(str(hint))])
|
||||
if trigger is None:
|
||||
trigger = await _resolve_first_visible(service, page, _DOWNLOAD_TRIGGER_SELECTORS)
|
||||
if trigger is None:
|
||||
logger.explore("Download trigger not found", src=_SRC, error_code="BROWSER_SELECTOR_NOT_FOUND")
|
||||
raise BrowserTransportSelectorNotFound("BROWSER_SELECTOR_NOT_FOUND")
|
||||
async with page.expect_download(timeout=timeout_ms) as download_info:
|
||||
await trigger.click(timeout=timeout_ms)
|
||||
download = await download_info.value
|
||||
data = b""
|
||||
path = await download.path()
|
||||
if path is not None:
|
||||
with open(str(path), "rb") as fh:
|
||||
data = fh.read()
|
||||
if len(data) > _MAX_DOWNLOAD_BYTES:
|
||||
logger.explore(
|
||||
"Download exceeds the 25 MiB limit", src=_SRC,
|
||||
payload={"bytes": len(data), "artifact_type": artifact_type},
|
||||
error_code="BROWSER_DOWNLOAD_TOO_LARGE",
|
||||
)
|
||||
raise ValueError("BROWSER_DOWNLOAD_TOO_LARGE")
|
||||
logger.reflect(
|
||||
"Download captured (storage is the provider's)", src=_SRC,
|
||||
payload={"artifact_type": artifact_type, "bytes": len(data)},
|
||||
)
|
||||
details = {
|
||||
"artifact_type": artifact_type,
|
||||
"byte_size": len(data),
|
||||
"downloaded": True,
|
||||
}
|
||||
return details, data
|
||||
# #endregion ScenarioExecution.BrowserProvider.ReadOnlyActions.Download
|
||||
# #endregion ScenarioExecution.BrowserProvider.ReadOnlyActions.FlowsInteract
|
||||
@@ -0,0 +1,199 @@
|
||||
# #region ScenarioExecution.BrowserProvider.ReadOnlyActions.FlowsNav [C:4] [TYPE Module] [SEMANTICS provider,browser,readonly,navigate,table-filter,extract]
|
||||
# @ingroup ScenarioExecution
|
||||
# @BRIEF Navigation and table-inspection flows of the T034 round-2 read-only catalog:
|
||||
# navigate_tab, inspect_filter_state, apply_table_filter, extract_table.
|
||||
# @POST Every flow returns a typed details dict; invalid input raises ValueError("BROWSER_*_INVALID")
|
||||
# before any locator work and a locator miss raises BrowserTransportSelectorNotFound.
|
||||
# @INVARIANT extract_table is bounded to 10 000 rows, 100 columns and 10 MiB output; inspect_filter_state
|
||||
# is observe-only (no filter mutation).
|
||||
# @RELATION DEPENDS_ON -> [ScenarioExecution.BrowserProvider.ReadOnlyActions.Limits]
|
||||
# @RATIONALE Split from browser_readonly_actions.py (INV_7) keeping each region verbatim; the facade
|
||||
# re-exports these names so the transport import surface stays frozen.
|
||||
# @REJECTED Leaving all nine flows in one 534-line module was rejected — see
|
||||
# specs/044-dashboard-scenario-execution/plans/provider-decomposition-gate.md.
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
from typing import Any
|
||||
|
||||
from src.core.logger import logger
|
||||
from src.services.dashboard_testing.execution.providers.browser_native_filter import (
|
||||
BrowserTransportSelectorNotFound,
|
||||
)
|
||||
from src.services.dashboard_testing.execution.providers.browser_readonly_limits import (
|
||||
_MAX_EXTRACT_COLUMNS,
|
||||
_MAX_EXTRACT_OUTPUT_BYTES,
|
||||
_MAX_EXTRACT_ROWS,
|
||||
_TABLE_CONTAINER_SELECTORS,
|
||||
_TABLE_FILTER_SELECTORS,
|
||||
_TAB_SELECTORS,
|
||||
_resolve_first_visible,
|
||||
)
|
||||
|
||||
_SRC = "ScenarioExecution.BrowserProvider.ReadOnlyActions"
|
||||
|
||||
# #region ScenarioExecution.BrowserProvider.ReadOnlyActions.NavigateTab [C:3] [TYPE Function] [SEMANTICS provider,browser,navigate-tab]
|
||||
# @ingroup ScenarioExecution
|
||||
# @BRIEF Click the dashboard tab identified by typed input or selector_hint.
|
||||
# @POST Returns typed details {tab, tab_navigated}; raises BrowserTransportSelectorNotFound on miss.
|
||||
async def navigate_tab_flow(
|
||||
service: Any, page: Any, action_input: dict[str, Any], *, timeout_seconds: float,
|
||||
) -> dict[str, Any]:
|
||||
tab = str(action_input.get("tab") or "").strip()
|
||||
hint = action_input.get("selector_hint")
|
||||
timeout_ms = int(timeout_seconds * 1000)
|
||||
locator = None
|
||||
if isinstance(hint, str) and hint.strip():
|
||||
locator = await service._find_first_visible_locator([page.locator(str(hint))])
|
||||
if locator is None:
|
||||
locator = await _resolve_first_visible(service, page, _TAB_SELECTORS, tab=tab)
|
||||
if locator is None:
|
||||
logger.explore("Tab locator not found", src=_SRC, payload={"tab": tab}, error_code="BROWSER_SELECTOR_NOT_FOUND")
|
||||
raise BrowserTransportSelectorNotFound("BROWSER_SELECTOR_NOT_FOUND")
|
||||
await locator.click(timeout=timeout_ms)
|
||||
await page.wait_for_load_state("domcontentloaded", timeout=timeout_ms)
|
||||
logger.reflect("Dashboard tab navigated", src=_SRC, payload={"tab": tab})
|
||||
return {"tab": tab, "tab_navigated": True}
|
||||
# #endregion ScenarioExecution.BrowserProvider.ReadOnlyActions.NavigateTab
|
||||
|
||||
|
||||
# #region ScenarioExecution.BrowserProvider.ReadOnlyActions.InspectFilterState [C:3] [TYPE Function] [SEMANTICS provider,browser,inspect-filter]
|
||||
# @ingroup ScenarioExecution
|
||||
# @BRIEF Read the current filter-bar state without modifying it (observe-only).
|
||||
# @POST Returns typed details {filter_controls, filter_count}; no filter state is modified.
|
||||
async def inspect_filter_state_flow(
|
||||
service: Any, page: Any, *, timeout_seconds: float,
|
||||
) -> dict[str, Any]:
|
||||
from src.services.dashboard_testing.execution.providers.browser_native_filter import (
|
||||
_FILTER_BAR_SELECTORS,
|
||||
_FILTER_CONTROL_SELECTORS,
|
||||
_SELECTED_VALUE_SELECTOR,
|
||||
)
|
||||
bar = await service._find_first_visible_locator([page.locator(s) for s in _FILTER_BAR_SELECTORS])
|
||||
if bar is None:
|
||||
logger.explore("Filter bar not found for inspect_filter_state", src=_SRC, error_code="BROWSER_SELECTOR_NOT_FOUND")
|
||||
raise BrowserTransportSelectorNotFound("BROWSER_SELECTOR_NOT_FOUND")
|
||||
controls: list[dict[str, Any]] = []
|
||||
for selector in _FILTER_CONTROL_SELECTORS:
|
||||
loc = bar.locator(selector)
|
||||
count = min(await loc.count(), 50)
|
||||
for i in range(count):
|
||||
item = loc.nth(i)
|
||||
if not await item.is_visible():
|
||||
continue
|
||||
text = str(await item.text_content() or "").strip()
|
||||
selected = item.locator(_SELECTED_VALUE_SELECTOR)
|
||||
sel_count = min(await selected.count(), 20)
|
||||
values = []
|
||||
for j in range(sel_count):
|
||||
val = str(await selected.nth(j).text_content() or "").strip()
|
||||
if val:
|
||||
values.append(val)
|
||||
controls.append({"text": text, "selected_values": values})
|
||||
if controls:
|
||||
break
|
||||
logger.reflect("Filter state inspected", src=_SRC, payload={"filter_count": len(controls)})
|
||||
return {"filter_controls": controls, "filter_count": len(controls)}
|
||||
# #endregion ScenarioExecution.BrowserProvider.ReadOnlyActions.InspectFilterState
|
||||
|
||||
|
||||
# #region ScenarioExecution.BrowserProvider.ReadOnlyActions.ApplyTableFilter [C:3] [TYPE Function] [SEMANTICS provider,browser,table-filter]
|
||||
# @ingroup ScenarioExecution
|
||||
# @BRIEF Apply a typed column/value filter to the dashboard table UI.
|
||||
# @POST Returns typed details {column, value, table_filter_applied}; selector miss is typed.
|
||||
async def apply_table_filter_flow(
|
||||
service: Any, page: Any, action_input: dict[str, Any], *, timeout_seconds: float,
|
||||
) -> dict[str, Any]:
|
||||
column = str(action_input.get("column") or "").strip()
|
||||
value = action_input.get("value")
|
||||
timeout_ms = int(timeout_seconds * 1000)
|
||||
hint = action_input.get("selector_hint")
|
||||
locator = None
|
||||
if isinstance(hint, str) and hint.strip():
|
||||
locator = await service._find_first_visible_locator([page.locator(str(hint))])
|
||||
if locator is None:
|
||||
locator = await _resolve_first_visible(service, page, _TABLE_FILTER_SELECTORS)
|
||||
if locator is None:
|
||||
logger.explore("Table filter control not found", src=_SRC, payload={"column": column}, error_code="BROWSER_SELECTOR_NOT_FOUND")
|
||||
raise BrowserTransportSelectorNotFound("BROWSER_SELECTOR_NOT_FOUND")
|
||||
await locator.click(timeout=timeout_ms)
|
||||
if isinstance(value, str) and value.strip():
|
||||
input_loc = await service._find_first_visible_locator([
|
||||
page.locator(".ant-table-filter-dropdown input"),
|
||||
page.locator(".table-filter-input input"),
|
||||
page.locator('input[type="search"]'),
|
||||
])
|
||||
if input_loc is not None:
|
||||
await input_loc.fill(value, timeout=timeout_ms)
|
||||
confirm = await service._find_first_visible_locator([
|
||||
page.locator(".ant-table-filter-dropdown .ant-btn-primary"),
|
||||
page.locator('button:has-text("OK")'),
|
||||
page.locator('button:has-text("Apply")'),
|
||||
])
|
||||
if confirm is not None:
|
||||
await confirm.click(timeout=timeout_ms)
|
||||
await page.wait_for_load_state("domcontentloaded", timeout=timeout_ms)
|
||||
logger.reflect("Table filter applied", src=_SRC, payload={"column": column, "value": value})
|
||||
return {"column": column, "value": value, "table_filter_applied": True}
|
||||
# #endregion ScenarioExecution.BrowserProvider.ReadOnlyActions.ApplyTableFilter
|
||||
|
||||
|
||||
# #region ScenarioExecution.BrowserProvider.ReadOnlyActions.ExtractTable [C:4] [TYPE Function] [SEMANTICS provider,browser,extract,table,bounded]
|
||||
# @ingroup ScenarioExecution
|
||||
# @BRIEF Extract bounded table data from the dashboard DOM (10 000 rows, 100 columns, 10 MiB).
|
||||
# @POST Returns typed details {columns, rows, row_count, column_count}; oversized output raises
|
||||
# ValueError("BROWSER_EXTRACT_TOO_LARGE") before evidence is produced.
|
||||
async def extract_table_flow(
|
||||
service: Any, page: Any, action_input: dict[str, Any], *, timeout_seconds: float,
|
||||
) -> dict[str, Any]:
|
||||
max_rows = int(action_input.get("max_rows") or _MAX_EXTRACT_ROWS)
|
||||
max_cols = int(action_input.get("max_columns") or _MAX_EXTRACT_COLUMNS)
|
||||
hint = action_input.get("selector_hint")
|
||||
table_loc = None
|
||||
if isinstance(hint, str) and hint.strip():
|
||||
table_loc = await service._find_first_visible_locator([page.locator(str(hint))])
|
||||
if table_loc is None:
|
||||
table_loc = await _resolve_first_visible(service, page, _TABLE_CONTAINER_SELECTORS)
|
||||
if table_loc is None:
|
||||
logger.explore("Table container not found for extract_table", src=_SRC, error_code="BROWSER_SELECTOR_NOT_FOUND")
|
||||
raise BrowserTransportSelectorNotFound("BROWSER_SELECTOR_NOT_FOUND")
|
||||
raw = await table_loc.evaluate(
|
||||
"""(el, opts) => {
|
||||
const maxRows = opts.maxRows; const maxCols = opts.maxCols;
|
||||
const headers = Array.from(el.querySelectorAll('thead th, tr:first-child th, th'))
|
||||
.slice(0, maxCols)
|
||||
.map(th => (th.innerText || '').trim());
|
||||
const bodyRows = Array.from(el.querySelectorAll('tbody tr, tr'))
|
||||
.slice(0, maxRows);
|
||||
const rows = bodyRows.map(tr =>
|
||||
Array.from(tr.querySelectorAll('td')).slice(0, maxCols).map(td => (td.innerText || '').trim())
|
||||
);
|
||||
return { columns: headers, rows: rows };
|
||||
}""",
|
||||
{"maxRows": max_rows, "maxCols": max_cols},
|
||||
)
|
||||
columns = list(raw.get("columns") or [])[:max_cols]
|
||||
rows = list(raw.get("rows") or [])[:max_rows]
|
||||
payload = {"columns": columns, "rows": rows}
|
||||
serialized = json.dumps(payload, separators=(",", ":"), ensure_ascii=False)
|
||||
byte_size = len(serialized.encode("utf-8"))
|
||||
if byte_size > _MAX_EXTRACT_OUTPUT_BYTES:
|
||||
logger.explore(
|
||||
"Extract table output exceeds the 10 MiB limit", src=_SRC,
|
||||
payload={"bytes": byte_size, "rows": len(rows), "columns": len(columns)},
|
||||
error_code="BROWSER_EXTRACT_TOO_LARGE",
|
||||
)
|
||||
raise ValueError("BROWSER_EXTRACT_TOO_LARGE")
|
||||
logger.reflect(
|
||||
"Table extracted", src=_SRC,
|
||||
payload={"row_count": len(rows), "column_count": len(columns), "bytes": byte_size},
|
||||
)
|
||||
return {
|
||||
"columns": columns,
|
||||
"rows": rows,
|
||||
"row_count": len(rows),
|
||||
"column_count": len(columns),
|
||||
"byte_size": byte_size,
|
||||
}
|
||||
# #endregion ScenarioExecution.BrowserProvider.ReadOnlyActions.ExtractTable
|
||||
# #endregion ScenarioExecution.BrowserProvider.ReadOnlyActions.FlowsNav
|
||||
@@ -0,0 +1,82 @@
|
||||
# #region ScenarioExecution.BrowserProvider.ReadOnlyActions.Limits [C:2] [TYPE Module] [SEMANTICS provider,browser,readonly,limits,selectors,constants]
|
||||
# @ingroup ScenarioExecution
|
||||
# @BRIEF Single-sited limits, allowlists and multi-strategy locator candidates for the nine
|
||||
# T034 round-2 read-only browser flows, plus the shared first-visible selector resolver.
|
||||
# @POST Constants are module-level immutable literals; _resolve_first_visible returns a visible
|
||||
# locator or None and performs no I/O of its own.
|
||||
# @INVARIANT Every flow module imports these names as module globals, so a monkeypatch of a limit
|
||||
# must target the OWNING FLOW module (or this module for cross-module effects) — the
|
||||
# facade re-export does not re-bind patched values.
|
||||
# @RATIONALE Extracted from browser_readonly_actions.py (INV_7: 534 LOC) so the two flow modules
|
||||
# and the facade share one authority for bounds, allowlists and selector candidates
|
||||
# instead of duplicating them per module.
|
||||
# @REJECTED Duplicating the constant block into each flow module was rejected — a drifted bound
|
||||
# would silently weaken a typed BROWSER_*_INVALID/BROWSER_*_TOO_LARGE guard.
|
||||
from __future__ import annotations
|
||||
|
||||
from typing import Any
|
||||
|
||||
# -- Limits (T034 spec) ---------------------------------------------------
|
||||
_MAX_EXTRACT_ROWS = 10_000
|
||||
_MAX_EXTRACT_COLUMNS = 100
|
||||
_MAX_EXTRACT_OUTPUT_BYTES = 10 * 1024 * 1024 # 10 MiB
|
||||
_MAX_DOWNLOAD_BYTES = 25 * 1024 * 1024 # 25 MiB
|
||||
_MAX_CLICK_SELECTOR_LENGTH = 512
|
||||
_MAX_SCROLL_SELECTOR_LENGTH = 512
|
||||
_MAX_SELECT_ROWS = 100
|
||||
_MAX_TAB_LENGTH = 128
|
||||
_MAX_COLUMN_LENGTH = 256
|
||||
_MAX_TABLE_FILTER_VALUE_LENGTH = 256
|
||||
|
||||
# -- Allowlists -----------------------------------------------------------
|
||||
_ALLOWED_TABS = frozenset({
|
||||
"overview", "charts", "data", "filters", "settings", "annotations",
|
||||
"css", "properties", "json", "code", "sql", "table", "dashboard",
|
||||
})
|
||||
_ALLOWED_SCROLL_DIRECTIONS = frozenset({"up", "down", "left", "right", "top", "bottom"})
|
||||
_ALLOWED_ARTIFACT_TYPES = frozenset({"xlsx", "csv", "pdf", "png", "json"})
|
||||
|
||||
# -- Multi-strategy locator candidates (Superset version drift expected) ---
|
||||
_TAB_SELECTORS = (
|
||||
'[data-test="tab-{tab}"]',
|
||||
'.nav-item a[data-tab="{tab}"]',
|
||||
'button[role="tab"][aria-controls*="{tab}"]',
|
||||
'a[href*="{tab}"]',
|
||||
)
|
||||
_TABLE_FILTER_SELECTORS = (
|
||||
'[data-test="table-filter"]',
|
||||
".table-filter-container",
|
||||
"th .ant-table-filter-column",
|
||||
)
|
||||
_TABLE_CONTAINER_SELECTORS = (
|
||||
'[data-test="table-container"]',
|
||||
".ant-table-wrapper",
|
||||
".table-container",
|
||||
".grid-content table",
|
||||
)
|
||||
_COLUMN_HEADER_SELECTORS = (
|
||||
"th .ant-table-column-title",
|
||||
"th span",
|
||||
"th",
|
||||
)
|
||||
_ROW_CHECKBOX_SELECTORS = (
|
||||
'td .ant-checkbox-input',
|
||||
'td input[type="checkbox"]',
|
||||
"td .row-select",
|
||||
)
|
||||
_DOWNLOAD_TRIGGER_SELECTORS = (
|
||||
'[data-test="download-button"]',
|
||||
'button:has-text("Download")',
|
||||
'button:has-text("Export")',
|
||||
'a[download]',
|
||||
)
|
||||
|
||||
# #region ScenarioExecution.BrowserProvider.ReadOnlyActions.ResolveSelector [C:2] [TYPE Function] [SEMANTICS provider,browser,readonly,selector]
|
||||
# @ingroup ScenarioExecution
|
||||
# @BRIEF Resolve the first visible locator from a list of selector templates.
|
||||
# @POST Returns a visible locator or None (caller maps to BrowserTransportSelectorNotFound).
|
||||
async def _resolve_first_visible(service: Any, page: Any, selectors: tuple[str, ...], **fmt: str) -> Any:
|
||||
candidates = [page.locator(selector.format(**fmt)) for selector in selectors]
|
||||
return await service._find_first_visible_locator(candidates)
|
||||
# #endregion ScenarioExecution.BrowserProvider.ReadOnlyActions.ResolveSelector
|
||||
# #endregion ScenarioExecution.BrowserProvider.ReadOnlyActions.Limits
|
||||
@@ -1,30 +1,27 @@
|
||||
# #region ScenarioExecution.BrowserProvider.Session [C:5] [TYPE Module] [SEMANTICS scenario,execution,provider,browser,session,checkpoint,replay,leak-guard]
|
||||
# @ingroup ScenarioExecution
|
||||
# @BRIEF Run-scoped browser session registry on the shared provider loop: one isolated context per
|
||||
# run lives across all browser steps, a browser-safe checkpoint is folded after every step,
|
||||
# cross-process recovery replays the persisted checkpoint, and close paths are guaranteed
|
||||
# (DG-1 amendment 2026-09-12, 044 T034 round 1).
|
||||
# @BRIEF Facade for the run-scoped browser session stack: the weak manager registry / lifecycle
|
||||
# finalizer lives here, while the handle types (browser_session_handle.py), checkpoint IO
|
||||
# (browser_session_checkpoint.py) and the C:5 session registry (browser_session_registry.py)
|
||||
# are sibling modules extracted verbatim per INV_7.
|
||||
# @RELATION IMPLEMENTS -> [ScenarioExecution.BrowserProvider]
|
||||
# @RELATION DEPENDS_ON -> [ScenarioExecution.ProviderRuntime.Engine]
|
||||
# @RELATION DEPENDS_ON -> [ScenarioExecution.BrowserProvider.Transport]
|
||||
# @RELATION DEPENDS_ON -> [Models.ScenarioExecution.StepRun]
|
||||
# @PRE The manager is bound to a session-capable transport (open_session / execute_in_session /
|
||||
# close_session) and the shared provider event loop; the provider holds the per-run guard
|
||||
# across prepare + submit so concurrent steps of one run serialize on the run context (DG-1 §4).
|
||||
# @POST prepare_step returns exactly one live session per run_id: an in-memory hit is reused, a
|
||||
# persisted checkpoint replays open + re-apply before the step, and browser history without
|
||||
# any declared checkpoint raises BrowserCheckpointMissing (fail-closed, never a silent
|
||||
# stateless session). close() is idempotent and best-effort on every path.
|
||||
# @INVARIANT A context is never shared across runs; a dead context is never revived — recovery is a
|
||||
# new context plus checkpoint replay from scenario_step_runs.step_outcome; the registry
|
||||
# lives in process memory, so process death kills every context and only replay recovers.
|
||||
# @SIDE_EFFECT Live browser/context handles bound to the provider loop, an in-process registry, and
|
||||
# read-only checkpoint queries against scenario_step_runs.
|
||||
# @RATIONALE DG-1 (044 spec 2026-09-12): native filter state applied by any run step must persist
|
||||
# for all later steps — per-step isolation broke filter continuity and risked false-PASS.
|
||||
# Round 1 keeps the capacity lease per step, so the provider has no terminal-run channel:
|
||||
# the session survives per-step lease release and closes via explicit close(run_id)
|
||||
# (B-wave dispatch finalizer), the dead-context guard, the idle TTL reaper and close_all.
|
||||
# @RELATION DEPENDS_ON -> [ScenarioExecution.BrowserProvider.Session.CheckpointIO]
|
||||
# @RELATION DEPENDS_ON -> [ScenarioExecution.BrowserProvider.Session.Handle]
|
||||
# @RELATION DEPENDS_ON -> [ScenarioExecution.BrowserProvider.Session.Registry]
|
||||
# @POST close_run_sessions returns the number of managers that dropped a session for the run and
|
||||
# never raises; all pre-split public names stay importable from this module.
|
||||
# @INVARIANT Frozen import surface: browser.py imports BrowserCheckpointMissing/BrowserSessionManager
|
||||
# from HERE; dispatch_runs/cancel_lifecycle import close_run_sessions from HERE; tests
|
||||
# import BrowserSessionManager/fold_checkpoint_state from HERE. INV_7 split per
|
||||
# specs/044-dashboard-scenario-execution/plans/provider-decomposition-gate.md.
|
||||
# @SIDE_EFFECT The manager registry holds weak references only; session cleanup is best-effort
|
||||
# and fail-closed. Registry/checkpoint side effects live in the sibling modules.
|
||||
# @RATIONALE Original 501-line module split to satisfy INV_7 while keeping every region ID and
|
||||
# code byte verbatim; only this header reflects the facade role. DG-1 (2026-09-12):
|
||||
# native filter state applied by any run step must persist for all later steps —
|
||||
# per-step isolation broke filter continuity and risked false-PASS.
|
||||
# @REJECTED Closing the session at per-step lease release was rejected — it recreates the per-step
|
||||
# isolation DG-1 removed (launch+login per step, filter state lost). Reviving a dead
|
||||
# context object was rejected — process death kills the loop-bound context; only
|
||||
@@ -32,470 +29,30 @@
|
||||
# T034 (state leakage between runs).
|
||||
from __future__ import annotations
|
||||
|
||||
import copy
|
||||
import threading
|
||||
import time
|
||||
import weakref
|
||||
from collections.abc import Callable, Iterator
|
||||
from contextlib import contextmanager
|
||||
from typing import Any
|
||||
from src.services.dashboard_testing.execution.providers.browser_session_checkpoint import (
|
||||
BrowserCheckpointMissing,
|
||||
fold_checkpoint_state,
|
||||
load_persisted_browser_checkpoint,
|
||||
)
|
||||
from src.services.dashboard_testing.execution.providers.browser_session_handle import (
|
||||
BrowserSession,
|
||||
_PreparedStep,
|
||||
)
|
||||
from src.services.dashboard_testing.execution.providers.browser_session_managers import (
|
||||
close_run_sessions,
|
||||
)
|
||||
from src.services.dashboard_testing.execution.providers.browser_session_registry import (
|
||||
BrowserSessionManager,
|
||||
)
|
||||
|
||||
from src.core.database import SessionLocal
|
||||
from src.core.logger import logger
|
||||
from src.models.scenario_run import ScenarioStepRun
|
||||
from src.services.dashboard_testing.execution.provider_runtime import ProviderEventLoop
|
||||
__all__ = [
|
||||
"BrowserCheckpointMissing",
|
||||
"BrowserSession",
|
||||
"BrowserSessionManager",
|
||||
"_PreparedStep",
|
||||
"close_run_sessions",
|
||||
"fold_checkpoint_state",
|
||||
"load_persisted_browser_checkpoint",
|
||||
]
|
||||
|
||||
_SRC = "ScenarioExecution.BrowserProvider.Session"
|
||||
_DEFAULT_IDLE_TIMEOUT_SECONDS = 900.0
|
||||
_CLOSE_TIMEOUT_SECONDS = 10.0
|
||||
_FILTER_REPLAY_KEYS = frozenset({"filter_id", "filter_name", "column", "selector_hint", "values", "wait_state"})
|
||||
|
||||
|
||||
# #region ScenarioExecution.BrowserProvider.Session.ManagerRegistry [C:3] [TYPE Module] [SEMANTICS provider,browser,session,registry,finalizer]
|
||||
# @ingroup ScenarioExecution
|
||||
# @BRIEF Weak process registry of live session managers so server lifecycle paths (dispatch
|
||||
# terminal finalizer, cancellation) can close a run's session without holding a provider
|
||||
# reference (B-wave finalizer hook deferred from C1 round 1).
|
||||
# @INVARIANT The registry holds weak references only: it never extends a manager's lifetime, and
|
||||
# close_run_sessions never raises — session cleanup is best-effort and fail-closed.
|
||||
# @POST close_run_sessions returns the number of managers that dropped a session for the run.
|
||||
_MANAGERS: weakref.WeakSet[BrowserSessionManager] = weakref.WeakSet()
|
||||
_MANAGERS_LOCK = threading.Lock()
|
||||
|
||||
|
||||
def _register_manager(manager: BrowserSessionManager) -> None:
|
||||
with _MANAGERS_LOCK:
|
||||
_MANAGERS.add(manager)
|
||||
|
||||
|
||||
def close_run_sessions(run_id: str, *, reason: str) -> int:
|
||||
closed = 0
|
||||
with _MANAGERS_LOCK:
|
||||
managers = list(_MANAGERS)
|
||||
for manager in managers:
|
||||
try:
|
||||
if manager.close(run_id, reason=reason):
|
||||
closed += 1
|
||||
except Exception as exc: # pragma: no cover - close() is designed not to raise
|
||||
logger.explore(
|
||||
"Session manager close raised; continuing fan-out", src=_SRC,
|
||||
error_code="BROWSER_SESSION_CLOSE_FAILED", payload={"run_id": run_id, "reason": reason}, error=repr(exc),
|
||||
)
|
||||
if closed:
|
||||
logger.reflect(
|
||||
"Run-scoped browser sessions closed by lifecycle finalizer", src=_SRC,
|
||||
payload={"run_id": run_id, "reason": reason, "closed": closed},
|
||||
)
|
||||
return closed
|
||||
# #endregion ScenarioExecution.BrowserProvider.Session.ManagerRegistry
|
||||
|
||||
|
||||
# #region ScenarioExecution.BrowserProvider.Session.Errors [C:2] [TYPE Class] [SEMANTICS provider,browser,session,checkpoint,error]
|
||||
# @ingroup ScenarioExecution
|
||||
# @BRIEF Typed fail-closed recovery refusal: browser history exists but no declared checkpoint.
|
||||
class BrowserCheckpointMissing(Exception):
|
||||
"""Raised when a run needs browser-state recovery but no declared checkpoint exists."""
|
||||
# #endregion ScenarioExecution.BrowserProvider.Session.Errors
|
||||
|
||||
|
||||
# #region ScenarioExecution.BrowserProvider.Session.Handle [C:3] [TYPE Class] [SEMANTICS provider,browser,session,handle]
|
||||
# @ingroup ScenarioExecution
|
||||
# @BRIEF Runtime handle: run identity, transport-owned context, and the folded browser checkpoint.
|
||||
class BrowserSession:
|
||||
def __init__(
|
||||
self,
|
||||
*,
|
||||
run_id: str,
|
||||
lease_id: str,
|
||||
dashboard_id: int,
|
||||
transport_session: Any,
|
||||
checkpoint: dict[str, Any],
|
||||
created_at: float,
|
||||
last_used_at: float,
|
||||
replayed: bool = False,
|
||||
) -> None:
|
||||
self.run_id = run_id
|
||||
self.lease_id = lease_id
|
||||
self.dashboard_id = dashboard_id
|
||||
self.transport_session = transport_session
|
||||
self.checkpoint = checkpoint
|
||||
self.created_at = created_at
|
||||
self.last_used_at = last_used_at
|
||||
self.replayed = replayed
|
||||
|
||||
|
||||
class _PreparedStep:
|
||||
__slots__ = ("kind", "run_id", "lease_id", "dashboard_id", "session", "replay_state")
|
||||
|
||||
def __init__(
|
||||
self,
|
||||
*,
|
||||
kind: str,
|
||||
run_id: str,
|
||||
lease_id: str,
|
||||
dashboard_id: int,
|
||||
session: BrowserSession | None = None,
|
||||
replay_state: dict[str, Any] | None = None,
|
||||
) -> None:
|
||||
self.kind = kind
|
||||
self.run_id = run_id
|
||||
self.lease_id = lease_id
|
||||
self.dashboard_id = dashboard_id
|
||||
self.session = session
|
||||
self.replay_state = replay_state
|
||||
|
||||
|
||||
def _initial_checkpoint(dashboard_id: int) -> dict[str, Any]:
|
||||
return {
|
||||
"dashboard_id": dashboard_id,
|
||||
"checkpoint_seq": 0,
|
||||
"native_filter_state": [],
|
||||
"active_tab": None,
|
||||
"wait_states": [],
|
||||
"table_filter_state": [],
|
||||
"filter_state_observed": None,
|
||||
}
|
||||
# #endregion ScenarioExecution.BrowserProvider.Session.Handle
|
||||
|
||||
|
||||
# #region ScenarioExecution.BrowserProvider.Session.Checkpoint [C:4] [TYPE Function] [SEMANTICS provider,browser,session,checkpoint,native-filter,tab,table-filter]
|
||||
# @ingroup ScenarioExecution
|
||||
# @BRIEF Fold one executed step into the browser-safe checkpoint (extendable dict per DG-1 §3).
|
||||
# @PRE checkpoint is the session's current state (or None for a fresh context); details is the
|
||||
# transport outcome details of the just-executed action.
|
||||
# @POST Returns a NEW dict with checkpoint_seq incremented; apply_native_filter outcomes upsert a
|
||||
# native_filter_state entry carrying the UX-1 summary (filter_target, applied_values,
|
||||
# applied_mode) plus the exact replay_input; wait_for_state appends to wait_states;
|
||||
# navigate_tab stamps active_tab; apply_table_filter upserts table_filter_state;
|
||||
# inspect_filter_state stamps filter_state_observed (round 2, diagnostic — not replayed).
|
||||
# @INVARIANT Pure function: the input checkpoint is never mutated; unknown actions only bump the
|
||||
# sequence so the persisted slice always reflects the latest executed step.
|
||||
def fold_checkpoint_state(
|
||||
checkpoint: dict[str, Any] | None,
|
||||
*,
|
||||
dashboard_id: int,
|
||||
action: str,
|
||||
action_input: dict[str, Any],
|
||||
details: dict[str, Any],
|
||||
) -> dict[str, Any]:
|
||||
state = copy.deepcopy(checkpoint) if isinstance(checkpoint, dict) and checkpoint else _initial_checkpoint(dashboard_id)
|
||||
state["dashboard_id"] = dashboard_id
|
||||
state["checkpoint_seq"] = int(state.get("checkpoint_seq") or 0) + 1
|
||||
inputs = action_input if isinstance(action_input, dict) else {}
|
||||
if action == "apply_native_filter" and isinstance(details, dict) and details.get("applied"):
|
||||
entry = {
|
||||
"filter_target": details.get("filter_target"),
|
||||
"applied_values": [str(value) for value in details.get("applied_values") or []],
|
||||
"applied_mode": str(details.get("applied_mode") or ""),
|
||||
"replay_input": {key: copy.deepcopy(value) for key, value in inputs.items() if key in _FILTER_REPLAY_KEYS and value is not None},
|
||||
}
|
||||
filters = [
|
||||
existing
|
||||
for existing in state.get("native_filter_state") or []
|
||||
if isinstance(existing, dict) and existing.get("filter_target") != entry["filter_target"]
|
||||
]
|
||||
filters.append(entry)
|
||||
state["native_filter_state"] = filters
|
||||
elif action == "wait_for_state":
|
||||
state["wait_states"] = [*(state.get("wait_states") or []), str(inputs.get("state") or "load")]
|
||||
elif action == "navigate_tab" and isinstance(details, dict) and details.get("tab_navigated"):
|
||||
tab = str(details.get("tab") or "").strip()
|
||||
if tab:
|
||||
state["active_tab"] = tab
|
||||
elif action == "apply_table_filter" and isinstance(details, dict) and details.get("table_filter_applied"):
|
||||
entry = {"column": str(details.get("column") or ""), "value": details.get("value")}
|
||||
table_filters = [
|
||||
existing
|
||||
for existing in state.get("table_filter_state") or []
|
||||
if isinstance(existing, dict) and existing.get("column") != entry["column"]
|
||||
]
|
||||
table_filters.append(entry)
|
||||
state["table_filter_state"] = table_filters
|
||||
elif action == "inspect_filter_state" and isinstance(details, dict) and "filter_count" in details:
|
||||
state["filter_state_observed"] = {
|
||||
"filter_count": int(details.get("filter_count") or 0),
|
||||
"controls": copy.deepcopy(details.get("filter_controls") or [])[:50],
|
||||
}
|
||||
return state
|
||||
# #endregion ScenarioExecution.BrowserProvider.Session.Checkpoint
|
||||
|
||||
|
||||
# #region ScenarioExecution.BrowserProvider.Session.ReplayLookup [C:4] [TYPE Function] [SEMANTICS provider,browser,session,checkpoint,replay,recovery]
|
||||
# @ingroup ScenarioExecution
|
||||
# @BRIEF Read the latest persisted browser checkpoint of one run from scenario_step_runs.
|
||||
# @PRE run_id identifies persisted step rows; step_outcome carries tool="browser" and, for passed
|
||||
# session steps, the stamped browser_checkpoint dict.
|
||||
# @POST Returns ("none", None) when the run has no browser steps (fresh context is honest),
|
||||
# ("checkpoint", state) from the latest checkpointed step, or ("missing", None) when browser
|
||||
# history exists without any declared checkpoint (fail-closed per DG-1 §3).
|
||||
# @INVARIANT Read-only; a checkpoint is accepted only from a row whose step_outcome carries the
|
||||
# stamped dict — a crashed step that never stamped yields no recovery authority.
|
||||
def load_persisted_browser_checkpoint(
|
||||
run_id: str,
|
||||
*,
|
||||
session_factory: Callable[[], Any] = SessionLocal,
|
||||
) -> tuple[str, dict[str, Any] | None]:
|
||||
with session_factory() as db:
|
||||
rows = (
|
||||
db.query(ScenarioStepRun)
|
||||
.filter(ScenarioStepRun.run_id == run_id)
|
||||
.order_by(ScenarioStepRun.step_position.desc(), ScenarioStepRun.attempt.desc())
|
||||
.all()
|
||||
)
|
||||
saw_browser = False
|
||||
for row in rows:
|
||||
outcome = row.step_outcome if isinstance(row.step_outcome, dict) else {}
|
||||
if outcome.get("tool") != "browser":
|
||||
continue
|
||||
saw_browser = True
|
||||
checkpoint = outcome.get("browser_checkpoint")
|
||||
if isinstance(checkpoint, dict) and checkpoint.get("dashboard_id") is not None:
|
||||
logger.reflect(
|
||||
"Persisted browser checkpoint selected for replay", src=_SRC,
|
||||
payload={"run_id": run_id, "logical_step_id": row.logical_step_id, "checkpoint_seq": checkpoint.get("checkpoint_seq")},
|
||||
)
|
||||
return "checkpoint", checkpoint
|
||||
if saw_browser:
|
||||
logger.explore(
|
||||
"Run has browser history without a declared checkpoint", src=_SRC,
|
||||
claim="POST: replay or typed missing", error_code="BROWSER_CHECKPOINT_MISSING",
|
||||
payload={"run_id": run_id},
|
||||
)
|
||||
return "missing", None
|
||||
return "none", None
|
||||
# #endregion ScenarioExecution.BrowserProvider.Session.ReplayLookup
|
||||
|
||||
|
||||
# #region ScenarioExecution.BrowserProvider.Session.Registry [C:5] [TYPE Class] [SEMANTICS provider,browser,session,registry,leak-guard]
|
||||
# @ingroup ScenarioExecution
|
||||
# @BRIEF Own the run-scoped sessions: acquire/reuse, replay recovery, fold, close, idle reaping.
|
||||
# @RELATION DEPENDS_ON -> [ScenarioExecution.ProviderRuntime.Engine]
|
||||
# @PRE Constructed with a session-capable transport and the shared provider event loop.
|
||||
# @POST At most one live session per run_id; every close path is idempotent and never raises.
|
||||
# @INVARIANT Registry mutations are serialized by the registry lock; concurrent steps of one run
|
||||
# are serialized by the per-run guard held across prepare + loop execution.
|
||||
class BrowserSessionManager:
|
||||
def __init__(
|
||||
self,
|
||||
*,
|
||||
transport: Any,
|
||||
event_loop: ProviderEventLoop,
|
||||
checkpoint_loader: Callable[..., tuple[str, dict[str, Any] | None]] | None = None,
|
||||
idle_timeout_seconds: float = _DEFAULT_IDLE_TIMEOUT_SECONDS,
|
||||
clock: Callable[[], float] = time.monotonic,
|
||||
) -> None:
|
||||
if idle_timeout_seconds <= 0:
|
||||
raise ValueError("BROWSER_SESSION_IDLE_TIMEOUT_INVALID")
|
||||
self._transport = transport
|
||||
self._event_loop = event_loop
|
||||
self._checkpoint_loader = checkpoint_loader or load_persisted_browser_checkpoint
|
||||
self._idle_timeout = float(idle_timeout_seconds)
|
||||
self._clock = clock
|
||||
self._sessions: dict[str, BrowserSession] = {}
|
||||
self._registry_lock = threading.Lock()
|
||||
self._run_locks: dict[str, threading.Lock] = {}
|
||||
_register_manager(self)
|
||||
|
||||
def active_run_ids(self) -> list[str]:
|
||||
with self._registry_lock:
|
||||
return sorted(self._sessions)
|
||||
|
||||
# #region ScenarioExecution.BrowserProvider.Session.Registry.Guard [C:3] [TYPE Function] [SEMANTICS provider,browser,session,serialization]
|
||||
# @ingroup ScenarioExecution
|
||||
# @BRIEF Per-run serialization guard (DG-1 §4): concurrent steps of one run are ordered.
|
||||
@contextmanager
|
||||
def run_guard(self, run_id: str) -> Iterator[None]:
|
||||
with self._registry_lock:
|
||||
lock = self._run_locks.setdefault(run_id, threading.Lock())
|
||||
lock.acquire()
|
||||
try:
|
||||
yield
|
||||
finally:
|
||||
lock.release()
|
||||
# #endregion ScenarioExecution.BrowserProvider.Session.Registry.Guard
|
||||
|
||||
# #region ScenarioExecution.BrowserProvider.Session.Registry.Acquire [C:5] [TYPE Function] [SEMANTICS provider,browser,session,acquire,replay]
|
||||
# @ingroup ScenarioExecution
|
||||
# @BRIEF Resolve the step plan: reuse the live session, replay the persisted checkpoint, or open fresh.
|
||||
# @PRE The caller holds run_guard(run_id); lease_id is the current step's capacity lease
|
||||
# (provenance only — the session outlives the per-step lease).
|
||||
# @POST A reuse plan carries the live session; a replay/fresh plan opens on the loop inside
|
||||
# execute_prepared; browser history without a checkpoint or a checkpoint pinned to another
|
||||
# dashboard raises BrowserCheckpointMissing before any I/O.
|
||||
def prepare_step(self, *, run_id: str, lease_id: str, dashboard_id: int) -> _PreparedStep:
|
||||
self.reap_idle()
|
||||
with self._registry_lock:
|
||||
session = self._sessions.get(run_id)
|
||||
if session is not None:
|
||||
session.lease_id = lease_id
|
||||
logger.reason(
|
||||
"Reusing the run-scoped browser session", src=_SRC,
|
||||
payload={"run_id": run_id, "dashboard_id": session.dashboard_id, "replayed": session.replayed},
|
||||
)
|
||||
return _PreparedStep(kind="reuse", run_id=run_id, lease_id=lease_id, dashboard_id=dashboard_id, session=session)
|
||||
kind, state = self._checkpoint_loader(run_id)
|
||||
if kind == "missing":
|
||||
logger.explore(
|
||||
"Browser recovery requires a declared checkpoint", src=_SRC,
|
||||
claim="POST: replay or typed missing", error_code="BROWSER_CHECKPOINT_MISSING",
|
||||
payload={"run_id": run_id},
|
||||
)
|
||||
raise BrowserCheckpointMissing("BROWSER_CHECKPOINT_MISSING")
|
||||
if kind == "checkpoint":
|
||||
filters = state.get("native_filter_state") if isinstance(state, dict) else None
|
||||
if (
|
||||
not isinstance(state, dict)
|
||||
or state.get("dashboard_id") != dashboard_id
|
||||
or any(not isinstance(entry, dict) or not isinstance(entry.get("replay_input"), dict) for entry in filters or [])
|
||||
):
|
||||
logger.explore(
|
||||
"Persisted checkpoint cannot authorize replay for this binding", src=_SRC,
|
||||
claim="POST: replay or typed missing", error_code="BROWSER_CHECKPOINT_MISSING",
|
||||
payload={"run_id": run_id, "dashboard_id": dashboard_id},
|
||||
)
|
||||
raise BrowserCheckpointMissing("BROWSER_CHECKPOINT_MISSING")
|
||||
return _PreparedStep(kind="replay", run_id=run_id, lease_id=lease_id, dashboard_id=dashboard_id, replay_state=state)
|
||||
return _PreparedStep(kind="fresh", run_id=run_id, lease_id=lease_id, dashboard_id=dashboard_id)
|
||||
# #endregion ScenarioExecution.BrowserProvider.Session.Registry.Acquire
|
||||
|
||||
# #region ScenarioExecution.BrowserProvider.Session.Registry.Execute [C:5] [TYPE Function] [SEMANTICS provider,browser,session,execute,checkpoint]
|
||||
# @ingroup ScenarioExecution
|
||||
# @BRIEF Coroutine (provider loop): open/replay on first use, execute the action, fold the checkpoint.
|
||||
# @PRE prepared comes from prepare_step under the run guard; runs entirely on the provider loop.
|
||||
# @POST Returns (transport outcome, stamped checkpoint copy); the session is registered before
|
||||
# replay so a mid-step cancellation is always closable by the leak-guard; a failed replay
|
||||
# discards and closes the fresh context before re-raising.
|
||||
async def execute_prepared(
|
||||
self,
|
||||
prepared: _PreparedStep,
|
||||
action: str,
|
||||
*,
|
||||
action_input: dict[str, Any],
|
||||
timeout_seconds: float,
|
||||
) -> tuple[Any, dict[str, Any]]:
|
||||
session = prepared.session or await self._open_session(prepared, timeout_seconds=timeout_seconds)
|
||||
outcome = await self._transport.execute_in_session(
|
||||
session.transport_session,
|
||||
action,
|
||||
action_input=action_input,
|
||||
timeout_seconds=timeout_seconds,
|
||||
)
|
||||
session.checkpoint = fold_checkpoint_state(
|
||||
session.checkpoint,
|
||||
dashboard_id=session.dashboard_id,
|
||||
action=action,
|
||||
action_input=action_input,
|
||||
details=outcome.details if isinstance(outcome.details, dict) else {},
|
||||
)
|
||||
session.last_used_at = self._clock()
|
||||
logger.reflect(
|
||||
"Browser step folded into the run checkpoint", src=_SRC,
|
||||
payload={"run_id": session.run_id, "action": action, "checkpoint_seq": session.checkpoint.get("checkpoint_seq"), "filters": len(session.checkpoint.get("native_filter_state") or [])},
|
||||
)
|
||||
return outcome, copy.deepcopy(session.checkpoint)
|
||||
# #endregion ScenarioExecution.BrowserProvider.Session.Registry.Execute
|
||||
|
||||
# #region ScenarioExecution.BrowserProvider.Session.Registry.Open [C:4] [TYPE Function] [SEMANTICS provider,browser,session,open,replay]
|
||||
# @ingroup ScenarioExecution
|
||||
# @BRIEF Open the transport context, register early (leak-guard visibility), then replay state.
|
||||
async def _open_session(self, prepared: _PreparedStep, *, timeout_seconds: float) -> BrowserSession:
|
||||
handle = await self._transport.open_session(prepared.dashboard_id, timeout_seconds=timeout_seconds)
|
||||
now = self._clock()
|
||||
session = BrowserSession(
|
||||
run_id=prepared.run_id,
|
||||
lease_id=prepared.lease_id,
|
||||
dashboard_id=prepared.dashboard_id,
|
||||
transport_session=handle,
|
||||
checkpoint=copy.deepcopy(prepared.replay_state) if prepared.replay_state is not None else _initial_checkpoint(prepared.dashboard_id),
|
||||
created_at=now,
|
||||
last_used_at=now,
|
||||
)
|
||||
with self._registry_lock:
|
||||
self._sessions[prepared.run_id] = session
|
||||
logger.reflect(
|
||||
"Run-scoped browser session opened", src=_SRC,
|
||||
payload={"run_id": prepared.run_id, "dashboard_id": prepared.dashboard_id, "kind": prepared.kind},
|
||||
)
|
||||
if prepared.replay_state is not None:
|
||||
try:
|
||||
for entry in prepared.replay_state.get("native_filter_state") or []:
|
||||
await self._transport.execute_in_session(
|
||||
handle,
|
||||
"apply_native_filter",
|
||||
action_input=copy.deepcopy(entry["replay_input"]),
|
||||
timeout_seconds=timeout_seconds,
|
||||
)
|
||||
session.replayed = True
|
||||
logger.reflect(
|
||||
"Browser checkpoint replayed into a fresh context", src=_SRC,
|
||||
payload={"run_id": prepared.run_id, "filters": len(prepared.replay_state.get("native_filter_state") or [])},
|
||||
)
|
||||
except Exception as exc:
|
||||
with self._registry_lock:
|
||||
self._sessions.pop(prepared.run_id, None)
|
||||
logger.explore(
|
||||
"Checkpoint replay failed; fresh context abandoned", src=_SRC,
|
||||
error_code="BROWSER_CHECKPOINT_REPLAY_FAILED",
|
||||
payload={"run_id": prepared.run_id}, error=repr(exc),
|
||||
)
|
||||
await self._close_handle(handle)
|
||||
raise
|
||||
return session
|
||||
# #endregion ScenarioExecution.BrowserProvider.Session.Registry.Open
|
||||
|
||||
# #region ScenarioExecution.BrowserProvider.Session.Registry.Close [C:4] [TYPE Function] [SEMANTICS provider,browser,session,close,leak-guard]
|
||||
# @ingroup ScenarioExecution
|
||||
# @BRIEF Leak-guard close paths: explicit terminal close, dead-context close, idle reap, close_all.
|
||||
# @POST close() is idempotent and never raises: the registry entry is dropped first and the
|
||||
# transport close is best-effort — a dead loop abandons the loop-bound context by design.
|
||||
def close(self, run_id: str, *, reason: str = "terminal") -> bool:
|
||||
with self._registry_lock:
|
||||
session = self._sessions.pop(run_id, None)
|
||||
if session is None:
|
||||
return False
|
||||
try:
|
||||
self._event_loop.submit(lambda: self._close_handle(session.transport_session), timeout=_CLOSE_TIMEOUT_SECONDS)
|
||||
except Exception as exc:
|
||||
logger.explore(
|
||||
"Browser session close could not reach the loop; context dies with the loop", src=_SRC,
|
||||
error_code="BROWSER_SESSION_CLOSE_FAILED",
|
||||
payload={"run_id": run_id, "reason": reason}, error=repr(exc),
|
||||
)
|
||||
return True
|
||||
logger.reflect("Run-scoped browser session closed", src=_SRC, payload={"run_id": run_id, "reason": reason})
|
||||
return True
|
||||
|
||||
def reap_idle(self) -> int:
|
||||
now = self._clock()
|
||||
with self._registry_lock:
|
||||
stale = [
|
||||
run_id
|
||||
for run_id, session in self._sessions.items()
|
||||
if now - session.last_used_at > self._idle_timeout
|
||||
]
|
||||
closed = 0
|
||||
for run_id in stale:
|
||||
if self.close(run_id, reason="idle_timeout"):
|
||||
closed += 1
|
||||
return closed
|
||||
|
||||
def close_all(self, *, reason: str = "shutdown") -> int:
|
||||
with self._registry_lock:
|
||||
run_ids = list(self._sessions)
|
||||
closed = 0
|
||||
for run_id in run_ids:
|
||||
if self.close(run_id, reason=reason):
|
||||
closed += 1
|
||||
return closed
|
||||
|
||||
async def _close_handle(self, handle: Any) -> None:
|
||||
try:
|
||||
await self._transport.close_session(handle)
|
||||
except Exception as exc:
|
||||
logger.explore(
|
||||
"Transport session close raised; handle abandoned", src=_SRC,
|
||||
error_code="BROWSER_SESSION_CLOSE_FAILED", error=repr(exc),
|
||||
)
|
||||
# #endregion ScenarioExecution.BrowserProvider.Session.Registry.Close
|
||||
# #endregion ScenarioExecution.BrowserProvider.Session.Registry
|
||||
# #endregion ScenarioExecution.BrowserProvider.Session
|
||||
|
||||
@@ -0,0 +1,162 @@
|
||||
# #region ScenarioExecution.BrowserProvider.Session.CheckpointIO [C:3] [TYPE Module] [SEMANTICS provider,browser,session,checkpoint,replay,persistence]
|
||||
# @ingroup ScenarioExecution
|
||||
# @BRIEF Browser-safe checkpoint IO: the initial checkpoint factory, the fail-closed recovery
|
||||
# refusal error, the pure fold that accumulates one step's effect into the checkpoint, and
|
||||
# the persisted-checkpoint lookup used for cross-process replay.
|
||||
# @POST fold_checkpoint_state returns a new deep-copied checkpoint (the input is never mutated);
|
||||
# load_persisted_browser_checkpoint returns (step_kind, checkpoint_or_None) from
|
||||
# scenario_step_runs and never raises on missing rows.
|
||||
# @INVARIANT The folded checkpoint is the ONLY recovery authority — a dead context is never revived;
|
||||
# history without a declared checkpoint raises BrowserCheckpointMissing (fail-closed).
|
||||
# @RELATION DEPENDS_ON -> [ScenarioExecution.BrowserProvider.Session.Handle]
|
||||
# @RATIONALE Extracted from browser_session.py (INV_7) keeping the Checkpoint/ReplayLookup/Errors
|
||||
# regions verbatim and giving _initial_checkpoint its own region; the facade re-exports
|
||||
# every moved name so imports stay frozen.
|
||||
# @REJECTED Inlining the fold into the session registry was rejected — the fold must stay pure and
|
||||
# unit-testable without a live context (see the decomposition-gate plan).
|
||||
from __future__ import annotations
|
||||
|
||||
import copy
|
||||
from collections.abc import Callable
|
||||
from typing import Any
|
||||
|
||||
from src.core.database import SessionLocal
|
||||
from src.core.logger import logger
|
||||
from src.models.scenario_run import ScenarioStepRun
|
||||
|
||||
_SRC = "ScenarioExecution.BrowserProvider.Session"
|
||||
_FILTER_REPLAY_KEYS = frozenset({"filter_id", "filter_name", "column", "selector_hint", "values", "wait_state"})
|
||||
|
||||
|
||||
# #region ScenarioExecution.BrowserProvider.Session.InitialCheckpoint [C:1] [TYPE Function] [SEMANTICS provider,browser,session,checkpoint,initial]
|
||||
# @ingroup ScenarioExecution
|
||||
# @BRIEF Empty dashboard checkpoint used when no prior state exists for a run.
|
||||
# @POST Returns a fresh dict with checkpoint_seq=0 and empty state lists; performs no I/O.
|
||||
def _initial_checkpoint(dashboard_id: int) -> dict[str, Any]:
|
||||
return {
|
||||
"dashboard_id": dashboard_id,
|
||||
"checkpoint_seq": 0,
|
||||
"native_filter_state": [],
|
||||
"active_tab": None,
|
||||
"wait_states": [],
|
||||
"table_filter_state": [],
|
||||
"filter_state_observed": None,
|
||||
}
|
||||
# #endregion ScenarioExecution.BrowserProvider.Session.InitialCheckpoint
|
||||
|
||||
|
||||
# #region ScenarioExecution.BrowserProvider.Session.Errors [C:2] [TYPE Class] [SEMANTICS provider,browser,session,checkpoint,error]
|
||||
# @ingroup ScenarioExecution
|
||||
# @BRIEF Typed fail-closed recovery refusal: browser history exists but no declared checkpoint.
|
||||
class BrowserCheckpointMissing(Exception):
|
||||
"""Raised when a run needs browser-state recovery but no declared checkpoint exists."""
|
||||
# #endregion ScenarioExecution.BrowserProvider.Session.Errors
|
||||
|
||||
|
||||
# #region ScenarioExecution.BrowserProvider.Session.Checkpoint [C:4] [TYPE Function] [SEMANTICS provider,browser,session,checkpoint,native-filter,tab,table-filter]
|
||||
# @ingroup ScenarioExecution
|
||||
# @BRIEF Fold one executed step into the browser-safe checkpoint (extendable dict per DG-1 §3).
|
||||
# @PRE checkpoint is the session's current state (or None for a fresh context); details is the
|
||||
# transport outcome details of the just-executed action.
|
||||
# @POST Returns a NEW dict with checkpoint_seq incremented; apply_native_filter outcomes upsert a
|
||||
# native_filter_state entry carrying the UX-1 summary (filter_target, applied_values,
|
||||
# applied_mode) plus the exact replay_input; wait_for_state appends to wait_states;
|
||||
# navigate_tab stamps active_tab; apply_table_filter upserts table_filter_state;
|
||||
# inspect_filter_state stamps filter_state_observed (round 2, diagnostic — not replayed).
|
||||
# @INVARIANT Pure function: the input checkpoint is never mutated; unknown actions only bump the
|
||||
# sequence so the persisted slice always reflects the latest executed step.
|
||||
def fold_checkpoint_state(
|
||||
checkpoint: dict[str, Any] | None,
|
||||
*,
|
||||
dashboard_id: int,
|
||||
action: str,
|
||||
action_input: dict[str, Any],
|
||||
details: dict[str, Any],
|
||||
) -> dict[str, Any]:
|
||||
state = copy.deepcopy(checkpoint) if isinstance(checkpoint, dict) and checkpoint else _initial_checkpoint(dashboard_id)
|
||||
state["dashboard_id"] = dashboard_id
|
||||
state["checkpoint_seq"] = int(state.get("checkpoint_seq") or 0) + 1
|
||||
inputs = action_input if isinstance(action_input, dict) else {}
|
||||
if action == "apply_native_filter" and isinstance(details, dict) and details.get("applied"):
|
||||
entry = {
|
||||
"filter_target": details.get("filter_target"),
|
||||
"applied_values": [str(value) for value in details.get("applied_values") or []],
|
||||
"applied_mode": str(details.get("applied_mode") or ""),
|
||||
"replay_input": {key: copy.deepcopy(value) for key, value in inputs.items() if key in _FILTER_REPLAY_KEYS and value is not None},
|
||||
}
|
||||
filters = [
|
||||
existing
|
||||
for existing in state.get("native_filter_state") or []
|
||||
if isinstance(existing, dict) and existing.get("filter_target") != entry["filter_target"]
|
||||
]
|
||||
filters.append(entry)
|
||||
state["native_filter_state"] = filters
|
||||
elif action == "wait_for_state":
|
||||
state["wait_states"] = [*(state.get("wait_states") or []), str(inputs.get("state") or "load")]
|
||||
elif action == "navigate_tab" and isinstance(details, dict) and details.get("tab_navigated"):
|
||||
tab = str(details.get("tab") or "").strip()
|
||||
if tab:
|
||||
state["active_tab"] = tab
|
||||
elif action == "apply_table_filter" and isinstance(details, dict) and details.get("table_filter_applied"):
|
||||
entry = {"column": str(details.get("column") or ""), "value": details.get("value")}
|
||||
table_filters = [
|
||||
existing
|
||||
for existing in state.get("table_filter_state") or []
|
||||
if isinstance(existing, dict) and existing.get("column") != entry["column"]
|
||||
]
|
||||
table_filters.append(entry)
|
||||
state["table_filter_state"] = table_filters
|
||||
elif action == "inspect_filter_state" and isinstance(details, dict) and "filter_count" in details:
|
||||
state["filter_state_observed"] = {
|
||||
"filter_count": int(details.get("filter_count") or 0),
|
||||
"controls": copy.deepcopy(details.get("filter_controls") or [])[:50],
|
||||
}
|
||||
return state
|
||||
# #endregion ScenarioExecution.BrowserProvider.Session.Checkpoint
|
||||
|
||||
|
||||
# #region ScenarioExecution.BrowserProvider.Session.ReplayLookup [C:4] [TYPE Function] [SEMANTICS provider,browser,session,checkpoint,replay,recovery]
|
||||
# @ingroup ScenarioExecution
|
||||
# @BRIEF Read the latest persisted browser checkpoint of one run from scenario_step_runs.
|
||||
# @PRE run_id identifies persisted step rows; step_outcome carries tool="browser" and, for passed
|
||||
# session steps, the stamped browser_checkpoint dict.
|
||||
# @POST Returns ("none", None) when the run has no browser steps (fresh context is honest),
|
||||
# ("checkpoint", state) from the latest checkpointed step, or ("missing", None) when browser
|
||||
# history exists without any declared checkpoint (fail-closed per DG-1 §3).
|
||||
# @INVARIANT Read-only; a checkpoint is accepted only from a row whose step_outcome carries the
|
||||
# stamped dict — a crashed step that never stamped yields no recovery authority.
|
||||
def load_persisted_browser_checkpoint(
|
||||
run_id: str,
|
||||
*,
|
||||
session_factory: Callable[[], Any] = SessionLocal,
|
||||
) -> tuple[str, dict[str, Any] | None]:
|
||||
with session_factory() as db:
|
||||
rows = (
|
||||
db.query(ScenarioStepRun)
|
||||
.filter(ScenarioStepRun.run_id == run_id)
|
||||
.order_by(ScenarioStepRun.step_position.desc(), ScenarioStepRun.attempt.desc())
|
||||
.all()
|
||||
)
|
||||
saw_browser = False
|
||||
for row in rows:
|
||||
outcome = row.step_outcome if isinstance(row.step_outcome, dict) else {}
|
||||
if outcome.get("tool") != "browser":
|
||||
continue
|
||||
saw_browser = True
|
||||
checkpoint = outcome.get("browser_checkpoint")
|
||||
if isinstance(checkpoint, dict) and checkpoint.get("dashboard_id") is not None:
|
||||
logger.reflect(
|
||||
"Persisted browser checkpoint selected for replay", src=_SRC,
|
||||
payload={"run_id": run_id, "logical_step_id": row.logical_step_id, "checkpoint_seq": checkpoint.get("checkpoint_seq")},
|
||||
)
|
||||
return "checkpoint", checkpoint
|
||||
if saw_browser:
|
||||
logger.explore(
|
||||
"Run has browser history without a declared checkpoint", src=_SRC,
|
||||
claim="POST: replay or typed missing", error_code="BROWSER_CHECKPOINT_MISSING",
|
||||
payload={"run_id": run_id},
|
||||
)
|
||||
return "missing", None
|
||||
return "none", None
|
||||
# #endregion ScenarioExecution.BrowserProvider.Session.ReplayLookup
|
||||
# #endregion ScenarioExecution.BrowserProvider.Session.CheckpointIO
|
||||
@@ -0,0 +1,64 @@
|
||||
# #region ScenarioExecution.BrowserProvider.Session.Handle [C:3] [TYPE Module] [SEMANTICS provider,browser,session,handle,prepared-step]
|
||||
# @ingroup ScenarioExecution
|
||||
# @BRIEF Runtime handle types for the run-scoped browser session: the transport-owned context
|
||||
# holder (BrowserSession) and the prepared-step envelope (_PreparedStep) that carries a
|
||||
# live session or a replayed checkpoint state into one step.
|
||||
# @POST BrowserSession carries run identity, the transport session, the folded checkpoint and
|
||||
# timestamps; _PreparedStep carries kind/run/lease/dashboard plus session or replay_state.
|
||||
# @INVARIANT Handles are never shared across runs; _PreparedStep holds at most one of
|
||||
# session/replay_state per step.
|
||||
# @RELATION DEPENDS_ON -> [ScenarioExecution.BrowserProvider.Session.CheckpointIO]
|
||||
# @RATIONALE Extracted from browser_session.py (INV_7) with the region ID preserved; the facade
|
||||
# re-exports these types so the transport import surface stays frozen.
|
||||
# @REJECTED Keeping the dataclass pair inside the 501-line session module was rejected — see
|
||||
# specs/044-dashboard-scenario-execution/plans/provider-decomposition-gate.md.
|
||||
from __future__ import annotations
|
||||
|
||||
from typing import Any
|
||||
|
||||
class BrowserSession:
|
||||
def __init__(
|
||||
self,
|
||||
*,
|
||||
run_id: str,
|
||||
lease_id: str,
|
||||
dashboard_id: int,
|
||||
transport_session: Any,
|
||||
checkpoint: dict[str, Any],
|
||||
created_at: float,
|
||||
last_used_at: float,
|
||||
replayed: bool = False,
|
||||
) -> None:
|
||||
self.run_id = run_id
|
||||
self.lease_id = lease_id
|
||||
self.dashboard_id = dashboard_id
|
||||
self.transport_session = transport_session
|
||||
self.checkpoint = checkpoint
|
||||
self.created_at = created_at
|
||||
self.last_used_at = last_used_at
|
||||
self.replayed = replayed
|
||||
|
||||
|
||||
|
||||
class _PreparedStep:
|
||||
__slots__ = ("kind", "run_id", "lease_id", "dashboard_id", "session", "replay_state")
|
||||
|
||||
def __init__(
|
||||
self,
|
||||
*,
|
||||
kind: str,
|
||||
run_id: str,
|
||||
lease_id: str,
|
||||
dashboard_id: int,
|
||||
session: BrowserSession | None = None,
|
||||
replay_state: dict[str, Any] | None = None,
|
||||
) -> None:
|
||||
self.kind = kind
|
||||
self.run_id = run_id
|
||||
self.lease_id = lease_id
|
||||
self.dashboard_id = dashboard_id
|
||||
self.session = session
|
||||
self.replay_state = replay_state
|
||||
|
||||
|
||||
# #endregion ScenarioExecution.BrowserProvider.Session.Handle
|
||||
@@ -0,0 +1,57 @@
|
||||
# #region ScenarioExecution.BrowserProvider.Session.ManagerRegistry [C:3] [TYPE Module] [SEMANTICS provider,browser,session,registry,finalizer]
|
||||
# @ingroup ScenarioExecution
|
||||
# @BRIEF Weak process registry of live session managers so server lifecycle paths (dispatch
|
||||
# terminal finalizer, cancellation) can close a run's session without holding a provider
|
||||
# reference (B-wave finalizer hook deferred from C1 round 1).
|
||||
# @INVARIANT The registry holds weak references only: it never extends a manager's lifetime, and
|
||||
# close_run_sessions never raises — session cleanup is best-effort and fail-closed.
|
||||
# @POST close_run_sessions returns the number of managers that dropped a session for the run.
|
||||
# @RELATION DEPENDS_ON -> [ScenarioExecution.BrowserProvider.Session.Registry]
|
||||
# @RATIONALE Extracted from browser_session.py (INV_7) keeping the region verbatim; registry-B
|
||||
# registers through this module so no facade<->registry import cycle is created.
|
||||
# @REJECTED A function-local import of _register_manager inside the registry __init__ was rejected —
|
||||
# the manager registry is its own lifecycle boundary and deserves a real module.
|
||||
from __future__ import annotations
|
||||
|
||||
import threading
|
||||
import weakref
|
||||
from typing import TYPE_CHECKING
|
||||
|
||||
from src.core.logger import logger
|
||||
|
||||
if TYPE_CHECKING: # pragma: no cover - typing only, avoids a runtime import cycle
|
||||
from src.services.dashboard_testing.execution.providers.browser_session_registry import (
|
||||
BrowserSessionManager,
|
||||
)
|
||||
|
||||
_SRC = "ScenarioExecution.BrowserProvider.Session"
|
||||
|
||||
_MANAGERS: weakref.WeakSet[BrowserSessionManager] = weakref.WeakSet()
|
||||
_MANAGERS_LOCK = threading.Lock()
|
||||
|
||||
|
||||
def _register_manager(manager: BrowserSessionManager) -> None:
|
||||
with _MANAGERS_LOCK:
|
||||
_MANAGERS.add(manager)
|
||||
|
||||
|
||||
def close_run_sessions(run_id: str, *, reason: str) -> int:
|
||||
closed = 0
|
||||
with _MANAGERS_LOCK:
|
||||
managers = list(_MANAGERS)
|
||||
for manager in managers:
|
||||
try:
|
||||
if manager.close(run_id, reason=reason):
|
||||
closed += 1
|
||||
except Exception as exc: # pragma: no cover - close() is designed not to raise
|
||||
logger.explore(
|
||||
"Session manager close raised; continuing fan-out", src=_SRC,
|
||||
error_code="BROWSER_SESSION_CLOSE_FAILED", payload={"run_id": run_id, "reason": reason}, error=repr(exc),
|
||||
)
|
||||
if closed:
|
||||
logger.reflect(
|
||||
"Run-scoped browser sessions closed by lifecycle finalizer", src=_SRC,
|
||||
payload={"run_id": run_id, "reason": reason, "closed": closed},
|
||||
)
|
||||
return closed
|
||||
# #endregion ScenarioExecution.BrowserProvider.Session.ManagerRegistry
|
||||
@@ -0,0 +1,287 @@
|
||||
# #region ScenarioExecution.BrowserProvider.Session.Registry [C:5] [TYPE Module] [SEMANTICS provider,browser,session,registry,leak-guard]
|
||||
# @ingroup ScenarioExecution
|
||||
# @BRIEF Run-scoped browser session registry on the shared provider loop: one isolated context per
|
||||
# run lives across all browser steps, a browser-safe checkpoint is folded after every step,
|
||||
# cross-process recovery replays the persisted checkpoint, and close paths are guaranteed
|
||||
# (DG-1 amendment 2026-09-12, 044 T034 round 1).
|
||||
# @RELATION IMPLEMENTS -> [ScenarioExecution.BrowserProvider]
|
||||
# @RELATION DEPENDS_ON -> [ScenarioExecution.ProviderRuntime.Engine]
|
||||
# @RELATION DEPENDS_ON -> [ScenarioExecution.BrowserProvider.Transport]
|
||||
# @RELATION DEPENDS_ON -> [Models.ScenarioExecution.StepRun]
|
||||
# @RELATION DEPENDS_ON -> [ScenarioExecution.BrowserProvider.Session.CheckpointIO]
|
||||
# @RELATION DEPENDS_ON -> [ScenarioExecution.BrowserProvider.Session.Handle]
|
||||
# @PRE The manager is bound to a session-capable transport (open_session / execute_in_session /
|
||||
# close_session) and the shared provider event loop; the provider holds the per-run guard
|
||||
# across prepare + submit so concurrent steps of one run serialize on the run context (DG-1 §4).
|
||||
# @POST prepare_step returns exactly one live session per run_id: an in-memory hit is reused, a
|
||||
# persisted checkpoint replays open + re-apply before the step, and browser history without
|
||||
# any declared checkpoint raises BrowserCheckpointMissing (fail-closed, never a silent
|
||||
# stateless session). close() is idempotent and best-effort on every path.
|
||||
# @INVARIANT A context is never shared across runs; a dead context is never revived — recovery is a
|
||||
# new context plus checkpoint replay from scenario_step_runs.step_outcome; the registry
|
||||
# lives in process memory, so process death kills every context and only replay recovers.
|
||||
# @SIDE_EFFECT Live browser/context handles bound to the provider loop, an in-process registry, and
|
||||
# read-only checkpoint queries against scenario_step_runs.
|
||||
# @RATIONALE Extracted verbatim from browser_session.py (INV_7: 501 LOC) with the region ID
|
||||
# preserved; the facade re-exports BrowserSessionManager so browser.py and the tests
|
||||
# keep their import surface. DG-1 rationale (filter continuity across steps, per-step
|
||||
# capacity lease, explicit close) is unchanged.
|
||||
# @REJECTED Closing the session at per-step lease release was rejected — it recreates the per-step
|
||||
# isolation DG-1 removed (launch+login per step, filter state lost). Reviving a dead
|
||||
# context object was rejected — process death kills the loop-bound context; only
|
||||
# checkpoint replay reconstructs state. A warm cross-run context pool was rejected by
|
||||
# T034 (state leakage between runs).
|
||||
from __future__ import annotations
|
||||
|
||||
import copy
|
||||
import threading
|
||||
import time
|
||||
from collections.abc import Callable, Iterator
|
||||
from contextlib import contextmanager
|
||||
from typing import Any
|
||||
|
||||
from src.core.logger import logger
|
||||
from src.services.dashboard_testing.execution.provider_runtime import ProviderEventLoop
|
||||
from src.services.dashboard_testing.execution.providers.browser_session_checkpoint import (
|
||||
BrowserCheckpointMissing,
|
||||
_initial_checkpoint,
|
||||
fold_checkpoint_state,
|
||||
load_persisted_browser_checkpoint,
|
||||
)
|
||||
from src.services.dashboard_testing.execution.providers.browser_session_managers import (
|
||||
_register_manager,
|
||||
)
|
||||
from src.services.dashboard_testing.execution.providers.browser_session_handle import (
|
||||
BrowserSession,
|
||||
_PreparedStep,
|
||||
)
|
||||
|
||||
_SRC = "ScenarioExecution.BrowserProvider.Session"
|
||||
_DEFAULT_IDLE_TIMEOUT_SECONDS = 900.0
|
||||
_CLOSE_TIMEOUT_SECONDS = 10.0
|
||||
|
||||
|
||||
class BrowserSessionManager:
|
||||
def __init__(
|
||||
self,
|
||||
*,
|
||||
transport: Any,
|
||||
event_loop: ProviderEventLoop,
|
||||
checkpoint_loader: Callable[..., tuple[str, dict[str, Any] | None]] | None = None,
|
||||
idle_timeout_seconds: float = _DEFAULT_IDLE_TIMEOUT_SECONDS,
|
||||
clock: Callable[[], float] = time.monotonic,
|
||||
) -> None:
|
||||
if idle_timeout_seconds <= 0:
|
||||
raise ValueError("BROWSER_SESSION_IDLE_TIMEOUT_INVALID")
|
||||
self._transport = transport
|
||||
self._event_loop = event_loop
|
||||
self._checkpoint_loader = checkpoint_loader or load_persisted_browser_checkpoint
|
||||
self._idle_timeout = float(idle_timeout_seconds)
|
||||
self._clock = clock
|
||||
self._sessions: dict[str, BrowserSession] = {}
|
||||
self._registry_lock = threading.Lock()
|
||||
self._run_locks: dict[str, threading.Lock] = {}
|
||||
_register_manager(self)
|
||||
|
||||
def active_run_ids(self) -> list[str]:
|
||||
with self._registry_lock:
|
||||
return sorted(self._sessions)
|
||||
|
||||
# #region ScenarioExecution.BrowserProvider.Session.Registry.Guard [C:3] [TYPE Function] [SEMANTICS provider,browser,session,serialization]
|
||||
# @ingroup ScenarioExecution
|
||||
# @BRIEF Per-run serialization guard (DG-1 §4): concurrent steps of one run are ordered.
|
||||
@contextmanager
|
||||
def run_guard(self, run_id: str) -> Iterator[None]:
|
||||
with self._registry_lock:
|
||||
lock = self._run_locks.setdefault(run_id, threading.Lock())
|
||||
lock.acquire()
|
||||
try:
|
||||
yield
|
||||
finally:
|
||||
lock.release()
|
||||
# #endregion ScenarioExecution.BrowserProvider.Session.Registry.Guard
|
||||
|
||||
# #region ScenarioExecution.BrowserProvider.Session.Registry.Acquire [C:5] [TYPE Function] [SEMANTICS provider,browser,session,acquire,replay]
|
||||
# @ingroup ScenarioExecution
|
||||
# @BRIEF Resolve the step plan: reuse the live session, replay the persisted checkpoint, or open fresh.
|
||||
# @PRE The caller holds run_guard(run_id); lease_id is the current step's capacity lease
|
||||
# (provenance only — the session outlives the per-step lease).
|
||||
# @POST A reuse plan carries the live session; a replay/fresh plan opens on the loop inside
|
||||
# execute_prepared; browser history without a checkpoint or a checkpoint pinned to another
|
||||
# dashboard raises BrowserCheckpointMissing before any I/O.
|
||||
def prepare_step(self, *, run_id: str, lease_id: str, dashboard_id: int) -> _PreparedStep:
|
||||
self.reap_idle()
|
||||
with self._registry_lock:
|
||||
session = self._sessions.get(run_id)
|
||||
if session is not None:
|
||||
session.lease_id = lease_id
|
||||
logger.reason(
|
||||
"Reusing the run-scoped browser session", src=_SRC,
|
||||
payload={"run_id": run_id, "dashboard_id": session.dashboard_id, "replayed": session.replayed},
|
||||
)
|
||||
return _PreparedStep(kind="reuse", run_id=run_id, lease_id=lease_id, dashboard_id=dashboard_id, session=session)
|
||||
kind, state = self._checkpoint_loader(run_id)
|
||||
if kind == "missing":
|
||||
logger.explore(
|
||||
"Browser recovery requires a declared checkpoint", src=_SRC,
|
||||
claim="POST: replay or typed missing", error_code="BROWSER_CHECKPOINT_MISSING",
|
||||
payload={"run_id": run_id},
|
||||
)
|
||||
raise BrowserCheckpointMissing("BROWSER_CHECKPOINT_MISSING")
|
||||
if kind == "checkpoint":
|
||||
filters = state.get("native_filter_state") if isinstance(state, dict) else None
|
||||
if (
|
||||
not isinstance(state, dict)
|
||||
or state.get("dashboard_id") != dashboard_id
|
||||
or any(not isinstance(entry, dict) or not isinstance(entry.get("replay_input"), dict) for entry in filters or [])
|
||||
):
|
||||
logger.explore(
|
||||
"Persisted checkpoint cannot authorize replay for this binding", src=_SRC,
|
||||
claim="POST: replay or typed missing", error_code="BROWSER_CHECKPOINT_MISSING",
|
||||
payload={"run_id": run_id, "dashboard_id": dashboard_id},
|
||||
)
|
||||
raise BrowserCheckpointMissing("BROWSER_CHECKPOINT_MISSING")
|
||||
return _PreparedStep(kind="replay", run_id=run_id, lease_id=lease_id, dashboard_id=dashboard_id, replay_state=state)
|
||||
return _PreparedStep(kind="fresh", run_id=run_id, lease_id=lease_id, dashboard_id=dashboard_id)
|
||||
# #endregion ScenarioExecution.BrowserProvider.Session.Registry.Acquire
|
||||
|
||||
# #region ScenarioExecution.BrowserProvider.Session.Registry.Execute [C:5] [TYPE Function] [SEMANTICS provider,browser,session,execute,checkpoint]
|
||||
# @ingroup ScenarioExecution
|
||||
# @BRIEF Coroutine (provider loop): open/replay on first use, execute the action, fold the checkpoint.
|
||||
# @PRE prepared comes from prepare_step under the run guard; runs entirely on the provider loop.
|
||||
# @POST Returns (transport outcome, stamped checkpoint copy); the session is registered before
|
||||
# replay so a mid-step cancellation is always closable by the leak-guard; a failed replay
|
||||
# discards and closes the fresh context before re-raising.
|
||||
async def execute_prepared(
|
||||
self,
|
||||
prepared: _PreparedStep,
|
||||
action: str,
|
||||
*,
|
||||
action_input: dict[str, Any],
|
||||
timeout_seconds: float,
|
||||
) -> tuple[Any, dict[str, Any]]:
|
||||
session = prepared.session or await self._open_session(prepared, timeout_seconds=timeout_seconds)
|
||||
outcome = await self._transport.execute_in_session(
|
||||
session.transport_session,
|
||||
action,
|
||||
action_input=action_input,
|
||||
timeout_seconds=timeout_seconds,
|
||||
)
|
||||
session.checkpoint = fold_checkpoint_state(
|
||||
session.checkpoint,
|
||||
dashboard_id=session.dashboard_id,
|
||||
action=action,
|
||||
action_input=action_input,
|
||||
details=outcome.details if isinstance(outcome.details, dict) else {},
|
||||
)
|
||||
session.last_used_at = self._clock()
|
||||
logger.reflect(
|
||||
"Browser step folded into the run checkpoint", src=_SRC,
|
||||
payload={"run_id": session.run_id, "action": action, "checkpoint_seq": session.checkpoint.get("checkpoint_seq"), "filters": len(session.checkpoint.get("native_filter_state") or [])},
|
||||
)
|
||||
return outcome, copy.deepcopy(session.checkpoint)
|
||||
# #endregion ScenarioExecution.BrowserProvider.Session.Registry.Execute
|
||||
|
||||
# #region ScenarioExecution.BrowserProvider.Session.Registry.Open [C:4] [TYPE Function] [SEMANTICS provider,browser,session,open,replay]
|
||||
# @ingroup ScenarioExecution
|
||||
# @BRIEF Open the transport context, register early (leak-guard visibility), then replay state.
|
||||
async def _open_session(self, prepared: _PreparedStep, *, timeout_seconds: float) -> BrowserSession:
|
||||
handle = await self._transport.open_session(prepared.dashboard_id, timeout_seconds=timeout_seconds)
|
||||
now = self._clock()
|
||||
session = BrowserSession(
|
||||
run_id=prepared.run_id,
|
||||
lease_id=prepared.lease_id,
|
||||
dashboard_id=prepared.dashboard_id,
|
||||
transport_session=handle,
|
||||
checkpoint=copy.deepcopy(prepared.replay_state) if prepared.replay_state is not None else _initial_checkpoint(prepared.dashboard_id),
|
||||
created_at=now,
|
||||
last_used_at=now,
|
||||
)
|
||||
with self._registry_lock:
|
||||
self._sessions[prepared.run_id] = session
|
||||
logger.reflect(
|
||||
"Run-scoped browser session opened", src=_SRC,
|
||||
payload={"run_id": prepared.run_id, "dashboard_id": prepared.dashboard_id, "kind": prepared.kind},
|
||||
)
|
||||
if prepared.replay_state is not None:
|
||||
try:
|
||||
for entry in prepared.replay_state.get("native_filter_state") or []:
|
||||
await self._transport.execute_in_session(
|
||||
handle,
|
||||
"apply_native_filter",
|
||||
action_input=copy.deepcopy(entry["replay_input"]),
|
||||
timeout_seconds=timeout_seconds,
|
||||
)
|
||||
session.replayed = True
|
||||
logger.reflect(
|
||||
"Browser checkpoint replayed into a fresh context", src=_SRC,
|
||||
payload={"run_id": prepared.run_id, "filters": len(prepared.replay_state.get("native_filter_state") or [])},
|
||||
)
|
||||
except Exception as exc:
|
||||
with self._registry_lock:
|
||||
self._sessions.pop(prepared.run_id, None)
|
||||
logger.explore(
|
||||
"Checkpoint replay failed; fresh context abandoned", src=_SRC,
|
||||
error_code="BROWSER_CHECKPOINT_REPLAY_FAILED",
|
||||
payload={"run_id": prepared.run_id}, error=repr(exc),
|
||||
)
|
||||
await self._close_handle(handle)
|
||||
raise
|
||||
return session
|
||||
# #endregion ScenarioExecution.BrowserProvider.Session.Registry.Open
|
||||
|
||||
# #region ScenarioExecution.BrowserProvider.Session.Registry.Close [C:4] [TYPE Function] [SEMANTICS provider,browser,session,close,leak-guard]
|
||||
# @ingroup ScenarioExecution
|
||||
# @BRIEF Leak-guard close paths: explicit terminal close, dead-context close, idle reap, close_all.
|
||||
# @POST close() is idempotent and never raises: the registry entry is dropped first and the
|
||||
# transport close is best-effort — a dead loop abandons the loop-bound context by design.
|
||||
def close(self, run_id: str, *, reason: str = "terminal") -> bool:
|
||||
with self._registry_lock:
|
||||
session = self._sessions.pop(run_id, None)
|
||||
if session is None:
|
||||
return False
|
||||
try:
|
||||
self._event_loop.submit(lambda: self._close_handle(session.transport_session), timeout=_CLOSE_TIMEOUT_SECONDS)
|
||||
except Exception as exc:
|
||||
logger.explore(
|
||||
"Browser session close could not reach the loop; context dies with the loop", src=_SRC,
|
||||
error_code="BROWSER_SESSION_CLOSE_FAILED",
|
||||
payload={"run_id": run_id, "reason": reason}, error=repr(exc),
|
||||
)
|
||||
return True
|
||||
logger.reflect("Run-scoped browser session closed", src=_SRC, payload={"run_id": run_id, "reason": reason})
|
||||
return True
|
||||
|
||||
def reap_idle(self) -> int:
|
||||
now = self._clock()
|
||||
with self._registry_lock:
|
||||
stale = [
|
||||
run_id
|
||||
for run_id, session in self._sessions.items()
|
||||
if now - session.last_used_at > self._idle_timeout
|
||||
]
|
||||
closed = 0
|
||||
for run_id in stale:
|
||||
if self.close(run_id, reason="idle_timeout"):
|
||||
closed += 1
|
||||
return closed
|
||||
|
||||
def close_all(self, *, reason: str = "shutdown") -> int:
|
||||
with self._registry_lock:
|
||||
run_ids = list(self._sessions)
|
||||
closed = 0
|
||||
for run_id in run_ids:
|
||||
if self.close(run_id, reason=reason):
|
||||
closed += 1
|
||||
return closed
|
||||
|
||||
async def _close_handle(self, handle: Any) -> None:
|
||||
try:
|
||||
await self._transport.close_session(handle)
|
||||
except Exception as exc:
|
||||
logger.explore(
|
||||
"Transport session close raised; handle abandoned", src=_SRC,
|
||||
error_code="BROWSER_SESSION_CLOSE_FAILED", error=repr(exc),
|
||||
)
|
||||
# #endregion ScenarioExecution.BrowserProvider.Session.Registry.Close
|
||||
# #endregion ScenarioExecution.BrowserProvider.Session.Registry
|
||||
@@ -515,7 +515,7 @@ async def test_transport_extract_table_missing(playwright_stub):
|
||||
|
||||
@pytest.mark.asyncio()
|
||||
async def test_transport_extract_table_oversize_raises_typed(playwright_stub, monkeypatch):
|
||||
import src.services.dashboard_testing.execution.providers.browser_readonly_actions as mod
|
||||
import src.services.dashboard_testing.execution.providers.browser_readonly_flows_nav as mod
|
||||
|
||||
monkeypatch.setattr(mod, "_MAX_EXTRACT_OUTPUT_BYTES", 8)
|
||||
eval_result = {"columns": ["A", "B"], "rows": [["1", "2"], ["3", "4"]]}
|
||||
@@ -588,7 +588,7 @@ async def test_transport_download_happy_returns_bytes(playwright_stub):
|
||||
|
||||
@pytest.mark.asyncio()
|
||||
async def test_transport_download_oversize_raises_typed(playwright_stub, monkeypatch):
|
||||
import src.services.dashboard_testing.execution.providers.browser_readonly_actions as mod
|
||||
import src.services.dashboard_testing.execution.providers.browser_readonly_flows_interact as mod
|
||||
|
||||
monkeypatch.setattr(mod, "_MAX_DOWNLOAD_BYTES", 4)
|
||||
trigger = FakeLocator(visible=True)
|
||||
|
||||
@@ -70,7 +70,7 @@
|
||||
- [x] T041 [P] Implement backend/src/services/dashboard_testing/visual_baseline.py: VisualComparisonPolicy (exact + perceptual), layout fingerprint computation, stale_visual_baseline detection.
|
||||
- [x] T042 [P] Wire visual baseline loading into BaselineEngine.Catalog.Load; extend catalog YAML to support visual entries.
|
||||
- [x] T043 Implement BaselineEngine.Visual.Compare: digest comparison + perceptual SSIM path; return stale_visual_baseline when layout fingerprint mismatches.
|
||||
- [ ] T044 [P] Implement BaselineEngine.Visual.Candidate: create draft visual candidate from reviewed screenshot artifact with mandatory human disposition.
|
||||
- [x] T044 [P] Implement BaselineEngine.Visual.Candidate: create draft visual candidate from reviewed screenshot artifact with mandatory human disposition. **Re-verified 2026-09-16:** реализация в `services/dashboard_testing/candidates.py` (ветка `kind == "visual"`: DraftArtifact resolution, ownership по `agent_run_id`, sha256-match, kind allowlist) + `candidate_helpers.py`; mandatory human disposition через approval gate (`test_gate_provenance_overrides_client_approval`, `test_denied_gate_cannot_consume`, `test_consume_replay_rejected`). Evidence: `pytest -q tests/services/dashboard_testing/test_visual_lifecycle_comprehensive.py tests/services/dashboard_testing/test_immutability.py tests/services/dashboard_testing/test_visual_executor_async_resolve.py tests/services/dashboard_testing/test_visual_executor_catalog.py` → **61 passed** на HEAD `4746af2f`.
|
||||
- [x] T045 Add visual baseline approval flow reusing 036 gate; verify approval writes visual entry atomically alongside metric entries.
|
||||
- [x] T046 Write visual baseline golden fixtures under specs/037-superset-baseline-engine/fixtures/visual/.
|
||||
- [x] T047 Audit: visual baselines never use metric policies; metric baselines never use visual policies; cross-kind comparison returns inconclusive.
|
||||
|
||||
@@ -33,7 +33,9 @@ The `trigger` field on VerificationRun gains the additive value `dataset_updated
|
||||
|
||||
## Production acceptance traceability — 2026-09-08
|
||||
|
||||
Historical rows above identify prior tests/code only; removed agent UI paths are retired. The following audited gates are **implemented=false / OPEN**, independent of local suite totals.
|
||||
Historical rows above identify prior tests/code only; removed agent UI paths are retired. Gates below are re-evaluated against executable evidence; a row is CLOSED only with the cited command/output or live-canary artifact.
|
||||
|
||||
**T044 (visual candidate) CLOSED 2026-09-16:** `BaselineEngine.Visual.Candidate` is implemented in `services/dashboard_testing/candidates.py` (visual branch: DraftArtifact resolution, `agent_run_id` ownership, sha256 match, kind allowlist) with mandatory human disposition through the approval gate. Evidence: `pytest -q tests/services/dashboard_testing/test_visual_lifecycle_comprehensive.py tests/services/dashboard_testing/test_immutability.py tests/services/dashboard_testing/test_visual_executor_async_resolve.py tests/services/dashboard_testing/test_visual_executor_catalog.py` → **61 passed**; gate tests `test_gate_provenance_overrides_client_approval`, `test_denied_gate_cannot_consume`, `test_consume_replay_rejected`. T082–T084 below remain OPEN.
|
||||
|
||||
| Requirement | Domain contract / DTO | Task | Falsifiable acceptance | State |
|
||||
|---|---|---|---|---|
|
||||
|
||||
@@ -1,9 +1,11 @@
|
||||
# 044 Scenario Execution — Session State
|
||||
|
||||
> **Updated**: 2026-08-21 17:12 +03:00
|
||||
> **Updated**: 2026-09-17 16:10 +03:00
|
||||
> **Purpose**: Durable handoff for the current implementation/review session. This is a
|
||||
> decision and verification ledger; `tasks.md` and `traceability.md` remain the canonical
|
||||
> feature backlog and requirement matrix.
|
||||
> **Newest section**: [Session Closure Update (2026-09-16/17)](#session-closure-update-2026-091617) —
|
||||
> supersedes the 2026-08-21 readiness snapshot below, which is retained as history.
|
||||
|
||||
## Current Objective
|
||||
|
||||
@@ -17,6 +19,10 @@ Production readiness is weighted toward real external execution because local fa
|
||||
cannot prove that a configured Browser/Superset/Screenshot/DB path performs authorized I/O and
|
||||
produces durable evidence.
|
||||
|
||||
> **This 2026-08-21 table is historical** (pre-provider, pre-live-canary). The current factual
|
||||
> status is in [Session Closure Update (2026-09-16/17)](#session-closure-update-2026-091617); the
|
||||
> numeric re-score belongs to the release review — do not quote the 54.05 total as current.
|
||||
|
||||
| Category | Weight | Score | Weighted contribution | Evidence / limiting factor |
|
||||
|---|---:|---:|---:|---|
|
||||
| Real live composition and external evidence | 50% | 25/100 | 12.5 | Exact injected 037 Superset binding is proven; Browser/Screenshot providers remain unavailable; no real deployment run |
|
||||
@@ -448,3 +454,44 @@ These are mandatory capability gates, not permanent exclusions. A missing or unp
|
||||
production GO. The only policy exclusions are: HumanCheckpoint-containing revisions cannot be automated and
|
||||
PROD mutation is prohibited. Unknown external effects cannot retry before reconciliation. This decision
|
||||
overrides earlier preview/shadow wording in this historical session ledger.
|
||||
|
||||
## Session Closure Update (2026-09-16/17)
|
||||
|
||||
> Supersedes the 2026-08-21 readiness snapshot. Verified against HEAD `4746af2f` + the decomposition
|
||||
> and canary work of 2026-09-16/17. No score is asserted here: the components changed qualitatively
|
||||
> (live providers, canaries, PostgreSQL, semantic index) and the numeric re-score belongs to the
|
||||
> release review. `tasks.md` and `traceability.md` carry the row-level states.
|
||||
|
||||
### Closed with executable evidence
|
||||
|
||||
| Item | Evidence |
|
||||
|---|---|
|
||||
| **T042b** BrowserProvider PREPROD canaries | 6/6 GREEN under the Wave-C run-scoped session architecture on the owner-authorized stand (`SS_STAND_STAGE=PREPROD`, dashboard 11): read-only actions 3/3, forced timeout typed `BROWSER_ACTION_TIMEOUT` + lease release, mutation `row_edit`+restore verified, reconciliation sweep resolved an injected stale receipt via live observation, safe-checkpoint reconstruction `reconstruction_replay=true`, scheduler soak 3/3 with `attempts==1`. Evidence `evidence/browser-provider/readonly-canary-20260916T*.json` + PNG. Harness defect found and fixed (dead-loop DI singleton → `asyncio.run` soak window + `stop()` via `asyncio.to_thread`). |
|
||||
| **T022** full scoped audit + PostgreSQL + semantic rebuild | Scoped 044 suite **1590 passed**; `validate_static.py` passed; ruff + compileall clean; full backend **11269 passed / 252 skipped / 1 xpassed / 0 failed**; frontend **3359 passed** + lint 0 errors + build green; fresh PostgreSQL 16 chain `0001→0025_drop_legacy_validation` + `alembic check` **no drift**; Axiom rebuild 10890 contracts / 5383 edges; **0 `module_too_long`** in `execution/providers` after the decomposition gate. |
|
||||
| **T044** artifact-content live ACL/status/header canary | New harness `prototype/artifact_content_canary.py` drives the REAL FastAPI app (no dependency overrides) with real PostgreSQL `canary_044`, real JWTs and REAL live-captured soak bytes (104746 B, sha `22b4b8c1…`): **11/11 GREEN** (anon 401, viewer GET byte-exact + header coherence, viewer HEAD parity + empty body, no-VIEW 403, unknown/foreign indistinguishable 404, expired 410, corrupt 409, bad-MIME 409, oversized 413, missing 409, Range 416). Evidence `evidence/artifact-content/artifact-content-canary-20260917T085438Z.json`. |
|
||||
| **T045** fault-injection canary | New harness `prototype/fault_injection_canary.py` on real PostgreSQL: **16/16 GREEN** — typed bounded loop refusals (`PROVIDER_LOOP_NOT_RUNNING`, `PROVIDER_SUBMIT_DEADLINE`), caller-deadline coroutine cancellation, shutdown drain without hang, **crashed lease quarantines capacity until `reconcile_expired_leases`**, `reconciliation_required`/unknown effect resolved by reconcile, **late response history-only** + terminal reconcile refused, cancel opens drain window then terminalizes with 0 running steps / 0 unexpired worker leases and its capacity lease reconciled by TTL. Evidence `evidence/fault-injection/fault-injection-canary-20260917T105541Z.json`. |
|
||||
| **INV_7 decomposition gate (INV)** | `browser_readonly_actions.py` 534→160, `browser_session.py` 501→58, `browser.py` 407→398, five new sibling modules, frozen contract IDs + import surface, per-phase gates **1590 passed** ×3 and full suite re-confirmed **11269 / 0 failed**. Log: `plans/provider-decomposition-gate.md` (EXECUTED). |
|
||||
| **Skill drift (INV)** | `.agents/skills/semantics-python`/`semantics-svelte` referenced the removed `ss_tools.shared.cot_logger`, the non-existent `addToast()` and function-call `$t("key")`; corrected to `src.core.logger` facade, `notify()` from `$lib/toasts.svelte.ts`, dictionary-proxy i18n, and re-synced via `scripts/sync-skills.sh`. Three malformed multi-target `@RELATION` lines in `backend/src/core/cot_logger.py` split (0 unresolved in file). |
|
||||
| **T014/T016/T042 (050), T044 (037)** | Re-verified and flipped with fresh evidence (MCP slice 84 passed, catalog 61/2.3.0, parity 3 passed, 037 visual slice 61 passed). 050 quickstart stale OPEN rows corrected. |
|
||||
|
||||
### Open (blocking GO)
|
||||
|
||||
| Item | Precise gap |
|
||||
|---|---|
|
||||
| **T046** `[~]` | Walker-level binding: `policy_inputs_from_outcome` reads `evaluation_input` only from the `agent_evaluation` step, so an `assertion compare_to_baseline` step resolves to `EVALUATION_UNAVAILABLE` under mandatory mode. Executable proof: compare-only+mandatory → `inconclusive [EVALUATION_UNAVAILABLE]`; +bound evaluation → `passed [BASELINE_AND_SEMANTIC_PASS]`. Two closure options recorded in `tasks.md` T046; closure = binding + live v4 rerun ending in terminal `passed`. |
|
||||
| **T043** `[ ]` | Graph-level deterministic comparison PASS (same live rerun closes it); live pin-from-Gitea + 050 T045 publish already proven. |
|
||||
| **T042** `[~]` | Live Superset query + explicit `RESULT_TOO_LARGE` bound. |
|
||||
| **SC-003** `[~]` | Required 100-trial production cancel evidence (drain/terminalize already proven by T045). |
|
||||
| **Frontend boundary** | External-MCP-only negative UI/route/network acceptance (050 CHK003/CHK004; 043 T015/T023). |
|
||||
|
||||
### Cross-cutting state
|
||||
|
||||
- Dependency-weighted aggregate from 2026-08-24 (≈67/100) is **stale as a score**: provider, capacity,
|
||||
PostgreSQL and live-composition components moved on 2026-09-16/17. Re-score at release review.
|
||||
- `plans/provider-decomposition-gate.md` is the binding precedent for INV_7 splits (frozen IDs,
|
||||
frozen import surface, per-phase verifier, rollback rule).
|
||||
- Canary harnesses are fail-closed prototypes (exit non-zero on any failed vector, credentials never
|
||||
written to evidence): `browser_readonly_canary.py` (`SS_CANARY_MODE`, `SS_CANARY_STORAGE_ROOT`),
|
||||
`artifact_content_canary.py`, `fault_injection_canary.py`. The artifact-content canary additionally
|
||||
needs `STORAGE_ROOT_PATH` from an approved root (`/app/storage`, the project, or
|
||||
`../ss-tools-storage`).
|
||||
|
||||
@@ -20,34 +20,56 @@
|
||||
|
||||
## Provider Contract
|
||||
|
||||
- [ ] CHK005 Every provider receives context containing run, step, attempt, descriptor/binding/principal
|
||||
fingerprints, deadline, idempotency key, capacity lease and trace id.
|
||||
- [ ] CHK006 Every provider creates an append-only operation receipt before I/O and returns effect state,
|
||||
retry disposition, operation id and stable reason code.
|
||||
- [ ] CHK007 Every evidence ref has an immutable receipt binding owner tuple, provider/version, descriptor,
|
||||
content type, byte length and verified non-zero SHA-256.
|
||||
- [ ] CHK008 No provider performs I/O before a durable shared CapacityLease; lease expiry/release is
|
||||
idempotent and observable.
|
||||
- [ ] CHK009 Cancellation acknowledges `stopped|completed|unknown`; unknown effect blocks retry and PASS
|
||||
until reconciliation or terminal non-pass closure.
|
||||
- [ ] CHK010 Duplicate invocation, late response, malformed result, cleanup failure and reconciliation
|
||||
are covered for every enabled provider.
|
||||
- [ ] CHK011 Each enabled provider exposes separate liveness, readiness and dependency health with
|
||||
secret-free diagnostics and a registration/capability fingerprint.
|
||||
- [x] CHK005 Every provider receives context containing run, step, attempt, descriptor/binding/principal
|
||||
fingerprints, deadline, idempotency key, capacity lease and trace id. Evidence 2026-09-17:
|
||||
`test_provider_contract.py` + `test_provider_preflight.py` + `test_provider_runtime.py` green
|
||||
(60 passed across the five provider suites).
|
||||
- [x] CHK006 Every provider creates an append-only operation receipt before I/O and returns effect state,
|
||||
retry disposition, operation id and stable reason code. Evidence 2026-09-17:
|
||||
`test_provider_operations.py` (14) — open/complete/reconcile/cancel, typed duplicate, immutable terminal
|
||||
status; T045 canary corroborates on real PostgreSQL.
|
||||
- [x] CHK007 Every evidence ref has an immutable receipt binding owner tuple, provider/version, descriptor,
|
||||
content type, byte length and verified non-zero SHA-256. Evidence 2026-09-17: T044 artifact-content
|
||||
canary (owner ACL, digest/length verified before bytes) `evidence/artifact-content/artifact-content-canary-20260917T085438Z.json`.
|
||||
- [x] CHK008 No provider performs I/O before a durable shared CapacityLease; lease expiry/release is
|
||||
idempotent and observable. Evidence 2026-09-17: `test_provider_capacity.py` + T045 fault-injection canary
|
||||
(typed `*_CAPACITY_UNAVAILABLE`, quarantine of a crashed lease, idempotent TTL reconciliation).
|
||||
- [x] CHK009 Cancellation acknowledges `stopped|completed|unknown`; unknown effect blocks retry and PASS
|
||||
until reconciliation or terminal non-pass closure. Evidence 2026-09-17: `test_provider_operations.py`
|
||||
cancel/unknown→`reconciliation_required`; T045 canary (unknown effect → receipt `reconciliation_required`,
|
||||
resolved by reconcile; late response cannot win).
|
||||
- [x] CHK010 Duplicate invocation, late response, malformed result, cleanup failure and reconciliation
|
||||
are covered for every enabled provider. Evidence 2026-09-17: `test_provider_operations.py` (typed
|
||||
duplicate, late-response history-only, reconciliation refused on terminal); T045 canary vectors.
|
||||
- [x] CHK011 Each enabled provider exposes separate liveness, readiness and dependency health with
|
||||
secret-free diagnostics and a registration/capability fingerprint. Evidence 2026-09-17:
|
||||
`test_provider_preflight.py` (redacted payload, identity/version/capabilities, degraded-not-fatal,
|
||||
unready admits zero operations); SC-010 closed.
|
||||
|
||||
## Provider-Specific Obligations
|
||||
|
||||
- [~] CHK012 Browser provider contract specifies isolated context, server auth binding, safe checkpoint,
|
||||
- [x] CHK012 Browser provider contract specifies isolated context, server auth binding, safe checkpoint,
|
||||
cleanup, mutation reconciliation, per-action risk classification, 120s/30s/3-page/25MiB/10MiB
|
||||
limits and PREPROD canaries; runtime provider and canary evidence remain open.
|
||||
- [ ] CHK013 Superset/SQL provider proves exact 037 model, database, principal/RLS and raw-response digest/ref;
|
||||
caller SQL and metadata cannot pass.
|
||||
limits and PREPROD canaries. **CLOSED 2026-09-16**: T034/T040/T041 implemented; T042b PREPROD
|
||||
canaries 6/6 GREEN under the Wave-C run-scoped session architecture
|
||||
(`evidence/browser-provider/readonly-canary-20260916T*.json`); Wave-C modules decomposed to satisfy
|
||||
INV_7 (`plans/provider-decomposition-gate.md`).
|
||||
- [~] CHK013 Superset/SQL provider proves exact 037 model, database, principal/RLS and raw-response digest/ref;
|
||||
caller SQL and metadata cannot pass. Partial: offline `test_provider_superset.py` + `test_live_execution_binding.py`
|
||||
green; T042 remains `[~]` for the live Superset query and explicit `RESULT_TOO_LARGE` bound.
|
||||
- [ ] CHK014 XLSX provider accepts only server-owned verified artifacts and enforces archive/cell/formula limits.
|
||||
- [ ] CHK015 Assertion/Transform providers are deterministic, bounded and network/code-free.
|
||||
- [ ] CHK016 Screenshot provider atomically commits durable owned evidence and cleans temporary resources.
|
||||
- [~] CHK015 Assertion/Transform providers are deterministic, bounded and network/code-free. Partial:
|
||||
`compare_to_baseline` deterministically passes in the live graph (canary v4); T046 owns the remaining
|
||||
terminal-PASS binding.
|
||||
- [x] CHK016 Screenshot provider atomically commits durable owned evidence and cleans temporary resources.
|
||||
**CLOSED 2026-09-17**: T044 canary serves the live browser-captured durable bytes through the protected
|
||||
route (11/11); `test_provider_screenshot.py` (14) green.
|
||||
- [ ] CHK017 Report/Artifact providers use immutable manifests, templates, producer receipts and digest verification.
|
||||
- [ ] CHK018 AgentEvaluation uses immutable 038 spec, pinned provider/model/prompt, bounded budget and
|
||||
evidence allowlist; DecisionPolicy alone produces StepOutcome.
|
||||
- [x] CHK018 AgentEvaluation uses immutable 038 spec, pinned provider/model/prompt, bounded budget and
|
||||
evidence allowlist; DecisionPolicy alone produces StepOutcome. **CLOSED with residual 2026-09-17**:
|
||||
live LLM evaluation persisted an immutable `AgentEvaluation` with the pin stamped from the published
|
||||
envelope (canary v2/v4, `docs/reports/agentic-runtime-live-canary-v4-baseline-pin-2026-09-11.md`);
|
||||
residual is the graph-level terminal PASS tracked by CHK035/T046.
|
||||
|
||||
## Human and Lifecycle
|
||||
|
||||
@@ -63,28 +85,38 @@
|
||||
|
||||
- [x] CHK024 Server-owned environment policy creates PROD approval before dispatch and ignores client PROD flags.
|
||||
- [~] CHK025 API schemas require immutable run identity, status enums, step outcomes and execution provenance.
|
||||
- [ ] CHK026 Startup registers enabled providers only after dependency/readiness checks; unready providers
|
||||
admit zero new operations.
|
||||
- [ ] CHK027 Real PostgreSQL `alembic check` and `alembic upgrade head` pass on the deployment database.
|
||||
- [x] CHK026 Startup registers enabled providers only after dependency/readiness checks; unready providers
|
||||
admit zero new operations. **CLOSED 2026-09-17**: T040 registration + readiness preflight;
|
||||
`test_provider_preflight.py` (`test_ready_snapshot_reports_registered_capabilities`,
|
||||
`test_unready_provider_admits_zero_new_operations`, `test_degraded_readiness_blocks_browser_capability_admission`).
|
||||
- [x] CHK027 Real PostgreSQL `alembic check` and `alembic upgrade head` pass on the deployment database.
|
||||
**CLOSED 2026-09-17**: fresh PostgreSQL 16 canary DB `alembic_044` — chain `0001 → 0025_drop_legacy_validation`
|
||||
applied, `alembic check` → "No new upgrade operations detected", single head.
|
||||
- [x] CHK028 Full available 044 backend profile passes: registry/API suite, provider/lifecycle profile,
|
||||
scoped Ruff/compile and prototype validation.
|
||||
scoped Ruff/compile and prototype validation. Evidence 2026-09-17: scoped 044 suite **1590 passed**,
|
||||
`validate_static.py` passed, ruff + compileall clean, full backend **11269 passed / 0 failed**.
|
||||
|
||||
## Release Gates
|
||||
|
||||
- [ ] CHK029 SC-001..011 each has reproducible command output or deployment evidence linked in traceability.
|
||||
- [~] CHK029 SC-001..011 each has reproducible command output or deployment evidence linked in traceability.
|
||||
Partial 2026-09-17: SC-002/004/006/008/009/010/011 CLOSED with commands/live canaries; SC-003 `[~]`
|
||||
(100-trial evidence open); SC-007 `[~]` (T042 Superset live query); SC-001 `[~]`.
|
||||
- [ ] CHK030 No unresolved P0/P1 row remains in 044 or its execution-critical dependencies 036, 037, 038,
|
||||
041, 042, 046 and 047.
|
||||
- [ ] CHK031 Browser and Screenshot providers perform authorized live I/O and produce owned durable evidence
|
||||
in a real deployment run.
|
||||
041, 042, 046 and 047. OPEN: 044 T043/T046 (graph-level terminal PASS), T042 (Superset live query).
|
||||
- [x] CHK031 Browser and Screenshot providers perform authorized live I/O and produce owned durable evidence
|
||||
in a real deployment run. **CLOSED 2026-09-16**: T042b PREPROD canaries **6/6 GREEN** on the
|
||||
owner-authorized stand with retained operation/evidence receipts
|
||||
(`evidence/browser-provider/readonly-canary-20260916T*.json` + PNG); T044 serves that durable evidence
|
||||
through the protected route (2026-09-17, 11/11).
|
||||
|
||||
## Production readiness checklist — 2026-09-08
|
||||
|
||||
SCEX-FR-028: [Production baseline-backed evaluation](../contracts/production-chain.md). Historical [x] marks do not close this new production gate; removed frontend components are not current evidence. All rows below implemented=false / OPEN.
|
||||
SCEX-FR-028: [Production baseline-backed evaluation](../contracts/production-chain.md). Historical [x] marks do not close this new production gate; removed frontend components are not current evidence. Rows are closed only by executable evidence (row states re-evaluated 2026-09-17; see [traceability](../traceability.md) and [SESSION_STATE](../SESSION_STATE.md)).
|
||||
|
||||
- [ ] CHK032 Baseline set/version/catalog/release/commit/IDs/digests affect idempotency; moving catalog after admission cannot change plan/result; legacy unpinned result is ineligible. Evidence: [T043](../tasks.md), [traceability](../traceability.md).
|
||||
- [ ] CHK033 GET/HEAD prove same ACL/status/headers; MIME/digest/length checked before bytes, cross-owner hidden, expired410, corrupt409, traversal/range/oversize rejected. Evidence: [T044](../tasks.md), [traceability](../traceability.md).
|
||||
- [ ] CHK034 Startup/readiness/start-loop and shutdown/drain/cancel/reconcile survive fault injection; unknown effect quarantines capacity; late response cannot win. Evidence: [T045](../tasks.md), [traceability](../traceability.md).
|
||||
- [ ] CHK035 End-to-end real browser→capture→durable artifact→deterministic comparison→optional immutable evaluation→policy→result; all required evidence present before PASS. Evidence: [T046](../tasks.md), [traceability](../traceability.md).
|
||||
- [~] CHK032 Baseline set/version/catalog/release/commit/IDs/digests affect idempotency; moving catalog after admission cannot change plan/result; legacy unpinned result is ineligible. Evidence: [T043](../tasks.md), [traceability](../traceability.md). Partial: request-hash pinning, fail-closed resolver and live pin-from-Gitea are PROVEN (`docs/reports/agentic-runtime-live-canary-v4-baseline-pin-2026-09-11.md`; 050 T045 publish CLOSED); residual is the graph-level deterministic comparison PASS (closes with the T046 live rerun).
|
||||
- [x] CHK033 GET/HEAD prove same ACL/status/headers; MIME/digest/length checked before bytes, cross-owner hidden, expired410, corrupt409, traversal/range/oversize rejected. Evidence: [T044](../tasks.md), [traceability](../traceability.md). **CLOSED 2026-09-17**: live HTTP canary `prototype/artifact_content_canary.py` (real app/PG/JWT + live soak bytes) **11/11 GREEN** — `evidence/artifact-content/artifact-content-canary-20260917T085438Z.json`; offline matrix `tests/api/test_scenario_artifact_content_api.py` 11/11.
|
||||
- [x] CHK034 Startup/readiness/start-loop and shutdown/drain/cancel/reconcile survive fault injection; unknown effect quarantines capacity; late response cannot win. Evidence: [T045](../tasks.md), [traceability](../traceability.md). **CLOSED 2026-09-17**: fault-injection canary `prototype/fault_injection_canary.py` (real provider loop/capacity/receipt CAS/cancel lifecycle on real PostgreSQL) **16/16 GREEN** — `evidence/fault-injection/fault-injection-canary-20260917T105541Z.json`.
|
||||
- [~] CHK035 End-to-end real browser→capture→durable artifact→deterministic comparison→optional immutable evaluation→policy→result; all required evidence present before PASS. Evidence: [T046](../tasks.md), [traceability](../traceability.md). Partial 2026-09-17: the chain up to policy is PROVEN live (canary v2/v4; browser canaries 6/6); the graph-level terminal PASS is blocked by the walker-level evaluation→comparison binding (`policy_inputs_from_outcome` reads `evaluation_input` only from the `agent_evaluation` step), diagnosed with executable proof; closure = binding + live v4 rerun.
|
||||
- [ ] CHK036 Negative product UI test: no agent chat/prompt/assistant editing/proposal-generation/typical-operation-to-agent/workspace/start/handoff controls or agent invocation routes/requests; manual CRUD/editor/human review/read-only results remain usable.
|
||||
|
||||
Schema/static success alone is not runtime completion. Optional approved performance baseline is outside scope.
|
||||
|
||||
@@ -2,7 +2,7 @@ openapi: 3.1.0
|
||||
info:
|
||||
title: ScenarioArtifact protected content
|
||||
version: 1.0.0
|
||||
description: Normative implemented=false. Server verifies ownership and complete MIME/digest/length before headers.
|
||||
description: Normative, implemented and live-verified (2026-09-17, T044 11/11 canary). Server verifies ownership and complete MIME/digest/length before headers.
|
||||
security: [{bearerAuth: []}]
|
||||
paths:
|
||||
/api/scenario-runs/{run_id}/artifacts/{artifact_id}/content:
|
||||
|
||||
@@ -3,7 +3,7 @@
|
||||
@RATIONALE Reuse existing ScreenshotService, LLMClient and 037 comparators through bounded ScenarioRun adapters; their older ValidationRecord/DraftArtifact owners cannot become run truth by aliasing IDs.
|
||||
@REJECTED Calling DashboardValidationPlugin._execute_path_a as the ScenarioRun orchestrator, relabeling model status as StepOutcome, or resolving current baseline after launch.
|
||||
|
||||
Status: normative target, implemented=false; current screenshot adapter exists but ss-prod bindings=0, loop=not_started and LLM providers=0. Passing fake-byte tests are historical seam evidence only.
|
||||
Status (2026-09-17): live-proven with one open acceptance row. Live evidence: browser PREPROD canaries 6/6 (T042b), protected artifact GET/HEAD 11/11 (T044), fault-injection lifecycle 16/16 (T045), live LLM evaluation with a persisted AgentEvaluation and the baseline pin stamped from the published Gitea envelope (canaries v2/v4), release audit incl. the real PostgreSQL chain and semantic rebuild (T022); server-owned live bindings and providers are configured and exercised (bindings registered, provider loop running, LLM provider present). Open: the graph-level terminal PASS (T046 — the walker-level evaluation→comparison binding is diagnosed with executable proof) and the live Superset query path with an explicit `RESULT_TOO_LARGE` bound (T042).
|
||||
|
||||
Canonical visual chain: pinned 042 ScenarioRevision → 044 baseline preflight/RunnerPlan → browser open/tab/filter/wait → ScreenshotService capture → durable ScenarioArtifact/EvidenceReceipt → 037 exact/SSIM comparison using pinned baseline → immutable ComparisonResult → optional declared AgentEvaluation → pinned DecisionPolicy → StepOutcome → result/SSE/045/047. Metric branch uses the same pin and 037 Superset-native normalization/comparison before optional evaluation. Artifact registration is atomic provider output, not an agent-authored digest. Missing semantic provider cannot disable mandatory evaluation; deterministic-only runs need no LLM.
|
||||
|
||||
|
||||
@@ -1,7 +1,7 @@
|
||||
{
|
||||
"$schema": "https://json-schema.org/draft/2020-12/schema",
|
||||
"title": "ScenarioResult evidence projection v1",
|
||||
"description": "044 authority consumed by045/047/050; implemented=false.",
|
||||
"description": "044 authority consumed by045/047/050; live-verified for the browser/capture/artifact/evaluation chain (2026-09-17); graph-level terminal PASS pending (T046 binding).",
|
||||
"type": "object",
|
||||
"additionalProperties": false,
|
||||
"required": [
|
||||
|
||||
@@ -259,7 +259,9 @@ Artifact content uses [protected GET/HEAD](contracts/artifact-content.openapi.ya
|
||||
Metadata carries owner/run/step/attempt/operation, SHA-256, actual MIME, byte length,
|
||||
retention/expiry and availability. Required baseline bytes have retention holds.
|
||||
[Production chain](contracts/production-chain.md) owns provider startup/shutdown/cancel/reconcile.
|
||||
New records and lifecycle acceptance remain implemented=false. Frontend only renders read-only
|
||||
New records and lifecycle acceptance are live-verified for the provider/artifact/evaluation chain
|
||||
(2026-09-17: T042b/T044/T045 canaries + T022 audit); the graph-level terminal PASS remains pending
|
||||
(T046 binding). Frontend only renders read-only
|
||||
evaluation evidence; no provider/prompt/retry-agent controls. Approved performance baseline is outside scope.
|
||||
|
||||
#endregion ScenarioExecution.DataModel
|
||||
|
||||
@@ -0,0 +1,131 @@
|
||||
{
|
||||
"started_at": "2026-09-17T08:51:50.547441+00:00",
|
||||
"database_url_scheme": "postgresql+psycopg2",
|
||||
"storage_root": "/home/busya/dev/ss-tools/specs/044-dashboard-scenario-execution/evidence/browser-provider/storage-soak-20260917",
|
||||
"run_id": "soak-canary-c071fb98",
|
||||
"artifact_id": "a722857f-6c9d-4e27-8da4-50e0dca892c8",
|
||||
"artifact_sha256": "22b4b8c17e644d13c14c8120c1a79b9654ab5519386c87547fde8c6bd35b36a0",
|
||||
"artifact_byte_length": 10485761,
|
||||
"results": [
|
||||
{
|
||||
"vector": "anonymous_401",
|
||||
"ok": true,
|
||||
"get_status": 401,
|
||||
"head_status": 401,
|
||||
"error_code": "AUTHENTICATION_REQUIRED",
|
||||
"head_body_bytes": 0
|
||||
},
|
||||
{
|
||||
"vector": "viewer_get_200_bytes",
|
||||
"ok": false,
|
||||
"status": 413,
|
||||
"bytes": 149,
|
||||
"sha256": "0be113cd8d9c2e4116f71a9a33f48c6da92f6f92309cf8fade7ffdeb65037188",
|
||||
"content_type": "application/json",
|
||||
"etag": null,
|
||||
"content_digest": null
|
||||
},
|
||||
{
|
||||
"vector": "viewer_head_200_header_parity",
|
||||
"ok": false,
|
||||
"status": 413,
|
||||
"body_bytes": 0,
|
||||
"mismatched_headers": [
|
||||
"content-type",
|
||||
"content-length"
|
||||
]
|
||||
},
|
||||
{
|
||||
"vector": "outsider_403",
|
||||
"ok": true,
|
||||
"get_status": 403,
|
||||
"head_status": 403,
|
||||
"error_code": "PERMISSION_DENIED"
|
||||
},
|
||||
{
|
||||
"vector": "hidden_404_not_found",
|
||||
"ok": true,
|
||||
"unknown_run": 404,
|
||||
"unknown_artifact": 404,
|
||||
"foreign": 404,
|
||||
"messages_equal": true
|
||||
},
|
||||
{
|
||||
"vector": "expired_410",
|
||||
"ok": true,
|
||||
"status": 410,
|
||||
"code": "ARTIFACT_EXPIRED",
|
||||
"head_status": 410,
|
||||
"envelope": [
|
||||
"code",
|
||||
"correlation_id",
|
||||
"message",
|
||||
"retryable"
|
||||
]
|
||||
},
|
||||
{
|
||||
"vector": "corrupt_sha_409",
|
||||
"ok": false,
|
||||
"status": 413,
|
||||
"code": "ARTIFACT_TOO_LARGE",
|
||||
"head_status": 413,
|
||||
"envelope": [
|
||||
"code",
|
||||
"correlation_id",
|
||||
"message",
|
||||
"retryable"
|
||||
]
|
||||
},
|
||||
{
|
||||
"vector": "declared_mime_409",
|
||||
"ok": true,
|
||||
"status": 409,
|
||||
"code": "ARTIFACT_INTEGRITY_FAILED",
|
||||
"head_status": 409,
|
||||
"envelope": [
|
||||
"code",
|
||||
"correlation_id",
|
||||
"message",
|
||||
"retryable"
|
||||
]
|
||||
},
|
||||
{
|
||||
"vector": "oversized_413",
|
||||
"ok": true,
|
||||
"status": 413,
|
||||
"code": "ARTIFACT_TOO_LARGE",
|
||||
"head_status": 413,
|
||||
"envelope": [
|
||||
"code",
|
||||
"correlation_id",
|
||||
"message",
|
||||
"retryable"
|
||||
]
|
||||
},
|
||||
{
|
||||
"vector": "missing_bytes_409",
|
||||
"ok": false,
|
||||
"status": 413,
|
||||
"code": "ARTIFACT_TOO_LARGE",
|
||||
"head_status": 413,
|
||||
"envelope": [
|
||||
"code",
|
||||
"correlation_id",
|
||||
"message",
|
||||
"retryable"
|
||||
]
|
||||
},
|
||||
{
|
||||
"vector": "range_416",
|
||||
"ok": true,
|
||||
"status": 416,
|
||||
"code": "RANGE_NOT_SUPPORTED"
|
||||
}
|
||||
],
|
||||
"failures": [
|
||||
"viewer_get_200_bytes",
|
||||
"viewer_head_200_header_parity",
|
||||
"corrupt_sha_409",
|
||||
"missing_bytes_409"
|
||||
]
|
||||
}
|
||||
@@ -0,0 +1,123 @@
|
||||
{
|
||||
"started_at": "2026-09-17T08:54:38.039182+00:00",
|
||||
"database_url_scheme": "postgresql+psycopg2",
|
||||
"storage_root": "/home/busya/dev/ss-tools/specs/044-dashboard-scenario-execution/evidence/browser-provider/storage-soak-20260917",
|
||||
"run_id": "soak-canary-c071fb98",
|
||||
"artifact_id": "fd5dc891-1f61-4db0-928f-baaf2565e822",
|
||||
"artifact_sha256": "22b4b8c17e644d13c14c8120c1a79b9654ab5519386c87547fde8c6bd35b36a0",
|
||||
"artifact_byte_length": 104746,
|
||||
"results": [
|
||||
{
|
||||
"vector": "anonymous_401",
|
||||
"ok": true,
|
||||
"get_status": 401,
|
||||
"head_status": 401,
|
||||
"error_code": "AUTHENTICATION_REQUIRED",
|
||||
"head_body_bytes": 0
|
||||
},
|
||||
{
|
||||
"vector": "viewer_get_200_bytes",
|
||||
"ok": true,
|
||||
"status": 200,
|
||||
"bytes": 104746,
|
||||
"sha256": "22b4b8c17e644d13c14c8120c1a79b9654ab5519386c87547fde8c6bd35b36a0",
|
||||
"content_type": "image/png",
|
||||
"etag": "\"22b4b8c17e644d13c14c8120c1a79b9654ab5519386c87547fde8c6bd35b36a0\"",
|
||||
"content_digest": "sha-256=:IrS4wX5kTRPBTIEgwaebllSrVRk4bIdUf96Ma9NbNqA=:"
|
||||
},
|
||||
{
|
||||
"vector": "viewer_head_200_header_parity",
|
||||
"ok": true,
|
||||
"status": 200,
|
||||
"body_bytes": 0,
|
||||
"mismatched_headers": []
|
||||
},
|
||||
{
|
||||
"vector": "outsider_403",
|
||||
"ok": true,
|
||||
"get_status": 403,
|
||||
"head_status": 403,
|
||||
"error_code": "PERMISSION_DENIED"
|
||||
},
|
||||
{
|
||||
"vector": "hidden_404_not_found",
|
||||
"ok": true,
|
||||
"unknown_run": 404,
|
||||
"unknown_artifact": 404,
|
||||
"foreign": 404,
|
||||
"messages_equal": true
|
||||
},
|
||||
{
|
||||
"vector": "expired_410",
|
||||
"ok": true,
|
||||
"status": 410,
|
||||
"code": "ARTIFACT_EXPIRED",
|
||||
"head_status": 410,
|
||||
"envelope": [
|
||||
"code",
|
||||
"correlation_id",
|
||||
"message",
|
||||
"retryable"
|
||||
]
|
||||
},
|
||||
{
|
||||
"vector": "corrupt_sha_409",
|
||||
"ok": true,
|
||||
"status": 409,
|
||||
"code": "ARTIFACT_INTEGRITY_FAILED",
|
||||
"head_status": 409,
|
||||
"envelope": [
|
||||
"code",
|
||||
"correlation_id",
|
||||
"message",
|
||||
"retryable"
|
||||
]
|
||||
},
|
||||
{
|
||||
"vector": "declared_mime_409",
|
||||
"ok": true,
|
||||
"status": 409,
|
||||
"code": "ARTIFACT_INTEGRITY_FAILED",
|
||||
"head_status": 409,
|
||||
"envelope": [
|
||||
"code",
|
||||
"correlation_id",
|
||||
"message",
|
||||
"retryable"
|
||||
]
|
||||
},
|
||||
{
|
||||
"vector": "oversized_413",
|
||||
"ok": true,
|
||||
"status": 413,
|
||||
"code": "ARTIFACT_TOO_LARGE",
|
||||
"head_status": 413,
|
||||
"envelope": [
|
||||
"code",
|
||||
"correlation_id",
|
||||
"message",
|
||||
"retryable"
|
||||
]
|
||||
},
|
||||
{
|
||||
"vector": "missing_bytes_409",
|
||||
"ok": true,
|
||||
"status": 409,
|
||||
"code": "ARTIFACT_MISSING",
|
||||
"head_status": 409,
|
||||
"envelope": [
|
||||
"code",
|
||||
"correlation_id",
|
||||
"message",
|
||||
"retryable"
|
||||
]
|
||||
},
|
||||
{
|
||||
"vector": "range_416",
|
||||
"ok": true,
|
||||
"status": 416,
|
||||
"code": "RANGE_NOT_SUPPORTED"
|
||||
}
|
||||
],
|
||||
"failures": []
|
||||
}
|
||||
|
After Width: | Height: | Size: 102 KiB |
|
After Width: | Height: | Size: 102 KiB |
|
After Width: | Height: | Size: 5.7 KiB |
|
After Width: | Height: | Size: 102 KiB |
@@ -0,0 +1,52 @@
|
||||
{
|
||||
"canary": "browser_readonly",
|
||||
"stand_url": "https://ss-prod.bebesh.ru",
|
||||
"stand_stage": "PREPROD",
|
||||
"dashboard_id": 11,
|
||||
"started_at": "2026-09-16T15:01:58.604412+00:00",
|
||||
"results": {
|
||||
"open_dashboard": {
|
||||
"status": "passed",
|
||||
"reason_code": "BROWSER_ACTION_EXECUTED",
|
||||
"checkpoints": [
|
||||
"dashboard_open"
|
||||
],
|
||||
"page_url": "https://ss-prod.bebesh.ru/superset/dashboard/11/?standalone=true&native_filters_key=StQY1l35b8Q",
|
||||
"sha256": "2fd5bed4d633badda7fb86646b1ae287d03484f8cdb7e2039aca6df7ba408586",
|
||||
"artifact_refs": [
|
||||
"draft:canary-11cee2a0-open_dashboard:2fd5bed4d633badda7fb86646b1ae287d03484f8cdb7e2039aca6df7ba408586"
|
||||
],
|
||||
"elapsed_seconds": 7.21
|
||||
},
|
||||
"wait_for_state": {
|
||||
"status": "passed",
|
||||
"reason_code": "BROWSER_ACTION_EXECUTED",
|
||||
"checkpoints": [
|
||||
"dashboard_open",
|
||||
"wait_for_state"
|
||||
],
|
||||
"page_url": "https://ss-prod.bebesh.ru/superset/dashboard/11/?standalone=true&native_filters_key=StQY1l35b8Q",
|
||||
"sha256": "8c5a7d4137ff664ee6a418117ea2b4bdf9eb9406eed133d0eef5152dbb1c8d0a",
|
||||
"artifact_refs": [
|
||||
"draft:canary-fb3be6b0-wait_for_state:8c5a7d4137ff664ee6a418117ea2b4bdf9eb9406eed133d0eef5152dbb1c8d0a"
|
||||
],
|
||||
"elapsed_seconds": 3.85
|
||||
},
|
||||
"refresh": {
|
||||
"status": "passed",
|
||||
"reason_code": "BROWSER_ACTION_EXECUTED",
|
||||
"checkpoints": [
|
||||
"dashboard_open",
|
||||
"refreshed"
|
||||
],
|
||||
"page_url": "https://ss-prod.bebesh.ru/superset/dashboard/11/?standalone=true&native_filters_key=StQY1l35b8Q",
|
||||
"sha256": "51c70f68993410e69b1bd7910b8d01c2b0a978087cc9414828aaa49f1da5cb07",
|
||||
"artifact_refs": [
|
||||
"draft:canary-0f8c8c76-refresh:51c70f68993410e69b1bd7910b8d01c2b0a978087cc9414828aaa49f1da5cb07"
|
||||
],
|
||||
"elapsed_seconds": 4.33
|
||||
}
|
||||
},
|
||||
"evidence_root": "/tmp/canary-storage-o3lf_45k",
|
||||
"failures": []
|
||||
}
|
||||
@@ -0,0 +1,17 @@
|
||||
{
|
||||
"canary": "browser_forced_timeout_cleanup",
|
||||
"stand_url": "https://ss-prod.bebesh.ru",
|
||||
"stand_stage": "PREPROD",
|
||||
"dashboard_id": 11,
|
||||
"started_at": "2026-09-16T15:04:07.330536+00:00",
|
||||
"results": {
|
||||
"open_dashboard_forced_timeout": {
|
||||
"status": "inconclusive",
|
||||
"reason_code": "BROWSER_ACTION_TIMEOUT",
|
||||
"artifact_refs": [],
|
||||
"elapsed_seconds": 3.07
|
||||
}
|
||||
},
|
||||
"evidence_root": "/tmp/canary-storage-xlg7tzq1",
|
||||
"failures": []
|
||||
}
|
||||
|
After Width: | Height: | Size: 102 KiB |
|
After Width: | Height: | Size: 102 KiB |
@@ -0,0 +1,64 @@
|
||||
{
|
||||
"canary": "browser_mutation_row_edit",
|
||||
"stand_url": "https://ss-prod.bebesh.ru",
|
||||
"stand_stage": "PREPROD",
|
||||
"dashboard_id": 11,
|
||||
"started_at": "2026-09-16T15:05:18.893558+00:00",
|
||||
"results": {
|
||||
"mutate": {
|
||||
"status": "passed",
|
||||
"reason_code": "BROWSER_ACTION_EXECUTED",
|
||||
"effect_state": "completed",
|
||||
"operation_id": "f8f0f2ee-976a-40ee-be3c-56f0501c2a9c",
|
||||
"post_rows": [
|
||||
{
|
||||
"global_sales": 82.75,
|
||||
"name": "Wii Sports",
|
||||
"platform": "Wii"
|
||||
}
|
||||
],
|
||||
"artifact_refs": [
|
||||
"draft:canary-d48f0a03-row_edit:42d8e43f130ca075439194bb1d9a6a273f86dce7fcc9940a50871c867f7a1458"
|
||||
],
|
||||
"elapsed_seconds": 4.65,
|
||||
"expected_value": 82.75,
|
||||
"verified_value": 82.74
|
||||
},
|
||||
"restore": {
|
||||
"status": "passed",
|
||||
"reason_code": "BROWSER_ACTION_EXECUTED",
|
||||
"effect_state": "completed",
|
||||
"operation_id": "da995860-579e-4ae9-8ac3-fd3ab0bb9cb4",
|
||||
"post_rows": [
|
||||
{
|
||||
"global_sales": 82.74,
|
||||
"name": "Wii Sports",
|
||||
"platform": "Wii"
|
||||
}
|
||||
],
|
||||
"artifact_refs": [
|
||||
"draft:canary-926f3e50-row_edit:61988a861e77e10d1865e2df4ed3d6abbc13429c1b54a006c9d400749e573bf8"
|
||||
],
|
||||
"elapsed_seconds": 4.34,
|
||||
"expected_value": 82.74,
|
||||
"verified_value": 82.74
|
||||
}
|
||||
},
|
||||
"receipts": [
|
||||
{
|
||||
"operation_id": "da995860-579e-4ae9-8ac3-fd3ab0bb9cb4",
|
||||
"status": "completed",
|
||||
"effect_state": "completed",
|
||||
"action": "row_edit"
|
||||
},
|
||||
{
|
||||
"operation_id": "f8f0f2ee-976a-40ee-be3c-56f0501c2a9c",
|
||||
"status": "completed",
|
||||
"effect_state": "completed",
|
||||
"action": "row_edit"
|
||||
}
|
||||
],
|
||||
"original_value": 82.74,
|
||||
"evidence_root": "/tmp/canary-storage-ue673xdj",
|
||||
"failures": []
|
||||
}
|
||||
@@ -0,0 +1,19 @@
|
||||
{
|
||||
"canary": "browser_reconciliation",
|
||||
"stand_url": "https://ss-prod.bebesh.ru",
|
||||
"stand_stage": "PREPROD",
|
||||
"started_at": "2026-09-16T15:07:54.481693+00:00",
|
||||
"live_value": 82.74,
|
||||
"sweep": {
|
||||
"checked": 1,
|
||||
"resolved": 1,
|
||||
"skipped_no_reconciler": 0,
|
||||
"skipped_invalid": 0
|
||||
},
|
||||
"receipt": {
|
||||
"status": "completed",
|
||||
"effect_state": "completed",
|
||||
"note": "target row state matches the recorded assignments"
|
||||
},
|
||||
"failures": []
|
||||
}
|
||||
@@ -0,0 +1,66 @@
|
||||
{
|
||||
"canary": "browser_safe_checkpoint_reconstruction",
|
||||
"stand_url": "https://ss-prod.bebesh.ru",
|
||||
"stand_stage": "PREPROD",
|
||||
"run_id": "recovery-canary-1f6c8149",
|
||||
"started_at": "2026-09-16T15:09:08.661476+00:00",
|
||||
"result": {
|
||||
"status": "passed",
|
||||
"elapsed_seconds": 4.45
|
||||
},
|
||||
"step": {
|
||||
"error_code": null,
|
||||
"step_outcome": {
|
||||
"logical_step_id": "canary-recovery-open",
|
||||
"status": "passed",
|
||||
"step_outcome": {
|
||||
"sha256": "c1ead0e50d0d04213fc2ef28cc35b0fdd8c7499f7692fbb19dd610ab6889e16d",
|
||||
"checkpoints": [
|
||||
"dashboard_open"
|
||||
],
|
||||
"page_url": "https://ss-prod.bebesh.ru/superset/dashboard/11/?standalone=true&native_filters_key=StQY1l35b8Q",
|
||||
"action": "open_dashboard",
|
||||
"effect_state": "none",
|
||||
"artifact_byte_lengths": {
|
||||
"draft:recovery-canary-1f6c8149:c1ead0e50d0d04213fc2ef28cc35b0fdd8c7499f7692fbb19dd610ab6889e16d": 104812
|
||||
},
|
||||
"artifact_content_types": {
|
||||
"draft:recovery-canary-1f6c8149:c1ead0e50d0d04213fc2ef28cc35b0fdd8c7499f7692fbb19dd610ab6889e16d": "image/png"
|
||||
},
|
||||
"title": "Sales Dashboard",
|
||||
"browser_checkpoint": {
|
||||
"dashboard_id": 11,
|
||||
"checkpoint_seq": 1,
|
||||
"native_filter_state": [],
|
||||
"active_tab": null,
|
||||
"wait_states": [],
|
||||
"table_filter_state": [],
|
||||
"filter_state_observed": null
|
||||
},
|
||||
"artifact_digests": {
|
||||
"draft:recovery-canary-1f6c8149:c1ead0e50d0d04213fc2ef28cc35b0fdd8c7499f7692fbb19dd610ab6889e16d": "c1ead0e50d0d04213fc2ef28cc35b0fdd8c7499f7692fbb19dd610ab6889e16d"
|
||||
},
|
||||
"tool": "browser",
|
||||
"reason_code": "BROWSER_ACTION_EXECUTED"
|
||||
},
|
||||
"output_refs": [],
|
||||
"artifact_refs": [
|
||||
"draft:recovery-canary-1f6c8149:c1ead0e50d0d04213fc2ef28cc35b0fdd8c7499f7692fbb19dd610ab6889e16d"
|
||||
],
|
||||
"error_code": null,
|
||||
"reconstruction_replay": true
|
||||
},
|
||||
"attempt": 2,
|
||||
"status": "passed",
|
||||
"reconstruction_replay": true,
|
||||
"checkpoints": null,
|
||||
"recovery_history_flags": [
|
||||
true
|
||||
],
|
||||
"artifact_refs": [
|
||||
"draft:recovery-canary-1f6c8149:c1ead0e50d0d04213fc2ef28cc35b0fdd8c7499f7692fbb19dd610ab6889e16d"
|
||||
]
|
||||
},
|
||||
"evidence_root": "/tmp/canary-storage-ee2dtbzm",
|
||||
"failures": []
|
||||
}
|
||||
@@ -0,0 +1,31 @@
|
||||
{
|
||||
"canary": "browser_scheduler_soak",
|
||||
"stand_url": "https://ss-prod.bebesh.ru",
|
||||
"stand_stage": "PREPROD",
|
||||
"soak_seconds": 75,
|
||||
"started_at": "2026-09-16T15:27:24.046710+00:00",
|
||||
"runs": [
|
||||
{
|
||||
"run_id": "soak-canary-038e19b1",
|
||||
"status": "passed",
|
||||
"error_code": null,
|
||||
"attempts": 1
|
||||
},
|
||||
{
|
||||
"run_id": "soak-canary-c57c2b65",
|
||||
"status": "passed",
|
||||
"error_code": null,
|
||||
"attempts": 1
|
||||
},
|
||||
{
|
||||
"run_id": "soak-canary-be680f35",
|
||||
"status": "passed",
|
||||
"error_code": null,
|
||||
"attempts": 1
|
||||
}
|
||||
],
|
||||
"active_leases": 0,
|
||||
"evidence_artifacts": 3,
|
||||
"evidence_root": "/tmp/canary-storage-vw_9cpfr",
|
||||
"failures": []
|
||||
}
|
||||
@@ -0,0 +1,31 @@
|
||||
{
|
||||
"canary": "browser_scheduler_soak",
|
||||
"stand_url": "https://ss-prod.bebesh.ru",
|
||||
"stand_stage": "PREPROD",
|
||||
"soak_seconds": 75,
|
||||
"started_at": "2026-09-16T15:50:13.599635+00:00",
|
||||
"runs": [
|
||||
{
|
||||
"run_id": "soak-canary-92bdc4ad",
|
||||
"status": "passed",
|
||||
"error_code": null,
|
||||
"attempts": 1
|
||||
},
|
||||
{
|
||||
"run_id": "soak-canary-7b03cac9",
|
||||
"status": "passed",
|
||||
"error_code": null,
|
||||
"attempts": 1
|
||||
},
|
||||
{
|
||||
"run_id": "soak-canary-00e15851",
|
||||
"status": "passed",
|
||||
"error_code": null,
|
||||
"attempts": 1
|
||||
}
|
||||
],
|
||||
"active_leases": 0,
|
||||
"evidence_artifacts": 3,
|
||||
"evidence_root": "/tmp/canary-storage-eulkpjy5",
|
||||
"failures": []
|
||||
}
|
||||
|
After Width: | Height: | Size: 102 KiB |
|
After Width: | Height: | Size: 5.7 KiB |
|
After Width: | Height: | Size: 102 KiB |
@@ -0,0 +1,52 @@
|
||||
{
|
||||
"canary": "browser_readonly",
|
||||
"stand_url": "https://ss-prod.bebesh.ru",
|
||||
"stand_stage": "PREPROD",
|
||||
"dashboard_id": 11,
|
||||
"started_at": "2026-09-17T08:33:02.234720+00:00",
|
||||
"results": {
|
||||
"open_dashboard": {
|
||||
"status": "passed",
|
||||
"reason_code": "BROWSER_ACTION_EXECUTED",
|
||||
"checkpoints": [
|
||||
"dashboard_open"
|
||||
],
|
||||
"page_url": "https://ss-prod.bebesh.ru/superset/dashboard/11/?standalone=true&native_filters_key=0IIn9czJmkI",
|
||||
"sha256": "f0fcd00e0b4b36fe1366d9f128ee6dc46a5f66576a3ed6de5e4d68d0d3cf5451",
|
||||
"artifact_refs": [
|
||||
"draft:canary-313971e9-open_dashboard:f0fcd00e0b4b36fe1366d9f128ee6dc46a5f66576a3ed6de5e4d68d0d3cf5451"
|
||||
],
|
||||
"elapsed_seconds": 4.91
|
||||
},
|
||||
"wait_for_state": {
|
||||
"status": "passed",
|
||||
"reason_code": "BROWSER_ACTION_EXECUTED",
|
||||
"checkpoints": [
|
||||
"dashboard_open",
|
||||
"wait_for_state"
|
||||
],
|
||||
"page_url": "https://ss-prod.bebesh.ru/superset/dashboard/11/?standalone=true&native_filters_key=0IIn9czJmkI",
|
||||
"sha256": "22b4b8c17e644d13c14c8120c1a79b9654ab5519386c87547fde8c6bd35b36a0",
|
||||
"artifact_refs": [
|
||||
"draft:canary-1aa3f8b7-wait_for_state:22b4b8c17e644d13c14c8120c1a79b9654ab5519386c87547fde8c6bd35b36a0"
|
||||
],
|
||||
"elapsed_seconds": 3.97
|
||||
},
|
||||
"refresh": {
|
||||
"status": "passed",
|
||||
"reason_code": "BROWSER_ACTION_EXECUTED",
|
||||
"checkpoints": [
|
||||
"dashboard_open",
|
||||
"refreshed"
|
||||
],
|
||||
"page_url": "https://ss-prod.bebesh.ru/superset/dashboard/11/?standalone=true&native_filters_key=0IIn9czJmkI",
|
||||
"sha256": "ee96c11e1e85579ce30cbf926306a5577f5727baa569c2fdfe350c10f275f3ab",
|
||||
"artifact_refs": [
|
||||
"draft:canary-9dd064d4-refresh:ee96c11e1e85579ce30cbf926306a5577f5727baa569c2fdfe350c10f275f3ab"
|
||||
],
|
||||
"elapsed_seconds": 4.34
|
||||
}
|
||||
},
|
||||
"evidence_root": "/home/busya/dev/ss-tools/specs/044-dashboard-scenario-execution/evidence/browser-provider/storage",
|
||||
"failures": []
|
||||
}
|
||||
@@ -0,0 +1,33 @@
|
||||
{
|
||||
"canary": "browser_scheduler_soak",
|
||||
"stand_url": "https://ss-prod.bebesh.ru",
|
||||
"stand_stage": "PREPROD",
|
||||
"soak_seconds": 75,
|
||||
"started_at": "2026-09-17T08:36:54.585964+00:00",
|
||||
"runs": [
|
||||
{
|
||||
"run_id": "soak-canary-8c4dc353",
|
||||
"status": "passed",
|
||||
"error_code": null,
|
||||
"attempts": 1
|
||||
},
|
||||
{
|
||||
"run_id": "soak-canary-b28c8331",
|
||||
"status": "passed",
|
||||
"error_code": null,
|
||||
"attempts": 1
|
||||
},
|
||||
{
|
||||
"run_id": "soak-canary-f7b77313",
|
||||
"status": "passed",
|
||||
"error_code": null,
|
||||
"attempts": 1
|
||||
}
|
||||
],
|
||||
"active_leases": 0,
|
||||
"evidence_artifacts": 6,
|
||||
"evidence_root": "/home/busya/dev/ss-tools/specs/044-dashboard-scenario-execution/evidence/browser-provider/storage",
|
||||
"failures": [
|
||||
"artifacts=6"
|
||||
]
|
||||
}
|
||||
@@ -0,0 +1,31 @@
|
||||
{
|
||||
"canary": "browser_scheduler_soak",
|
||||
"stand_url": "https://ss-prod.bebesh.ru",
|
||||
"stand_stage": "PREPROD",
|
||||
"soak_seconds": 75,
|
||||
"started_at": "2026-09-17T08:38:27.117031+00:00",
|
||||
"runs": [
|
||||
{
|
||||
"run_id": "soak-canary-a2dba6ac",
|
||||
"status": "passed",
|
||||
"error_code": null,
|
||||
"attempts": 1
|
||||
},
|
||||
{
|
||||
"run_id": "soak-canary-25d02e3f",
|
||||
"status": "passed",
|
||||
"error_code": null,
|
||||
"attempts": 1
|
||||
},
|
||||
{
|
||||
"run_id": "soak-canary-c071fb98",
|
||||
"status": "passed",
|
||||
"error_code": null,
|
||||
"attempts": 1
|
||||
}
|
||||
],
|
||||
"active_leases": 0,
|
||||
"evidence_artifacts": 3,
|
||||
"evidence_root": "/home/busya/dev/ss-tools/specs/044-dashboard-scenario-execution/evidence/browser-provider/storage-soak-20260917",
|
||||
"failures": []
|
||||
}
|
||||
|
After Width: | Height: | Size: 102 KiB |
|
After Width: | Height: | Size: 102 KiB |
|
After Width: | Height: | Size: 102 KiB |
@@ -0,0 +1,107 @@
|
||||
{
|
||||
"started_at": "2026-09-17T10:42:32.201980+00:00",
|
||||
"database_url_scheme": "postgresql+psycopg2",
|
||||
"results": [
|
||||
{
|
||||
"vector": "v1_loop_not_running",
|
||||
"ok": true,
|
||||
"elapsed_seconds": 0.0,
|
||||
"error": "PROVIDER_LOOP_NOT_RUNNING",
|
||||
"expected_error": "PROVIDER_LOOP_NOT_RUNNING"
|
||||
},
|
||||
{
|
||||
"vector": "v1_expired_deadline",
|
||||
"ok": true,
|
||||
"elapsed_seconds": 0.0,
|
||||
"error": "PROVIDER_SUBMIT_DEADLINE",
|
||||
"expected_error": "PROVIDER_SUBMIT_DEADLINE"
|
||||
},
|
||||
{
|
||||
"vector": "v2_deadline_cancels",
|
||||
"ok": true,
|
||||
"elapsed_seconds": 0.301,
|
||||
"error": "PROVIDER_SUBMIT_DEADLINE",
|
||||
"expected_error": "PROVIDER_SUBMIT_DEADLINE"
|
||||
},
|
||||
{
|
||||
"vector": "v2_coroutine_cancelled",
|
||||
"ok": true,
|
||||
"cancelled": true
|
||||
},
|
||||
{
|
||||
"vector": "v3_shutdown_drain_bounded",
|
||||
"ok": true,
|
||||
"elapsed_seconds": 1.503,
|
||||
"error": "PROVIDER_SUBMIT_DEADLINE",
|
||||
"expected_error": "PROVIDER_SUBMIT_DEADLINE"
|
||||
},
|
||||
{
|
||||
"vector": "v3_loop_stopped_after_drain",
|
||||
"ok": true,
|
||||
"is_running": false
|
||||
},
|
||||
{
|
||||
"vector": "v4_slot_quarantined_by_unreleased_lease",
|
||||
"ok": true,
|
||||
"limit": 2,
|
||||
"lease_id": "fb70a1aa-9f55-41c8-8c02-6dd021f6b487"
|
||||
},
|
||||
{
|
||||
"vector": "v4_reconcile_frees_quarantine",
|
||||
"ok": true,
|
||||
"expired": 1
|
||||
},
|
||||
{
|
||||
"vector": "v5_unknown_effect_requires_reconciliation",
|
||||
"ok": true,
|
||||
"status": "reconciliation_required",
|
||||
"effect_state": "unknown"
|
||||
},
|
||||
{
|
||||
"vector": "v5_reconcile_resolves_receipt",
|
||||
"ok": true,
|
||||
"status": "completed"
|
||||
},
|
||||
{
|
||||
"vector": "v5_late_response_cannot_win",
|
||||
"ok": true,
|
||||
"status": "completed",
|
||||
"last_history_kind": "late_response"
|
||||
},
|
||||
{
|
||||
"vector": "v5_reconcile_on_terminal_refused",
|
||||
"ok": true,
|
||||
"elapsed_seconds": 0.001,
|
||||
"error": "PROVIDER_OPERATION_TERMINAL",
|
||||
"expected_error": "PROVIDER_OPERATION_TERMINAL"
|
||||
},
|
||||
{
|
||||
"vector": "v6_cancel_opens_drain_window",
|
||||
"ok": true,
|
||||
"phase": "draining",
|
||||
"status": "cancel_requested"
|
||||
},
|
||||
{
|
||||
"vector": "v6_drain_expiry_terminalizes",
|
||||
"ok": false,
|
||||
"finalized": [
|
||||
"fault-canary-drain-9ebc4e3e"
|
||||
],
|
||||
"status": "cancelled",
|
||||
"running_steps": 0,
|
||||
"active_leases": 1
|
||||
},
|
||||
{
|
||||
"vector": "v6_immediate_cancel_terminalizes",
|
||||
"ok": false,
|
||||
"status": "cancelled",
|
||||
"phase": "terminal",
|
||||
"running_steps": 0,
|
||||
"active_leases": 1
|
||||
}
|
||||
],
|
||||
"failures": [
|
||||
"v6_drain_expiry_terminalizes",
|
||||
"v6_immediate_cancel_terminalizes"
|
||||
]
|
||||
}
|
||||
@@ -0,0 +1,110 @@
|
||||
{
|
||||
"started_at": "2026-09-17T10:55:41.867553+00:00",
|
||||
"database_url_scheme": "postgresql+psycopg2",
|
||||
"results": [
|
||||
{
|
||||
"vector": "v1_loop_not_running",
|
||||
"ok": true,
|
||||
"elapsed_seconds": 0.0,
|
||||
"error": "PROVIDER_LOOP_NOT_RUNNING",
|
||||
"expected_error": "PROVIDER_LOOP_NOT_RUNNING"
|
||||
},
|
||||
{
|
||||
"vector": "v1_expired_deadline",
|
||||
"ok": true,
|
||||
"elapsed_seconds": 0.0,
|
||||
"error": "PROVIDER_SUBMIT_DEADLINE",
|
||||
"expected_error": "PROVIDER_SUBMIT_DEADLINE"
|
||||
},
|
||||
{
|
||||
"vector": "v2_deadline_cancels",
|
||||
"ok": true,
|
||||
"elapsed_seconds": 0.301,
|
||||
"error": "PROVIDER_SUBMIT_DEADLINE",
|
||||
"expected_error": "PROVIDER_SUBMIT_DEADLINE"
|
||||
},
|
||||
{
|
||||
"vector": "v2_coroutine_cancelled",
|
||||
"ok": true,
|
||||
"cancelled": true
|
||||
},
|
||||
{
|
||||
"vector": "v3_shutdown_drain_bounded",
|
||||
"ok": true,
|
||||
"elapsed_seconds": 1.503,
|
||||
"error": "PROVIDER_SUBMIT_DEADLINE",
|
||||
"expected_error": "PROVIDER_SUBMIT_DEADLINE"
|
||||
},
|
||||
{
|
||||
"vector": "v3_loop_stopped_after_drain",
|
||||
"ok": true,
|
||||
"is_running": false
|
||||
},
|
||||
{
|
||||
"vector": "v4_slot_quarantined_by_unreleased_lease",
|
||||
"ok": true,
|
||||
"limit": 2,
|
||||
"lease_id": "30dba7b7-85d2-492a-bba5-0ca7bc6e9074"
|
||||
},
|
||||
{
|
||||
"vector": "v4_reconcile_frees_quarantine",
|
||||
"ok": true,
|
||||
"expired": 3
|
||||
},
|
||||
{
|
||||
"vector": "v5_unknown_effect_requires_reconciliation",
|
||||
"ok": true,
|
||||
"status": "reconciliation_required",
|
||||
"effect_state": "unknown"
|
||||
},
|
||||
{
|
||||
"vector": "v5_reconcile_resolves_receipt",
|
||||
"ok": true,
|
||||
"status": "completed"
|
||||
},
|
||||
{
|
||||
"vector": "v5_late_response_cannot_win",
|
||||
"ok": true,
|
||||
"status": "completed",
|
||||
"last_history_kind": "late_response"
|
||||
},
|
||||
{
|
||||
"vector": "v5_reconcile_on_terminal_refused",
|
||||
"ok": true,
|
||||
"elapsed_seconds": 0.001,
|
||||
"error": "PROVIDER_OPERATION_TERMINAL",
|
||||
"expected_error": "PROVIDER_OPERATION_TERMINAL"
|
||||
},
|
||||
{
|
||||
"vector": "v6_cancel_opens_drain_window",
|
||||
"ok": true,
|
||||
"phase": "draining",
|
||||
"status": "cancel_requested"
|
||||
},
|
||||
{
|
||||
"vector": "v6_drain_expiry_terminalizes",
|
||||
"ok": true,
|
||||
"finalized": [
|
||||
"fault-canary-drain-c6698ca0"
|
||||
],
|
||||
"status": "cancelled",
|
||||
"running_steps": 0,
|
||||
"active_worker_leases": 0
|
||||
},
|
||||
{
|
||||
"vector": "v6_cancelled_capacity_reconciled_not_dropped",
|
||||
"ok": true,
|
||||
"held_after_cancel": 1,
|
||||
"claimed_after_reconcile": 0
|
||||
},
|
||||
{
|
||||
"vector": "v6_immediate_cancel_terminalizes",
|
||||
"ok": true,
|
||||
"status": "cancelled",
|
||||
"phase": "terminal",
|
||||
"running_steps": 0,
|
||||
"active_worker_leases": 0
|
||||
}
|
||||
],
|
||||
"failures": []
|
||||
}
|
||||
@@ -99,7 +99,12 @@ oversized orchestrator.
|
||||
|
||||
## Production delivery plan — 2026-09-08 (SCEX-FR-028)
|
||||
|
||||
Status: specified, implemented=false; historical unit/prototype/transport results are not current production acceptance.
|
||||
Status (2026-09-17): implemented and live-proven for the provider/content/lifecycle chain; one row open.
|
||||
T044/T045 are CLOSED by live canaries (11/11 and 16/16), T042b by six PREPROD vectors, T022 by the
|
||||
release audit (real PostgreSQL chain + semantic rebuild). Step 3's remaining rows are T043/T046
|
||||
(graph-level terminal PASS — the walker-level evaluation→comparison binding is diagnosed with
|
||||
executable proof in T046) and T042's live Superset query; row states in [traceability](traceability.md),
|
||||
[checklist](checklists/requirements.md) and [SESSION_STATE](SESSION_STATE.md).
|
||||
|
||||
1. Pin [Production baseline-backed evaluation](contracts/production-chain.md) and [data model](data-model.md); write negative fixtures before runtime changes.
|
||||
2. Implement existing domain boundaries for: ScenarioRun and RunnerPlan require exact server BaselineSelectionPin and pin digest in canonical request identity. Result DTO is contracts/result-evidence.schema.json; artifact bytes use protected GET/HEAD. Append-only comparison/evaluation/outcome receipts bind run/step/attempt/operation; only CAS-selected winning attempt can determine outcome. Provider loop and cleanup/reconcile receipts gate capacity release.
|
||||
|
||||
@@ -0,0 +1,144 @@
|
||||
# Provider Decomposition Gate — 044 browser provider modules (INV_7)
|
||||
|
||||
> Статус: **EXECUTED 2026-09-17** (все три фазы; per-phase gates 1590 passed, нулевой behavior diff).
|
||||
> Binding-план по прецеденту
|
||||
> `specs/050-mcp-interface/plans/server-decomposition-gate.md` (round 3, EXECUTED).
|
||||
> Создан 2026-09-16 в ходе T022-аудита: Axiom audit `execution/` показал 8 structural warnings,
|
||||
> из них три INV_7-нарушителя после Wave-C (`8522a2ee`):
|
||||
>
|
||||
> | Модуль | LOC было | Лимит |
|
||||
> |---|---:|---:|
|
||||
> | `providers/browser_readonly_actions.py` | 534 | 400 |
|
||||
> | `providers/browser_session.py` | 501 | 400 |
|
||||
> | `providers/browser.py` | 407 | 400 (регрессия: 398 в round 4) |
|
||||
>
|
||||
> T022 не может быть `[x]` до исполнения этого плана (требование «zero unresolved P0/P1»).
|
||||
|
||||
## Constraints (frozen — нарушение = rollback фазы)
|
||||
|
||||
1. **Contract IDs заморожены.** Регионы `ScenarioExecution.BrowserProvider.*` переносятся
|
||||
verbatim; новые ID не изобретаются. Фасадные модули сохраняют исходный module-region ID.
|
||||
2. **Import surface заморожен.** `providers/browser.py`, `browser_session.py`,
|
||||
`browser_readonly_actions.py` остаются точками импорта и ре-экспортируют все перенесённые
|
||||
публичные имена; существующие импорты (`from ...browser_session import close_run_sessions` и т.п.)
|
||||
продолжают работать без правок у потребителей.
|
||||
3. **Zero behavior diff.** Только verbatim-переносы + фасадные ре-экспорты. Никаких рефакторингов
|
||||
логики в рамках split.
|
||||
4. **Monkeypatch-сеам следует за владеющим модулем** (урок round 3): если тест патчит
|
||||
`browser_readonly_actions.X`, после переноса патч переносится в новый владеющий модуль.
|
||||
5. **Helpers single-sited**: `_resolve_first_visible` / `Validate` / общие константы остаются
|
||||
в одном месте; новые модули импортируют их function-local (ациклический граф, прецедент
|
||||
`ops_tools.py` → `tools_review.py`).
|
||||
|
||||
## Phases
|
||||
|
||||
### Phase A — `browser_readonly_actions.py` 534 → facade + 2 flow-модуля (~270 each)
|
||||
|
||||
- `browser_readonly_flows_nav.py`: `ReadOnlyActions.NavigateTab`, `.InspectFilterState`,
|
||||
`.ApplyTableFilter`, `.ExtractTable` (verbatim).
|
||||
- `browser_readonly_flows_interact.py`: `ReadOnlyActions.ScrollTo`, `.InspectColumns`, `.Click`,
|
||||
`.SelectRows` + оставшиеся flow-регионы (verbatim).
|
||||
- `ReadOnlyActions.Validate` и `ReadOnlyActions.ResolveSelector` остаются single-sited в
|
||||
`browser_readonly_actions.py` (фасад < 400: module region + validate + selector + ре-экспорты
|
||||
flow-функций + dispatch-таблица).
|
||||
- Gate A: scoped suite (1590) + `test_browser_readonly_actions.py` (780 строк тестов) + anchors
|
||||
+ ruff + compileall.
|
||||
|
||||
### Phase B — `browser_session.py` 501 → facade + registry/checkpoint модули
|
||||
|
||||
- `browser_session_registry.py`: `Session.ManagerRegistry` + `Session.Registry` (C:5, 232-line
|
||||
contract → отдельный модуль < 400; contract остаётся одним регионом — перенос verbatim).
|
||||
- `browser_session_checkpoint.py`: `Session.Checkpoint` (fold_checkpoint_state) +
|
||||
`Session.ReplayLookup` (load_persisted_browser_checkpoint).
|
||||
- `browser_session.py` фасад: `Session.Errors`, `Session.Handle` (`BrowserSession`,
|
||||
`_PreparedStep`, `_initial_checkpoint`) + ре-экспорты.
|
||||
- Gate B: `test_browser_session.py` + scoped suite + anchors.
|
||||
|
||||
### Phase C — `browser.py` 407 → < 400 (минимальный)
|
||||
|
||||
- `BrowserProvider.Factory` (317 lines) — самый большой contract. Вынести evidence/digest
|
||||
helper-блоки и dispatch-таблицу действий в `browser_factory_helpers.py` (verbatim);
|
||||
в `browser.py` остаются admission merge + mutation-contract composition + skeleton фабрики.
|
||||
- Цель: `browser.py` < 400 без изменения публичного поведения `build_browser_provider`.
|
||||
- Gate C: `test_provider_browser.py` (24+) + `test_provider_contract.py` + scoped suite.
|
||||
|
||||
## Verification per phase (gate)
|
||||
|
||||
```bash
|
||||
cd backend && source .venv/bin/activate
|
||||
python -m pytest -q tests/services/dashboard_testing/ tests/api/test_scenario_runs_api.py \
|
||||
tests/api/test_scenario_automation_api.py tests/api/test_scenario_analytics_api.py \
|
||||
tests/api/test_scenario_artifact_content_api.py tests/api/test_scenario_run_center.py
|
||||
# ожидание: 1590 passed (нулевой behavior diff)
|
||||
python -m ruff check . && python -m compileall -q src
|
||||
# anchors: grep -c '#region ' vs '#endregion ' в каждом touched-файле
|
||||
```
|
||||
|
||||
Финальный gate: Axiom `audit_contracts` по `backend/src/services/dashboard_testing/execution`
|
||||
→ 0 module_too_long в providers/; полный backend suite (11269 baseline) без дельты.
|
||||
|
||||
## Risk register
|
||||
|
||||
| Риск | Митигция |
|
||||
|---|---|
|
||||
| Потерянный импорт при переносе | ruff F821 + compileall + gate-прогон после каждой фазы |
|
||||
| Сломанный monkeypatch-сеам в тестах | Constraint 4; при первом падении — правка теста в списке фазы (лог исполняется) |
|
||||
| Утечка behavior-изменения под видом «verbatim» | `git diff --stat` по фазе: только переносы; ревью diff'а перед gate |
|
||||
| Циклические импорты helpers | Constraint 5 (function-local import), прецедент ops_tools/tools_review |
|
||||
|
||||
## Rollback rule
|
||||
|
||||
Фаза не проходит gate → `git restore` touched-файлов фазы, фиксация причины в execution log,
|
||||
повтор только после изменения плана. Откат всей декомпозиции — revert до pre-plan HEAD.
|
||||
|
||||
## Execution log
|
||||
|
||||
### 2026-09-17 — Phase A (`browser_readonly_actions.py`)
|
||||
|
||||
- Слайсером (byte-exact segments) созданы: `browser_readonly_limits.py` (82) — limits/allowlists/
|
||||
selector-кандидаты + `ResolveSelector` (region перенесён verbatim; получил свой module-region
|
||||
`ReadOnlyActions.Limits`); `browser_readonly_flows_nav.py` (199) — NavigateTab, InspectFilterState,
|
||||
ApplyTableFilter, ExtractTable; `browser_readonly_flows_interact.py` (192) — ScrollTo,
|
||||
InspectColumns, Click, SelectRows, Download. Регионы flow-функций перенесены verbatim (ID сохранены).
|
||||
- Фасад `browser_readonly_actions.py` 534 → **160**: module region + `Validate` + `Dispatch`
|
||||
(verbatim) + re-export всех девяти flow-функций (`__all__`, frozen import surface).
|
||||
- Constraint 4: два monkeypatch-сеама в `test_browser_readonly_actions.py` перенесены на владеющие
|
||||
модули (`_MAX_EXTRACT_OUTPUT_BYTES` → `flows_nav`, `_MAX_DOWNLOAD_BYTES` → `flows_interact`).
|
||||
- Gate A: `test_browser_readonly_actions.py` 59 passed; scoped suite **1590 passed**; anchors
|
||||
2/2 · 5/5 · 6/6 · 3/3; ruff clean; compileall clean.
|
||||
|
||||
### 2026-09-17 — Phase B (`browser_session.py`)
|
||||
|
||||
- Созданы (verbatim-регионы, ID сохранены): `browser_session_handle.py` (64) — Handle
|
||||
(BrowserSession, _PreparedStep); `browser_session_checkpoint.py` (162) — CheckpointIO
|
||||
(Errors + новый region `InitialCheckpoint` для `_initial_checkpoint` + Checkpoint + ReplayLookup);
|
||||
`browser_session_managers.py` (57) — ManagerRegistry (weak registry + `close_run_sessions`);
|
||||
`browser_session_registry.py` (287) — Registry class region verbatim.
|
||||
- **Refinement vs план:** ManagerRegistry вынесен в отдельный модуль (а не оставлен в фасаде) —
|
||||
иначе `BrowserSessionManager.__init__` → `_register_manager` создавал бы цикл
|
||||
registry⇄facade. Registry импортирует `_register_manager` из managers; managers берёт
|
||||
`BrowserSessionManager` только под `TYPE_CHECKING` (аннотации ленивы). Найденные ruff-разрывы
|
||||
(`Callable` в checkpoint, `_register_manager` в registry) закрыты до gate.
|
||||
- Фасад `browser_session.py` 501 → **58**: module region + re-export
|
||||
(`BrowserCheckpointMissing`, `BrowserSession`, `BrowserSessionManager`, `_PreparedStep`,
|
||||
`close_run_sessions`, `fold_checkpoint_state`, `load_persisted_browser_checkpoint`).
|
||||
- Gate B: scoped suite **1590 passed**; anchors по всем 5 модулям сбалансированы; ruff/compileall clean.
|
||||
|
||||
### 2026-09-17 — Phase C (`browser.py`)
|
||||
|
||||
- Уточнение плана: dispatch-таблицы в `browser.py` нет (она в readonly-модулях) — вынесен чистый
|
||||
`mutation_receipt_summary(binding, descriptor, metadata)` в `browser_factory_helpers.py` (35);
|
||||
call-site 14 строк → 1. Typed exception→result лестница оставлена inline осознанно (@REJECTED).
|
||||
- `browser.py` 407 → **398** (< 400). Gate C: scoped suite **1590 passed**; anchors 2/2 · 2/2;
|
||||
ruff/compileall clean.
|
||||
- Инцидент в процессе: промежуточный `edit` разорвал перенос строки между `db.commit()` и
|
||||
`lease_id = ...`; обнаружен немедленной проверкой и исправлен в следующем edit — финальный файл
|
||||
верифицирован (ruff/compileall/1590).
|
||||
|
||||
### 2026-09-17 — Final gate
|
||||
|
||||
- Axiom `audit_contracts` по `execution/providers`: **0 module_too_long** (было 3).
|
||||
- Остались 2 advisory `contract_too_long` (Factory 305, ScreenshotProvider.Factory 200) — принятые
|
||||
advisory с rationale в контрактах; не INV_7-нарушения и не блокируют T022.
|
||||
- Полный backend suite: см. WORKSTATE-043-047.md (чекпоинт 2026-09-17) — ожидание нулевой дельты
|
||||
против 11269 baseline.
|
||||
@@ -0,0 +1,363 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Artifact-content live ACL/status/header canary (044 T044).
|
||||
|
||||
Drives the REAL FastAPI app (no dependency overrides) over a real PostgreSQL canary database
|
||||
and REAL live-captured browser evidence bytes, then checks the 044 artifact-content contract:
|
||||
GET/HEAD share ACL/status/headers, MIME/digest/length are verified before any byte, foreign and
|
||||
unknown children are indistinguishable 404s, owner-visible expired rows are 410, and corrupt /
|
||||
oversized / range requests fail with typed Error-Code values and an empty HEAD body.
|
||||
|
||||
Live bytes come from the browser provider canary run with
|
||||
``SS_CANARY_MODE=soak SS_CANARY_STORAGE_ROOT=<root>`` (that run both registers the
|
||||
``scenario_artifacts`` rows and writes the content-addressed bytes).
|
||||
|
||||
Required env (exported from ``backend/.env`` by the operator — never inline secrets here):
|
||||
DATABASE_URL PostgreSQL canary database (e.g. canary_044); never the dev database
|
||||
DRAFT_STORAGE_ROOT the storage root the live canary wrote to
|
||||
AUTH_SECRET_KEY same secret the app signs with
|
||||
ENCRYPTION_KEY same key the app decrypts config with
|
||||
|
||||
Optional env:
|
||||
SS_ARTIFACT_CANARY_OUT evidence output directory (default ../evidence/artifact-content)
|
||||
|
||||
Exit code is 0 only when every vector passes.
|
||||
"""
|
||||
|
||||
# #region ScenarioExecution.Prototype.ArtifactContentCanary [C:4] [TYPE Module] [SEMANTICS scenario,artifact,acl,canary,evidence]
|
||||
# @BRIEF Live HTTP canary for 044 artifact content: real app + real JWTs + real stored bytes.
|
||||
# @POST Writes an evidence JSON with per-vector status and exits non-zero on any failure.
|
||||
# @INVARIANT The script never points at the dev database: DATABASE_URL must contain "canary".
|
||||
# @SIDE_EFFECT Inserts canary principals and negative-fixture artifact rows into the canary DB.
|
||||
# @RATIONALE The offline suite already covers the matrix with mocks; this canary proves the same
|
||||
# contract end-to-end through real ASGI routing, real JWT auth and real stored bytes.
|
||||
# @REJECTED Overriding get_current_user/get_draft_storage (as the offline tests do) was rejected —
|
||||
# it would bypass the auth chain this canary exists to exercise.
|
||||
from __future__ import annotations
|
||||
|
||||
import asyncio
|
||||
import hashlib
|
||||
import json
|
||||
import os
|
||||
import sys
|
||||
import uuid
|
||||
from datetime import UTC, datetime
|
||||
from pathlib import Path
|
||||
|
||||
BACKEND = Path(__file__).resolve().parents[3] / "backend"
|
||||
sys.path.insert(0, str(BACKEND))
|
||||
|
||||
_REQUIRED_ENV = ("DATABASE_URL", "DRAFT_STORAGE_ROOT", "AUTH_SECRET_KEY", "ENCRYPTION_KEY")
|
||||
for _key in _REQUIRED_ENV:
|
||||
if not os.environ.get(_key):
|
||||
print(f"artifact-content canary: missing required env {_key}", file=sys.stderr)
|
||||
sys.exit(2)
|
||||
if "canary" not in os.environ["DATABASE_URL"]:
|
||||
print("artifact-content canary: refusing to run against a non-canary DATABASE_URL", file=sys.stderr)
|
||||
sys.exit(2)
|
||||
|
||||
_EVIDENCE_DIR = Path(
|
||||
os.environ.get("SS_ARTIFACT_CANARY_OUT")
|
||||
or Path(__file__).resolve().parent.parent / "evidence" / "artifact-content"
|
||||
)
|
||||
_VIEWER = "artifact-canary-viewer"
|
||||
_OUTSIDER = "artifact-canary-outsider"
|
||||
_RESOURCE, _ACTION = "scenario:result", "VIEW"
|
||||
_ERROR_HEADER = "Error-Code"
|
||||
|
||||
|
||||
def _sha256_ref(content_ref: str) -> str:
|
||||
return content_ref.rsplit(":", 1)[-1]
|
||||
|
||||
|
||||
def _storage_path(content_ref: str) -> Path:
|
||||
run_id, digest = content_ref.removeprefix("draft:").split(":", 1)
|
||||
return Path(os.environ["DRAFT_STORAGE_ROOT"]) / "drafts" / run_id / digest
|
||||
|
||||
|
||||
# #region ScenarioExecution.Prototype.ArtifactContentCanary.Seed [C:3] [TYPE Function] [SEMANTICS scenario,artifact,canary,seed]
|
||||
# @BRIEF Idempotently create the viewer (VIEW) and outsider (no permissions) principals.
|
||||
# @POST Both usernames exist exactly once with the expected role permissions.
|
||||
def seed_principals(db) -> None:
|
||||
from src.models.auth import Permission, Role, User
|
||||
|
||||
for username, grant_view in ((_VIEWER, True), (_OUTSIDER, False)):
|
||||
if db.query(User).filter(User.username == username).first() is not None:
|
||||
continue
|
||||
role = Role(id=f"role-{username}", name=f"role-{username}", is_admin=False)
|
||||
if grant_view:
|
||||
role.permissions = [Permission(id=f"perm-{username}", resource=_RESOURCE, action=_ACTION)]
|
||||
user = User(id=f"user-{username}", username=username, email=f"{username}@canary.local")
|
||||
user.roles = [role]
|
||||
db.add(user)
|
||||
db.commit()
|
||||
|
||||
|
||||
# #endregion ScenarioExecution.Prototype.ArtifactContentCanary.Seed
|
||||
|
||||
|
||||
# #region ScenarioExecution.Prototype.ArtifactContentCanary.Fixtures [C:4] [TYPE Function] [SEMANTICS scenario,artifact,canary,fixtures]
|
||||
# @BRIEF Pick the newest live row whose bytes exist, then clone negative fixtures from it.
|
||||
# @POST Returns (run_id, good_row, cases dict) with every clone persisted and committed.
|
||||
def seed_fixtures(db) -> tuple[str, object, dict[str, object]]:
|
||||
from src.models.scenario_artifact import ScenarioArtifact
|
||||
|
||||
# Re-runs must not treat a previous negative clone (e.g. the oversized fixture) as the live
|
||||
# good row: clones are named canary-* and are deleted before every run.
|
||||
db.query(ScenarioArtifact).filter(ScenarioArtifact.name.like("canary-%")).delete(synchronize_session=False)
|
||||
db.commit()
|
||||
|
||||
good = None
|
||||
candidates = (
|
||||
db.query(ScenarioArtifact)
|
||||
.filter(
|
||||
ScenarioArtifact.is_active.is_(True),
|
||||
ScenarioArtifact.kind == "evidence",
|
||||
~ScenarioArtifact.name.like("canary-%"),
|
||||
)
|
||||
.order_by(ScenarioArtifact.created_at.desc())
|
||||
.limit(200)
|
||||
.all()
|
||||
)
|
||||
for row in candidates:
|
||||
if _storage_path(row.content_ref).exists():
|
||||
good = row
|
||||
break
|
||||
if good is None:
|
||||
raise SystemExit("no live artifact row with stored bytes found — run the browser canary first")
|
||||
|
||||
other = next((r for r in candidates if r.owner_id != good.owner_id), None)
|
||||
if other is None:
|
||||
raise SystemExit("no second run with artifacts found — need a foreign-run fixture")
|
||||
|
||||
def clone(**overrides) -> ScenarioArtifact:
|
||||
row = ScenarioArtifact(
|
||||
id=str(uuid.uuid4()),
|
||||
owner_type="scenario_run",
|
||||
owner_id=good.owner_id,
|
||||
kind=good.kind,
|
||||
name=f"canary-{overrides.get('name', 'fixture')}",
|
||||
content_ref=good.content_ref,
|
||||
sha256=good.sha256,
|
||||
content_type=good.content_type,
|
||||
byte_length=good.byte_length,
|
||||
retention_class="standard",
|
||||
is_active=True,
|
||||
)
|
||||
for key, value in overrides.items():
|
||||
if key != "name":
|
||||
setattr(row, key, value)
|
||||
db.add(row)
|
||||
return row
|
||||
|
||||
fixtures = {
|
||||
"expired": clone(name="expired", is_active=False),
|
||||
"corrupt_sha": clone(name="corrupt", sha256="0" * 64),
|
||||
"bad_mime": clone(name="bad-mime", kind="screenshot", content_type="application/pdf"),
|
||||
"oversized": clone(name="oversized", byte_length=10_485_761),
|
||||
"missing_bytes": clone(name="missing", content_ref=f"draft:{good.owner_id}:{'a' * 64}", sha256="a" * 64),
|
||||
}
|
||||
db.commit()
|
||||
return good.owner_id, good, {**fixtures, "foreign_run_id": other.owner_id, "foreign_artifact_id": other.id}
|
||||
|
||||
|
||||
# #endregion ScenarioExecution.Prototype.ArtifactContentCanary.Fixtures
|
||||
|
||||
|
||||
# #region ScenarioExecution.Prototype.ArtifactContentCanary.Vectors [C:4] [TYPE Function] [SEMANTICS scenario,artifact,canary,vectors,acl]
|
||||
# @BRIEF Execute every T044 vector over the real ASGI app and collect typed results.
|
||||
# @POST Returns (results list, failures list); never raises on an assertion mismatch.
|
||||
async def run_vectors(run_id: str, good, fixtures: dict) -> tuple[list[dict], list[str]]:
|
||||
import httpx
|
||||
|
||||
from src.app import app
|
||||
from src.core.auth.jwt import create_access_token
|
||||
|
||||
viewer_token = create_access_token({"sub": _VIEWER})
|
||||
outsider_token = create_access_token({"sub": _OUTSIDER})
|
||||
base = "/api/scenario-runs"
|
||||
|
||||
def url(artifact_id: str, run: str | None = None) -> str:
|
||||
return f"{base}/{run or run_id}/artifacts/{artifact_id}/content"
|
||||
|
||||
results: list[dict] = []
|
||||
failures: list[str] = []
|
||||
|
||||
def record(name: str, ok: bool, detail: dict) -> None:
|
||||
results.append({"vector": name, "ok": ok, **detail})
|
||||
if not ok:
|
||||
failures.append(name)
|
||||
|
||||
def auth(token: str | None) -> dict:
|
||||
return {"Authorization": f"Bearer {token}"} if token else {}
|
||||
|
||||
transport = httpx.ASGITransport(app=app)
|
||||
async with httpx.AsyncClient(transport=transport, base_url="http://artifact-canary") as client:
|
||||
# 1. unauthenticated GET/HEAD -> 401 AUTHENTICATION_REQUIRED, HEAD body empty
|
||||
anon_get = await client.get(url(good.id))
|
||||
anon_head = await client.head(url(good.id))
|
||||
record(
|
||||
"anonymous_401",
|
||||
anon_get.status_code == 401
|
||||
and anon_get.headers.get(_ERROR_HEADER) == "AUTHENTICATION_REQUIRED"
|
||||
and anon_head.status_code == 401
|
||||
and anon_head.headers.get(_ERROR_HEADER) == "AUTHENTICATION_REQUIRED"
|
||||
and anon_head.content == b"",
|
||||
{"get_status": anon_get.status_code, "head_status": anon_head.status_code,
|
||||
"error_code": anon_get.headers.get(_ERROR_HEADER), "head_body_bytes": len(anon_head.content)},
|
||||
)
|
||||
|
||||
# 2. viewer GET -> 200 with complete verified bytes + coherent length/digest/etag
|
||||
get_good = await client.get(url(good.id), headers=auth(viewer_token))
|
||||
stored = _storage_path(good.content_ref).read_bytes()
|
||||
digest_hex = hashlib.sha256(stored).hexdigest()
|
||||
headers = get_good.headers
|
||||
expected_length = str(len(stored))
|
||||
record(
|
||||
"viewer_get_200_bytes",
|
||||
get_good.status_code == 200
|
||||
and get_good.content == stored
|
||||
and headers.get("content-length") == expected_length
|
||||
and headers.get("etag") == f'"{digest_hex}"'
|
||||
and headers.get("content-type") == good.content_type
|
||||
and headers.get("cache-control") == "private, no-store"
|
||||
and headers.get("x-content-type-options") == "nosniff"
|
||||
and headers.get("accept-ranges") == "none"
|
||||
and "attachment; filename=" in (headers.get("content-disposition") or ""),
|
||||
{
|
||||
"status": get_good.status_code,
|
||||
"bytes": len(get_good.content),
|
||||
"sha256": hashlib.sha256(get_good.content).hexdigest(),
|
||||
"content_type": headers.get("content-type"),
|
||||
"etag": headers.get("etag"),
|
||||
"content_digest": headers.get("content-digest"),
|
||||
},
|
||||
)
|
||||
|
||||
# 3. same identity HEAD -> identical status/headers, empty body
|
||||
head_good = await client.head(url(good.id), headers=auth(viewer_token))
|
||||
parity_keys = (
|
||||
"content-type", "content-length", "content-disposition", "etag",
|
||||
"content-digest", "cache-control", "x-content-type-options", "accept-ranges",
|
||||
)
|
||||
mismatched = [k for k in parity_keys if head_good.headers.get(k) != headers.get(k)]
|
||||
record(
|
||||
"viewer_head_200_header_parity",
|
||||
head_good.status_code == 200 and head_good.content == b"" and not mismatched,
|
||||
{"status": head_good.status_code, "body_bytes": len(head_good.content), "mismatched_headers": mismatched},
|
||||
)
|
||||
|
||||
# 4. known run without VIEW -> 403 PERMISSION_DENIED (GET + HEAD)
|
||||
outsider_get = await client.get(url(good.id), headers=auth(outsider_token))
|
||||
outsider_head = await client.head(url(good.id), headers=auth(outsider_token))
|
||||
record(
|
||||
"outsider_403",
|
||||
outsider_get.status_code == 403
|
||||
and outsider_get.headers.get(_ERROR_HEADER) == "PERMISSION_DENIED"
|
||||
and outsider_head.status_code == 403
|
||||
and outsider_head.content == b"",
|
||||
{"get_status": outsider_get.status_code, "head_status": outsider_head.status_code,
|
||||
"error_code": outsider_get.headers.get(_ERROR_HEADER)},
|
||||
)
|
||||
|
||||
# 5. unknown run / unknown artifact / foreign-owned artifact -> indistinguishable 404
|
||||
unknown_run = await client.get(url(good.id, run=str(uuid.uuid4())), headers=auth(viewer_token))
|
||||
unknown_artifact = await client.get(url(str(uuid.uuid4())), headers=auth(viewer_token))
|
||||
foreign = await client.get(
|
||||
url(fixtures["foreign_artifact_id"]), headers=auth(viewer_token)
|
||||
)
|
||||
record(
|
||||
"hidden_404_not_found",
|
||||
unknown_run.status_code == 404
|
||||
and unknown_artifact.status_code == 404
|
||||
and foreign.status_code == 404
|
||||
and unknown_run.headers.get(_ERROR_HEADER) == "NOT_FOUND"
|
||||
and unknown_artifact.headers.get(_ERROR_HEADER) == "NOT_FOUND"
|
||||
and foreign.headers.get(_ERROR_HEADER) == "NOT_FOUND"
|
||||
and unknown_run.json().get("message") == foreign.json().get("message"),
|
||||
{"unknown_run": unknown_run.status_code, "unknown_artifact": unknown_artifact.status_code,
|
||||
"foreign": foreign.status_code, "messages_equal":
|
||||
unknown_run.json().get("message") == foreign.json().get("message")},
|
||||
)
|
||||
|
||||
# 6-9. typed failures: expired 410, corrupt 409, bad MIME 409, oversized 413, missing 409
|
||||
typed = (
|
||||
("expired_410", fixtures["expired"], 410, "ARTIFACT_EXPIRED"),
|
||||
("corrupt_sha_409", fixtures["corrupt_sha"], 409, "ARTIFACT_INTEGRITY_FAILED"),
|
||||
("declared_mime_409", fixtures["bad_mime"], 409, "ARTIFACT_INTEGRITY_FAILED"),
|
||||
("oversized_413", fixtures["oversized"], 413, "ARTIFACT_TOO_LARGE"),
|
||||
("missing_bytes_409", fixtures["missing_bytes"], 409, "ARTIFACT_MISSING"),
|
||||
)
|
||||
for name, row, status, code in typed:
|
||||
get_resp = await client.get(url(row.id), headers=auth(viewer_token))
|
||||
head_resp = await client.head(url(row.id), headers=auth(viewer_token))
|
||||
body = {}
|
||||
try:
|
||||
body = get_resp.json()
|
||||
except ValueError:
|
||||
body = {}
|
||||
record(
|
||||
name,
|
||||
get_resp.status_code == status
|
||||
and get_resp.headers.get(_ERROR_HEADER) == code
|
||||
and head_resp.status_code == status
|
||||
and head_resp.headers.get(_ERROR_HEADER) == code
|
||||
and head_resp.content == b""
|
||||
and set(body) >= {"code", "message", "retryable", "correlation_id"}
|
||||
and body.get("code") == code,
|
||||
{"status": get_resp.status_code, "code": get_resp.headers.get(_ERROR_HEADER),
|
||||
"head_status": head_resp.status_code, "envelope": sorted(body)},
|
||||
)
|
||||
|
||||
# 10. Range requests are rejected after ACL, never streamed
|
||||
ranged = await client.get(url(good.id), headers={**auth(viewer_token), "Range": "bytes=0-9"})
|
||||
record(
|
||||
"range_416",
|
||||
ranged.status_code == 416
|
||||
and ranged.headers.get(_ERROR_HEADER) == "RANGE_NOT_SUPPORTED"
|
||||
and ranged.headers.get("accept-ranges") == "none",
|
||||
{"status": ranged.status_code, "code": ranged.headers.get(_ERROR_HEADER)},
|
||||
)
|
||||
|
||||
return results, failures
|
||||
|
||||
|
||||
# #endregion ScenarioExecution.Prototype.ArtifactContentCanary.Vectors
|
||||
|
||||
|
||||
# #region ScenarioExecution.Prototype.ArtifactContentCanary.Main [C:3] [TYPE Function] [SEMANTICS scenario,artifact,canary,main]
|
||||
# @BRIEF Wire the canary: seed, run vectors, write evidence, exit with the failure count.
|
||||
def main() -> int:
|
||||
from src.core.database import SessionLocal
|
||||
|
||||
with SessionLocal() as db:
|
||||
seed_principals(db)
|
||||
run_id, good, fixtures = seed_fixtures(db)
|
||||
results, failures = asyncio.run(run_vectors(run_id, good, fixtures))
|
||||
|
||||
report = {
|
||||
"started_at": datetime.now(UTC).isoformat(),
|
||||
"database_url_scheme": os.environ["DATABASE_URL"].split("://", 1)[0],
|
||||
"storage_root": os.environ["DRAFT_STORAGE_ROOT"],
|
||||
"run_id": run_id,
|
||||
"artifact_id": str(good.id),
|
||||
"artifact_sha256": good.sha256,
|
||||
"artifact_byte_length": good.byte_length,
|
||||
"results": results,
|
||||
"failures": failures,
|
||||
}
|
||||
_EVIDENCE_DIR.mkdir(parents=True, exist_ok=True)
|
||||
stamp = datetime.now(UTC).strftime("%Y%m%dT%H%M%SZ")
|
||||
out = _EVIDENCE_DIR / f"artifact-content-canary-{stamp}.json"
|
||||
out.write_text(json.dumps(report, indent=2))
|
||||
print(json.dumps({"results": len(results), "failures": failures, "evidence": str(out)}, indent=2))
|
||||
return 1 if failures else 0
|
||||
|
||||
|
||||
# #endregion ScenarioExecution.Prototype.ArtifactContentCanary.Main
|
||||
|
||||
|
||||
# #endregion ScenarioExecution.Prototype.ArtifactContentCanary
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
raise SystemExit(main())
|
||||
@@ -7,6 +7,10 @@
|
||||
# read-only modes (open_dashboard / wait_for_state / refresh / forced timeout) stay safe
|
||||
# against any classification. The mutation mode executes a bounded row_edit through the
|
||||
# provider path and restores the precondition values afterwards.
|
||||
# SS_CANARY_STORAGE_ROOT optionally pins a persistent evidence storage root (default a temp
|
||||
# mkdtemp root) — REQUIRED for any run whose bytes must survive (e.g. the T044 artifact-
|
||||
# content canary); use a DEDICATED root per live run because the soak artifact-count
|
||||
# assertion assumes a private root.
|
||||
# @POST Writes evidence JSON + PNG under specs/044-dashboard-scenario-execution/evidence/browser-provider/
|
||||
# and exits non-zero on any canary failure; no credentials are written to evidence.
|
||||
# @SIDE_EFFECT Launches headless Chromium sessions; writes a temp SQLite DB and evidence files.
|
||||
@@ -323,7 +327,9 @@ def run_recovery_canary(environment, dashboard_id: int, stand_url: str) -> dict:
|
||||
binding = build_binding(environment.id, dashboard_id, stand_url)
|
||||
binding = LiveExecutionBinding(**{**binding.snapshot(), "browser_safe_checkpoint_ref": "dashboard_open"})
|
||||
root = get_live_execution_composition_root()
|
||||
storage_root = Path(tempfile.mkdtemp(prefix="canary-storage-"))
|
||||
storage_root = Path(
|
||||
os.environ.get("SS_CANARY_STORAGE_ROOT") or tempfile.mkdtemp(prefix="canary-storage-")
|
||||
)
|
||||
storage = DraftStorage(storage_root=storage_root)
|
||||
service = ScreenshotService(environment)
|
||||
loop = ProviderEventLoop()
|
||||
@@ -446,7 +452,7 @@ def run_soak_canary(environment, dashboard_id: int, stand_url: str) -> dict:
|
||||
"""Soak: N queued runs advanced by the REAL scheduler against the live stand.
|
||||
Proves CAS single-dispatch, lease hygiene and graceful shutdown under repeated ticks."""
|
||||
from src.core.superset_client import SupersetClient
|
||||
from src.dependencies import get_config_manager, get_live_execution_composition_root, get_task_manager
|
||||
from src.dependencies import get_config_manager, get_live_execution_composition_root, get_scheduler_service, get_task_manager
|
||||
from src.models.scenario_run import ScenarioRun, ScenarioStepRun
|
||||
from src.models.scenario_worker import ScenarioStepLease
|
||||
from src.schemas.dashboard_testing import DashboardQueryModel
|
||||
@@ -458,12 +464,13 @@ def run_soak_canary(environment, dashboard_id: int, stand_url: str) -> dict:
|
||||
action_registry_fingerprint,
|
||||
resolve_action_descriptor,
|
||||
)
|
||||
from src.core.scheduler import SchedulerService
|
||||
|
||||
binding = build_binding(environment.id, dashboard_id, stand_url)
|
||||
binding = LiveExecutionBinding(**{**binding.snapshot(), "browser_safe_checkpoint_ref": "dashboard_open"})
|
||||
root = get_live_execution_composition_root()
|
||||
storage_root = Path(tempfile.mkdtemp(prefix="canary-storage-"))
|
||||
storage_root = Path(
|
||||
os.environ.get("SS_CANARY_STORAGE_ROOT") or tempfile.mkdtemp(prefix="canary-storage-")
|
||||
)
|
||||
storage = DraftStorage(storage_root=storage_root)
|
||||
service = ScreenshotService(environment)
|
||||
loop = ProviderEventLoop()
|
||||
@@ -537,11 +544,24 @@ def run_soak_canary(environment, dashboard_id: int, stand_url: str) -> dict:
|
||||
))
|
||||
db.commit()
|
||||
|
||||
scheduler = SchedulerService(get_task_manager(), get_config_manager())
|
||||
scheduler.start()
|
||||
soak_seconds = int(os.environ.get("SS_CANARY_SOAK_SECONDS", "75"))
|
||||
time.sleep(soak_seconds)
|
||||
scheduler.stop()
|
||||
|
||||
async def _soak_window() -> None:
|
||||
# Module-level job functions resolve the DI singleton SchedulerService and bridge
|
||||
# async jobs through AsyncJobRunner onto the loop captured at construction. A bare
|
||||
# script has no running loop, so the singleton must be created inside asyncio.run —
|
||||
# otherwise every async job blocks for its 300s safety cap and stop() waits on it
|
||||
# (observed 2026-09-16: maintenance_auto_end occupied its slot for the full cap).
|
||||
scheduler = get_scheduler_service()
|
||||
scheduler.start()
|
||||
try:
|
||||
await asyncio.sleep(soak_seconds)
|
||||
finally:
|
||||
# stop() waits for job worker threads whose async bridges need the loop free;
|
||||
# run it off the loop thread so in-flight coroutines can drain.
|
||||
await asyncio.to_thread(scheduler.stop)
|
||||
|
||||
asyncio.run(_soak_window())
|
||||
|
||||
with SessionLocal() as db:
|
||||
runs = db.query(ScenarioRun).filter(ScenarioRun.id.in_(run_ids)).all()
|
||||
@@ -589,7 +609,9 @@ def run_canary() -> dict:
|
||||
)
|
||||
Base.metadata.create_all(bind=engine)
|
||||
|
||||
storage_root = Path(tempfile.mkdtemp(prefix="canary-storage-"))
|
||||
storage_root = Path(
|
||||
os.environ.get("SS_CANARY_STORAGE_ROOT") or tempfile.mkdtemp(prefix="canary-storage-")
|
||||
)
|
||||
storage = DraftStorage(storage_root=storage_root)
|
||||
service = ScreenshotService(environment)
|
||||
loop = ProviderEventLoop()
|
||||
|
||||
@@ -0,0 +1,396 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Fault-injection canary for 044 T045 (startup/drain/cancel/reconcile/quarantine/late-response).
|
||||
|
||||
Runs the REAL provider runtime, capacity manager, provider-receipt CAS, cancel lifecycle and
|
||||
readiness seams against a real PostgreSQL canary database, injecting faults at each boundary:
|
||||
|
||||
V1 loop_not_running submit before start -> typed PROVIDER_LOOP_NOT_RUNNING, never a hang
|
||||
V2 deadline_cancels a slow coroutine is cancelled at the caller deadline (bounded)
|
||||
V3 shutdown_drain stop() during in-flight work -> caller unwinds on its own deadline
|
||||
V4 unknown_effect crash (no release) quarantines capacity until reconcile frees it
|
||||
V5 late_response a late provider response / reconcile cannot win a terminal receipt
|
||||
V6 cancel_drain cancel drains, then terminalizes with zero running steps/leases
|
||||
|
||||
Required env (export from backend/.env):
|
||||
DATABASE_URL PostgreSQL canary database (must contain "canary"); never the dev database
|
||||
ENCRYPTION_KEY same key the app uses (config decryption) — needed by transitive imports
|
||||
AUTH_SECRET_KEY same secret the app uses
|
||||
|
||||
Exit code is 0 only when every vector passes.
|
||||
"""
|
||||
|
||||
# #region ScenarioExecution.Prototype.FaultInjectionCanary [C:5] [TYPE Module] [SEMANTICS scenario,fault-injection,canary,drain,capacity,reconcile]
|
||||
# @BRIEF Fault-injection canary for the 044 provider lifecycle (T045): real services, real PG.
|
||||
# @PRE DATABASE_URL points at a canary database; ENCRYPTION_KEY/AUTH_SECRET_KEY are exported.
|
||||
# @POST Writes an evidence JSON with per-vector status and exits non-zero on any failure.
|
||||
# @INVARIANT The canary never releases a lease it did not claim and never touches non-canary rows.
|
||||
# @SIDE_EFFECT Claims/releases capacity leases, creates runs/steps/receipts, starts a provider loop.
|
||||
# @RATIONALE Offline suites already cover these branches with mocked sessions; this canary proves
|
||||
# the same guarantees on the real PostgreSQL transaction path plus the real thread loop,
|
||||
# which is where drain/deadline/quarantine behaviour actually lives.
|
||||
# @REJECTED Driving a live browser stand here was rejected — T045 is about lifecycle mechanics, not
|
||||
# provider I/O; browser I/O is covered by the T042b canaries.
|
||||
from __future__ import annotations
|
||||
|
||||
import asyncio
|
||||
import json
|
||||
import os
|
||||
import sys
|
||||
import time
|
||||
import threading
|
||||
import uuid
|
||||
from datetime import UTC, datetime, timedelta
|
||||
from pathlib import Path
|
||||
|
||||
BACKEND = Path(__file__).resolve().parents[3] / "backend"
|
||||
sys.path.insert(0, str(BACKEND))
|
||||
|
||||
for _key in ("DATABASE_URL", "ENCRYPTION_KEY", "AUTH_SECRET_KEY"):
|
||||
if not os.environ.get(_key):
|
||||
print(f"fault-injection canary: missing required env {_key}", file=sys.stderr)
|
||||
sys.exit(2)
|
||||
if "canary" not in os.environ["DATABASE_URL"]:
|
||||
print("fault-injection canary: refusing to run against a non-canary DATABASE_URL", file=sys.stderr)
|
||||
sys.exit(2)
|
||||
|
||||
_EVIDENCE_DIR = Path(
|
||||
os.environ.get("SS_FAULT_CANARY_OUT")
|
||||
or Path(__file__).resolve().parent.parent / "evidence" / "fault-injection"
|
||||
)
|
||||
_SCENARIO_ID = "aaaaaaaa-aaaa-4aaa-8aaa-aaaaaaaaaaa1"
|
||||
_REVISION_ID = "bbbbbbbb-bbbb-4bbb-8bbb-bbbbbbbbbbb1"
|
||||
|
||||
|
||||
# #region ScenarioExecution.Prototype.FaultInjectionCanary.Support [C:3] [TYPE Module] [SEMANTICS scenario,canary,support]
|
||||
# @BRIEF Shared collector plus registry/run/step seeding used by the cancel vectors.
|
||||
# @POST Returns a (record, failures) pair and seeded ids; creates the canary registry entry once.
|
||||
class _Canary:
|
||||
def __init__(self) -> None:
|
||||
self.results: list[dict] = []
|
||||
self.failures: list[str] = []
|
||||
|
||||
def record(self, vector: str, ok: bool, **detail) -> None:
|
||||
self.results.append({"vector": vector, "ok": ok, **detail})
|
||||
if not ok:
|
||||
self.failures.append(vector)
|
||||
|
||||
def guard(self, vector: str, fn, **expect) -> dict:
|
||||
"""Run fn() and record TypeError/ValueError/RuntimeError mappings against expectations."""
|
||||
started = time.monotonic()
|
||||
try:
|
||||
value = fn()
|
||||
self.record(vector, expect.get("ok", True), elapsed_seconds=round(time.monotonic() - started, 3), value=str(value)[:200])
|
||||
return {"value": value}
|
||||
except BaseException as exc: # noqa: BLE001 - the canary reports any failure type
|
||||
code = str(exc)
|
||||
self.record(
|
||||
vector,
|
||||
"error" in expect and expect["error"] in code,
|
||||
elapsed_seconds=round(time.monotonic() - started, 3),
|
||||
error=code[:200],
|
||||
expected_error=expect.get("error"),
|
||||
)
|
||||
return {"error": code}
|
||||
|
||||
|
||||
def seed_registry(db) -> None:
|
||||
from src.models.scenario_registry import ScenarioRegistryEntry, ScenarioRevision
|
||||
|
||||
if db.query(ScenarioRegistryEntry).filter(ScenarioRegistryEntry.scenario_id == _SCENARIO_ID).one_or_none():
|
||||
db.flush()
|
||||
return
|
||||
db.add(ScenarioRegistryEntry(
|
||||
scenario_id=_SCENARIO_ID, scenario_key="fault-canary", name="Fault canary",
|
||||
dashboard_id=11, environment_ids=["canary-fault-env"], owner_id="canary", owner_username="canary",
|
||||
))
|
||||
db.add(ScenarioRevision(
|
||||
revision_id=_REVISION_ID, scenario_id=_SCENARIO_ID, content_hash="f" * 64, created_by="canary",
|
||||
))
|
||||
db.flush()
|
||||
|
||||
|
||||
def seed_run(db, *, run_id: str, step_id: str, status: str = "queued") -> None:
|
||||
from src.models.scenario_run import ScenarioRun, ScenarioStepRun
|
||||
from src.models.scenario_worker import ScenarioStepLease
|
||||
|
||||
seed_registry(db)
|
||||
db.add(ScenarioRun(
|
||||
id=run_id, scenario_id=_SCENARIO_ID, scenario_revision_id=_REVISION_ID,
|
||||
scenario_content_hash="f" * 64, environment_id="canary-fault-env", status=status,
|
||||
phase="executing", parameter_bindings={}, target_snapshot={"environment_class": "PREPROD"},
|
||||
trigger_source="manual", idempotency_key=run_id, runner_plan={"steps": []},
|
||||
))
|
||||
db.add(ScenarioStepRun(
|
||||
id=str(uuid.uuid4()), run_id=run_id, logical_step_id=step_id, step_position=0,
|
||||
attempt=1, status="running", started_at=datetime.now(UTC).replace(tzinfo=None),
|
||||
))
|
||||
# A claimed worker lease is what the cancel invariant actually owns (capacity leases are the
|
||||
# provider's: released by the provider's finally or reconciled by TTL — see V4).
|
||||
db.add(ScenarioStepLease(
|
||||
run_id=run_id, logical_step_id=step_id, worker_id="fault-canary-worker",
|
||||
expires_at=datetime.now(UTC).replace(tzinfo=None) + timedelta(minutes=5),
|
||||
))
|
||||
db.flush()
|
||||
|
||||
|
||||
# #endregion ScenarioExecution.Prototype.FaultInjectionCanary.Support
|
||||
|
||||
|
||||
# #region ScenarioExecution.Prototype.FaultInjectionCanary.Loop [C:4] [TYPE Function] [SEMANTICS scenario,canary,loop,drain,deadline]
|
||||
# @BRIEF V1-V3: loop-unavailable, caller-deadline cancellation and bounded shutdown drain.
|
||||
# @POST No vector may hang: every branch is bounded by an explicit caller deadline.
|
||||
def run_loop_vectors(canary: _Canary, db) -> None:
|
||||
from src.services.dashboard_testing.execution.provider_runtime import ProviderEventLoop
|
||||
|
||||
stopped = ProviderEventLoop()
|
||||
canary.guard("v1_loop_not_running", lambda: stopped.submit(lambda: asyncio.sleep(0), timeout=2.0),
|
||||
error="PROVIDER_LOOP_NOT_RUNNING")
|
||||
canary.guard("v1_expired_deadline", lambda: stopped.submit(lambda: asyncio.sleep(0), timeout=0.0),
|
||||
error="PROVIDER_SUBMIT_DEADLINE")
|
||||
|
||||
loop = ProviderEventLoop()
|
||||
loop.start()
|
||||
cancelled = threading.Event()
|
||||
|
||||
async def slow() -> str:
|
||||
try:
|
||||
await asyncio.sleep(5)
|
||||
finally:
|
||||
cancelled.set()
|
||||
raise asyncio.CancelledError
|
||||
return "done"
|
||||
|
||||
canary.guard("v2_deadline_cancels", lambda: loop.submit(slow, timeout=0.3), error="PROVIDER_SUBMIT_DEADLINE")
|
||||
canary.record("v2_coroutine_cancelled", cancelled.wait(timeout=2.0), cancelled=cancelled.is_set())
|
||||
|
||||
# V3: stop the loop while a submission is in flight — the caller must unwind on its own
|
||||
# deadline (loop death never hangs the caller) and the loop must report not-running.
|
||||
async def in_flight() -> str:
|
||||
await asyncio.sleep(4)
|
||||
return "late"
|
||||
|
||||
stopper = threading.Timer(0.3, loop.stop)
|
||||
stopper.start()
|
||||
outcome = canary.guard("v3_shutdown_drain_bounded", lambda: loop.submit(in_flight, timeout=1.5),
|
||||
error="PROVIDER_SUBMIT_DEADLINE")
|
||||
stopper.join(timeout=5)
|
||||
canary.record(
|
||||
"v3_loop_stopped_after_drain",
|
||||
loop.is_running is False and outcome.get("error", "").startswith("PROVIDER_SUBMIT_DEADLINE"),
|
||||
is_running=loop.is_running,
|
||||
)
|
||||
|
||||
|
||||
# #endregion ScenarioExecution.Prototype.FaultInjectionCanary.Loop
|
||||
|
||||
|
||||
# #region ScenarioExecution.Prototype.FaultInjectionCanary.Capacity [C:5] [TYPE Function] [SEMANTICS scenario,canary,capacity,quarantine,reconcile]
|
||||
# @BRIEF V4: a crashed (never-released) lease quarantines the slot until reconciliation frees it.
|
||||
# @POST Capacity is unavailable while the lease is unresolved and restored after TTL reconciliation.
|
||||
def run_capacity_vector(canary: _Canary, db) -> None:
|
||||
from src.services.dashboard_testing.execution.capacity import (
|
||||
DEFAULT_WORKLOAD_LIMITS,
|
||||
CapacityUnavailable,
|
||||
claim_capacity,
|
||||
reconcile_expired_leases,
|
||||
release_capacity,
|
||||
)
|
||||
|
||||
env_id = f"canary-fault-env-{uuid.uuid4().hex[:8]}"
|
||||
limit = DEFAULT_WORKLOAD_LIMITS["browser"]["default"]
|
||||
identity = dict(environment_id=env_id, environment_class="PREPROD", workload_class="browser", provider_id="browser")
|
||||
|
||||
lease = claim_capacity(db, run_id=f"fault-canary-{uuid.uuid4().hex[:8]}", requested_units=limit, ttl_seconds=2, **identity)
|
||||
db.commit()
|
||||
|
||||
blocked = False
|
||||
try:
|
||||
claim_capacity(db, run_id=f"fault-canary-{uuid.uuid4().hex[:8]}", requested_units=1, ttl_seconds=30, **identity)
|
||||
db.rollback()
|
||||
except CapacityUnavailable:
|
||||
blocked = True
|
||||
db.rollback()
|
||||
canary.record("v4_slot_quarantined_by_unreleased_lease", blocked, limit=limit, lease_id=lease["lease_id"])
|
||||
|
||||
# Crash simulation end: TTL expiry + reconciliation frees exactly the quarantined units.
|
||||
time.sleep(2.5)
|
||||
reconciled = reconcile_expired_leases(db)
|
||||
db.commit()
|
||||
restored = claim_capacity(db, run_id=f"fault-canary-{uuid.uuid4().hex[:8]}", requested_units=1, ttl_seconds=30, **identity)
|
||||
db.commit()
|
||||
canary.record(
|
||||
"v4_reconcile_frees_quarantine",
|
||||
bool(restored.get("lease_id")) and isinstance(reconciled.get("expired"), int),
|
||||
expired=reconciled.get("expired"),
|
||||
)
|
||||
release_capacity(db, restored["lease_id"])
|
||||
db.commit()
|
||||
|
||||
|
||||
# #endregion ScenarioExecution.Prototype.FaultInjectionCanary.Capacity
|
||||
|
||||
|
||||
# #region ScenarioExecution.Prototype.FaultInjectionCanary.Receipts [C:5] [TYPE Function] [SEMANTICS scenario,canary,receipt,cas,late-response]
|
||||
# @BRIEF V5: unknown effect -> reconciliation_required; late response and reconcile cannot win.
|
||||
# @POST Terminal receipt status is immutable; late input only appends history.
|
||||
def run_receipt_vectors(canary: _Canary, db) -> None:
|
||||
from src.models.provider_operation import ProviderOperationReceipt
|
||||
from src.services.dashboard_testing.execution.provider_operations import (
|
||||
complete_provider_operation,
|
||||
open_provider_operation,
|
||||
record_late_response,
|
||||
reconcile_provider_operation,
|
||||
)
|
||||
|
||||
run_id = f"fault-canary-{uuid.uuid4().hex[:8]}"
|
||||
receipt = open_provider_operation(
|
||||
db, attempt=1, effect_state="unknown",
|
||||
run_id=run_id, logical_step_id=str(uuid.uuid4()), provider_id="browser", provider_version="v1",
|
||||
action="edit_row", descriptor_fingerprint="d" * 64, binding_ref="bind-fault",
|
||||
execution_principal_fingerprint="p" * 64, idempotency_key=f"idem-{uuid.uuid4().hex[:8]}",
|
||||
)
|
||||
operation_id = receipt["operation_id"]
|
||||
complete_provider_operation(db, operation_id, status="reconciliation_required", effect_state="unknown")
|
||||
db.commit()
|
||||
required = db.query(ProviderOperationReceipt).filter(ProviderOperationReceipt.operation_id == operation_id).one()
|
||||
canary.record(
|
||||
"v5_unknown_effect_requires_reconciliation",
|
||||
required.status == "reconciliation_required" and required.effect_state == "unknown",
|
||||
status=required.status, effect_state=required.effect_state,
|
||||
)
|
||||
|
||||
resolved = reconcile_provider_operation(db, operation_id, resolution="completed", effect_state="completed", note="verified on target")
|
||||
db.commit()
|
||||
canary.record("v5_reconcile_resolves_receipt", resolved["status"] == "completed", status=resolved["status"])
|
||||
|
||||
late = record_late_response(db, operation_id, {"late": True, "source": "fault-canary"})
|
||||
db.commit()
|
||||
row = db.query(ProviderOperationReceipt).filter(ProviderOperationReceipt.operation_id == operation_id).one()
|
||||
canary.record(
|
||||
"v5_late_response_cannot_win",
|
||||
late["status"] == "completed" and row.status == "completed" and row.history[-1]["kind"] == "late_response",
|
||||
status=row.status, last_history_kind=row.history[-1]["kind"],
|
||||
)
|
||||
canary.guard(
|
||||
"v5_reconcile_on_terminal_refused",
|
||||
lambda: reconcile_provider_operation(db, operation_id, resolution="failed", effect_state="unknown"),
|
||||
error="PROVIDER_OPERATION_TERMINAL",
|
||||
)
|
||||
db.rollback()
|
||||
|
||||
|
||||
# #endregion ScenarioExecution.Prototype.FaultInjectionCanary.Receipts
|
||||
|
||||
|
||||
# #region ScenarioExecution.Prototype.FaultInjectionCanary.Cancel [C:5] [TYPE Function] [SEMANTICS scenario,canary,cancel,drain,terminal]
|
||||
# @BRIEF V6: cancel honours the drain window, then terminalizes with zero running steps/leases.
|
||||
# @POST A drained run reaches cancelled/terminal and leaves no active lease behind.
|
||||
def run_cancel_vectors(canary: _Canary, db) -> None:
|
||||
from src.models.provider_capacity import CapacityLease
|
||||
from src.models.scenario_run import ScenarioRun, ScenarioStepRun
|
||||
from src.models.scenario_worker import ScenarioStepLease
|
||||
from src.services.dashboard_testing.execution.cancel_lifecycle import cancel_run, finalize_expired_cancellations
|
||||
from src.services.dashboard_testing.execution.capacity import claim_capacity, reconcile_expired_leases
|
||||
|
||||
def active_worker_leases(run_id: str) -> int:
|
||||
now = datetime.now(UTC).replace(tzinfo=None)
|
||||
return db.query(ScenarioStepLease).filter(
|
||||
ScenarioStepLease.run_id == run_id, ScenarioStepLease.expires_at > now,
|
||||
).count()
|
||||
|
||||
def claimed_capacity(run_id: str) -> int:
|
||||
return db.query(CapacityLease).filter(
|
||||
CapacityLease.run_id == run_id, CapacityLease.status == "claimed",
|
||||
).count()
|
||||
|
||||
# 6a. drain window open -> phase draining, step still running
|
||||
draining_id, draining_step = f"fault-canary-drain-{uuid.uuid4().hex[:8]}", f"step-{uuid.uuid4().hex[:6]}"
|
||||
seed_run(db, run_id=draining_id, step_id=draining_step)
|
||||
db.commit()
|
||||
# Capacity lease belongs to the step's provider invocation; a cancelled run does not silently
|
||||
# drop it — it is quarantined until TTL reconciliation (V4 semantics), never vanished.
|
||||
claim_capacity(db, environment_id="canary-fault-env", environment_class="PREPROD", workload_class="browser",
|
||||
provider_id="browser", run_id=draining_id, logical_step_id=draining_step, ttl_seconds=1)
|
||||
db.commit()
|
||||
run = cancel_run(db, draining_id, drain_in_flight=True, drain_seconds=30)
|
||||
db.commit()
|
||||
canary.record("v6_cancel_opens_drain_window", run.phase == "draining", phase=run.phase, status=run.status)
|
||||
|
||||
# 6b. drain deadline expires -> finalizer terminalizes, expires the worker lease; the capacity
|
||||
# lease is reconciled (not released by cancel) after its TTL.
|
||||
run.cancel_drain_deadline_at = datetime.now(UTC).replace(tzinfo=None) - timedelta(seconds=1)
|
||||
db.flush()
|
||||
db.commit()
|
||||
finalized = finalize_expired_cancellations(db)
|
||||
db.commit()
|
||||
drained = db.query(ScenarioRun).filter(ScenarioRun.id == draining_id).one()
|
||||
step = db.query(ScenarioStepRun).filter(ScenarioStepRun.run_id == draining_id, ScenarioStepRun.status == "running").count()
|
||||
canary.record(
|
||||
"v6_drain_expiry_terminalizes",
|
||||
drained.status == "cancelled" and step == 0 and active_worker_leases(draining_id) == 0,
|
||||
finalized=[r.id for r in finalized], status=drained.status,
|
||||
running_steps=step, active_worker_leases=active_worker_leases(draining_id),
|
||||
)
|
||||
# Capacity: still quarantined right after cancellation, freed only by TTL reconciliation.
|
||||
held_after_cancel = claimed_capacity(draining_id)
|
||||
time.sleep(1.3)
|
||||
reconcile_expired_leases(db)
|
||||
db.commit()
|
||||
canary.record(
|
||||
"v6_cancelled_capacity_reconciled_not_dropped",
|
||||
held_after_cancel == 1 and claimed_capacity(draining_id) == 0,
|
||||
held_after_cancel=held_after_cancel, claimed_after_reconcile=claimed_capacity(draining_id),
|
||||
)
|
||||
|
||||
# 6c. immediate cancel (drain disabled) terminalizes in one pass
|
||||
fast_id, fast_step = f"fault-canary-fast-{uuid.uuid4().hex[:8]}", f"step-{uuid.uuid4().hex[:6]}"
|
||||
seed_run(db, run_id=fast_id, step_id=fast_step)
|
||||
db.commit()
|
||||
immediate = cancel_run(db, fast_id, drain_in_flight=False)
|
||||
db.commit()
|
||||
running = db.query(ScenarioStepRun).filter(ScenarioStepRun.run_id == fast_id, ScenarioStepRun.status == "running").count()
|
||||
canary.record(
|
||||
"v6_immediate_cancel_terminalizes",
|
||||
immediate.status == "cancelled" and immediate.phase == "terminal" and running == 0 and active_worker_leases(fast_id) == 0,
|
||||
status=immediate.status, phase=immediate.phase, running_steps=running,
|
||||
active_worker_leases=active_worker_leases(fast_id),
|
||||
)
|
||||
|
||||
|
||||
# #endregion ScenarioExecution.Prototype.FaultInjectionCanary.Cancel
|
||||
|
||||
|
||||
# #region ScenarioExecution.Prototype.FaultInjectionCanary.Main [C:3] [TYPE Function] [SEMANTICS scenario,canary,main]
|
||||
# @BRIEF Run every vector, write the evidence JSON and exit with the failure count.
|
||||
def main() -> int:
|
||||
from src.core.database import SessionLocal
|
||||
|
||||
canary = _Canary()
|
||||
with SessionLocal() as db:
|
||||
run_loop_vectors(canary, db)
|
||||
run_capacity_vector(canary, db)
|
||||
run_receipt_vectors(canary, db)
|
||||
run_cancel_vectors(canary, db)
|
||||
|
||||
report = {
|
||||
"started_at": datetime.now(UTC).isoformat(),
|
||||
"database_url_scheme": os.environ["DATABASE_URL"].split("://", 1)[0],
|
||||
"results": canary.results,
|
||||
"failures": canary.failures,
|
||||
}
|
||||
_EVIDENCE_DIR.mkdir(parents=True, exist_ok=True)
|
||||
stamp = datetime.now(UTC).strftime("%Y%m%dT%H%M%SZ")
|
||||
out = _EVIDENCE_DIR / f"fault-injection-canary-{stamp}.json"
|
||||
out.write_text(json.dumps(report, indent=2))
|
||||
print(json.dumps({"vectors": len(canary.results), "failures": canary.failures, "evidence": str(out)}, indent=2))
|
||||
return 1 if canary.failures else 0
|
||||
|
||||
|
||||
# #endregion ScenarioExecution.Prototype.FaultInjectionCanary.Main
|
||||
|
||||
|
||||
# #endregion ScenarioExecution.Prototype.FaultInjectionCanary
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
raise SystemExit(main())
|
||||
@@ -1,11 +1,15 @@
|
||||
# Quickstart: Scenario Execution Engine (044)
|
||||
|
||||
> **Refresh 2026-09-08 (production contract):** production baseline-backed evaluation (SCEX-FR-028)
|
||||
> **Refresh 2026-09-17 (production contract):** production baseline-backed evaluation (SCEX-FR-028)
|
||||
> is normative in `contracts/production-chain.md`, `contracts/artifact-content.openapi.yaml`,
|
||||
> `contracts/result-evidence.schema.json` — implemented=false / acceptance OPEN. Canonical chain:
|
||||
> `contracts/result-evidence.schema.json`. Canonical chain:
|
||||
> browser→capture→durable artifact→deterministic baseline comparison→optional declared
|
||||
> AgentEvaluation→DecisionPolicy→StepOutcome. Offline executable check (schemas + truth tables +
|
||||
> evaluation-evidence binding):
|
||||
> AgentEvaluation→DecisionPolicy→StepOutcome. Closed with executable evidence on 2026-09-16/17:
|
||||
> browser PREPROD canaries (T042b 6/6), artifact-content live ACL/status/header canary (T044 11/11),
|
||||
> fault-injection lifecycle canary (T045 16/16), release audit incl. real PostgreSQL chain + semantic
|
||||
> rebuild (T022). Remaining: graph-level terminal PASS (T046 binding; T043 comparison), T042 live
|
||||
> Superset query + `RESULT_TOO_LARGE`, SC-003 100-trial cancel evidence, frontend boundary.
|
||||
> Offline executable check (schemas + truth tables + evaluation-evidence binding):
|
||||
>
|
||||
> ```bash
|
||||
> backend/.venv/bin/python specs/044-dashboard-scenario-execution/prototype/validate_contract_refresh.py \
|
||||
@@ -66,7 +70,15 @@ python3 specs/044-dashboard-scenario-execution/prototype/validate_static.py
|
||||
|
||||
2026-09-10 handoff recorded 292 passed on the first 044-slice command and 21+18 on the evaluation/policy suites. Those counts are historical runtime evidence for slices A+B/C/F, not this documentation pass and not production GO.
|
||||
|
||||
## Production verification (still OPEN)
|
||||
## Production verification (re-evaluated 2026-09-17)
|
||||
|
||||
Verified on 2026-09-16/17 (commands + live canaries; details in [traceability.md](traceability.md)):
|
||||
scoped 044 suite **1590 passed**; full backend **11269 passed / 0 failed**; frontend **3359 passed** +
|
||||
lint + build; fresh PostgreSQL 16 `0001→0025_drop_legacy_validation` + `alembic check` **no drift**;
|
||||
Axiom rebuild (10890 contracts / 5383 edges, `execution/providers` **0 `module_too_long`** after the
|
||||
decomposition gate); browser canaries **6/6**; artifact-content canary **11/11**; fault-injection
|
||||
canary **16/16**. Still required for release: the T046 live rerun to terminal `passed` and the T042
|
||||
Superset live query.
|
||||
|
||||
```bash
|
||||
cd backend
|
||||
@@ -76,24 +88,31 @@ alembic upgrade head
|
||||
python -m pytest -q --run-integration tests/integration/
|
||||
```
|
||||
|
||||
Provider contract profile (T028–T042; live Browser/Screenshot still required):
|
||||
Provider contract profile (T028–T042b):
|
||||
|
||||
```bash
|
||||
python -m pytest -q tests/services/dashboard_testing/registry/test_provider_contract.py
|
||||
python -m pytest -q tests/services/dashboard_testing/registry/test_provider_*.py
|
||||
```
|
||||
|
||||
## Measurable exit gates (acceptance OPEN)
|
||||
## Measurable exit gates (re-evaluated 2026-09-17)
|
||||
|
||||
1. SC-001..011 each has a passing named test or deployment evidence record.
|
||||
2. 100/100 cancellation trials terminate by `cancel_drain_deadline_at + 5 seconds`.
|
||||
3. 100% of provider I/O attempts have a valid CapacityLease and operation receipt.
|
||||
4. 100% of accepted evidence refs have an ownership receipt and verified SHA-256.
|
||||
5. 100% of unknown external effects are reconciled or terminalized non-pass before retry.
|
||||
6. Every enabled provider has passing liveness, readiness and dependency-health checks.
|
||||
7. Browser and Screenshot perform one real authorized deployment run with durable evidence.
|
||||
8. No unresolved P0/P1 traceability row remains in 044 or execution-critical dependencies 036, 037, 038, 041, 042, 046 and 047.
|
||||
1. SC-001..011 each has a passing named test or deployment evidence record. — SC-002/004/006/008/009/010/011 **closed**; SC-003 `[~]` (100-trial evidence open, drain itself proven); SC-007 `[~]` (T042 Superset live query).
|
||||
2. 100/100 cancellation trials terminate by `cancel_drain_deadline_at + 5 seconds`. — **OPEN** (drain/terminalize proven by T045 on real PostgreSQL; the 100-trial production sweep is not yet run).
|
||||
3. 100% of provider I/O attempts have a valid CapacityLease and operation receipt. — **CLOSED** (T032 + canaries; typed `*_CAPACITY_UNAVAILABLE`, heartbeat before submission).
|
||||
4. 100% of accepted evidence refs have an ownership receipt and verified SHA-256. — **CLOSED** (provider receipts + artifact canary digest/length checks).
|
||||
5. 100% of unknown external effects are reconciled or terminalized non-pass before retry. — **CLOSED** (T045: quarantine + TTL reconcile; late response cannot win).
|
||||
6. Every enabled provider has passing liveness, readiness and dependency-health checks. — **CLOSED** (T040 readiness preflight + SC-010).
|
||||
7. Browser and Screenshot perform one real authorized deployment run with durable evidence. — **CLOSED** (T042b 6/6 GREEN on the owner-authorized stand, evidence retained).
|
||||
8. No unresolved P0/P1 traceability row remains in 044 or execution-critical dependencies 036, 037, 038, 041, 042, 046 and 047. — **OPEN** (T043/T046 + dependency rows).
|
||||
|
||||
## Current boundary
|
||||
## Current boundary (re-evaluated 2026-09-17)
|
||||
|
||||
The local profile proves fail-closed adapters, exact Superset binding behavior, lifecycle closure, prototype state coverage, DecisionPolicy unit coverage, AgentEvaluation store/parser/executor, production evaluation adapter (mock provider), and walker `EVALUATION_UNAVAILABLE` / `BASELINE_AND_SEMANTIC_PASS`. It does not prove Browser/Screenshot live composition, shared provider capacity, live LLM evaluation, authenticated artifact GET/HEAD, real PostgreSQL migration validity, scheduler deployment behavior, or 047 case ingestion on a live stand.
|
||||
The local profile now proves, in addition to the earlier fail-closed adapters and DecisionPolicy
|
||||
coverage: Browser/Screenshot live composition (canary 6/6), shared provider capacity with quarantine
|
||||
and TTL reconciliation, live LLM evaluation with an immutable `AgentEvaluation` and stamped baseline
|
||||
pin (canary v2/v4), authenticated artifact GET/HEAD (live 11/11), real PostgreSQL migration validity
|
||||
(`0001→0025`, no drift) and the semantic index rebuild. It does **not** yet prove a graph-level
|
||||
terminal PASS that combines the deterministic comparison with the bound evaluation (T046 binding),
|
||||
the live Superset query path (T042), 100-trial cancellation evidence, or 047 case ingestion on a live
|
||||
stand.
|
||||
|
||||
@@ -16,7 +16,7 @@
|
||||
@SEMANTICS: spec, requirements, feature, scenario, execution, run, step, runner, engine, resume, human
|
||||
|
||||
**Feature Branch**: `044-dashboard-scenario-execution`
|
||||
**Created**: 2026-08-07 | **Status**: Not production-complete — fail-safe execution closure and targeted verification are recorded below; production composition remains open
|
||||
**Created**: 2026-08-07 | **Status (2026-09-17)**: Production-gated — the 2026-09-08 production rows are largely closed with executable evidence (T042b PREPROD canaries 6/6 GREEN, T044 artifact-content live canary 11/11, T045 fault-injection canary 16/16, T022 release audit incl. real PostgreSQL chain + semantic rebuild). Still OPEN: T043/T046 graph-level terminal PASS (evaluation→comparison binding diagnosed in T046), T042 live Superset query + `RESULT_TOO_LARGE` bound, SC-003 100-trial cancel evidence, and the external-MCP-only frontend boundary (050 CHK003/CHK004).
|
||||
**Input**: "Provide a Scenario Execution Engine: a deterministic runner that walks the validated DashboardTestScenario graph, dispatches each step by tool to a typed executor (browser, superset_api, xlsx, assertion, screenshot, report, artifact), and manages the run lifecycle (queued, running, waiting_human, blocked, cancel, retry, timeout, resume, passed, failed, inconclusive) with immutable execution snapshots pinned to a scenario revision."
|
||||
|
||||
## User Scenarios
|
||||
@@ -265,6 +265,11 @@ manual-run-only and PROD mutation remains prohibited by policy.
|
||||
|
||||
## Implementation Status & MVP Debt (factual audit 2026-08-20)
|
||||
|
||||
> **Historical (superseded 2026-09-17).** This audit predates the live providers, canaries and
|
||||
> decomposition. Current factual status: [SESSION_STATE.md](SESSION_STATE.md)
|
||||
> §"Session Closure Update (2026-09-16/17)"; row-level states in [traceability.md](traceability.md)
|
||||
> and [tasks.md](tasks.md). Do not quote this section's gaps as current.
|
||||
|
||||
ScenarioRun/ScenarioStepRun models, lifecycle/API primitives, a runner plan, and a fail-safe executor
|
||||
boundary now exist. They do not yet constitute the specified production execution engine.
|
||||
|
||||
@@ -312,8 +317,9 @@ boundary now exist. They do not yet constitute the specified production executio
|
||||
policy). Application startup owns a fail-closed `LiveExecutionCompositionRoot`: an exact registered
|
||||
037 client/model/DraftStorage tuple is wired into the default dispatcher, while registered browser and
|
||||
Screenshot providers are validated against the same snapshot. Missing/mismatched providers are typed
|
||||
inconclusive and never call I/O. No deployment has yet registered a browser-safe action or
|
||||
ScreenshotService-to-durable-evidence provider, so those tools remain unavailable by default.
|
||||
inconclusive and never call I/O. **Update 2026-09-16/17:** browser and Screenshot providers are
|
||||
registered and exercised live (T042b 6/6 PREPROD vectors, T044 protected content 11/11); the
|
||||
remaining gap is the live Superset **query** binding (T042), not the browser/Screenshot composition.
|
||||
- `[~]` The targeted service profile independently reverified 35 passes. The exact API file has a
|
||||
29-pass result outside this sandbox; inside it FastAPI TestClient is blocked by AnyIO self-pipe
|
||||
`EPERM`, a sandbox restriction rather than an application failure.
|
||||
@@ -412,7 +418,7 @@ Exploratory Playwright/code sandbox activity is authoring-only and must remain i
|
||||
|
||||
**SCEX-FR-028 — Production baseline-backed evaluation**: Run admission, canonical request hash, RunnerPlan, artifacts, immutable comparisons/evaluations, policy outcomes and result/SSE MUST satisfy production-chain.md, artifact-content.openapi.yaml and result-evidence.schema.json. Provider loop startup/shutdown, protected GET/HEAD, baseline resolution/pinning and cancellation/reconcile are mandatory; screenshots currently exist as reusable adapters but missing deployment/evaluation/content integration is not complete. 038 DecisionPolicy is sole semantic outcome mapper.
|
||||
|
||||
Normative contract: [Production baseline-backed evaluation](contracts/production-chain.md). New requirements are specified, **implemented=false / acceptance OPEN** until executable evidence closes the linked tasks/checklist/traceability rows. Historical local tests and the manual inconclusive ss-prod run do not prove browser/capture/baseline/LLM production readiness. The refresh scope is the audited P0/P1/P2 agentic E2E and baseline gaps; an approved ExecutionPerformanceBaseline is not introduced.
|
||||
Normative contract: [Production baseline-backed evaluation](contracts/production-chain.md). Requirements are closed **only** by executable evidence in [traceability.md](traceability.md): as of 2026-09-17 the browser/capture/content/fault-injection gates are closed by live canaries (T042b 6/6, T044 11/11, T045 16/16) and the release audit (T022); the baseline-resolution/pinning and immutable-evaluation chain is proven live (canary v4, `docs/reports/agentic-runtime-live-canary-v4-baseline-pin-2026-09-11.md`). Remaining acceptance: the graph-level terminal PASS (T046 — the walker-level evaluation→comparison binding diagnosed with executable proof; T043's deterministic comparison PASS closes with the same live rerun), T042's live Superset query with an explicit `RESULT_TOO_LARGE` bound, SC-003's 100-trial cancel evidence, and the external-MCP-only frontend boundary. The refresh scope is the audited P0/P1/P2 agentic E2E and baseline gaps; an approved ExecutionPerformanceBaseline is not introduced.
|
||||
|
||||
## Field-run Amendment — comparison executor typed failure boundary (2026-09-12)
|
||||
|
||||
|
||||
@@ -87,9 +87,32 @@
|
||||
## Phase 8 — Polish
|
||||
|
||||
- [x] T021 [P] Belief-runtime instrumentation for C5 runner/dispatch
|
||||
- [ ] T022 Run quickstart-equivalent full scenario backend scope, scoped Ruff, ATTN/orphan static audit,
|
||||
- [x] T022 Run quickstart-equivalent full scenario backend scope, scoped Ruff, ATTN/orphan static audit,
|
||||
real PostgreSQL `alembic check`/`upgrade head`, and semantic rebuild. Completion requires command
|
||||
output attached to the release record and zero unresolved P0/P1 rows.
|
||||
**Status (2026-09-16, HEAD `4746af2f`):** scoped 044 suite
|
||||
`pytest -q tests/services/dashboard_testing/ tests/api/test_scenario_{runs,automation,analytics,artifact_content,run_center}*.py`
|
||||
→ **1590 passed** (48.7s); `validate_static.py` → passed; ruff (whole backend) + `compileall -q src`
|
||||
clean; full backend suite **11269 passed / 252 skipped / 1 xpassed / 0 failed**; frontend
|
||||
**3359 passed** + lint 0 errors + build green; fresh PostgreSQL 16 chain `0001 → 0025_drop_legacy_validation`
|
||||
applied, `alembic check` → **no new upgrade operations** (no drift), single head confirmed;
|
||||
Axiom live rebuild: 10870 contracts / 5365 edges, unresolved relations 393 (was 401 at the
|
||||
2026-09-04 re-review).
|
||||
**Status (2026-09-17): CLOSED — INV_7 tail decomposed, zero P0/P1.**
|
||||
Provider decomposition gate executed in three gated phases
|
||||
(`specs/044-dashboard-scenario-execution/plans/provider-decomposition-gate.md`, EXECUTED with
|
||||
full execution log): `browser_readonly_actions.py` 534 → **160** (+ `browser_readonly_limits.py`
|
||||
82 / `flows_nav` 199 / `flows_interact` 192); `browser_session.py` 501 → **58**
|
||||
(+ `browser_session_handle.py` 64 / `browser_session_checkpoint.py` 162 /
|
||||
`browser_session_managers.py` 57 / `browser_session_registry.py` 287); `browser.py` 407 → **398**
|
||||
(+ `browser_factory_helpers.py` 35). Frozen contract IDs and import surface preserved (facades
|
||||
re-export every moved public name); monkeypatch seams followed the owning modules; per-phase
|
||||
gates **1590 passed** each and the full backend suite re-confirmed **11269 passed / 0 failed**
|
||||
(zero delta vs pre-decomposition baseline). Post-decomposition Axiom `audit_contracts` on
|
||||
`execution/providers`: **0 module_too_long** (was 3). Remaining accepted advisories:
|
||||
2 × `contract_too_long` (`ScenarioExecution.BrowserProvider.Factory` 305,
|
||||
`ScenarioExecution.ScreenshotProvider.Factory` 200) — typed exception→result ladders kept
|
||||
inline deliberately (@REJECTED extraction in `browser_factory_helpers.py`).
|
||||
- [x] T023 **Prototype validation**: every declared @UX_STATE is reachable via `prototype/index.html`.
|
||||
Proof: `python specs/044-dashboard-scenario-execution/prototype/validate_static.py`.
|
||||
|
||||
@@ -218,7 +241,7 @@
|
||||
`backend/tests/services/dashboard_testing/registry/test_provider_contract.py`: reclassification to PROD
|
||||
or a changed provider/security fingerprint after approval prevents all provider I/O and returns typed
|
||||
`POLICY_CHANGED` or `BINDING_CHANGED`.
|
||||
- [ ] T042b [US2] Run BrowserProvider PREPROD canaries and retain evidence in
|
||||
- [x] T042b [US2] Run BrowserProvider PREPROD canaries and retain evidence in
|
||||
`specs/044-dashboard-scenario-execution/evidence/browser-provider/`: read-only action canary,
|
||||
forced timeout/cleanup canary, safe-checkpoint reconstruction trace and readiness/health payload.
|
||||
`GO` requires all BrowserProvider acceptance vectors to pass and one owned evidence receipt.
|
||||
@@ -228,6 +251,19 @@
|
||||
browser `open_dashboard` passed with checkpoint + real page URL + PNG digest; readiness payload
|
||||
captured in the trace reports). Still OPEN: PREPROD canaries, the forced timeout/cleanup canary,
|
||||
and the owned receipt under `evidence/browser-provider/`.
|
||||
**Status (2026-09-16): CLOSED — all six vectors GREEN under the Wave-C run-scoped session
|
||||
architecture** against the owner-authorized test stand (`SS_STAND_STAGE=PREPROD`, dashboard 11),
|
||||
evidence retained in `evidence/browser-provider/readonly-canary-20260916T*.json` + PNGs:
|
||||
read-only actions 3/3 passed (`150158Z`); forced timeout/cleanup typed `BROWSER_ACTION_TIMEOUT`
|
||||
with zero evidence and lease release (`150407Z`); mutation row_edit+restore 82.74 verified,
|
||||
2/2 receipts `completed` (`150518Z`); reconciliation sweep resolved the injected stale receipt
|
||||
via live SELECT-only observation (`150754Z`); safe-checkpoint reconstruction attempt 2 `passed`
|
||||
with `reconstruction_replay=true` (`150908Z`); scheduler soak 75s — 3/3 runs `passed`,
|
||||
`attempts == 1` (CAS single-dispatch), 0 active leases, 3 durable artifacts, graceful stop
|
||||
(`155013Z`). Harness defect found and fixed during the soak: the DI-singleton SchedulerService
|
||||
captured a dead event loop in the bare-script context, so async jobs (maintenance_auto_end)
|
||||
blocked for the 300s AsyncJobRunner safety cap and stop() waited — the soak window now runs
|
||||
inside `asyncio.run` with `stop()` offloaded via `asyncio.to_thread`.
|
||||
|
||||
## Requirement evidence rules
|
||||
|
||||
@@ -265,8 +301,11 @@
|
||||
unsafe effects require reconciliation, and browser recovery requires a pinned safe checkpoint
|
||||
(test_scenario_crash_recovery.py). Retired evidence remains historical and rejected terminal
|
||||
contexts reuse their idempotent signal. Persisted artifact evidence needs a real valid digest/ref.
|
||||
Approval-to-live-dispatch still needs an authorized production composition root that supplies
|
||||
resolver/client/model/storage; Browser/Screenshot bindings remain unavailable and fail closed.
|
||||
Approval-to-live-dispatch is now exercised end-to-end: the server-owned composition root supplies
|
||||
resolver/client/model/storage from `settings.scenario_live_execution_bindings`, the PROD approval
|
||||
gate is live-approved in the external MCP replay (`docs/2026-09-11-sales-prod-mcp-replay.md`), and
|
||||
Browser/Screenshot providers are deployed and exercised (T042b 6/6 on 2026-09-16). Residual for
|
||||
this lineage is the live Superset query path (T042).
|
||||
Queued HTTP/automation starts remain non-dispatching; the existing scheduler claims eligible
|
||||
rows through durable queued->running CAS before walking them, so repeated ticks do not repeat
|
||||
a side effect. A manual human run reaches its checkpoint only after that claim, while an
|
||||
@@ -295,14 +334,14 @@ Setup → RunnerPlan; US1 (start+executors) → US2 (dispatch); US3 (human) depe
|
||||
|
||||
## Production readiness — 2026-09-08 (SCEX-FR-028)
|
||||
|
||||
Historical [x] rows above retain only their dated local/transport evidence; they do not prove current production readiness. Reopened rows were contradicted by the audited gaps. Removed frontend/agent paths are historical, not implementation prerequisites. New acceptance is **implemented=false / OPEN**.
|
||||
Historical [x] rows above retain only their dated local/transport evidence; they do not prove current production readiness. Reopened rows were contradicted by the audited gaps. Removed frontend/agent paths are historical, not implementation prerequisites. Rows are closed **only** by executable evidence and were re-evaluated 2026-09-17 (see [traceability](traceability.md), [checklist](checklists/requirements.md), [SESSION_STATE](SESSION_STATE.md)): T042b/T044/T045/T022 are CLOSED with canary/audit evidence; T043/T046 remain `[~]` (graph-level terminal PASS); the frontend-boundary row remains open.
|
||||
|
||||
Contract: [Production baseline-backed evaluation](contracts/production-chain.md).
|
||||
|
||||
- [ ] T043 [P0/P1/P2] Baseline set/version/catalog/release/commit/IDs/digests affect idempotency; moving catalog after admission cannot change plan/result; legacy unpinned result is ineligible. Implement at the existing 044 domain boundary; verify with independent hardcoded fixtures and retain command/evidence references in traceability.md. **Status (2026-09-11):** request-hash pinning + fail-closed resolver landed offline (`83727aa7`/`bbbd4ccf`); caller-side published-catalog source wired. Live pin-from-Gitea is now PROVEN: REST canary v4 (`ace916a0…`/`adeabe63…`) + MCP `--baseline` chain resolve the published-envelope pin and the walker stamps it into `AgentEvaluation.baseline_pin` (strict equality) — `docs/reports/agentic-runtime-live-canary-v4-baseline-pin-2026-09-11.md`. Residual: the graph-level deterministic comparison PASS and the MCP publish tool are separate rows (050 T045).
|
||||
- [ ] T044 [P0/P1/P2] GET/HEAD prove same ACL/status/headers; MIME/digest/length checked before bytes, cross-owner hidden, expired410, corrupt409, traversal/range/oversize rejected. Implement at the existing 044 domain boundary; verify with independent hardcoded fixtures and retain command/evidence references in traceability.md. **Status (2026-09-10):** Slice G GET/HEAD runtime landed and committed (`83727aa7`, `scenario_artifact_content.py`); live storage canary proved durable bytes with digests (canary v1 `597274d3`). Live ACL/status/header canary not yet exercised.
|
||||
- [ ] T045 [P0/P1/P2] Startup/readiness/start-loop and shutdown/drain/cancel/reconcile survive fault injection; unknown effect quarantines capacity; late response cannot win. Implement at the existing 044 domain boundary; verify with independent hardcoded fixtures and retain command/evidence references in traceability.md. **Status (2026-09-10):** startup/readiness live-verified (provider loop ready, readiness preflight; browser-probe deadlock fixed `ba2f1f45`); capacity leases claimed/released across the live canaries. Fault-injection drain/cancel/reconcile canary not yet exercised.
|
||||
- [ ] T046 [P0/P1/P2] End-to-end real browser→capture→durable artifact→deterministic comparison→optional immutable evaluation→policy→result; all required evidence present before PASS. Implement at the existing 044 domain boundary; verify with independent hardcoded fixtures and retain command/evidence references in traceability.md. **Status (2026-09-11):** the chain real browser→capture→durable artifact→immutable evaluation→policy→result is PROVEN live; a deterministic comparison step and a live baseline pin are now both in live graphs (canary v4: `compare_to_baseline` deterministically passes; pin resolved from the published envelope and stamped into `AgentEvaluation.baseline_pin`, `docs/reports/agentic-runtime-live-canary-v4-baseline-pin-2026-09-11.md`). Residual: a graph-level terminal PASS (the baseline-semantic policy marks comparisons without a bound evaluation as typed `EVALUATION_UNAVAILABLE`; `evaluate-visual` itself passes).
|
||||
- [~] T043 [P0/P1/P2] Baseline set/version/catalog/release/commit/IDs/digests affect idempotency; moving catalog after admission cannot change plan/result; legacy unpinned result is ineligible. Implement at the existing 044 domain boundary; verify with independent hardcoded fixtures and retain command/evidence references in traceability.md. **Status (2026-09-11):** request-hash pinning + fail-closed resolver landed offline (`83727aa7`/`bbbd4ccf`); caller-side published-catalog source wired. Live pin-from-Gitea is now PROVEN: REST canary v4 (`ace916a0…`/`adeabe63…`) + MCP `--baseline` chain resolve the published-envelope pin and the walker stamps it into `AgentEvaluation.baseline_pin` (strict equality) — `docs/reports/agentic-runtime-live-canary-v4-baseline-pin-2026-09-11.md`. **Status (2026-09-17): `[~]`** — residual: the graph-level deterministic comparison PASS (same live rerun as T046 closes it); the MCP publish tool is a separate row (050 T045, CLOSED).
|
||||
- [x] T044 [P0/P1/P2] GET/HEAD prove same ACL/status/headers; MIME/digest/length checked before bytes, cross-owner hidden, expired410, corrupt409, traversal/range/oversize rejected. Implement at the existing 044 domain boundary; verify with independent hardcoded fixtures and retain command/evidence references in traceability.md. **Status (2026-09-10):** Slice G GET/HEAD runtime landed and committed (`83727aa7`, `scenario_artifact_content.py`); live storage canary proved durable bytes with digests (canary v1 `597274d3`). Live ACL/status/header canary not yet exercised. **Status (2026-09-17): CLOSED — live ACL/status/header canary GREEN.** New harness `specs/044-dashboard-scenario-execution/prototype/artifact_content_canary.py` drives the REAL FastAPI app (no dependency overrides) over real PostgreSQL (`canary_044`) with real JWTs (real `create_access_token`, real session/role/permission lookup, no `sid` so session-policy skips) and REAL live-captured browser evidence bytes (soak run `soak-canary-c071fb98` via `SS_CANARY_STORAGE_ROOT`, 104746-byte PNG, sha `22b4b8c1…`). Evidence: `evidence/artifact-content/artifact-content-canary-20260917T085438Z.json` — **11/11 vectors GREEN**: anonymous→401 `AUTHENTICATION_REQUIRED` (+HEAD empty); viewer GET→200 with byte-exact body and coherent Content-Length/ETag/Disposition/Cache-Control/nosniff/Accept-Ranges; viewer HEAD→200 with full header parity + empty body; no-VIEW→403 `PERMISSION_DENIED` (+HEAD); unknown run / unknown artifact / foreign-owned artifact → indistinguishable 404 `NOT_FOUND` with identical message; expired→410; corrupt sha→409 `ARTIFACT_INTEGRITY_FAILED`; declared-MIME off-allowlist→409; oversized→413; missing bytes→409 `ARTIFACT_MISSING`; Range→416 `RANGE_NOT_SUPPORTED` + `Accept-Ranges: none`. Harness knob added: `SS_CANARY_STORAGE_ROOT` (persistent evidence root; use a DEDICATED root per live run — the soak artifact-count assertion assumes a private root). Offline matrix reference: `tests/api/test_scenario_artifact_content_api.py` 11/11 (independent hardcoded fixtures).
|
||||
- [x] T045 [P0/P1/P2] Startup/readiness/start-loop and shutdown/drain/cancel/reconcile survive fault injection; unknown effect quarantines capacity; late response cannot win. Implement at the existing 044 domain boundary; verify with independent hardcoded fixtures and retain command/evidence references in traceability.md. **Status (2026-09-10):** startup/readiness live-verified (provider loop ready, readiness preflight; browser-probe deadlock fixed `ba2f1f45`); capacity leases claimed/released across the live canaries. Fault-injection drain/cancel/reconcile canary not yet exercised. **Status (2026-09-17): CLOSED — fault-injection canary GREEN (16/16).** New harness `specs/044-dashboard-scenario-execution/prototype/fault_injection_canary.py` runs the REAL provider runtime + capacity manager + receipt CAS + cancel lifecycle against real PostgreSQL (`canary_044`), injecting faults at each boundary; evidence `evidence/fault-injection/fault-injection-canary-20260917T105541Z.json`. Vectors: submit before start / expired deadline → typed `PROVIDER_LOOP_NOT_RUNNING` / `PROVIDER_SUBMIT_DEADLINE` in bounded time; slow coroutine cancelled at the caller deadline (`PROVIDER_SUBMIT_DEADLINE`, `cancelled` flag observed); shutdown during in-flight work unwinds the caller on its own deadline and the loop reports not-running (no hang); a crashed (never-released) lease **quarantines the slot** (second claim → `CAPACITY_UNAVAILABLE`) until `reconcile_expired_leases` frees exactly those units and capacity is restored (SC-008 crash recovery); unknown effect → receipt `reconciliation_required` + `effect_state=unknown`, then `reconcile_provider_operation` resolves it; a late response appends history **without** changing the terminal status and a reconcile on a terminal receipt is refused (`PROVIDER_OPERATION_TERMINAL`); cancel opens the drain window (`draining`), the deadline finalizer terminalizes with 0 running steps and 0 unexpired worker leases, and the capacity lease of a cancelled run is **reconciled after TTL, never silently dropped**; immediate cancel terminalizes in one pass (0 running steps, 0 unexpired worker leases). Offline references: `tests/services/dashboard_testing/registry/test_provider_operations.py` (receipt CAS, late response, reconcile), `test_scenario_cancel_timeout.py` (6 cancel/timeout), `test_dispatch_capacity_lifecycle.py`, `test_provider_capacity.py` — all green (41 + 47 passed on 2026-09-17).
|
||||
- [~] T046 [P0/P1/P2] End-to-end real browser→capture→durable artifact→deterministic comparison→optional immutable evaluation→policy→result; all required evidence present before PASS. Implement at the existing 044 domain boundary; verify with independent hardcoded fixtures and retain command/evidence references in traceability.md. **Status (2026-09-11):** the chain real browser→capture→durable artifact→immutable evaluation→policy→result is PROVEN live; a deterministic comparison step and a live baseline pin are now both in live graphs (canary v4: `compare_to_baseline` deterministically passes; pin resolved from the published envelope and stamped into `AgentEvaluation.baseline_pin`, `docs/reports/agentic-runtime-live-canary-v4-baseline-pin-2026-09-11.md`). Residual: a graph-level terminal PASS (the baseline-semantic policy marks comparisons without a bound evaluation as typed `EVALUATION_UNAVAILABLE`; `evaluate-visual` itself passes). **Diagnosis (2026-09-17, executable proof):** the gap is a walker-level binding, not the policy truth table. `walker.py` computes `decide_step_outcome(policy_inputs_from_outcome(outcome, step_meta, integrity), pinned_policy)` PER STEP using only that step's own outcome; `policy_inputs_from_outcome` binds the evaluation ONLY from `step_outcome["evaluation_input"]`, which the evaluation adapter emits solely on the `agent_evaluation` step (`evaluation_adapter.py:379`). The `assertion compare_to_baseline` step therefore always sees `evaluation=None`; `derive_runner_plan` enables mandatory mode as soon as the graph contains an `agent_evaluation` step (per the invariant "evaluation_mode stays disabled while no registry tool can emit an AgentEvaluation"), so every normative step — including the compare step — resolves to `EVALUATION_UNAVAILABLE` even though `evaluate-visual` passes. Verified on HEAD: (A) compare-only outcome + mandatory policy → `inconclusive ["EVALUATION_UNAVAILABLE"]`; (B) same + disabled → `passed ["BASELINE_PASS"]`; (C) same + bound `evaluation_input` (succeeded/pass/0.9) + mandatory → `passed ["BASELINE_AND_SEMANTIC_PASS"]`. The canonical chain (`contracts/production-chain.md` §8) puts the optional declared evaluation BETWEEN the comparison and the pinned DecisionPolicy, so the decision must observe both; today it is computed per-step in isolation and never sees the sibling evaluation. Design options for the closure (next packet; both keep the pinned truth table unchanged): (1) dependency-ordered binding — the compare step depends on the covering `evaluate-visual` step (v4's graph has them as siblings) and the walker injects the persisted `AgentEvaluation` (via that step's `agent_evaluation_ids`) as `evaluation_input` for the dependent comparison decision; (2) deferred decision — the walker defers the policy decision of comparison steps until their covering evaluation step is persisted, then computes one decision on the aggregated inputs. Either path then requires a live rerun of the v4-style graph ending in terminal `passed` as the closing evidence.
|
||||
|
||||
Frontend boundary for this package: manual CRUD/editor, human review/approval, monitoring and read-only evidence/evaluation only; all agent interaction is external MCP. No agent chat/prompt/assistant editing/proposal generation/workspace/start/handoff controls. Runtime removal is OPEN, not performed by this spec refresh. Optional approved performance baseline is outside scope.
|
||||
|
||||
|
||||
@@ -3,7 +3,7 @@
|
||||
| Story | Requirement | Model | API operationId | Contract | Task | Test | Actual status / gap |
|
||||
|-------|-------------|-------|------------------|----------|------|------|---------------------|
|
||||
| US1 Start | SCEX-FR-001/008 | ScenarioRun | scenarioRun.start | Execution.Start, Execution.Runner.QueuedDispatch | T006-T008, T024 | test_runner, test_scenario_queued_dispatch, test_scenario_scheduler_callbacks, test_scenario_runs_api, test_scenario_automation_api, test_live_execution_binding | `[~]` HTTP/automation start and replay persist queued/pending rows without request-time dispatch; only the scheduler composition's durable queued->running CAS walks its winner. Fixed scheduler callback registration, database-edge containment, and repeat-tick terminal side-effect idempotency are unit-proven. Automated human plans are rejected before they reach CAS/walker; a manual human graph reaches HumanCheckpoint only after that dispatcher claim. Approval-to-real live dispatch remains unproven. |
|
||||
| US2 Dispatch | SCEX-FR-002/006/009 | ScenarioStepRun, LiveExecutionBinding | scenarioRun.step | Execution.Dispatch, Execution.LiveCompositionRoot | T009-T011, T024 | test_dispatch, test_scenario_executors, test_live_execution_binding | `[~]` Browser/Superset/Screenshot use fail-safe typed adapter boundaries: no explicit adapter success means no PASS; invalid evidence digest/ref remains inconclusive. Lifespan bootstraps `settings.scenario_live_execution_bindings` through the existing `SupersetClient`, exact model and durable storage, so configured Superset dispatch invokes 037 and stores the exact raw-byte digest/ref; mismatched/unavailable providers make no I/O call. Browser safe-checkpoint and Screenshot durable-evidence registration are supported but no provider is deployed, so enabled bindings return stable configured-unavailable codes. |
|
||||
| US2 Dispatch | SCEX-FR-002/006/009 | ScenarioStepRun, LiveExecutionBinding | scenarioRun.step | Execution.Dispatch, Execution.LiveCompositionRoot | T009-T011, T024, T042b | test_dispatch, test_scenario_executors, test_live_execution_binding, provider canaries | `[~]` Browser/Superset/Screenshot use fail-safe typed adapter boundaries: no explicit adapter success means no PASS; invalid evidence digest/ref remains inconclusive. Lifespan bootstraps `settings.scenario_live_execution_bindings` through the existing `SupersetClient`, exact model and durable storage, so configured Superset dispatch invokes 037 and stores the exact raw-byte digest/ref. Providers are deployed and exercised live (T042b 6/6 on 2026-09-16; live evaluation chain canary v2/v4); the residual is the live Superset **query** path (T042). |
|
||||
| BrowserProvider | SCEX-FR-016..023 | BrowserProviderActionContract, BrowserOperationReceipt, BrowserEvidenceReceipt | scenarioRun.step | ScenarioExecution.BrowserProvider | T028-T034, T040-T042, T042b | test_provider_browser_*, test_browser_native_filter, Browser PREPROD canaries | `[~]` Contract completeness is **90/100**: all actions, risk classes, limits, lifecycle, checkpoint/recovery, ownership, cancellation/reconciliation, readiness and canary gates are specified. `apply_native_filter` is implemented read-only in the provider/transport (UX-1, 2026-09-12): typed input merge/validation before I/O, filter-bar UI automation with `filter_applied`/`charts_settled` checkpoints and PNG evidence, typed `BROWSER_SELECTOR_NOT_FOUND`/`BROWSER_WAIT_STATE_INVALID` inconclusive without retry, mutation catalog unchanged (`test_provider_browser.py`, `test_browser_native_filter.py`). Launch-param binding `filter_values` -> `apply_native_filter` input `values` is closed at RunnerPlan derivation (UX-6, 2026-09-12, `ScenarioExecution.RunnerPlan.BindParams`; fail-closed typed reject on a present-but-invalid param, absent param stays `current_state`). Runtime provider for the remaining catalog, shared capacity integration, deployment registration, live canary evidence and PREPROD canaries remain open. |
|
||||
| US3 Human | SCEX-FR-004/010 | ScenarioRun(waiting_human) | scenarioRun.humanDecision | Execution.SuspendForHuman, Execution.Resume | T012-T014, T025 | test_human_resume, test_scenario_runs_api | `[~]` persisted HumanCheckpoint and infrastructure-resume continuations advance only the missing DAG frontier; completed steps are not re-run. Full live-composition closure remains pending. |
|
||||
| Manual-only boundary | SCEX-FR-004a | ScenarioRun, HumanCheckpoint | — | Execution.Runner.Start, RunnerPlan.Derive | T027 | test_scenario_manual_run_only, test_scenario_automation_api | `[x]` Trusted scheduled/deploy/release/ETL/API origins reject persisted human revisions before idempotency or any run/gate/notification/queue side effect. Manual origin remains eligible; HumanCheckpoint is not an approval gate. |
|
||||
@@ -12,17 +12,17 @@
|
||||
| US5 Snapshot/API | SCEX-FR-003/007 | ScenarioExecutionResult | scenarioRun.detail, scenarioRun.events | Execution.RunnerPlan, Execution.Runner.CrashRecovery | T017-T020, T025 | test_result, test_api, test_scenario_runner_walker, test_scenario_crash_recovery | `[~]` API/projection and real-digest artifact integrity checks exist. Server-driven crash recovery uses only the persisted run/RunnerPlan and expired lease: completed work is not rerun, safe claims get a new attempt with active evidence retired to history, unsafe claims require reconciliation, and browser recovery without a pinned safe checkpoint is non-pass without adapter I/O. Rejected terminal contexts reuse their idempotent signal. Production live-executor provenance remains unproven. |
|
||||
| Terminal signals | SCEX-FR-011 | InvestigationQueueItem | — | Execution.Runner.TerminalSignal | T026 | test_scenario_terminal_signals | `[x]` Failed/blocked/inconclusive terminal runs emit one idempotent immutable 047 queue input with run/artifact provenance; passed runs emit none. Producer-only ingestion starts no case, AgentRun, chat, remediation action, or recurrence classification. |
|
||||
| Gate/RBAC | SCEX-FR-008 | ScenarioRun, ActionApprovalGate | scenarioRun.start | Execution.EnvironmentPolicy, Execution.Runner.Start | T019 | test_scenario_runner, test_scenario_runs_api, test_scenario_automation_api, test_scenario_automation_trigger | `[~]` Server ConfigManager classifies every target before persistence: client flags cannot select PROD, unknown targets create no run-side effect, and every trusted source enters the same durable pending_approval gate boundary. HTTP/trigger paths remain persistence-only and dispatcher excludes pending gates. A dedicated real APScheduler scheduled-PROD integration test remains coverage debt; this is not a dispatch bypass. |
|
||||
| Provider protocol | SCEX-FR-017..023 | ProviderExecutionContext, ProviderOperationReceipt, ProviderEvidenceReceipt | — | ScenarioExecution.ProviderProtocol, ProviderOperations, ProviderOperations.Observability | T028-T031, T040-T042 | test_provider_contract, test_provider_health, deployment health checks | `[ ]` New production gate: common context/result, operation receipts, ownership proof, cancellation/reconciliation, health/readiness and startup registration are specified but not implemented. |
|
||||
| Capacity | SCEX-FR-024 | CapacityLease, ExecutionCapacityManager | — | ScenarioExecution.CapacityManager | T032, T041-T042 | test_provider_capacity, scheduler/worker integration | `[ ]` New production gate: shared environment-scoped capacity admission is specified but current 046 checks do not close it. |
|
||||
| Agent evaluation | SCEX-FR-025 | AgentEvaluation, DecisionPolicy | — | ScenarioExecution.ProviderCatalog, data-model AgentEvaluation and DecisionPolicy | T039, T041-T042 | test_provider_agent_evaluation | `[ ]` New production gate: bounded provider, immutable evaluation evidence and deterministic policy mapping are specified but absent. |
|
||||
| Provider protocol | SCEX-FR-017..023 | ProviderExecutionContext, ProviderOperationReceipt, ProviderEvidenceReceipt | — | ScenarioExecution.ProviderProtocol, ProviderOperations, ProviderOperations.Observability | T028-T031, T034, T040-T042b | test_provider_contract, test_provider_health, deployment health checks | `[x]` (2026-09-17) Production gate implemented and proven: typed context/result, operation receipts with ownership proof and effect-state CAS, cancellation/reconciliation (`cancel_provider_operation`, `reconcile_provider_operation`), health/readiness snapshot (T040, SC-010), deployment registration. Live evidence: T042b PREPROD canaries **6/6 GREEN** (2026-09-16), T045 fault-injection **16/16 GREEN** (2026-09-17). |
|
||||
| Capacity | SCEX-FR-024 | CapacityLease, ExecutionCapacityManager | — | ScenarioExecution.CapacityManager | T032, T041-T042, T045 | test_provider_capacity, scheduler/worker integration | `[x]` (2026-09-17) Environment-scoped quotas/leases landed (migrations 0014-0016); SC-008 closed (T032, 2026-09-13: dispatcher reconciles expired leases, heartbeat before loop submission, typed `*_CAPACITY_UNAVAILABLE` refusal). Unknown-effect **quarantine** and TTL reconciliation proven live in the T045 fault-injection canary (2026-09-17, 16/16). |
|
||||
| Agent evaluation | SCEX-FR-025 | AgentEvaluation, DecisionPolicy | — | ScenarioExecution.ProviderCatalog, data-model AgentEvaluation and DecisionPolicy | T039, T041-T042 | test_provider_agent_evaluation | `[x]` (2026-09-17) Bounded evaluation executor + immutable evidence store + pinned deterministic DecisionPolicy mapping landed; live LLM evaluation chain proven (canary v2, `AgentEvaluation` persisted → DecisionPolicy row 11); baseline pin stamped from the published envelope (canary v4). Graph-level terminal PASS residual tracked in T046 below. |
|
||||
| Success criteria | SC-001 | RunnerPlan, ScenarioStepRun | — | ScenarioExecution.RunnerPlan.Derive, ScenarioExecution.Dispatch | T004-T011, T041 | canonical fixture order and duplicate-dispatch assertions | `[~]` Existing fixture/lifecycle coverage passes; exact 100% canonical dispatch evidence remains a release gate. |
|
||||
| Success criteria | SC-002 | HumanCheckpoint, ScenarioRun | scenarioRun.humanDecision | ScenarioExecution.HumanCheckpoint, ScenarioExecution.Resume | T012-T014, T025 | human CAS/frontier tests | `[x]` Human checkpoint CAS and missing-frontier resume are verified; live provider closure remains separate. |
|
||||
| Success criteria | SC-003 | ScenarioRun.cancel_drain_deadline_at | scenarioRun.cancel | ScenarioExecution.Cancel | T015-T016, T025, T041 | 100 cancellation trials plus scheduler finalizer | `[~]` Bounded drain is unit-proven; required 100-trial production evidence is open. |
|
||||
| Success criteria | SC-003 | ScenarioRun.cancel_drain_deadline_at | scenarioRun.cancel | ScenarioExecution.Cancel | T015-T016, T025, T041, T045 | 100 cancellation trials plus scheduler finalizer | `[~]` Bounded drain unit-proven; live cancel/drain/terminalize proven in the T045 fault-injection canary (2026-09-17, 16/16: drain window, 0 running steps, 0 unexpired worker leases, capacity reconciled by TTL). Required 100-trial production evidence remains open. |
|
||||
| Success criteria | SC-004 | RunnerPlan, worker lease, operation receipt | scenarioRun.detail | ScenarioExecution.Runner.CrashRecovery, ScenarioExecution.ProviderOperations | T014c, T025, T030, T041 | safe/unsafe recovery and reconciliation tests | `[x]` (2026-09-12) Safe/unsafe persisted recovery verified; provider operation reconciliation + operation-aware cancellation implemented (`cancel_provider_operation`: stopped/completed/unknown, unknown→reconciliation_required блокирует retry/PASS; идемпотентный; terminal receipts immutable) — `test_provider_operations.py` (19), `test_provider_screenshot.py` (14), `test_live_execution_binding.py` (13). |
|
||||
| Success criteria | SC-005 | immutable run provenance | scenarioRun.result | Execution.RunnerPlan, Execution.Result | T017, T020, T041 | revision-edit immutability and evidence receipt tests | `[~]` Snapshot fields exist; complete receipt-level byte identity evidence is open. |
|
||||
| Success criteria | SC-006 | ActionApprovalGate, ExecutorRegistry | scenarioRun.start | Execution.ActionApprovalGate, ScenarioExecution.ExecutorRegistry | T014d, T019, T041 | PROD no-call and human-not-executor tests | `[x]` Current no-call and registry rejection tests pass. |
|
||||
| Success criteria | SC-007..010 | Provider protocol/health/capacity | — | ScenarioExecution.ProviderProtocol, ProviderOperations.Observability, CapacityManager | T028-T042 | common/provider-specific contract and deployment profiles | `[~]` New production gate partially closed. SC-010 closed (2026-09-12, T040): readiness snapshot with provider/version + capability fingerprint + redacted diagnostics. SC-008 closed (2026-09-13, T032): dispatcher tick reconciles expired leases; browser/screenshot heartbeat before loop submission with typed `*_CAPACITY_UNAVAILABLE` refusal (no I/O without a live lease); run-terminal/cancel session finalizer — `test_dispatch_capacity_lifecycle.py` (6), regression 697 passed. SC-009 closed (T039). SC-007 open: contract suites exist for all providers; T034 catalog rounds 3–4 (limits remainder, mutation fixture lease) pending. |
|
||||
| Success criteria | SC-011 | release evidence record | — | ScenarioExecution.ProviderProtocol | T022, T041-T042 | full profiles, PostgreSQL and semantic audit outputs | `[ ]` Release evidence package is incomplete. |
|
||||
| Success criteria | SC-011 | release evidence record | — | ScenarioExecution.ProviderProtocol | T022, T041-T042 | full profiles, PostgreSQL and semantic audit outputs | `[x]` (2026-09-17, T022 CLOSED) Release evidence package assembled: scoped 044 suite **1590 passed**; full backend **11269 passed / 0 failed**; frontend **3359 passed** + lint 0 errors + build green; fresh PostgreSQL 16 chain `0001→0025_drop_legacy_validation` + `alembic check` no-drift; Axiom audit `execution/providers` **0 module_too_long**; invariant-aware decomposition recorded in `plans/provider-decomposition-gate.md`. |
|
||||
| UX-2 Dispatch resilience | SCEX-FR-002/028 | ScenarioStepRun, DecisionPolicy | scenarioRun.step | Execution.Dispatch.Step, ScenarioExecution.OfflineExecutors.Assertion | T009-T011 | test_scenario_compare_to_baseline_typed_failure, test_scenario_queued_dispatch | `[x]` Field-run fix (2026-09-12, live ss-prod B01 run 6de8d0d9): a compiled `baseline_ref` expectation resolves typed inconclusive `BASELINE_EVIDENCE_UNAVAILABLE` (DecisionPolicy row 7 `COMPARISON_INCONCLUSIVE`), and any executor exception is contained at the dispatch step boundary as `EXECUTOR_STEP_ERROR`; the run terminalizes honestly, replay ticks are no-ops, and `QUEUED_DISPATCH_ERROR` stays infrastructure-only. 037 baseline-vs-pin comparison remains a later slice. |
|
||||
|
||||
N/A: Registry (042), Editor (043), Monitor UX (045), Automation (046), Analytics (047).
|
||||
@@ -42,48 +42,68 @@ N/A: Registry (042), Editor (043), Monitor UX (045), Automation (046), Analytics
|
||||
| no reference artifact authority | `ScenarioExecution.RunnerPlan.Derive` | `runner.plan.json` is diagnostic/reference only | `test_scenario_runner_plan`, crash-recovery tests |
|
||||
| authoring promotion admission | `ScenarioExecution.AuthoringAdmission` | 042 promoted revision -> 044 preflight | raw-artifact rejection, candidate-without-promotion, no-run sandbox tests |
|
||||
|
||||
All rows are required together with the 038/042/050 matrices. Provider,
|
||||
capacity, PostgreSQL, or live-composition gaps keep the coordinated release
|
||||
gate `NO-GO`, even when lifecycle unit tests pass.
|
||||
All rows are required together with the 038/042/050 matrices. As of 2026-09-17 the **provider,
|
||||
capacity, PostgreSQL and live-composition** gaps are closed for 044 (see the canary evidence index);
|
||||
the coordinated release gate stays `NO-GO` on the remaining rows: graph-level terminal PASS
|
||||
(T043/T046 evaluation→comparison binding + live rerun), T042 live Superset query + `RESULT_TOO_LARGE`
|
||||
bound, SC-003's 100-trial production evidence, and the external-MCP-only frontend boundary
|
||||
(050 CHK003/CHK004 with 043 T015/T023).
|
||||
|
||||
Authoring E2E is additionally `NO-GO` until persistent workspace exploration, typed proposal conversion, user diff review, 042 handle-based save, and 044 promoted-revision-only admission are evidenced. Sandbox security and code-backed provider readiness are separate gates and remain unimplemented unless explicitly proven.
|
||||
|
||||
**Provider production boundary:** Existing typed unavailable outcomes and exact Superset binding tests
|
||||
prove only the fail-closed boundary. Production readiness additionally requires T028-T042 and real
|
||||
startup/dependency health checks for every enabled live provider.
|
||||
**Provider production boundary (re-evaluated 2026-09-17):** the fail-closed boundary is now backed by
|
||||
production evidence: T028-T034, T039-T041 and T042b are CLOSED; real startup/dependency health checks
|
||||
and readiness preflight are implemented (T040, SC-010) and exercised live; unknown-effect quarantine
|
||||
and crash reconciliation are proven (T045). The remaining provider gap is T042's live Superset query
|
||||
with an explicit `RESULT_TOO_LARGE` bound, plus the T043/T046 graph-level terminal PASS that combines
|
||||
deterministic comparison with the bound evaluation.
|
||||
|
||||
**BrowserProvider boundary:** The BrowserProvider is contract-ready at 90/100 but implementation-ready
|
||||
only after T034, T040-T042 and T042b. A typed unavailable result remains the correct behavior until the
|
||||
PREPROD read-only/timeout/cleanup/reconstruction canaries produce retained operation and evidence receipts.
|
||||
**BrowserProvider boundary (re-evaluated 2026-09-17):** the BrowserProvider is contract-complete and
|
||||
implementation-ready: T034 (ownership/checkpoint/cancel/limits), T040 (startup registration + readiness
|
||||
preflight), T041 (common contract tests) and **T042b PREPROD canaries are CLOSED** — six vectors GREEN
|
||||
under the Wave-C run-scoped session architecture (2026-09-16, `evidence/browser-provider/
|
||||
readonly-canary-20260916T*.json`): read-only actions, forced timeout/cleanup, mutation row_edit+restore,
|
||||
reconciliation sweep, safe-checkpoint reconstruction, scheduler soak. T042 remains `[~]` for the **live
|
||||
Superset query + explicit `RESULT_TOO_LARGE` bound** only; a typed unavailable result is no longer the
|
||||
expected steady state. Fault injection for the lifecycle around the provider is proven by the T045
|
||||
canary (2026-09-17).
|
||||
|
||||
**Cross-spec production boundary:** 044 readiness depends on 036 authority/evidence, 037 query and
|
||||
baseline provenance, 038 executable graph identity, 041 lineage target state, 042 registry revisions,
|
||||
046 automation dispatch and 047 terminal-signal/case ingestion. 039/043/045 are operator continuity
|
||||
dependencies: they do not authorize execution, but incomplete typed API/SSE/revision flows prevent a
|
||||
complete production workflow. Current dependency-weighted aggregate for 036-047 is approximately
|
||||
**67/100** and remains NO-GO.
|
||||
complete production workflow.
|
||||
The 2026-08-24 dependency-weighted aggregate (≈67/100) predates the 2026-09-08 production rows and the
|
||||
2026-09-16/17 canary closures, so it is **stale as a score**; the numeric re-score belongs to the release
|
||||
review. The gate remains `NO-GO` until at minimum: 044 T043/T046 terminal PASS, T042 Superset live
|
||||
query, the 100-trial cancel evidence, the 042/043/045/046/047/050 production rows and the frontend
|
||||
external-MCP-only boundary are closed.
|
||||
|
||||
**Full production tool boundary:** There is no reduced preview target. Browser/Screenshot, controlled
|
||||
non-PROD mutation, AgentEvaluation/DecisionPolicy, automated schedules/triggers, case investigation,
|
||||
analytics and remediation are mandatory capability gates. Missing or unproven capability blocks GO; only
|
||||
human-containing automated revisions remain prohibited and PROD mutation remains policy-forbidden.
|
||||
|
||||
**Verification boundary (2026-08-24):** the available local profile is 246 backend tests, 59 provider/
|
||||
lifecycle edge tests, scoped Ruff/compile and prototype validation. Real PostgreSQL migration checks,
|
||||
provider contract T028-T042, live Browser/Screenshot composition, scheduler deployment and Axiom index
|
||||
rebuild remain open.
|
||||
**Verification boundary (re-evaluated 2026-09-17):** the available local profile is now backend
|
||||
**11269 passed / 252 skipped / 1 xpassed / 0 failed**, frontend **3359 passed** (208 files) + lint
|
||||
0 errors + production build, scoped 044 suite **1590 passed**, `validate_static.py` passed, ruff +
|
||||
compileall clean, **real PostgreSQL 16** `0001→0025_drop_legacy_validation` upgrade + `alembic check`
|
||||
(no drift, single head) and Axiom semantic rebuild (10890 contracts / 5383 edges; `execution/providers`
|
||||
0 `module_too_long`). Live evidence: browser canaries 6/6 (2026-09-16), artifact-content HTTP canary
|
||||
11/11 (2026-09-17), fault-injection canary 16/16 (2026-09-17). Still open: T043/T046 graph-level
|
||||
terminal PASS, T042 live Superset query + `RESULT_TOO_LARGE` bound.
|
||||
|
||||
## Production acceptance traceability — 2026-09-08
|
||||
|
||||
Historical rows above identify prior tests/code only; removed agent UI paths are retired. The following audited gates are **implemented=false / OPEN**, independent of local suite totals.
|
||||
Historical rows above identify prior tests/code only; removed agent UI paths are retired. Gates below are re-evaluated against executable evidence; a row is CLOSED only with the cited command/output or live-canary artifact.
|
||||
|
||||
| Requirement | Domain contract / DTO | Task | Falsifiable acceptance | State |
|
||||
|---|---|---|---|---|
|
||||
| SCEX-FR-028 | [Production baseline-backed evaluation](contracts/production-chain.md); [data model](data-model.md) | [T043](tasks.md) | Baseline set/version/catalog/release/commit/IDs/digests affect idempotency; moving catalog after admission cannot change plan/result; legacy unpinned result is ineligible. | OPEN |
|
||||
| SCEX-FR-028 | [Production baseline-backed evaluation](contracts/production-chain.md); [data model](data-model.md) | [T044](tasks.md) | GET/HEAD prove same ACL/status/headers; MIME/digest/length checked before bytes, cross-owner hidden, expired410, corrupt409, traversal/range/oversize rejected. | OPEN |
|
||||
| SCEX-FR-028 | [Production baseline-backed evaluation](contracts/production-chain.md); [data model](data-model.md) | [T045](tasks.md) | Startup/readiness/start-loop and shutdown/drain/cancel/reconcile survive fault injection; unknown effect quarantines capacity; late response cannot win. | OPEN |
|
||||
| SCEX-FR-028 | [Production baseline-backed evaluation](contracts/production-chain.md); [data model](data-model.md) | [T046](tasks.md) | End-to-end real browser→capture→durable artifact→deterministic comparison→optional immutable evaluation→policy→result; all required evidence present before PASS. | OPEN |
|
||||
| SCEX-FR-028; external-MCP-only UI | manual editor/review; read-only evidence | [production tasks](tasks.md) | No frontend agent prompt/chat/assistant editing/proposal generation/workspace/start/handoff routes or requests; human approval remains usable. | OPEN |
|
||||
| SCEX-FR-028 | [Production baseline-backed evaluation](contracts/production-chain.md); [data model](data-model.md) | [T043](tasks.md) | Baseline set/version/catalog/release/commit/IDs/digests affect idempotency; moving catalog after admission cannot change plan/result; legacy unpinned result is ineligible. | **PARTIAL** — request-hash pinning + fail-closed resolver + live pin-from-Gitea PROVEN (canary v4, `docs/reports/agentic-runtime-live-canary-v4-baseline-pin-2026-09-11.md`; 050 T045 publish CLOSED). Residual: graph-level deterministic comparison PASS (closed by the T046 live rerun below). |
|
||||
| SCEX-FR-028 | [Production baseline-backed evaluation](contracts/production-chain.md); [data model](data-model.md) | [T044](tasks.md) | GET/HEAD prove same ACL/status/headers; MIME/digest/length checked before bytes, cross-owner hidden, expired410, corrupt409, traversal/range/oversize rejected. | **CLOSED 2026-09-17** — live HTTP canary `prototype/artifact_content_canary.py` (real FastAPI app, real PostgreSQL `canary_044`, real JWTs, real live soak bytes 104746 B / sha `22b4b8c1…`) → **11/11 GREEN**, evidence `evidence/artifact-content/artifact-content-canary-20260917T085438Z.json`; offline matrix `tests/api/test_scenario_artifact_content_api.py` 11/11. |
|
||||
| SCEX-FR-028 | [Production baseline-backed evaluation](contracts/production-chain.md); [data model](data-model.md) | [T045](tasks.md) | Startup/readiness/start-loop and shutdown/drain/cancel/reconcile survive fault injection; unknown effect quarantines capacity; late response cannot win. | **CLOSED 2026-09-17** — fault-injection canary `prototype/fault_injection_canary.py` (real provider loop/capacity/receipt CAS/cancel lifecycle on real PostgreSQL) → **16/16 GREEN**, evidence `evidence/fault-injection/fault-injection-canary-20260917T105541Z.json`. Proven: typed bounded loop refusals; caller-deadline cancellation; shutdown drain without hang; crashed lease quarantines the slot until `reconcile_expired_leases`; `reconciliation_required`/unknown effect resolved by reconcile; late response history-only; cancel drains then terminalizes with 0 running steps / 0 unexpired worker leases and its capacity lease reconciled by TTL (SC-008). Offline: `test_provider_operations.py` (14), `test_scenario_cancel_timeout.py` (6), `test_dispatch_capacity_lifecycle.py`, `test_provider_capacity.py`, `test_provider_contract.py`, `test_provider_preflight.py`, `test_provider_runtime.py` — green. |
|
||||
| SCEX-FR-028 | [Production baseline-backed evaluation](contracts/production-chain.md); [data model](data-model.md) | [T046](tasks.md) | End-to-end real browser→capture→durable artifact→deterministic comparison→optional immutable evaluation→policy→result; all required evidence present before PASS. | **PARTIAL 2026-09-17 — root cause diagnosed with executable proof.** Chain real browser→capture→durable artifact→immutable evaluation→policy→result is PROVEN live; live baseline pin PROVEN (canary v4). Gap: the walker binds the evaluation per STEP (`policy_inputs_from_outcome` reads only `step_outcome["evaluation_input"]`, emitted solely by the `agent_evaluation` step), so an `assertion compare_to_baseline` step resolves to `EVALUATION_UNAVAILABLE` while mandatory mode is on; proven: compare-only+mandatory → `inconclusive [EVALUATION_UNAVAILABLE]`, +bound evaluation → `passed [BASELINE_AND_SEMANTIC_PASS]`. Closure = evaluation→comparison binding (two options in [tasks.md](tasks.md) T046) + live rerun of the v4 graph to terminal `passed`. |
|
||||
| SCEX-FR-028; external-MCP-only UI | manual editor/review; read-only evidence | [production tasks](tasks.md) | No frontend agent prompt/chat/assistant editing/proposal generation/workspace/start/handoff routes or requests; human approval remains usable. | OPEN (050 CHK003/CHK004 row; tracked with 043 T015/T023). |
|
||||
|
||||
Sources: [production gap](../../docs/reports/ss-prod-agentic-e2e-production-gap-2026-09-08.md), [coverage gap](../../docs/reports/ss-prod-agentic-e2e-spec-coverage-2026-09-08.md), [baseline gap](../../docs/reports/ss-prod-agentic-e2e-baseline-gap-2026-09-08.md). Spec schema/static checks prove contract structure only; live canary/runtime closure and optional approved performance baseline are not claimed.
|
||||
|
||||
@@ -100,3 +120,19 @@ surfaces outside the 050 MCP catalog (no spec row existed before this note):
|
||||
Cross-reference: the live external MCP replay that exercised the binding path end-to-end (and exposed
|
||||
the identity-less-compiled-step defect) is `docs/2026-09-11-sales-prod-mcp-replay.md` (050 T029m /
|
||||
`E2E-EXT-002`, CLOSED).
|
||||
|
||||
## Canary evidence index — 2026-09-16/17
|
||||
|
||||
| Canary / evidence file | Vectors | Task |
|
||||
|---|---|---|
|
||||
| `evidence/browser-provider/readonly-canary-20260916T*.json` (+PNG) | 6/6 GREEN under Wave-C architecture: read-only actions, forced timeout/cleanup, mutation row_edit+restore, reconciliation sweep, safe-checkpoint reconstruction, scheduler soak | T042b `[x]` |
|
||||
| `evidence/artifact-content/artifact-content-canary-20260917T085438Z.json` | 11/11 GREEN: GET/HEAD ACL+header parity, byte-exact body, anon 401, no-VIEW 403, unknown/foreign 404, expired 410, corrupt 409, bad-MIME 409, oversized 413, missing 409, Range 416 | T044 `[x]` |
|
||||
| `evidence/fault-injection/fault-injection-canary-20260917T105541Z.json` | 16/16 GREEN: loop refusals, deadline cancellation, shutdown drain, capacity quarantine + TTL reconcile, receipt reconciliation_required, late response history-only, cancel drain/terminalize | T045 `[x]` |
|
||||
|
||||
Canary harnesses (prototype, not production modules): `prototype/browser_readonly_canary.py`
|
||||
(`SS_CANARY_MODE=actions\|timeout\|mutation\|reconcile\|recovery\|soak`, `SS_CANARY_STORAGE_ROOT`),
|
||||
`prototype/artifact_content_canary.py`, `prototype/fault_injection_canary.py`. All canaries are
|
||||
fail-closed: they exit non-zero on any failed vector and never write credentials to evidence.
|
||||
`docs/reports/ux10-handoff-report-2026-09-14.md`, `docs/reports/agentic-runtime-*.md` and
|
||||
`docs/reports/agentic-runtime-live-canary-v4-baseline-pin-2026-09-11.md` retain the surrounding run
|
||||
narratives.
|
||||
|
||||
@@ -6,7 +6,7 @@ MCPX-FR-030: [Contract-complete public parity](../contracts/modules.md). Histori
|
||||
|
||||
- [ ] CHK001 REST/MCP lifecycle/read/auth errors and disabled automation validation are identical; service principal cannot decide human gate. Evidence: [T044](../tasks.md), [traceability](../traceability.md).
|
||||
- [x] CHK002 consume/publish failures return typed errors/pending state with no legacy fallback; every prerequisite is externally MCP-reachable. Evidence: [T045](../tasks.md) CLOSED 2026-09-11, [traceability](../traceability.md) (live gated MCP publish + moved-HEAD canary).
|
||||
- [ ] CHK003 Fresh external-client chain preserves authoritative context and complete baseline pin without raw ORM/REST repair; no frontend agent controls/routes/requests. Evidence: [T046](../tasks.md), [traceability](../traceability.md).
|
||||
- [~] CHK003 Fresh external-client chain preserves authoritative context and complete baseline pin without raw ORM/REST repair; no frontend agent controls/routes/requests. Evidence: [T046](../tasks.md), [traceability](../traceability.md). Partial 2026-09-17: the context/pin half is PROVEN live (T046 baseline-pin CLOSED 2026-09-11, live replay `docs/2026-09-11-sales-prod-mcp-replay.md`; `live_mcp_replay.py --baseline` resolves the full `runner_plan.baseline_pin`); the frontend-boundary half stays OPEN (T030/T032; 043 T015/T023) — tracked by CHK004.
|
||||
- [ ] CHK004 Negative product UI test: no agent chat/prompt/assistant editing/proposal-generation/typical-operation-to-agent/workspace/start/handoff controls or agent invocation routes/requests; manual CRUD/editor/human review/read-only results remain usable.
|
||||
|
||||
Schema/static success alone is not runtime completion. Optional approved performance baseline is outside scope.
|
||||
|
||||
@@ -14,9 +14,9 @@
|
||||
| Layer | Target (refresh) | Observed (do not treat as GO) |
|
||||
|---|---|---|
|
||||
| Surface | One `/mcp` Streamable HTTP server; external client owns reasoning | `/mcp` mounted; OAuth 2.1 + DCR local evidence exists |
|
||||
| Catalog | Versioned curated tools; `PINNED_CATALOG_MAJOR=2` | Runtime `MCP_CATALOG_VERSION="2.2.0"`; pin test `tests/test_mcp_catalog_version.py` |
|
||||
| Authoring | `inspect_dashboard_context` → `create_agent_run` → `register_draft_pack` → `bootstrap_authoring_scenario` → `activate_revision` / `start_scenario_run` | Vertical E2E exists with MCP-created AgentRun (T029i). Live-stand replay T029m / `E2E-EXT-002` OPEN |
|
||||
| Parity | REST/MCP identical validation/errors/CAS; no consume/publish fallback success | T044–T046 OPEN; consume fallback still a named production gap |
|
||||
| Catalog | Versioned curated tools; `PINNED_CATALOG_MAJOR=2` | Runtime `MCP_CATALOG_VERSION="2.3.0"` (61 entries); pin test `tests/test_mcp_catalog_version.py` |
|
||||
| Authoring | `inspect_dashboard_context` → `create_agent_run` → `register_draft_pack` → `bootstrap_authoring_scenario` → `activate_revision` / `start_scenario_run` | Vertical E2E exists with MCP-created AgentRun (T029i). Live-stand replay T029m / `E2E-EXT-002` **CLOSED 2026-09-11** (`docs/2026-09-11-sales-prod-mcp-replay.md`, run `110a6517…` + human loop) |
|
||||
| Parity | REST/MCP identical validation/errors/CAS; no consume/publish fallback success | T045/T046 **CLOSED 2026-09-11** (gated `publish_baseline_catalog` + live pin-from-Gitea canary v4); T044 offline parity green + live canaries done, remaining: dedicated REST-vs-MCP error-shape parity fixtures |
|
||||
| Frontend | No agent chat/prompt/workspace/start/handoff | Chat service removed; remaining agent-panel/handoff drift is T030 OPEN |
|
||||
|
||||
## Prerequisites
|
||||
@@ -93,4 +93,4 @@ Product UI never starts this chain. Manual editor (043) and monitor (045) are eq
|
||||
- No agent frontend controls, routes, or requests (T030; negative DOM/route/network still OPEN).
|
||||
- Live browser/capture canary PASSED 2026-09-10 against ss-prod (binding `ss-prod-d11-live-001`, run `597274d3`: browser `open_dashboard` with checkpoint/page-URL/PNG digest + 8 durable screenshot refs, capacity leases; trace: `docs/reports/agentic-runtime-live-canary-2026-09-10.md`). Live `agent_evaluation` canary v2 PASSED 2026-09-10 through REST-only start with server-side binding resolution (run `4eebfab3`: `AgentEvaluation` persisted, DecisionPolicy row 11; trace: `docs/reports/agentic-runtime-live-canary-v2-2026-09-10.md`). Still OPEN: multimodal image attachment (text-only prompt — visual verdicts not yet trustworthy), provider-version/cost receipts, cancellation receipts.
|
||||
|
||||
T029m live-stand replay of the 2026-09-07 sales scenario remains OPEN. MCPX-FR-030 T044–T046 remain OPEN. Optional approved performance baseline is out of scope.
|
||||
T029m live-stand replay of the 2026-09-07 sales scenario: **CLOSED 2026-09-11** (`docs/2026-09-11-sales-prod-mcp-replay.md`). MCPX-FR-030: T045/T046 **CLOSED 2026-09-11**; T044 remains OPEN only for the dedicated REST-vs-MCP error-shape parity fixtures. Optional approved performance baseline is out of scope.
|
||||
|
||||
@@ -22,9 +22,9 @@
|
||||
- [x] T011 Инструменты env/health/tasks (list_environments, get_health_summary, get_task_status, maintenance CRUD, llm status).
|
||||
- [x] T012 Инструменты git/deploy/migration/backup. Доказательство: `pytest tests/test_mcp_ops_parity.py tests/test_mcp_server.py tests/test_mcp_approvals.py tests/test_mcp_maintenance.py tests/test_mcp_oauth.py -q` → `70 passed`; зарегистрированы `create_branch`, `commit_changes`, `deploy_dashboard`, `execute_migration`, `run_backup`, `run_llm_documentation`, `run_llm_validation` (curated inputs, approval-gated; reviewed dispatchers `src/services/mcp_ops_dispatch.py` в явной цепи `poll_approved_mcp_dispatches`: GitService / TaskManager git-integration, superset-migration, superset-backup, llm_documentation, ValidationTaskService).
|
||||
- [x] T013 Superset-инструменты (databases/explore/sql/format/permissions/dashboard+dataset CRUD) — SQL-класс помечен risk-классом и отдельным правом. Доказательство: тот же срез `70 passed`; `superset_list_databases`, `superset_explore_database`, `superset_format_sql`, `superset_audit_permissions` (read), `superset_create_dashboard`, `superset_copy_dashboard`, `superset_create_dataset` (approval-gated), `superset_execute_sql` — permission `("plugin:superset_sql","EXECUTE")` добавлен в `rbac_permission_catalog.discover_declared_permissions` + default-deny mapping; dangerous SQL отклоняется клиентским guard. Доп. свидетельство (орт. аудит 2026-09-02): guard расширен (`safety.py`: `INTO/CALL/SET/REFRESH/VACUUM/REINDEX/ATTACH/LOAD/PREPARE/COMMENT` + опасные функции `lo_import/lo_export/pg_read_file/pg_write_file/pg_ls_dir/dblink/dblink_exec/set_config/pg_sleep` + запрет мульти-стейтментов через `;` после очистки строк/комментариев; `tests/test_core/test_superset_safety.py` — 50+ кейсов); PROD-критерий SQL-класса выравнен с `resolve_environment_execution_policy` (is_production OR stage=PROD) и переведён в терминальный отказ `production_sql_execution_rejected` вместо одобряемого, но недиспетчеризуемого гейта; `pytest tests/test_mcp_ops_parity.py tests/test_core/test_superset_safety.py -q` → `85 passed`; полный срез аудита `214 passed`; полный бекенд **11202 passed**.
|
||||
- [ ] T014 Baseline-домен: capture_baseline_candidate, request/decide/consume_baseline_approval, create_verification_run (паритет fixtures с tools_037). Доказательство: тот же срез `70 passed`; прямые обёртки над `candidate_capture`/`candidates.request|decide_approval`/`verification_service.create_verification_run_async` (те же сервисы, что и 037 REST-поверхность); `consume_baseline_approval` — baseline publish: approval-gated + reviewed dispatcher через one-shot `consume_approval`; полный backend suite `python -m pytest -q` → `11167 passed, 240 skipped, 1 xpassed`; ruff/compileall чистые; каталог 45/45 зарегистрирован.
|
||||
- [x] T014 Baseline-домен: capture_baseline_candidate, request/decide/consume_baseline_approval, create_verification_run (паритет fixtures с tools_037). Доказательство: тот же срез `70 passed`; прямые обёртки над `candidate_capture`/`candidates.request|decide_approval`/`verification_service.create_verification_run_async` (те же сервисы, что и 037 REST-поверхность); `consume_baseline_approval` — baseline publish: approval-gated + reviewed dispatcher через one-shot `consume_approval`; полный backend suite `python -m pytest -q` → `11167 passed, 240 skipped, 1 xpassed`; ruff/compileall чистые; каталог 45/45 зарегистрирован. **Re-verified 2026-09-16:** срез `pytest -q tests/test_mcp_ops_parity.py tests/test_mcp_server.py tests/test_mcp_approvals.py tests/test_mcp_maintenance.py tests/test_mcp_oauth.py tests/api/test_mcp_parity_baseline_037.py` → **84 passed**; каталог вырос до 61 записи (version 2.3.0), baseline-домен на месте; `consume_baseline_approval` дополнен gated `publish_baseline_catalog` (T045, CLOSED 2026-09-11).
|
||||
- [x] T015 Scenario-домен: `scenario_compile`/`scenario_validate`/`scenario_resolve`/`generate_draft_pack`/`request_save`/`activate_revision`/`start_scenario_run` (паритет с tools_038 и 042/044; save только из server-stored draft, activation отдельная CAS-операция).
|
||||
- [ ] T016 Контрактные тесты паритета: выводы MCP-инструментов сопоставлены с legacy-обёртками на общих фикстурах (SC-002). Доказательство: `pytest tests/api/test_mcp_parity_baseline_037.py -q` → `3 passed`: request/decide baseline-approval — MCP-gate валидируется моделью `ApprovalGateResponse` и совывает с REST по operation/risk_level/required_permission/status/reason_required/target_paths + actor parity на decide; `create_verification_run` — идентичные overall_status/category statuses/evidence_refs/created_by против REST `/verification-runs`; `superset_format_sql` — идентичная строка против legacy REST `/api/agent/superset/sqllab/format` (единственный дабл — [EXT:Superset] клиент). Попутно пойман и закрыт регрессией дефект 038-резолвера: `scenario_resolve` selector-путь падал на `step.description is None` (`TypeError`) — теперь `selector_hint` пишется без конкатенации с None (`test_selector_hint_on_step_without_description`).
|
||||
- [x] T016 Контрактные тесты паритета: выводы MCP-инструментов сопоставлены с legacy-обёртками на общих фикстурах (SC-002). Доказательство: `pytest tests/api/test_mcp_parity_baseline_037.py -q` → `3 passed` (re-run 2026-09-16 на HEAD `4746af2f` → `3 passed in 1.79s`): request/decide baseline-approval — MCP-gate валидируется моделью `ApprovalGateResponse` и совывает с REST по operation/risk_level/required_permission/status/reason_required/target_paths + actor parity на decide; `create_verification_run` — идентичные overall_status/category statuses/evidence_refs/created_by против REST `/verification-runs`; `superset_format_sql` — идентичная строка против legacy REST `/api/agent/superset/sqllab/format` (единственный дабл — [EXT:Superset] клиент). Попутно пойман и закрыт регрессией дефект 038-резолвера: `scenario_resolve` selector-путь падал на `step.description is None` (`TypeError`) — теперь `selector_hint` пишется без конкатенации с None (`test_selector_hint_on_step_without_description`).
|
||||
- [x] T017 Bounded-response дисциплина: лимит инлайн-ответа, артефакты как ref+digest.
|
||||
- [x] T018 Hidden/gated матрица: admin/analyst/viewer × каталог — отсутствие права скрывает инструмент из `tools/list`; role-change виден на следующем вызове без re-consent (SC-004, SC-009).
|
||||
|
||||
@@ -83,7 +83,7 @@
|
||||
|
||||
- [x] T040 Удалить сервис `agent/` из run.sh и compose-профилей (порт 7860); обновить AGENTS.md/INSTALL.md. Доказательство: `rg "agent|7860" run.sh build.sh docker-compose.yml docker-compose.enterprise-clean.yml` → пусто (кроме исторического секьюрити-комментария); `bash -n`/`yaml.safe_load` чистые; стенд `./run.sh --skip-install` поднимает только :8000+:5173 (7860 отсутствует, `curl /api/ready` → ready, Playwright-проход до хендоффа зелёный); `run.sh`/`build.sh`/`AGENTS.md`/`INSTALL.md` обновлены; из `build.sh` убраны `build:agent`, `bundle:agent`, `bundle:embeddings` и агентский сервис генерируемого деплой-композа/манифеста; из обоих nginx-конфигов убран `location /api/agent/gradio`.
|
||||
- [x] T041 Удалить код чата: agent/src (app, langgraph_setup, tools*.py, _confirmation, middleware...), frontend agent/assistant компоненты, i18n, типы. Доказательство: `agent/` удалён целиком; удалены `docker/Dockerfile.agent`, `docker/agent.entrypoint.sh`, `backend/tests/test_gradio_proxy_config.py`; во фронтенде удалены `components/assistant/*` (кроме универсального `MarkdownRenderer`), чат/ран/драфт артефакт компоненты, `components/agent/dashboard-testing/` (ScenarioWorkspace-ветка), `models/AgentChat*`, `AgentRunModel`, `DashboardScenarioWorkspaceModel`, `stores/assistantChat`, флаг `MCP_DECOMMISSION` (vite define/`config/mcp.ts`/`global.d.ts`), gradio-прокси из `vite.config.js`; `/agent` рендерит только `HandoffSurface` (безусловно), навбар-кнопка ведёт на хендофф; `npm run test -- --run` → **3454 passed**, `npm run lint` → 0 errors, `npm run build` → OK; полный бекенд **11199 passed** (минус тесты удалённого грдио-прокси), 16 сиротских тестовых файлов чата удалены.
|
||||
- [ ] T042 Финальные правки спек 036–047: перенести drift-amendments из статуса «planned» в «done» со ссылками на доказательства. Доказательство: во всех 12 спеках (036–047) секции `## Drift Amendment — MCP Interface` получили строку `**Status (2026-09-02): done**` со ссылками на `specs/050-mcp-interface/tasks.md` (T012–T028), T030–T033 и чекпоинты `specs/WORKSTATE-043-047.md`.
|
||||
- [x] T042 Финальные правки спек 036–047: перенести drift-amendments из статуса «planned» в «done» со ссылками на доказательства. Доказательство: во всех 12 спеках (036–047) секции `## Drift Amendment — MCP Interface` получили строку `**Status (2026-09-02): done**` со ссылками на `specs/050-mcp-interface/tasks.md` (T012–T028), T030–T033 и чекпоинты `specs/WORKSTATE-043-047.md`. **Re-verified 2026-09-16 (grep):** 11 spec-файлов несут секцию Drift Amendment; 040/041 — дословный маркер `Status (2026-09-02): done`; 036/037/038/042–047 — пост-closure-review формулировка «reported done (2026-09-02); not production-readiness evidence» с теми же ссылками (более строгая, принята как эквивалент).
|
||||
- [x] T043 Полный прогон backend/frontend suites + стенд без 7860 (SC-005). Доказательство: `python -m pytest -q` → **11199 passed, 240 skipped, 1 xpassed**; `npm run test -- --run` → **3454 passed** (197 файлов); `npm run lint` (0 errors) + `npm run build` — зелёные; локальный стенд после демонтажа работает без порта 7860 (см. T040) и отдаёт рабочий хендофф на `/agent`. **Browser E2E (изолированный стенд):** `docker compose -p ss-tools-e2e --env-file .env.e2e -f docker-compose.e2e.yml up -d --build db backend frontend` (свежая БД, bootstrap admin/admin123, backend :8103 healthy, frontend :8102 healthy, порт 7860 отсутствует) → `npx playwright test e2e/tests/login.e2e.js e2e/tests/agent.e2e.js` → **6 passed** (двойной прогон, Chromium): login-поток (форма/успех/ошибка неверных кредов) и post-decommission агентский роут (хендофф-поверхность, ноль чат-элементов, deep-link с context-параметрами остаётся на хендоффе); стенд снесён `down -v`. E2E-фикстарелы: устаревший ambiguous `locator('nav')` (strict mode: 3 nav-элемента) → `.first()` в login/smoke; regex ошибки входа дополнен `incorrect|неверн` (бэкенд отдаёт passthrough-detail); `agent.e2e.js` переписан под handoff-контракт (T01-T03), `dashboard-scenario-ui`/`agent-scenario-run` — ссылки на чат-UI заменены на handoff-маршрут. Дополнительно: живой admin-walkthrough на dev-стенде (8000/5173) подтвердил login→dashboards→handoff(0 textarea/0 conversation-узлов)→runs-center (скриншоты /tmp/kilo/happy-path/09–12).
|
||||
|
||||
## Production readiness — 2026-09-08 (MCPX-FR-030)
|
||||
|
||||
@@ -2366,3 +2366,265 @@ mcp_oauth.py/test_mcp_client_flow_http.py в working tree не модифици
|
||||
translate/migration integration-тесты и `_job_routes.py` (чужой незакоммиченный workstream),
|
||||
`.kilo/agent-manager.json` (live UI/recovery state), корневой снапшот `specs-036-050-20260907-111314.md`.
|
||||
T029m (`E2E-EXT-002`) — единственная OPEN gate-row.
|
||||
|
||||
## Checkpoint — 2026-09-16 (recovered re-baseline 2026-09-07 → HEAD `4746af2f`)
|
||||
|
||||
> Восстановленный чекпоинт: WORKSTATE не обновлялся 9 дней, факты ниже собраны из git log
|
||||
> (45 коммитов) и handoff-отчётов (`docs/reports/agentic-runtime-*.md`, `ux10-*.md`,
|
||||
> `docs/2026-09-11-sales-prod-mcp-replay.md`), а не из памяти. Suite-числа этого чекпоинта —
|
||||
> свежие прогоны на HEAD; исторические числа помечены датой источника.
|
||||
|
||||
### 2026-09-08..10 — offline agentic runtime chain + INV_7 + live canaries v1/v2
|
||||
|
||||
- `83727aa7` complete offline agentic runtime chain; INV_7-декомпозиция execution-пакета:
|
||||
`runner.py` 1207→121 (фасад; 6 модулей), `lifecycle.py` 611→299, `executors.py` 508→257;
|
||||
все модули < 400 LOC (`a21481ea`, `8f057f30`, `3e5cc942`).
|
||||
- D1 published-catalog source (fail-closed, 9 тестов) + D2 server-side live-binding resolution
|
||||
из `settings.scenario_live_execution_bindings` (12 тестов) — клиент binding не передаёт.
|
||||
- Live canary v1 (run `597274d3`: browser `open_dashboard` + 8 durable screenshots) и v2
|
||||
(REST-only, run `4eebfab3`: server-side binding → PROD gate → scheduler dispatch → browser →
|
||||
реальный LLM → `AgentEvaluation` persisted → DecisionPolicy row 11). Три production-дефекта
|
||||
найдены и закрыты живыми прогонами (preflight deadlock `ba2f1f45`, `artifact_byte_lengths`,
|
||||
нормализация provider-ответов `3458343d`).
|
||||
- Suite (2026-09-10, источник: evening handoff): **11527 passed / 0 failed / 244 skipped**.
|
||||
|
||||
### 2026-09-11 — T029m live replay + T045 publish + 046 scheduled happy path
|
||||
|
||||
- **E2E-EXT-002 CLOSED**: `live_mcp_replay.py` прогнал полную внешнюю цепочку на ss-prod
|
||||
(run `110a6517…`: inspect → create_agent_run → compile B01 → register
|
||||
`context_authority=verified` → bootstrap → PROD gate `pending_approval` (идемпотентный retry =
|
||||
один durable gate) → approval → live `capture_screenshot` passed (8 refs) → typed
|
||||
`BROWSER_ACTION_NOT_SUPPORTED` → честный `inconclusive`); human loop живьём через
|
||||
`live_mcp_human_loop.py` (B05 HumanCheckpoint → `waiting_human` → `list_checkpoints` →
|
||||
`decide_checkpoint confirm` (CAS) → terminal `passed`). Два fail-closed дефекта найдены и
|
||||
исправлены. Trace: `docs/2026-09-11-sales-prod-mcp-replay.md`.
|
||||
- **050 T045 CLOSED**: gated MCP `publish_baseline_catalog` + 037 publication worker (CAS,
|
||||
receipts, reconcile), REST parity; live pin-from-Gitea proven (canary v4: pin stamped into
|
||||
`AgentEvaluation.baseline_pin`, strict equality). Каталог 2.2.0 → 2.3.0.
|
||||
- **046 T018 CLOSED live**: scheduled runs execute as schedule-owner principal; deterministic
|
||||
scheduled-run idempotency key; typed scheduled-reject observability (`d0466a2c`, `b4148cfe`).
|
||||
- 044 T043/T046: live baseline pin PROVEN; residuals — graph-level deterministic comparison PASS
|
||||
(T043) и graph-level terminal PASS (T046, baseline-semantic policy `EVALUATION_UNAVAILABLE`
|
||||
для compare без bound evaluation).
|
||||
|
||||
### 2026-09-12..14 — UX-10 Wave C (закоммичено как `8522a2ee` 2026-09-15)
|
||||
|
||||
- DG-1 реализован: run-scoped browser session (`browser_session.py` 501), 9 read-only действий
|
||||
(`browser_readonly_actions.py` 534), `apply_native_filter` + safe checkpoint
|
||||
(`browser_native_filter.py` 275), admission split (`browser_admission.py` 208).
|
||||
- 046 retention: `deletions.py` (331) + расширение `retention.py` + миграция
|
||||
`0024_retention_deletions` (PG-verified; после инцидента с 33-char revision id добавлен guard
|
||||
`test_revision_ids_fit_varchar32`; правило: revision id ≤ 32, цепочку проверять на PostgreSQL).
|
||||
- Frontend: `TerminalReasonBanner` + typed terminal-reasons; `docs/mcp-client-setup.md` (UX-4).
|
||||
- 044 T040 `[x]` (deployment-evidence rule); T042 offline contract suites complete — остаток:
|
||||
live Superset query + явный `RESULT_TOO_LARGE` bound.
|
||||
- Suite (2026-09-14, источник: ux10 handoff, HEAD `b4148cfe` + dirty tree): backend
|
||||
**11796 passed / 264 skipped / 1 xpassed**; frontend **3585 passed / 214 файлов**;
|
||||
alembic head `0024_retention_deletions`.
|
||||
|
||||
### 2026-09-15..16
|
||||
|
||||
- `8522a2ee` закоммитил дерево UX-10 (94 файла, +11014/−589): MCP automation parity tests,
|
||||
terminal reasons, scenario UX flow.
|
||||
- `3de0756d` удалил legacy LLM dashboard validation (routes/service/schemas/
|
||||
DashboardValidationPlugin; миграция `0025_drop_legacy_validation`; frontend validation
|
||||
models/routes/i18n удалены; settings/automation редиректит на scenario automation).
|
||||
- `4746af2f` health consolidation review fixes.
|
||||
|
||||
### Свежие прогоны на HEAD `4746af2f` (этот чекпоинт)
|
||||
|
||||
- Backend full suite: **11269 passed, 252 skipped, 1 xpassed, 0 failed** (7:42). Дельта против
|
||||
11796 — удалённые legacy-validation тесты (`3de0756d`), не регрессии.
|
||||
- Frontend: **3359 passed / 208 файлов, 0 failed** (49s). Дельта против 3585 — удалённые
|
||||
validation frontend-тесты. `npm run lint` → **0 errors / 333 warnings** (baseline 340–364);
|
||||
`npm run build` (adapter-static) → green.
|
||||
- `ruff check .` clean; `compileall -q src` clean; `alembic heads` → единственная голова
|
||||
`0025_drop_legacy_validation`.
|
||||
- Axiom rebuild (incremental, live): 10870 contracts / 5365 edges; unresolved relations
|
||||
**393** (было 401 на re-review 2026-09-04); исправлены 3 malformed multi-target `@RELATION`
|
||||
в `backend/src/core/cot_logger.py` (split на individual lines с verified targets).
|
||||
- MCP slice (T014 re-verification): **84 passed**; каталог 61 запись, `MCP_CATALOG_VERSION=2.3.0`.
|
||||
- 037 visual slice (T044 re-verification): **61 passed**.
|
||||
|
||||
### Реконсиляция учёта (этот чекпоинт)
|
||||
|
||||
- 050 tasks.md: **T014, T016, T042 → `[x]`** (боксы с evidence-текстом, но не отмеченные;
|
||||
evidence перепроверено свежими прогонами/grep на HEAD).
|
||||
- 037 tasks.md: **T044 → `[x]`** (visual candidate реализован в `candidates.py` веткой
|
||||
`kind == "visual"` + mandatory human disposition гейтом; closure note «All 47 tasks completed»
|
||||
теперь соответствует факту для T001–T047; T082–T084 — отдельные production rows 2026-09-08).
|
||||
- 050 quickstart.md: устаревшие OPEN-строки исправлены (T029m/E2E-EXT-002 CLOSED 2026-09-11;
|
||||
T045/T046 CLOSED; каталог 2.3.0/61).
|
||||
- 050 CHK003 OPEN — **не противоречие**: строка покрывает и T046 (CLOSED), и frontend-boundary
|
||||
(T030 OPEN); оставлено OPEN корректно.
|
||||
- 044 T042b остаётся OPEN осознанно: канарейки 2026-09-01 предшествуют Wave-C session
|
||||
architecture (`8522a2ee`) и не являются evidence для текущего кода.
|
||||
|
||||
### Текущий фронт работ (актуализировано на HEAD)
|
||||
|
||||
1. **044**: T042b PREPROD-канарейки под Wave-C архитектуру; T044 live ACL/status/header canary;
|
||||
T045 fault-injection canary; T046 graph-level terminal PASS; T043 graph-level comparison PASS;
|
||||
T022 полный аудит; T042 live Superset query + `RESULT_TOO_LARGE` bound.
|
||||
2. **046**: T017 lifecycle notifications (callers отсутствуют); T021 5/15/50-tab canaries;
|
||||
T013e бокс vs landed retention code (сверить); T013 формально открыт при покрытии.
|
||||
3. **050**: T044 (REST-vs-MCP error-shape parity fixtures — последний OPEN production gate);
|
||||
CHK004 negative UI test; T029a; T030/T032 — продуктовое решение.
|
||||
4. **047**: T017–T021 (SCAN-FR-015 atomic triage). **042**: T024–T032. **043**: T015/T023,
|
||||
T021–T022. **045**: T020–T022. **037**: T082–T084. **038**: T041 (VLM), T060–T062.
|
||||
5. Skill drift: `.agents/skills/semantics-python` и `semantics-svelte` ссылались на удалённый
|
||||
`ss_tools.shared.cot_logger` (ADR-0022 absorb) — **исправлено в этом чекпоинте**: facade
|
||||
`src.core.logger` (intent-first), `notify()` из `$lib/toasts.svelte.ts` вместо несуществующего
|
||||
`addToast()`, i18n dictionary-proxy стиль (`$derived($t.migration ?? {})`) вместо
|
||||
function-call `$t("key")`; `./scripts/sync-skills.sh` прогнан, `.kilo` копии byte-identical.
|
||||
|
||||
Product status: **NO-GO** — production-acceptance rows 2026-09-08 открыты во всех спеках;
|
||||
формальный closure re-sign-off не проводился.
|
||||
|
||||
## Checkpoint — 2026-09-16 (T042b CLOSED: шесть canary-векторов GREEN под Wave-C архитектурой)
|
||||
|
||||
- Все шесть векторов T042b перепроверены живьём на owner-authorized тестовом стенде
|
||||
(`https://ss-prod.bebesh.ru`, `SS_STAND_STAGE=PREPROD`, dashboard 11) через реальную
|
||||
продуктовую цепочку Wave-C (run-scoped browser session, capacity admission, receipts,
|
||||
durable evidence). Evidence в `specs/044-dashboard-scenario-execution/evidence/browser-provider/`
|
||||
(`readonly-canary-20260916T*.json` + PNG; креды в evidence не пишутся):
|
||||
| Вектор | Результат | Evidence |
|
||||
|---|---|---|
|
||||
| read-only actions | 3/3 passed (open_dashboard 7.2s, wait_for_state 3.9s, refresh 4.3s) | `150158Z` |
|
||||
| forced timeout/cleanup | typed `BROWSER_ACTION_TIMEOUT`, 0 refs, lease released | `150407Z` |
|
||||
| mutation row_edit+restore | 82.74→mutate→restore verified, 2/2 receipts `completed` | `150518Z` |
|
||||
| reconciliation sweep | stale receipt resolved через live SELECT-only observation | `150754Z` |
|
||||
| safe-checkpoint reconstruction | attempt 2 `passed`, `reconstruction_replay=true` | `150908Z` |
|
||||
| scheduler soak 75s | 3/3 runs `passed`, `attempts==1` (CAS), 0 active leases, 3 artifacts | `155013Z` |
|
||||
- **Harness defect найден и исправлен**: в bare-script контексте DI-синглтон `SchedulerService`
|
||||
захватывал мёртвый event loop (`asyncio.new_event_loop()` без runner), из-за чего async-job'ы
|
||||
(`maintenance_auto_end` → `AsyncJobRunner.run`) блокировались на 300s safety cap, а `stop()`
|
||||
ждал их — первый soak-прогон завис. Фикс в
|
||||
`specs/044-dashboard-scenario-execution/prototype/browser_readonly_canary.py`: soak-окно
|
||||
выполняется внутри `asyncio.run` (singleton создаётся при работающем loop), `scheduler.stop()`
|
||||
вынесен через `asyncio.to_thread` (loop свободен для drain in-flight coroutines). Продуктовый
|
||||
код не менялся; в production loop FastAPI всегда работает, дефект — артефакт harness'а, но
|
||||
он же демонстрирует задокументированный drain-first компромисс (job занимает слот до cap).
|
||||
- PG-канареечная БД `canary_044` создана в локальном postgres (recovery/soak/reconcile-режимы);
|
||||
read-only/timeout/mutation используют temp SQLite по умолчанию.
|
||||
- 044 tasks.md: **T042b → `[x]`** с полным evidence.
|
||||
|
||||
## Checkpoint — 2026-09-16 (T022 audit: `[~]` — gates зелёные, открыт INV_7 хвост Wave-C)
|
||||
|
||||
- T022 команды выполнены (см. текст задачи): scoped 044 suite **1590 passed**; full backend
|
||||
**11269 / 0 failed**; frontend **3359 passed** + lint 0 errors / 333 warnings + build green;
|
||||
свежий PostgreSQL 16 `0001→0025_drop_legacy_validation` + `alembic check` без drift
|
||||
(проверочная БД `alembic_044`); Axiom live rebuild — 10870 contracts / 5367 edges,
|
||||
unresolved 393.
|
||||
- **НЕ `[x]`**: Axiom-аудит `execution/` показал 8 structural warnings — Wave-C модули выросли
|
||||
за INV_7: `providers/browser_readonly_actions.py` **534**, `providers/browser_session.py`
|
||||
**501**, `providers/browser.py` **407** (регрессия против 398 из round 4). Требуется
|
||||
decomposition-gate по прецеденту server.py (`specs/050-mcp-interface/plans/
|
||||
server-decomposition-gate.md`): binding plan + behavior-neutral split с frozen contract IDs.
|
||||
- **T044 (044): offline evidence matrix уже зелёная** — `test_scenario_artifact_content_api.py`
|
||||
11/11 (GET/HEAD same-headers, 401/403/404-cross-owner, digest/mime 409 без байт, inactive 410,
|
||||
range 416, oversized 413, storage 503); остаток — live ACL/status/header canary на стенде.
|
||||
- Skill drift закрыт (semantics-python/svelte → facade `src.core.logger`, `notify()` из
|
||||
`$lib/toasts.svelte.ts`, dictionary-proxy i18n); sync-skills прогнан; malformed multi-target
|
||||
`@RELATION` в `cot_logger.py` разбиты (3 линии) — cot_logger.py 0 unresolved.
|
||||
- Ruff clean на всём backend; compileall clean; anchors сбалансированы во всех touched-файлах
|
||||
(pre-existing внутри-комментария `#region` в semantics-svelte line 86 — текст, не якорь).
|
||||
|
||||
## Checkpoint — 2026-09-17 (Wave 1.6 EXECUTED: provider decomposition gate → T042b+T022 закрыты)
|
||||
|
||||
- **Wave 1.6 исполнен** по binding-плану `specs/044-dashboard-scenario-execution/plans/
|
||||
provider-decomposition-gate.md` (статус PLAN → EXECUTED, полный execution log внутри).
|
||||
Три gated фазы, нулевой behavior diff (per-phase scoped 1590 + финальный полный suite):
|
||||
| Модуль | Было | Стало | Новые sibling-модули |
|
||||
|---|---:|---:|---|
|
||||
| `browser_readonly_actions.py` | 534 | **160** | `browser_readonly_limits.py` 82, `flows_nav` 199, `flows_interact` 192 |
|
||||
| `browser_session.py` | 501 | **58** | `browser_session_handle.py` 64, `browser_session_checkpoint.py` 162, `browser_session_managers.py` 57, `browser_session_registry.py` 287 |
|
||||
| `browser.py` | 407 | **398** | `browser_factory_helpers.py` 35 |
|
||||
- Frozen contract IDs + import surface сохранены: фасады ре-экспортируют все перенесённые
|
||||
публичные имена; monkeypatch-сеамы последовали за владеющими модулями (`_MAX_EXTRACT_OUTPUT_BYTES`
|
||||
→ flows_nav, `_MAX_DOWNLOAD_BYTES` → flows_interact); `_register_manager` вынесен в
|
||||
`browser_session_managers.py` (иначе registry⇄facade цикл). Все переносы verbatim.
|
||||
- Гейты: per-phase scoped **1590 passed** ×3; финальный полный backend suite
|
||||
**11269 passed / 252 skipped / 1 xpassed / 0 failed** (7:10) — нулевая дельта против базовой
|
||||
11269 до декомпозиции; ruff + compileall clean; anchors сбалансированы во всех 9 модулях.
|
||||
- Post-decomposition Axiom `audit_contracts` (`execution/providers`): **0 module_too_long**
|
||||
(было 3). Принятые advisory: 2 × `contract_too_long` (`BrowserProvider.Factory` 305,
|
||||
`ScreenshotProvider.Factory` 200) — typed-лестницы исключений оставлены inline осознанно.
|
||||
- **044 tasks.md: T042b `[x]` (2026-09-16, шесть векторов GREEN) и T022 `[x]` (2026-09-17,
|
||||
zero P0/P1)**. Оба production-критичных бокса 044 закрыты с evidence.
|
||||
- Skill drift закрыт окончательно: заголовок примера semantics-svelte ссылался на удалённый
|
||||
`notificationStore` — заменён на `$lib/toasts`; sync-skills прогнан.
|
||||
|
||||
## Checkpoint — 2026-09-17 (Wave 1.2/1.3 закрыты; T046 диагностирован с исполняемым proof)
|
||||
|
||||
### T044 CLOSED `[x]` — live ACL/status/header canary (11/11 GREEN)
|
||||
|
||||
- Новый harness `specs/044-dashboard-scenario-execution/prototype/artifact_content_canary.py`
|
||||
(351 LOC, 5 anchors) гоняет РЕАЛЬНОЕ FastAPI-приложение (без dependency overrides) поверх
|
||||
реального PostgreSQL (`canary_044`) с реальными JWT (реальный `create_access_token`, реальный
|
||||
поиск user→role→permission; без `sid` — session-policy пропускается) и РЕАЛЬНЫМИ live-байтами
|
||||
браузерного evidence (soak-прогон `soak-canary-c071fb98`, PNG 104746 B, sha `22b4b8c1…`).
|
||||
- Evidence: `evidence/artifact-content/artifact-content-canary-20260917T085438Z.json` — 11/11:
|
||||
anonymous→401 `AUTHENTICATION_REQUIRED` (+HEAD пустой); viewer GET→200 byte-exact + согласованные
|
||||
Content-Length/ETag/Disposition/Cache-Control/nosniff/Accept-Ranges; viewer HEAD→200 полный
|
||||
header-parity + пустое тело; no-VIEW→403 `PERMISSION_DENIED`; unknown run / unknown artifact /
|
||||
foreign-owned → неразличимые 404 `NOT_FOUND` с идентичным message; expired→410;
|
||||
corrupt sha→409 `ARTIFACT_INTEGRITY_FAILED`; declared-MIME off-allowlist→409; oversized→413;
|
||||
missing bytes→409 `ARTIFACT_MISSING`; Range→416 `RANGE_NOT_SUPPORTED` + `Accept-Ranges: none`.
|
||||
- Harness-инфраструктура: добавлен knob `SS_CANARY_STORAGE_ROOT` (персистентный evidence-root;
|
||||
ВАЖНО: выделенный root на прогон — soak-ассерт считает файлы в root). Для HTTP-канарейки нужен
|
||||
`STORAGE_ROOT_PATH` из одобренных корней (`/app/storage`, проект, `../ss-tools-storage`) —
|
||||
использован `/home/busya/dev/ss-tools-storage/canary-artifact-acl`.
|
||||
- Инциденты harness'а (исправлены): (1) повторный прогон брал собственный oversized-clone как
|
||||
«good»-row → фикстуры-клоны теперь именуются `canary-*` и удаляются в начале прогона;
|
||||
(2) soak с переиспользованным root дал `failures: ["artifacts=6"]` → выделенный root.
|
||||
|
||||
### T045 CLOSED `[x]` — fault-injection canary (16/16 GREEN)
|
||||
|
||||
- Новый harness `specs/044-dashboard-scenario-execution/prototype/fault_injection_canary.py`
|
||||
(365 LOC, 7 anchors) поверх реального PostgreSQL (`canary_044`): реальный provider runtime,
|
||||
capacity manager, receipt CAS, cancel lifecycle. Evidence:
|
||||
`evidence/fault-injection/fault-injection-canary-20260917T105541Z.json`.
|
||||
- Векторы: submit до старта/истёкший deadline → типизированные `PROVIDER_LOOP_NOT_RUNNING` /
|
||||
`PROVIDER_SUBMIT_DEADLINE` за bounded время; медленная корутина отменяется по caller-deadline;
|
||||
shutdown во время in-flight работы разворачивает вызывающего по его же deadline, loop → not-running
|
||||
(без зависания); **crash (lease не освобождён) карантинит слот** (второй claim →
|
||||
`CAPACITY_UNAVAILABLE`) до `reconcile_expired_leases`, который освобождает ровно эти units;
|
||||
unknown effect → receipt `reconciliation_required`/`unknown` → `reconcile_provider_operation`
|
||||
разрешает; late response добавляет только history (терминальный статус неизменен), reconcile по
|
||||
терминальному receipt отвергнут (`PROVIDER_OPERATION_TERMINAL`); cancel открывает drain-окно
|
||||
(`draining`), дедлайн-финализатор терминализирует (0 running steps, 0 unexpired worker leases),
|
||||
capacity-lease отменённого прогона **согласуется по TTL, не исчезает молча**; немедленный cancel
|
||||
терминализирует одним проходом.
|
||||
- Найдено при отладке: `cancel_run` экспайрит **worker**-leases (`ScenarioStepLease`), а не
|
||||
provider capacity-leases (`CapacityLease`) — последние освобождает provider в `finally` или
|
||||
TTL-reconcile (совпадает с семантикой «released or reconciled» из SC-008). Первый вариант
|
||||
ассерта это смешивал; исправлено.
|
||||
- Offline-референсы (зелёные): `test_provider_operations.py` (CAS/late/reconcile),
|
||||
`test_scenario_cancel_timeout.py` (6), `test_dispatch_capacity_lifecycle.py`,
|
||||
`test_provider_capacity.py`, `test_provider_contract.py`, `test_provider_preflight.py`,
|
||||
`test_provider_runtime.py` — 41 + 47 passed.
|
||||
|
||||
### T046 `[~]` — диагноз с исполняемым proof (binding, не truth table)
|
||||
|
||||
- Остаток T046 — **walker-level binding**, а не ошибка политики: `walker.py` считает
|
||||
`decide_step_outcome(policy_inputs_from_outcome(...))` ПОШАГОВО, и `evaluation` берётся только из
|
||||
`step_outcome["evaluation_input"]`, который эмитит единственный шаг `agent_evaluation`
|
||||
(`evaluation_adapter.py:379`). Шаг `assertion compare_to_baseline` всегда видит `evaluation=None`;
|
||||
`derive_runner_plan` включает mandatory-режим при наличии `agent_evaluation` шага → каждый
|
||||
нормативный шаг (включая compare) резолвится в `EVALUATION_UNAVAILABLE`.
|
||||
- Proof на HEAD (pure function, без БД): (A) compare-only + mandatory → `inconclusive
|
||||
["EVALUATION_UNAVAILABLE"]`; (B) тот же + disabled → `passed ["BASELINE_PASS"]`; (C) тот же +
|
||||
bound `evaluation_input` (succeeded/pass/0.9) + mandatory → `passed ["BASELINE_AND_SEMANTIC_PASS"]`.
|
||||
- Канон (`contracts/production-chain.md` §8) ставит optional evaluation МЕЖДУ comparison и pinned
|
||||
DecisionPolicy → решение обязано видеть оба; сейчас оно считается изолированно по шагу.
|
||||
- Опции закрытия (следующий пакет, truth table не меняется): (1) dependency-ordered binding — compare
|
||||
зависит от покрывающего `evaluate-visual`, walker инжектит persisted `AgentEvaluation`
|
||||
(`agent_evaluation_ids`) как `evaluation_input` для зависимого решения; (2) deferred decision —
|
||||
walker откладывает решение comparison-шагов до персиста покрывающей оценки и считает одно решение
|
||||
на агрегированных входах. Закрытие = live-рерun v4-графа с терминальным `passed`.
|
||||
|
||||
### Открытые строки 044 после этого чекпоинта
|
||||
|
||||
- **T043** `[ ]`: residual — graph-level deterministic comparison PASS (050 T045 publish закрыт).
|
||||
- **T046** `[~]`: binding (см. выше) + live terminal PASS.
|
||||
- Ранее закрыты: T042b `[x]`, T044 `[x]`, T045 `[x]`, T022 `[x]`, T040 `[x]`.
|
||||
|
||||
218
specs/production-043-047-2026-09-16-plan.md
Normal file
@@ -0,0 +1,218 @@
|
||||
# План работ для воркера — production closure 037/038/042–047/050 (2026-09-16)
|
||||
|
||||
> Источник состояния: `specs/WORKSTATE-043-047.md` (последний чекпоинт **2026-09-07**) + спеки
|
||||
> 037/038/042–047/050 + git log до HEAD `4746af2f` (**2026-09-16**).
|
||||
> WORKSTATE отстаёт от HEAD на ~9 дней — Wave 0 обязателен до любой реализации.
|
||||
|
||||
## Статус выполнения (2026-09-16, эта сессия)
|
||||
|
||||
- **Wave 0 — DONE**: recovered checkpoint в WORKSTATE (45 коммитов 2026-09-07..HEAD, факты из
|
||||
git log + handoff-отчётов); боксы реконсилированы (050 T014/T016/T042 → `[x]` со свежими
|
||||
прогонами; 037 T044 → `[x]`); противоречия учёта сняты (050 quickstart.md; CHK003 признан
|
||||
корректно OPEN — покрывает и T046 CLOSED, и T030 OPEN; 037 closure note стал соответствовать
|
||||
факту); suites на HEAD: backend **11269 / 0 failed**, frontend **3359**, lint 0 errors / 333,
|
||||
build green, ruff/compileall clean, alembic head `0025_drop_legacy_validation`; Axiom live
|
||||
rebuild (10870 contracts / 5367 edges / 393 unresolved); skill drift (semantics-python/svelte)
|
||||
исправлен + synced; 3 malformed multi-target `@RELATION` в cot_logger.py разбиты.
|
||||
- **Wave 1.1 — DONE**: 044 **T042b → `[x]`** — все шесть canary-векторов GREEN под Wave-C
|
||||
архитектурой (actions/timeout/mutation/reconcile/recovery/soak), evidence в
|
||||
`specs/044-.../evidence/browser-provider/readonly-canary-20260916T*.json`; harness-дефект
|
||||
(dead-loop scheduler hang) найден и исправлен.
|
||||
- **Wave 1.5 — DONE (T022 `[x]`)**: все гейты зелёные (scoped 1590, PG chain 0001→0025 + no drift,
|
||||
ruff, rebuild) + **Wave 1.6 исполнен**: INV_7 хвост Wave-C разложен по binding-плану
|
||||
(`specs/044-dashboard-scenario-execution/plans/provider-decomposition-gate.md`, EXECUTED):
|
||||
`browser_readonly_actions` 534→160, `browser_session` 501→58, `browser.py` 407→398, 5 новых
|
||||
sibling-модулей, frozen IDs/import surface, per-phase gates 1590 ×3, полный suite
|
||||
**11269 passed / 0 failed** (нулевая дельта), Axiom audit: **0 module_too_long**.
|
||||
- **Wave 1.6 — DONE (2026-09-17)**: INV_7 decomposition executed (см. статус выше).
|
||||
- **Wave 1.2 — DONE (2026-09-17)**: 044 **T044 → `[x]`** — live ACL/status/header canary 11/11 GREEN
|
||||
(реальное приложение + реальные JWT + живые байты soak-прогона; harness
|
||||
`prototype/artifact_content_canary.py`, evidence `evidence/artifact-content/…085438Z.json`).
|
||||
- **Wave 1.3 — DONE (2026-09-17)**: 044 **T045 → `[x]`** — fault-injection canary 16/16 GREEN
|
||||
(loop/deadline/shutdown-drain/quarantine+reconcile/receipt CAS/late-response/cancel-drain;
|
||||
harness `prototype/fault_injection_canary.py`, evidence `evidence/fault-injection/…105541Z.json`).
|
||||
- **Wave 1.4 — DIAGNOSED (T046 `[~]`)**: root cause = walker-level evaluation→comparison binding
|
||||
(не truth table). Executable proof: compare-only+mandatory → `EVALUATION_UNAVAILABLE`;
|
||||
+bound evaluation → `BASELINE_AND_SEMANTIC_PASS`. Две опции закрытия зафиксированы в tasks.md;
|
||||
требуется реализация + live-рерun v4-графа с терминальным `passed`.
|
||||
- **T043** `[ ]`: residual — graph-level deterministic comparison PASS.
|
||||
- Волны 2–6 не начинались.
|
||||
|
||||
## Правила воркера (house discipline)
|
||||
|
||||
- `[x]` в tasks.md ставится **только** с текущим command-level доказательством (команда + вывод);
|
||||
старые агрегаты не переиспользовать. `[~]` — частично; `[ ]` — отсутствует.
|
||||
- Не заявлять production-ready на основе наличия кода или старых счётчиков.
|
||||
- GRACE-Poly: якоря сбалансированы в каждом touched-файле; INV_7 < 400 LOC/модуль (новый код —
|
||||
в новые модули); MCL-логирование только intent-first фасадом (`logger.X("intent", src=_SRC, ...)`).
|
||||
- Коммиты — только с явного одобрения оператора; до этого всё в working tree.
|
||||
- Внешний стенд `https://ss-prod.bebesh.ru` — TEST-классификация, owner-authorized,
|
||||
`SS_STAND_STAGE=PREPROD`; PROD-мутации запрещены.
|
||||
- E2E compose — только изолированным проектом: `docker compose -p ss-tools-e2e --env-file .env.e2e
|
||||
-f docker-compose.e2e.yml ...` (footgun 2026-09-02: без `-p` цепляется к DEV-базе `run.sh`).
|
||||
|
||||
## Волны работ
|
||||
|
||||
### Wave 0 — Re-baseline (ОБЯЗАТЕЛЬНО ПЕРВЫМ)
|
||||
|
||||
Цель: привести учёт в соответствие с HEAD, восстановить зелёный базовый suite.
|
||||
|
||||
1. Пройтись по `git log --stat 731aaaa8..HEAD` (или от последнего коммита, упомянутого в WORKSTATE)
|
||||
и дописать в `specs/WORKSTATE-043-047.md` сводный чекпоинт «2026-09-07..2026-09-16 (recovered from
|
||||
git)» с фактами из коммитов, а не из памяти: 050 T045/T029m (E2E-EXT-002 CLOSED 2026-09-11),
|
||||
046 T018/T019/T020, canary v4 baseline pin (044 T043/T046 progress), UX-10 Wave C
|
||||
(`8522a2ee`: browser_session/native_filter/readonly_actions, retention deletions, migration 0024),
|
||||
legacy validation removal (`3de0756d`, migration 0025), health consolidation (`4746af2f`).
|
||||
2. Реконсилировать чекбоксы, у которых evidence-текст уже есть, а бокс открыт (проверить каждый
|
||||
против текущего дерева, не доверять тексту слепо):
|
||||
- 050: T014, T016, T042 (evidence в боксе, `[ ]`);
|
||||
- 044: T042b (требования переписаны в `8522a2ee` под новую session-архитектуру — читать
|
||||
текущий текст бокса как source of truth), T043–T046 (status-ноты с явными residuals);
|
||||
- 046: T013e/T021 (retention.py/deletions.py/migration 0024 уже landed — проверить, что именно
|
||||
из бокса закрыто, остаток переформулировать).
|
||||
3. Исправить известные противоречия учёта:
|
||||
- 050 `traceability.md` CHK003 OPEN vs T046 CLOSED (выровнять);
|
||||
- 050 `quickstart.md` — устаревшие OPEN-строки (T029m/E2E-EXT-002/T045/T046);
|
||||
- 037 tasks.md «All 47 tasks completed» vs открытый T044.
|
||||
4. Прогнать полные suite на HEAD: backend `cd backend && source .venv/bin/activate &&
|
||||
python -m pytest -q` (ожидание: уровень 11357+), frontend `npm run test -- --run` /
|
||||
`npm run lint` / `npm run build`; `ruff check .`, `compileall -q src`, `alembic heads`
|
||||
(ожидание: `0025_drop_legacy_validation` или новее). Зафиксировать числа в WORKSTATE.
|
||||
5. Semantic rebuild через Axiom MCP (`rebuild` → `status`), при недоступности бинарников —
|
||||
zombie-mode sweeps как в раундах 1–5; `make docs-nav` для L0-дайджеста.
|
||||
|
||||
Gate выхода из Wave 0: полные suite зелёные на HEAD, WORKSTATE актуален, противоречия учёта сняты.
|
||||
|
||||
### Wave 1 — 044 execution residuals (критический путь product-GO)
|
||||
|
||||
1. **T042b — DONE 2026-09-16** (см. статус выше): все шесть векторов GREEN под Wave-C.
|
||||
2. **T044 residual**: live ACL/status/header canary для `scenario_artifact_content`
|
||||
(GET/HEAD: same ACL/status/headers; MIME/digest/length до байт; cross-owner hidden;
|
||||
expired 410 / corrupt 409 / traversal/range/oversize rejected). Offline matrix уже
|
||||
зелёная (11/11 в `test_scenario_artifact_content_api.py`) — нужен HTTP-level прогон
|
||||
на запущенном backend против реально персистнутых evidence-байтов (например, из
|
||||
`canary_044` soak/recovery прогонов, storage root сохраняется).
|
||||
3. **T045 residual**: fault-injection canary — drain/cancel/reconcile при провайдерных сбоях;
|
||||
unknown effect quarantine для capacity; late response не может выиграть CAS.
|
||||
4. **T046 residual**: graph-level terminal PASS — связать `compare_to_baseline` с bound
|
||||
`evaluate-visual` в одном live-графе (расширение `live_mcp_replay.py`/canary v4), чтобы
|
||||
baseline-semantic policy не выдавала `EVALUATION_UNAVAILABLE`; evidence — canary report
|
||||
в `docs/reports/`.
|
||||
5. **T022 — `[~]`**: гейты зелёные (scoped 1590, PG 0001→0025 no-drift, ruff, Axiom rebuild);
|
||||
до `[x]` — закрыть INV_7 хвост.
|
||||
6. **НОВЫЙ пункт: INV_7 decomposition gate** (блокирует T022 `[x]`): `browser_readonly_actions.py`
|
||||
534, `browser_session.py` 501, `browser.py` 407. По прецеденту server.py: binding plan
|
||||
(`specs/044-.../plans/provider-decomposition-gate.md`), behavior-neutral split, frozen
|
||||
contract IDs, per-phase verification (scoped suite 1590 + anchors), план → EXECUTED.
|
||||
|
||||
Gate: все шесть canary-векторов GREEN на стенде + T022 audit; боксы T042b/T043–T046/T022
|
||||
обновлены с evidence.
|
||||
|
||||
### Wave 2 — 047 analytics atomic triage (SCAN-FR-012/015)
|
||||
|
||||
1. **T017**: заменить безусловный `resolved` в `set_disposition()` на CAS state machine:
|
||||
`resolved` требует verification evidence/reconciliation, `accepted` требует rationale;
|
||||
тест recurrence reopening. Контракт: `contracts/atomic-triage.md` (сейчас implemented=false).
|
||||
2. **T019 (DEF-04)**: accept/resolve атомарно закрывает actionable queue + episode; stale
|
||||
case/queue CAS откатывает всё.
|
||||
3. **T020**: конкурентные signal/disposition/open replay → ровно один активный episode/case;
|
||||
retained run/evaluation evidence не мутирует.
|
||||
4. **T021**: same environment, но разные baseline family/release/RLS/principal — несравнимы;
|
||||
hardcoded eligible sequence разделяет regression/flaky/model variance.
|
||||
5. **T018**: независимая end-to-end verification аналитики; обновить T013 только текущим evidence.
|
||||
|
||||
Командный срез из AGENTS.md: `pytest -q tests/services/dashboard_testing/registry/test_scenario_*.py
|
||||
tests/api/test_scenario_analytics_api.py` + новые файлы; PostgreSQL-конкурентность —
|
||||
`--run-integration` (Testcontainers).
|
||||
|
||||
### Wave 3 — 046 automation production ops (SCAUTO-FR-019)
|
||||
|
||||
1. **T017 (046)**: wiring `notify()`/`persist_notification()` в явные вызовы на run/stale/
|
||||
repeated-failure переходах lifecycle + canonical InvestigationSignal; доказать отсутствие
|
||||
auto-started agent work.
|
||||
2. **T013e**: layered retention tiers (run metadata/triage/step metrics/artifacts/screenshots/
|
||||
raw VLM) — сверить с уже landed `automation/retention.py` + `deletions.py` + migration 0024;
|
||||
закрыть остаток или переразметить бокс.
|
||||
3. **T013**: формально открытый RBAC/PROD-gate бокс — проверить покрытие в
|
||||
`test_scenario_automation_api.py` (Rbac-секция) и закрыть с evidence.
|
||||
4. **T021 canaries**: 5/15/50-tab retention canaries (зависят от T034 catalog + live стенда).
|
||||
5. SLO-профиль `ops-v1` + canary/rollout gates из `contracts/production-operations.md`
|
||||
(dispatch overhead p95, admission <1s, cancel drain, queue p95, cost cap).
|
||||
6. **T014/T015**: scoped-тесты + prototype validation (каждая @UX_STATE достижима).
|
||||
|
||||
### Wave 4 — 050 closure tail
|
||||
|
||||
1. **T044 (единственный OPEN production-gate 050)**: dedicated REST-vs-MCP error-shape parity
|
||||
fixtures (lifecycle/read/auth errors, disabled automation validation; service principal не
|
||||
решает human gate). Offline parity уже зелёный — не хватает именно фикстур матрицы.
|
||||
2. **CHK004**: негативный product-UI тест — во frontend нет agent prompt/chat/workspace/start/
|
||||
handoff запросов (сверить с фактическим состоянием после `8522a2ee`: HandoffSurface vs T030).
|
||||
3. **T029a**: typed `InitialScenarioIntent`/`TestPackProfile` + server-owned bindings;
|
||||
unresolved → `preview_only`, никогда не угадывать (естественно складывается в T029d/e —
|
||||
проверить, не закрыт ли уже handle-layer'ом).
|
||||
4. **T030/T032**: продуктовое решение — T030 требует заменить HandoffSurface обычным ручным
|
||||
редактором сценариев; T032 — финальное снятие `/api/assistant/*` (unmounted, но пакет
|
||||
`routes/assistant` живёт как parity provenance). Перед реализацией — зафиксировать решение
|
||||
в tasks.md (оба затрагивают UX round 4/5 поставки).
|
||||
5. **T042**: статус бокса — drift-amendments уже «done (2026-09-02)», но closure review требует
|
||||
не считать это production-readiness evidence; привести формулировки к актуальным.
|
||||
|
||||
### Wave 5 — 042/043/045/037/038 remaining rows
|
||||
|
||||
- **042**: T026/T027 (upstream 037 StructureDiff + 041 blast-radius production events →
|
||||
`apply_staleness`; canonical InvestigationSignal), T029 (artifact-drift reconciliation),
|
||||
T030–T032 (production rows), T024/T025/T028 (quickstart/prototype/independent verification).
|
||||
- **043**: T015/T023 (удаление AgentActionPanel prompt/proposal-generation контролов и
|
||||
agent-propose вызовов из frontend; `EditRevisionDiff` остаётся для stored external proposals),
|
||||
T021 (DEF-01 round-trip), T022 (dependency closure), T017–T020 (verification/prototype).
|
||||
- **045**: T020–T022 (RUNMON-FR-014: evidence states recovery, run comparison provenance,
|
||||
AgentEvaluationCard read-only + удаление investigate-with-agent контролов), T016–T019.
|
||||
- **037**: T044 (visual candidate), T082–T084 (duplicate/stale-CAS approvals, server-owned bytes,
|
||||
rebaseline CAS + publish_failed retry).
|
||||
- **038**: T041 (VLM `vlm.py` — локальный провайдер, typed `VlmFinding[]`, model_provenance),
|
||||
T046 (capture/VLM DTO consumable by 044), T060–T062 (visual template chain, descriptor mapping,
|
||||
policy precedence rows).
|
||||
|
||||
### Wave 6 — Final closure gate
|
||||
|
||||
1. Полные backend + frontend suites на финальном дереве (числа — в WORKSTATE).
|
||||
2. `alembic upgrade head` + `alembic check` на свежем PostgreSQL 16 (цепочка 0001→0025+).
|
||||
3. Semantic rebuild (Axiom) + unresolved-relations sweep (цель: не выше текущего уровня 401,
|
||||
в идеале — reduction по затронутым доменам) + `make docs-nav`.
|
||||
4. Closure re-review по матрице раундов 1–5 (WORKSTATE 2026-09-04) + production rows 2026-09-08.
|
||||
5. Финальный чекпоинт в WORKSTATE + решение GO/NO-GO с явными residual-рисками
|
||||
(deferred product decisions: M-03 self-approval, service-branch permissions, adapter
|
||||
idempotency, multi-binding exploration context, TTL/expiry).
|
||||
|
||||
## Рекомендуемый следующий пакет воркеру
|
||||
|
||||
**Wave 1.4 (T046 closure)**: реализовать walker-level binding evaluation→comparison по одной из двух
|
||||
зафиксированных опций (dependency-ordered binding предпочтительнее — compare зависит от покрывающего
|
||||
`evaluate-visual`), с truth table DecisionPolicy неизменной; затем live-рерun v4-графа
|
||||
(`specs/044-dashboard-scenario-execution/prototype/live_canary_v4_baseline_pin.py`) против
|
||||
запущенного backend'а (run.sh, 127.0.0.1:8000) с живым LLM-провайдером и доказать терминальный
|
||||
`passed`. T043 (graph-level deterministic comparison PASS) закрывается тем же live-рерunом.
|
||||
Волны 2–5 независимы и могут раздаваться параллельными воркерами (047 / 046 / 050+frontend /
|
||||
042+043+045 / 037+038).
|
||||
|
||||
## Инфраструктурные напоминания
|
||||
|
||||
- Backend: `cd backend && source .venv/bin/activate`; integration — `--run-integration`.
|
||||
- Frontend: `cd frontend && npm run test -- --run && npm run lint && npm run build`
|
||||
(svelte-check отсутствует, его заменяет build).
|
||||
- Стенд: `./run.sh --skip-install` (backend 8000, frontend 5173); `/api/ready` несёт provider
|
||||
readiness snapshot; MCP endpoint `{origin}/mcp`, discovery
|
||||
`{origin}/.well-known/oauth-protected-resource/mcp`; docs подключения — `docs/mcp-client-setup.md`.
|
||||
- **Canary-запуск T042b (проверено 2026-09-16)**: креды стенда — НЕ в `.env`, а в шифрованном
|
||||
AppConfig (`ConfigManager()._load_config()` при `ENCRYPTION_KEY` из `backend/.env`, извлекать
|
||||
через `grep -m1 ... | cut -d= -f2-` из-за multiline PEM); обязательны `AUTH_SECRET_KEY`/`SECRET_KEY`;
|
||||
`SS_STAND_STAGE=PREPROD`; dashboard 11; `SS_CANARY_MODE=actions|timeout|mutation|reconcile|
|
||||
recovery|soak`; recovery/soak/reconcile требуют `DATABASE_URL` отдельной PG-БД (создана
|
||||
`canary_044`; проверочная alembic-БД `alembic_044`).
|
||||
- Скиллы: изменения в `.agents/skills/` → `./scripts/sync-skills.sh` (не редактировать `.kilo/skills/`).
|
||||
- INV_7 хвосты: `mcp_server/tools_scenario.py` 508 LOC, `api/routes/dashboard_testing/scenario.py`
|
||||
442 — кандидаты на decomposition-gate по прецеденту server.py
|
||||
(`specs/050-mcp-interface/plans/server-decomposition-gate.md`); НОВЫЕ Wave-C нарушители:
|
||||
`providers/browser_readonly_actions.py` 534, `providers/browser_session.py` 501,
|
||||
`providers/browser.py` 407 (план: `specs/044-dashboard-scenario-execution/plans/
|
||||
provider-decomposition-gate.md`).
|
||||