Files
ss-tools/specs/WORKSTATE-043-047.md
busya bcc69f4bbe feat(dashboard-testing): add sampled traversal and ClickHouse test lab
Add versioned metric graph authority, owned browser evidence, paginated and all-tab traversal, deterministic sampling policies, and analyst-facing run inspection.

Provision the DEV/PREPROD/PROD Superset, Gitea and million-row ClickHouse lab; retain reproducible lifecycle evidence and explicit incomplete-traversal limits.
2026-10-02 10:54:42 +03:00

281 KiB
Raw Blame History

WORKSTATE 043–047 (intermediate 2026-09-24 06:52 +03:00; supersedes the 2026-08-20 header)

Статический + targeted-runtime checkpoint. Не является feature-completion. Правило: [x] только с текущим доказательством; [~] частичная реализация; [ ] отсутствует.

Current dashboard 11 release and T029a checkpoint — 2026-09-30

  • User-supplied Superset credentials enabled an authenticated, read-only lookup: /superset/dashboard/sales/ is Sales Dashboard, ID 11. The earlier ss-prod login redirect and missing local test credential describe an earlier attempt, not the current read-only discovery state.
  • The current local ss-tools release ledger has a dashboard 11 repository, but no deployment or DashboardRelease rows for it. The running configuration exposes DEV (ss-dev) and PROD (ss-prod) and no unique Git release PREPROD stage. create_release requires the latest successful, validated PREPROD deployment; PROD publication requires its exact named release and a human-authenticated ss-tools session. No release or PROD deployment was made.
  • The earlier SS_STAND_STAGE=PREPROD canary classification describes test policy for that run. It does not configure a Git release PREPROD deployment target. The accepted dashboard-release → baseline-capture dependency remains open for dashboard 11/ss-prod.
  • T029a has a CAS-persisted, server-derived baseline selection snapshot, with the URL absent from durable resolutions and responses. It remains needs_baseline/preview_only: compiled execute_metric lacks typed coordinate identity, so graph/receipt binding and runnable comparison cannot be inferred from a case and depends_on edge. Next: a 038 v2 producer coordinate contract and validator, then exact graph/receipt/launch admission and external replay. G-MCP-TEST-PACK-PROFILE is PARTIAL / NO-GO.

The operational release finding is detailed in the live dashboard 11 handoff and the profile contract gap in the T029a binding plan.

T029a MCP profile update — 2026-09-29 (post-acceptance checkpoint)

External MCP scenario authoring now has bounded profile-preview handlers in accepted runtime catalog 2.6.0 with dashboard:testing:READ (current unverified working tree: 2.7.0). The server re-inspects the selected dashboard and environment, returns complete selected-case coverage, stable chart→dataset→metric coordinates and native-filter applicability, and keeps unresolved questions typed. Selector answers are restricted to one exact compiler step and guarded by profile-digest CAS. Metric-coordinate selection changes needs_metric to needs_baseline; it never supplies expected numeric truth or a baseline pin.

Bootstrap now checks selected environment and case coverage against digest-bound server handles before any registry/revision/workspace write. T029a acceptance commits 2f094915, ce3dc56b, 1b6edaf7, 1a3d3958, and 3406433b add a server-owned eligibility receipt, durable owner-scoped CAS/idempotency state, receipt-bound handle admission, environment allow-list checks, and a linear, idempotent Alembic chain. Profile-focused tests pass 9/9. The selected MCP/profile run is 9 passed, 2 failed in the external-client and unresolved preview regression vectors listed in the handoff. This is PARTIAL, not production closure. Domain context and published-baseline resolution, fully proven save-eligible handle minting, fresh external-client profile→bootstrap E2E and catalog/schema parity remain open. Overall status remains NO-GO.

Working-tree continuation on 2026-09-29 (unverified): sequential profile CAS, receipt forwarding at bootstrap, scoped metric questions, fresh owner-scoped profile-session registration without a caller graph, and catalog 2.7.0 are implemented locally. Required unbound launch parameters now produce warnings and do not block authoring eligibility, matching the 038/044 contract. The two named regression fixtures were corrected, but their results above are still the latest executed evidence. Ruff, compileall, diff-check and AXIOM semantic verification passed; no behavioral tests or external E2E were run. T029a remains PARTIAL until targeted and fresh external MCP execution proves this slice and the remaining baseline and domain-context paths.

Current UX handoff

2026-09-24 architecture amendment: every new baseline is defined by an explicit same-dashboard Superset URL and its server-parsed immutable native-filter snapshot (037 ReferenceUrl). This is a P0 production invariant, tracked as G-BSL-REFERENCE-URL in the unified matrix, 037 T091 and 050 T052. Both supplied ss-dev dashboard 10 URLs have live read-only parser/query evidence; the native_filters_key=A4_rwhfQta0 link was rerun on 2026-09-25 with platform=3DS, 10/10 metrics executed and 6 distinct metric formulas suggested. A separate, removed-after-test Compose project with a dashboard 10 scenario and test-only synthetic release then proved browser URL entry → filter/metric preview → six draft candidates (Chromium 2/2, including a different-dashboard field error). Its isolated DB retained the exact URL hash and platform IN ["3DS"] snapshot across six captures. The ordinary local database still has no dashboard 10 scenario/release fixture; a later isolated Chromium pass proved source-byte verification, reasoned approval and local catalog materialization for one URL candidate, with the exact URL/filter snapshot retained in its revision. Git publication has no receipt; real release→Git publication→run pin→scheduled replay remains unproven. Do not infer the full invariant or analyst acceptance from the seeded 4/5 review score.

The dashboard-testing candidate is now 9acd9bad (UX slice on top of fa531225) and has live smoke proof for route loading, queue→case navigation, protected evidence, schedule preview and run-center rendering. The UX slice adds grouped queue states, business-language case summary, evidence availability cards and collapsed technical disclosures; focused frontend proof is 13 tests, with lint/build green and Axiom parse warnings at zero. A business-analyst expert review nevertheless scores the seven critical tasks 2, 1, 2, 2, 1, 3, 2 respectively, mean 1.9/5; this is not moderated five-analyst acceptance. See specs/HANDOFF-DASHBOARD-ANALYST-UX-2026-09-24.md. G-BA-UX-READINESS remains OPEN.

The most important remaining product work is live analyst validation of investigation/case and baseline workflows, period-change/curated baseline selection UX, business-language explanations for launch/result states, and the five-analyst study with retained browser/network evidence. The uncommitted seeded baseline review slice reaches a provisional expert 4/5 on its scoped happy path; it does not imply readiness acceptance.

Overall

Spec Status Notes
042 Registry [~] CRUD/lifecycle есть. Upstream 037/041 ingestion + queue signal wired. Full quickstart/T024 unverified.
043 Editor [~] Code/route present. No current E2E/a11y proof. Depends on 042/044.
044 Execution [~] not production-complete Walker/lifecycle/API работают. Executors fail-safe. Live Playwright/Superset session still missing.
045 Run Monitor [~] Typed launch contract aligned (release/baseline/toggles). Full route E2E open.
046 Automation [~] Scheduler reload + trigger dispatch + notifications persist. Live due-job E2E open.
047 Analytics [~] not production-complete Signal ingest, CAS closure, recurrence, owner ACL. Chat/AgentRun workspace open.

Verified facts (this checkpoint)

044 Execution

  • DAG walker _advance_run walks topological_order, claims steps, registers artifacts, suspends on human.
  • Failed/blocked producers block descendants including human; walker does not suspend a terminal failed run.
  • API POST /scenario-runs creates durable queued run (auto_advance=False); worker/tests advance separately.
  • Launch contract: typed dashboard_release_id, baseline_set, execution_toggles on StartRunRequest / start_run; frontend no longer hides them in params.launch_config.
  • Executors (execution/executors.py):
    • assertion → 037 compare_values (can fail; never hardcoded PASS).
    • xlsx → openpyxl parse of real workbook bytes + artifact ref.
    • screenshot / report / artifact bind only with real bytes/refs/digest.
    • browser / superset_api without session/query envelope → typed inconclusive.
  • Cancel: queued steps → skipped, run → cancelled (via cancel_requested/draining).
  • Timeout: apply_step_timeout → step/run inconclusive + STEP_TIMEOUT.
  • Terminal failed/blocked/inconclusive → auto_queue_failed_run + persisted notification.

047 Analytics

  • ingest_investigation_signal() idempotent queue projection; no auto agent start.
  • Queue/case persist evidence_snapshot + linked_run_ids.
  • Closure CAS: resolved needs reconciled verification evidence; accepted needs rationale.
  • Recurrence after resolved episode opens new episode + queue item.
  • Object ACL: InvestigationCase.owner_id; can_access_case; GET /cases/{id} 403 for non-owner view.
  • Migration b9c0d1e2f3a4 adds evidence snapshot, linked_run_ids, owner_id.

046 Automation

  • load_schedules() reloads enabled ScenarioSchedule and keeps scenario_* jobs in desired set.
  • add_scenario_job maps persisted misfire_grace_time, max_instances, missed-execution/coalesce.
  • dispatch_trigger_event resolves current revision and calls start_run.
  • POST /api/scenario-automation/events/dispatch is the HTTP boundary.
  • persist_notification writes lifecycle events (completed/failed/blocked).

042 / 045

  • ingest_upstream_staleness_event() for normalized structure_diff / lineage_blast_radius.
  • Staleness apply emits investigation queue signal.
  • RunMonitorModel posts typed 044 start fields (Vitest: launch body contains dashboard_release_id, not launch_config).

Targeted evidence (do not treat as suite-close)

Recorded in this session with AUTH_SECRET_KEY / SECRET_KEY set:

  • Investigation + staleness + executors: 12 passed (later expanded).
  • Automation services/API: 32 passed; trigger dispatch pinning: 9 passed.
  • Runner/walker/dispatch after fail-safe executors: 11 passed.
  • Run API after queued-start restore: 30 passed.
  • RunMonitorModel Vitest: 2 passed.
  • Core scheduler regression: 17 passed.
  • Consolidated 042/044/046/047 slice: 81 passed.
  • Recurrence + walker fail-path + executors: 47 passed, ruff clean on touched modules.
  • Lifecycle cancel/timeout + investigation ACL: 20 passed.
  • Analytics API after resolved-disposition contract: 21 passed.

Full backend/frontend suites, Playwright, semantic rebuild: not rerun.

Remaining (priority)

  1. 044 live adapters — Playwright session replay and catalog-backed Superset query when a real session/query envelope exists. Current inconclusive is fail-safe, not SCEX-FR-002/009 complete.
  2. 044 T025 remainder — approval-to-live-dispatch, infra resume continuation of remaining DAG, timeout during in-flight executor I/O.
  3. 047 T016 remainder — durable chat / linked AgentRun workspace on InvestigationCase.
  4. 047 T018 / 044 T022 — independent full scoped verification; do not reuse old aggregate counts.
  5. 045/043 UI — mount remaining production panels; browser/a11y E2E.
  6. 039 T058 — wire _trigger_release_verification into preprod deploy (out of 042–047 core, still listed).

Key files (this checkpoint)

  • backend/src/services/dashboard_testing/execution/{runner,executors,lifecycle,dispatch}.py
  • backend/src/services/dashboard_testing/analytics/{investigation,recurring}.py
  • backend/src/services/dashboard_testing/automation/{trigger,notify}.py
  • backend/src/services/dashboard_testing/registry/staleness.py
  • backend/src/core/scheduler.py
  • backend/src/api/routes/dashboard_testing/{scenario_runs,scenario_analytics,scenario_automation}.py
  • backend/src/models/scenario_investigation.py
  • backend/alembic/versions/b9c0d1e2f3a4_add_investigation_evidence_snapshot.py
  • frontend/src/lib/models/RunMonitorModel.svelte.ts

MCP Handoff (2026-08-28 08:00 +03:00)

Это промежуточный handoff для следующего агента. MCP infrastructure реализована частично. Maintenance adapters ниже являются только техническим доказательством approval/scheduler/task boundary и не являются продуктовым happy path dashboard testing.

Product priority

Основной продуктовый vertical path:

MCP client
  -> 042 current ScenarioRevision / WorkingDraft
  -> 038 compile + deterministic validation
  -> 037 baseline/lineage resolution
  -> 044 ScenarioRun creation
  -> 045 approval/monitoring
  -> 044 ExecutionCapacityManager + live providers
  -> 037/044 evidence and StepOutcome
  -> 047 InvestigationSignal / Queue / Analytics

Следующий агент НЕ должен продолжать расширять maintenance как самостоятельную feature. Приоритет: 038 -> 042 -> 044 -> 045 -> 047 MCP vertical integration.

Implemented MCP infrastructure

  • Official Python SDK mcp==1.29.1; Streamable HTTP mounted at /mcp.
  • FastMCP lifespan bridge and local host/origin allowlists (MCP_ALLOWED_HOSTS, MCP_ALLOWED_ORIGINS).
  • Request byte limit for both Content-Length and chunked bodies; buffered ASGI messages are replayed.
  • Explicit catalog with safe, guarded, approval, and read-only tools; no generic api_call.
  • Live RBAC catalog filtering and invocation guard; service principal is read-only.
  • RS256 MCP JWT, kid, JWKS /oauth/jwks.json, current/previous public-key overlap.
  • PKCE S256 authorization code flow, refresh rotation/reuse revocation, DCR rate limit, client revocation.
  • McpToolInvocationRecord: argument hash, approval/dispatch state, continuation payload, lease/fencing, result hash; raw bearer tokens, arguments, results, and credentials are not persisted.
  • MCP-owned ActionApprovalGate, human-only list_pending_approvals / decide_approval, CAS decision.
  • Payload hash and operation binding revalidation before dispatch.
  • Scheduler poller with explicit dispatcher registry, lease renewal, stale claim recovery and fencing.
  • Current reviewed adapters: start_maintenance and end_maintenance only; they enqueue existing TaskManager/MaintenanceBannerPlugin work and do not call Superset directly.

Alembic head

0001 baseline
  -> 0002 MCP OAuth
  -> 0003 MCP provenance
  -> 0004 continuation payload
  -> 0005 dispatch receipt
  -> 0006 dispatch lease
  -> 0007 DCR rate limit
  -> 0008_mcp_reconcile maintenance ownership/recovery

Evidence

  • Focused MCP/OAuth/approval/maintenance unit suite: 41 passed at the latest reconciliation checkpoint.
  • PostgreSQL Testcontainers concurrency suite: 4 passed for claim, stale-worker fencing, renewal, and MCP-owned orphan reconciliation.
  • Real MCP protocol test covers initialize -> mcp-session-id -> notifications/initialized -> tools/list -> tools/call.
  • ruff, compileall, git diff --check, and clean Alembic upgrade through head passed at the latest checkpoints.
  • Full backend suite, full frontend suite, Playwright product E2E, and semantic index rebuild were not rerun after all MCP changes.

Known MCP/product gaps

  1. OAuth browser consent/session flow and native FastMCP OAuth middleware are not production-complete.
  2. DCR PostgreSQL concurrent rate-limit evidence, admin client management, and automated key-rotation expiry remain open.
  3. MCP catalog is not parity-complete with the intended 37 tool surface; current explicit catalog is a staged subset.
  4. Approved MCP dispatch is wired only to maintenance adapters; no real 044 ScenarioRun dispatcher exists yet.
  5. 044 live Browser/Superset/Screenshot provider composition and ExecutionCapacityManager are not proven.
  6. No MCP vertical E2E currently proves ScenarioRevision -> compile -> validate -> ScenarioRun -> evidence -> analytics.
  7. Multi-process scheduler restart/soak, retry/dead-letter policy, and provider I/O failure injection remain open.
  8. PII is accepted for enterprise-local MCP/LLM/VLM providers. Credentials, cookies, access/refresh/service tokens, secrets, and raw storage paths remain prohibited in payloads, telemetry, provenance, and artifact refs.

Next implementation packet

Start with a dashboard-testing adapter, not maintenance:

  1. Find the existing 042 revision/WorkingDraft load and 038 compile/validate service APIs.
  2. Add a typed MCP inspect_scenario/validate_scenario read path using those services.
  3. Add a typed gated start_scenario_run request that creates the existing 044 queued ScenarioRun only after validation and approval; never accept a client-supplied compiled graph as canonical.
  4. Bind approval continuation to scenario_revision_id, revision_digest, validation digest, environment policy fingerprint, and execution toggles.
  5. Use existing 044 queued dispatcher and capacity boundary; do not create a second runner.
  6. Add PostgreSQL and MCP protocol tests for stale revision, digest drift, unknown lineage, duplicate run, approval CAS, cancellation, timeout and evidence receipt.

Working-tree note

All current changes are uncommitted. Existing unrelated modifications/deletions in .agents/, .kilo/, AGENTS.md, and earlier spec files were present during the work and must not be reverted. Review git status and git diff before any commit.

MCP Checkpoint (2026-08-28)

Dated implementation checkpoint. Claims below are limited to evidence collected in this round.

Verified changes

  • MCP catalog removed show_capabilities and unbound approval stubs; list_maintenance_events is read-only and non-gated.
  • Repeated direct RbacFastMCP gated calls reuse the invocation hash, gate hash, and request hash.
  • Fixed the gate-id handling that caused DetachedInstanceError.
  • Frontend exposes the MCP_DECOMMISSION handoff surface. Default and flag-on builds passed; the focused contract test passed.

Verification evidence

  • Backend MCP focused suite: 44 passed.
  • ruff and compileall: passed.
  • Frontend lint: 0 errors / 373 warnings.
  • Semantic anchor repair for mcp_approvals: 13/13.
  • Browser E2E was unavailable because the stack at ports 8101/8102 was unreachable.

Product decision

Product remains NO-GO. Open proof and delivery areas are:

  • live providers and execution capacity;
  • the full 038 -> 042 -> 044 MCP vertical;
  • parity with the intended 37-tool surface;
  • full backend and frontend suites;
  • semantic rebuild;
  • browser E2E.

MCP Checkpoint (2026-08-28, this round)

Dated implementation checkpoint. Claims below are limited to evidence collected in this round.

Verified changes

  • Added typed MCP inspect_scenario, validate_scenario, and start_scenario_run using existing 038/042/044 boundaries.
  • start_scenario_run has strict, server-owned revision input and does not accept a client-supplied graph.
  • Permission classification is RUN_PROD for PROD and RUN for non-PROD.
  • Removed duplicate scenario MCP approval ownership and the stale dynamic dispatcher/allowlist; maintenance-only explicit adapters remain.

Verification evidence

  • Focused MCP suites after the final fix: 43 passed.
  • ruff, compileall, and scoped git diff --check passed; anchors are balanced.

Product decision

Scenario MCP start remains intentionally queued/not-ready because no reviewed scenario dispatcher or live provider adapter exists.

Product remains NO-GO. The following remain open:

  • full scenario authoring persistence;
  • MCP 37-tool parity;
  • live providers and capacity;
  • PostgreSQL integration;
  • browser E2E;
  • full suites;
  • semantic rebuild.

Architectural Amendment Checkpoint (2026-08-28)

User-approved amendment to the 043–047 authoring and execution architecture. This checkpoint records the governing boundary; it does not change the completion status or convert unverified implementation into evidence.

Approved architecture

  • The agent is a full co-author of intent and diagnostics through the persistent, server-owned AgentAuthoringWorkspace and governed MCP operations. It is not a production authority.

  • Exploratory Playwright/code is allowed only inside an isolated authoring sandbox, with explicit time, size and network limits, cancellation, durable receipts and bounded artifact ownership. Credential, filesystem and shell escape is forbidden, as are network-policy escapes and production side effects.

  • AuthoringArtifact is distinct from ExecutionProgram: source, patch, trace, screenshot, diagnostic and operation-receipt artifacts are review material only; ExecutionProgram is the typed, deterministically validated 038 graph/program.

  • Promotion is strictly:

    sandbox
      -> typed action candidates / graph proposal
      -> deterministic 038 compile + validate
      -> user diff review
      -> 042 handle-based save (candidate)
      -> separate activate_revision CAS
      -> 044 execution
    
  • Raw code is never direct ScenarioRun authority. A code-backed production provider remains a future, separate and unimplemented contract; exploratory code support does not imply production code execution.

Canonical MCP operation names

The exact canonical MCP names are the following 15 snake_case operations. REST operationId values are transport aliases only and are distinct from MCP names; REST aliases MUST NOT be exposed as canonical MCP names in tools/list.

  1. inspect_dashboard_context
  2. create_authoring_session
  3. propose_test_plan
  4. start_exploration
  5. get_exploration_result
  6. propose_graph_revision
  7. get_graph_diff
  8. promote_to_scenario
  9. scenario_compile
  10. scenario_validate
  11. scenario_resolve
  12. generate_draft_pack
  13. request_save
  14. activate_revision
  15. start_scenario_run

The corrected 042 rule is explicit: ordinary save creates an immutable candidate; only the separate activate_revision operation may advance current_revision, subject to eligibility, materialization, policy, approval and CAS checks. Initial creation may be the sole explicit exception when its eligible initial revision is created as current.

Verification and release decision

  • Scoped git diff --check passed.
  • Required semantic anchors remained balanced after the append.
  • Product remains NO-GO. Open proof/delivery areas are authoring promotion E2E, sandbox runtime isolation and limits, MCP implementation and REST/MCP parity, live 044 providers and execution capacity, PostgreSQL integration evidence, and complete end-to-end/full-suite evidence.

Checkpoint — 2026-08-28: AgentAuthoringWorkspace foundation implementation

  • Added the AgentAuthoringWorkspace model and migration chain 0009 -> 0010 -> 0011.
  • The exact normative workspace states are: draft, exploring, exploration_failed, exploration_passed, proposal_ready, validation_blocked, awaiting_user_review, save_eligible, pending_approval, candidate, and current.
  • Create, load, transition, attach, and expire operations are owner-bound and use CAS; reads of expired workspaces fail read-only without a hidden mutation.
  • Durable per-operation idempotency and provenance receipts support replay after a later mutation and reject a changed-hash conflict.
  • The registry remains unmodified by the workspace foundation.
  • Twelve tests passed, including migration smoke; Ruff and compileall passed; one Alembic head is present at 0011.
  • PostgreSQL upgrade was not executed or verified; the scoped diff operation is absent.
  • Sandbox runtime, MCP authoring tools, graph promotion/activation, authoring E2E, provider/capacity evidence, and full production evidence remain open.
  • Product remains NO-GO.

Checkpoint — 2026-08-28: MCP authoring session implementation

  • create_authoring_session MCP tool is implemented with strict bounded input, authenticated user-derived ownership, service-principal denial, transactional workspace creation, durable idempotency replay/conflict handling, and no registry mutation.
  • propose_test_plan is intentionally deferred because no safe canonical persistence shape exists.
  • Twenty-five MCP/workspace-scope tests passed; Ruff and compileall passed; the Alembic head is 0011.
  • Canonical authoring operations and the scenario chain remain incomplete. Sandbox, proposal/diff/promotion/activation, MCP parity, response bounds, live 044 providers/capacity/PostgreSQL/full E2E remain open.
  • Product remains NO-GO.

Checkpoint — 2026-08-28: bounded test-plan intent artifact

  • Implemented the propose_test_plan MCP operation as a bounded, server-owned intent artifact.
  • Added immutable AgentAuthoringWorkspaceTestPlan, migration 0012 after 0011, and a canonical server-computed digest.
  • Enforced owner binding, CAS, expiry, idempotency, provenance, and dangerous content bounds. The operation accepts no raw executable graph or code and does not mutate the registry, revision, or run.
  • Fixed the fresh SQLite migration-chain issue by making guards in 0010–0012 idempotent; 0001 was not rewritten.
  • Fixed the owner-before-replay disclosure vulnerability.
  • Verification completed: fresh SQLite upgrade twice; 28 targeted tests; Ruff and compileall; one Alembic head at 0012; semantic anchors and git diff --check passed.
  • PostgreSQL concurrency and partial-schema migration branches were not verified. Sandbox, exploration, proposal/diff/promotion/activation, MCP parity, and live 044 evidence remain open.
  • Product remains NO-GO.

Checkpoint — 2026-08-28: bounded exploration request/result boundaries

  • Implemented start_exploration as a bounded, server-owned, persistence-only request boundary with the exact 038 REGISTERED_ACTIONS allowlist, dangerous-content rejection, CAS/idempotency/provenance, and a typed sandbox_unavailable outcome.
  • The workspace remains draft. The operation performs no provider or subprocess invocation and makes no registry, revision, run, or artifact mutation.
  • Implemented read-only get_exploration_result with owner-before-disclosure authorization, service-principal denial, a typed not-found outcome, and a bounded projection only.
  • Fixed the SQLite timezone CAS defect and closed the affected semantic anchor. A fresh migration chain now upgrades twice successfully.
  • Verification completed: 34 targeted tests passed; fresh SQLite upgrade twice; Ruff and compileall; Alembic head 0013; semantic anchors and git diff --check passed.
  • Remaining: no real sandbox/provider; no exploration result or artifact persistence from a provider; no graph proposal/diff/promotion/activation; no PostgreSQL concurrency, browser E2E, MCP parity, or full production verification.
  • Product remains NO-GO.

Checkpoint — 2026-09-01 (graph revision boundary)

  • Implemented propose_graph_revision MCP operation: persists one server-derived ScenarioEditProposal from the closed typed EditOperation union (via existing agent_propose + apply_ops) and attaches it to the workspace. Enforces owner-before-replay disclosure, draft-only state, scenario/revision binding, CAS, and durable idempotency (same key+ops replays the same proposal; changed ops raise a typed conflict). No candidate or current revision is created or activated; the client never supplies a graph.
  • Implemented read-only get_graph_diff MCP operation with owner-before- disclosure and bounded projection (proposal_id, base_revision_id, digest, diff, validation, proposal_status, cas_version). No mutation, provider I/O, or graph-snapshot disclosure.
  • Renamed the editor's graph-diff helper _graph_diff to a public graph_diff and reused it as the single diff authority across agent_propose and the read projection.
  • Verification: 47 targeted tests passed (workspace service, editor agent, MCP server); Ruff and compileall pass; Alembic head remains 0013; anchor pairing balanced (fixed a previously unclosed GetExplorationResult region); scoped git diff --check passes.
  • Remaining: exploration sandbox runtime; promote_to_scenario / request_save / activate_revision; PostgreSQL concurrency; browser E2E; full MCP authoring E2E; 044 live providers/capacity. Product remains NO-GO.

Checkpoint — 2026-09-01 (promotion boundary promote_to_scenario)

  • propose_graph_revision now advances the workspace draft -> proposal_ready (in addition to attaching the server-derived proposal), aligning with the normative promotion machine.
  • Implemented promote_to_scenario: deterministic promotion boundary that reaches proposal_ready -> awaiting_user_review (or validation_blocked), recomputes the proposal digest server-side, returns the computed diff and validation, and records a durable idempotent operation receipt. It never creates or activates a revision and never accepts a caller digest.
  • Verification: 50 targeted tests passed (workspace service + editor agent + MCP server); Ruff and compileall pass; Alembic head 0013; anchor pairing balanced (10/10 service, 26/26 MCP); scoped git diff --check passes.
  • Remaining: request_save (candidate save via save_proposal with policy authorization), activate_revision service + MCP, isolated sandbox runtime, 038 compile/validate parity names, PostgreSQL concurrency, browser E2E, full MCP authoring E2E, 044 live providers/capacity. Product remains NO-GO.

Checkpoint — 2026-09-01 (candidate save request_save)

  • Implemented request_save MCP operation: saves the owner-reviewed proposal into one immutable candidate ScenarioRevision through the existing guarded save_proposal -> save_revision -> create_revision path. The digest is recomputed server-side; the caller never supplies one. The workspace advances awaiting_user_review -> candidate; current_revision is never advanced.
  • request_save is the first authoring operation with a real registry mutation, so its catalog entry requires ("scenario", "EDIT") and denies service principals (risk_level="guarded").
  • Verification: 53 targeted tests passed (workspace service + editor agent + MCP server); Ruff and compileall pass; Alembic head 0013; anchor pairing balanced (11/11 service, 28/28 MCP); scoped git diff --check passes.
  • Remaining: activate_revision (separate CAS; activate_current_revision is still spec-only), isolated sandbox runtime, 038 compile/validate parity names, PostgreSQL concurrency, browser E2E, full MCP authoring E2E, 044 live providers/capacity. Product remains NO-GO.

Checkpoint — 2026-09-01 (activation activate_revision)

  • Implemented the spec-only activate_current_revision in registry/revisions.py: atomically promotes an eligible candidate to current, demotes the incumbent, advances entry.current_revision_id, enforces a compatibility-family check, and writes a ScenarioLifecycleAudit row.
  • Implemented the workspace activate_revision (and MCP tool): requires an owner workspace in candidate, an explicit revision_id, CAS, and idempotent replay. It calls activate_current_revision and advances the workspace candidate -> current. A ScenarioRun can now be pinned to a genuinely promoted/current revision_id.
  • activate_revision requires ("scenario", "EDIT") and denies service principals (risk_level="guarded"), matching request_save.
  • Verification: 286 targeted tests passed (workspace service + MCP server + the full registry service suite); Ruff and compileall pass; Alembic head 0013; anchor pairing balanced (12/12 service, 30/30 MCP, 6/6 revisions); scoped git diff --check passes.
  • The authoring promotion chain session -> plan -> proposal -> promote -> diff -> request_save -> activate_revision is now implemented end-to-end (exploration/sandbox aside).
  • Remaining: isolated sandbox runtime, 038 compile/validate parity names, PostgreSQL concurrency, browser E2E, full external MCP authoring E2E, 044 live providers/capacity. Product remains NO-GO.

Checkpoint — 2026-09-01 (038 parity: scenario_resolve + generate_draft_pack)

  • Added read-only MCP tools scenario_resolve and generate_draft_pack wrapping the existing 038 resolve_scenario and generate_draft_pack. scenario_resolve applies the typed ResolveChange union (parameter|selector|manual_conversion| remove_step) with a bounded value (<=2048 canonical chars), re-validates the resolved graph, and returns the new revision_hash + validation. generate_draft_pack returns the server-owned manifest only (no artifact bytes), save_eligible or preview_only. Both are read paths and persist nothing.
  • Catalog now: inspect_scenario (compile), validate_scenario, scenario_resolve, generate_draft_pack, start_scenario_run (plus the authoring chain).
  • Verification: 288 targeted tests passed (MCP server + workspace service + registry service suite); Ruff and compileall pass; Alembic head 0013; anchor pairing balanced (32/32 server); scoped git diff --check passes.
  • The 038 authoring pipeline compile -> validate -> resolve -> generate_draft_pack is now reachable over MCP, complementing the implemented promotion chain session -> plan -> proposal -> promote -> diff -> request_save -> activate.
  • Remaining: isolated sandbox runtime (start_exploration real provider), PostgreSQL concurrency, browser E2E, full external MCP authoring E2E, 044 live providers/ capacity. Product remains NO-GO.

Checkpoint — 2026-09-01 (PostgreSQL migration + concurrency evidence)

  • Fixed a PostgreSQL-only migration defect: revision IDs 0010..0013 exceeded the alembic_version.version_num VARCHAR(32) width and failed with StringDataRightTruncation. Shortened the four authoring revision IDs (0010_authoring_contract, 0011_authoring_operations, 0012_authoring_test_plan, 0013_authoring_exploration) and renamed the files; the linear chain is intact.
  • Fixed schema drift surfaced by alembic check: AgentAuthoringWorkspaceOperation used index=True on workspace_id, auto-generating an index name that differed from the 0011 migration. The model now declares the explicit ix_authoring_workspace_operations_workspace index, matching the migration.
  • Evidence captured against a fresh real PostgreSQL (postgres:16):
    • alembic upgrade head applies the full 0001 -> 0013 chain;
    • alembic check reports no new upgrade operations (no schema drift);
    • MCP PostgreSQL concurrency suite passes 4/4 (--run-integration): one-winner CAS claim, stale-worker fencing, lease renewal fencing, scoped/idempotent orphan reconciliation.
  • Verification: 54 targeted tests passed (migration smoke + workspace + MCP); Ruff and compileall pass; Alembic head 0013_authoring_exploration.
  • Remaining: isolated sandbox runtime, browser E2E, full external MCP authoring E2E, 044 live providers/capacity. Product remains NO-GO.

Checkpoint — 2026-09-01 (044 ProviderRuntime spec amendment)

  • Amended spec 044 to close the three provider-spec gaps identified during the live-provider effort assessment (contracts/modules.md, data-model.md, tasks.md; +43/-4 lines):
    1. Async/sync bridge pinned — new ScenarioExecution.ProviderRuntime ADR contract: exactly one application-owned long-lived provider event-loop thread; sync executors submit via run_coroutine_threadsafe and block under the step deadline; context/session objects live only on that loop and remain usable across all steps of one run. @REJECTED per-step asyncio.run (cannot hold live contexts), warm context pools for Phase 1; dedicated browser worker process deferred with adapter-Protocol survival requirement.
    2. Concurrency defaults pinned — 2 concurrent browser contexts for DEV/PREPROD, 1 for PROD per environment; screenshot capture admitted through the same CapacityManager lease accounting; exhausted capacity keeps the run queued with CAPACITY_BLOCKED; bounded submit queue (overflow waits under the step deadline, never starts I/O).
    3. Trace/video policy pinned — @REJECTED Playwright trace/video as durable evidence (non-deterministic, outside descriptor-declared evidence); operator debug flag keeps them local only, never registered as Artifact/EvidenceReceipt.
  • Implementation tasks amended accordingly: T032 (provider_runtime.py: loop thread + bounded submissions + defaults), T034 (loop-bound contexts, no per-step loop, no warm pool), T035 (wrap existing llm_analysis Playwright capture stack, shared loop + admission).
  • Anchor balance after edit: modules.md 19/19 region pairs, 2/2 brace contracts.

Checkpoint — 2026-09-01 (T028 provider protocol + ProviderRuntime loop bridge)

  • Implemented the first provider-production slice per the amended 044 spec:
    • backend/src/services/dashboard_testing/execution/provider_runtime.py — ScenarioExecution.ProviderRuntime.Engine: one long-lived application-owned event-loop thread; submit() is deadline-aware with bounded slots (semaphore); overflow raises ProviderSubmissionOverflow before the coroutine is created (no I/O started); deadline expiry cancels the in-flight coroutine and keeps the loop reusable; idempotent start(), draining stop(); process-wide accessor with test/shutdown reset.
    • backend/src/services/dashboard_testing/execution/provider_protocol.py — ScenarioExecution.ProviderProtocol.Schemas (T028): frozen ProviderExecutionContext / ProviderExecutionResult with closed enums (passed|failed|inconclusive|blocked, effect states, retry dispositions), stable PROVIDER_CONTEXT_INVALID / PROVIDER_RESULT_INVALID taxonomy, no-manufactured-PASS invariant (passed requires operation_id; non-pass requires reason_code), and pre-I/O create_operation_receipt with terminal-status predicate.
  • Tests: backend/tests/services/dashboard_testing/registry/test_provider_runtime.py (11), test_provider_protocol.py (19) — deadline cancellation, overflow-never-starts-IO, precondition rejections, lifecycle join, schema/receipt shapes.
  • Verification: 30 passed (new) + 63 passed (live-binding + MCP + workspace regression); Ruff clean; compileall clean; anchor nesting verified in all 4 new files.
  • Next: T032 full capacity.py lease table (migration 0014), then T035 ScreenshotProvider wrapping the llm_analysis Playwright stack on this loop. Product remains NO-GO.

Checkpoint — 2026-09-01 (T032 ExecutionCapacityManager)

  • Implemented the durable capacity slice per the amended 044 spec:
    • backend/src/models/provider_capacity.py — ScenarioExecution.Capacity: provider_capacity_quotas (one server-owned counter per environment+workload pair, unique index) and provider_capacity_leases (claimed|released|expired|reconciled).
    • backend/alembic/versions/0014_provider_capacity.py — guarded table creation, idempotent index creation, clean downgrade; revision id 22 chars (within VARCHAR(32)).
    • backend/src/services/dashboard_testing/execution/capacity.py — ScenarioExecution.CapacityManager.Service: atomic admission via one CAS increment (active_units + units <= limit_units), typed CapacityUnavailable (CAPACITY_UNAVAILABLE | CAPACITY_LEASE_NOT_ACTIVE | CAPACITY_LEASE_UNKNOWN), pinned defaults (browser/screenshot 1 PROD / 2 otherwise; other classes server-owned), idempotent release with zero floor, bounded expiry reconciliation (reconcile_expired_leases), heartbeat, read-only snapshot; naive/UTC-aware datetime normalization for SQLite and PostgreSQL TIMESTAMP columns; belief-runtime reason/reflect/explore on every branch.
    • Fixed during slice: heartbeat discovering an expired lease now frees quota units via the shared CAS _expire_lease_and_free (no double-free on concurrent paths).
  • Tests: backend/tests/services/dashboard_testing/registry/test_provider_capacity.py (17) — pinned defaults, exhaustion keeps counter unchanged, environment independence, heartbeat extension/rejection, idempotent release, expiry-once reconciliation, precondition taxonomy.
  • Verification: 17 passed (capacity) + 98 passed (provider runtime/protocol + MCP + workspace regression); Ruff clean; compileall clean; anchor nesting OK (4 files); migration smoke 3 passed; alembic heads = 0014_provider_capacity; full chain 0001→0014 applied on fresh postgres:16 with alembic check — no drift.
  • Next: T035 ScreenshotProvider wrapping the llm_analysis Playwright stack on the shared loop with capacity admission; then dispatcher integration (lease into ProviderExecutionContext) and T034 BrowserProvider. Product remains NO-GO.

Checkpoint — 2026-09-01 (T035 ScreenshotProvider)

  • Implemented the first real live provider per the amended 044 spec:
    • backend/src/services/dashboard_testing/execution/providers/screenshot.py — ScenarioExecution.ScreenshotProvider: registered LiveProvider wrapping the existing llm_analysis ScreenshotService transport; capture submitted to the single shared provider event loop (never per-call run_async); capacity admitted per attempt through the ExecutionCapacityManager (workload_class=screenshot, PROD limit 1 / otherwise 2) and always released in finally; evidence re-digested after capture, size-capped at 10 MiB, stored only as opaque draft:{run_id}:{digest} refs; typed non-pass taxonomy (SCREENSHOT_BINDING_INVALID/MISMATCH, RUN_UNRESOLVED, TARGET_MISMATCH, CAPACITY_UNAVAILABLE, CAPTURE_TIMEOUT/EMPTY/FAILED, TOO_LARGE, EVIDENCE_REF_INVALID, LOOP_UNAVAILABLE/OVERFLOW).
    • live_composition.py bootstrap now registers the deployment-owned screenshot provider per enabled binding (falls back to typed-unavailable); browser-safe providers remain explicitly unavailable. Amended the bootstrap @INVARIANT.
    • Fail-closed ordering fixed after the pre-existing lifespan test caught it: the provider checks loop availability BEFORE any DB/capacity I/O, so a not-running runtime performs zero side effects (SCREENSHOT_LOOP_UNAVAILABLE).
  • Tests: test_provider_screenshot.py (9) — happy path with lease release evidence, binding/identity rejections before I/O, capacity exhaustion refuses capture, PROD single-context limit, empty/oversize/multi-image evidence shapes, and a loop-binding probe proving capture runs on the shared loop. Transport ([EXT:Browser]) is the only mocked boundary; capacity, loop bridge, storage and DB are real.
  • Verification: 122 passed scoped regression (providers + live binding + MCP + workspace + migration smoke); Ruff clean; compileall clean; anchor nesting OK; git diff --check clean.
  • Next: dispatcher integration (capacity lease into ProviderExecutionContext) and T034 BrowserProvider on the same loop; then T040 startup readiness preflight. Product remains NO-GO (browser provider, sandbox, E2E, capacity soak open).

Checkpoint — 2026-09-01 (T034 read-only BrowserProvider slice)

  • Implemented the browser provider read-only core per the amended 044 spec:
    • backend/src/services/dashboard_testing/execution/providers/browser.py — ScenarioExecution.BrowserProvider: read-only action set open_dashboard, wait_for_state, refresh; descriptor validation before any I/O (presence, tool match, mutating rejection, registry fingerprint when present); capture via a BrowserActionTransport seam whose real implementation reuses the deployment-owned ScreenshotService _launch_and_login flow (server-owned auth, isolated headless Chromium context per attempt, context/browser closed on every path); evidence screenshot re-digested, size-capped at 10 MiB, stored as opaque draft:{run_id}:{digest}; capacity admitted per attempt through the ExecutionCapacityManager (workload_class=browser, PROD 1 / otherwise 2) with guaranteed release; typed taxonomy (BROWSER_BINDING_*, RUN_UNRESOLVED, TARGET_MISMATCH, ACTION_DESCRIPTOR_REQUIRED, ACTION_TOOL_MISMATCH, ACTION_NOT_SUPPORTED, REGISTRY_MISMATCH, CAPACITY_UNAVAILABLE, ACTION_TIMEOUT, LOOP_UNAVAILABLE/OVERFLOW, EVIDENCE_REQUIRED/TOO_LARGE/REF_INVALID, ACTION_FAILED).
    • live_composition.py bootstrap registers both providers (screenshot + read-only browser) from one deployment-owned ScreenshotService per enabled binding, falling back to typed-unavailable; mutating browser capability stays unregistered.
  • Tests: test_provider_browser.py (11) — happy path with checkpoints/lease release, descriptor rejections before I/O (missing/tool-mismatch/mutating), binding mismatch, capacity exhaustion refuses transport, PROD single-context limit, evidence missing/oversize, shared-loop probe. Only [EXT:Browser] is a test double.
  • Verification: 133 passed scoped regression (browser + screenshot + capacity + runtime + protocol + live binding + MCP + workspace + migration smoke); Ruff, compileall, anchors, git diff --check clean.
  • Next: wait_for_state/refresh real-stand canary, mutating-action flow (fixture lease + reconciliation), dispatcher lease integration, T040 readiness preflight, T042b PREPROD canaries. Product remains NO-GO.

Checkpoint — 2026-09-01 (capacity-aware dispatch)

  • Wired the capacity policy into the dispatch lifecycle per the amended 044 spec:
    • runner.py new region ScenarioExecution.Runner.CapacityBlock: a provider step refused by the shared allocator (BROWSER_CAPACITY_UNAVAILABLE / SCREENSHOT_CAPACITY_UNAVAILABLE) no longer terminalizes the run. The step returns to queued without consuming its attempt (outcome flagged capacity_retry), and the run parks as queued with CAPACITY_BLOCKED. reconcile_capacity_blocked_runs clears the parking code (bounded, idempotent) and dispatch_queued_runs calls it at cycle start, so parked runs re-attempt with a fresh capacity claim on the next cycle. @REJECTED: terminalizing capacity-refused runs as inconclusive.
  • Tests: test_scenario_capacity_block.py (3) — requeue semantics (attempt preserved, no terminal side effects), reconcile-then-retry reaching a real terminal passed with the same attempt, and scoping (non-capacity error codes untouched).
  • Verification: 3 passed new + 361 passed full registry suite regression (queued dispatch, crash recovery, providers, live binding, approvals, automation); Ruff, compileall, anchors, git diff --check clean.
  • Next: mutating browser-action flow (fixture lease + precondition hash + reconciliation), T040 startup readiness preflight, T042b PREPROD canaries. Product remains NO-GO.

Checkpoint — 2026-09-01 (T040 provider readiness preflight)

  • Implemented the startup readiness surface per the amended 044 spec:
    • backend/src/services/dashboard_testing/execution/providers/preflight.py — ScenarioExecution.ProviderPreflight: JSON-safe readiness snapshot per capability (provider loop, evidence storage, screenshot, browser, bindings) built after bootstrap registration; bounded Chromium-executable probe (resolves the executable without launching a browser) via the caller's run_async; evidence storage construction check (no bytes written); every check failure/crash becomes a typed degraded reason (BROWSER_EXECUTABLE_MISSING, BROWSER_PROBE_TIMEOUT, BROWSER_PROBE_FAILED, EVIDENCE_STORAGE_UNAVAILABLE, PROVIDER_CHECK_CRASHED, *_PROVIDER_UNREGISTERED), never an exception and never a startup blocker. @REJECTED: blocking startup on readiness — typed-unavailable registration already fails closed at dispatch.
    • live_composition.py bootstrap now logs the readiness snapshot at the end of composition (Provider readiness snapshot).
  • Tests: test_provider_preflight.py (5) — ready states with bookkeeping of unavailable bindings, probe failure degradation, check-crash safety, unregistered capability reasons, and a real-checks smoke (returns None or stable codes).
  • Verification: 5 passed new + 366 passed full registry regression; Ruff, compileall, anchors, git diff --check clean.
  • Next: mutating browser-action flow (fixture lease + precondition hash + reconciliation), T042/T042c contract test matrix, T042b PREPROD canaries. Product remains NO-GO.

Checkpoint — 2026-09-01 (T034 mutation admission + reconciliation semantics)

  • Extended the BrowserProvider with the mutation contract gate and unknown-effect semantics per the 044 BrowserProvider contract:
    • Admission split into ScenarioExecution.BrowserProvider.Admission (loop, binding, run identity, target identity, descriptor gating) and .MutationContract (per-key validation: fixture_lease_id, non-empty bounded target_keys <= 100, field_allowlist, 64-hex precondition_hash, cleanup_policy, retry_safe=false — typed codes BROWSER_MUTATION_CONTRACT_INVALID / PRECONDITION_INVALID / TARGETS_INVALID, all before any browser I/O).
    • Outcome semantics: mutating PASS carries operation_id + effect_state=completed; mutating timeout/crash -> BROWSER_MUTATION_RECONCILE_REQUIRED with effect_state unknown, reconciliation_required=true, retry after_reconciliation (never a plain timeout or PASS); transport-unsupported -> effect_state not_started (nothing started, fail-closed); read-only PASS carries effect_state=none without operation_id.
    • Real transport explicitly raises BrowserTransportUnsupported for mutating actions before any I/O until the DOM-write flow lands; real read-only flow unchanged.
  • Tests: test_provider_browser.py grown to 24 — valid-mutation receipt, 8 contract rejection shapes, timeout/crash reconciliation semantics, unsupported not_started, read-only effect_state none.
  • Verification: 24 passed browser + 379 passed full registry regression; Ruff, compileall, anchors, git diff --check clean.
  • Next: real DOM-write flow for edit_row/bulk_edit (fixture lease claim in transport, precondition verification, cleanup execution), durable provider operation receipts (T030), T042/T042c contract matrix, T042b canaries. Product remains NO-GO.

Checkpoint — 2026-09-01 (T030 durable provider operation receipts)

  • Implemented durable write-once operation receipts and wired them into the browser mutation flow:
    • backend/src/models/provider_operation.py — ScenarioExecution.ProviderOperation: provider_operation_receipts with pinned identity (run/step/attempt, provider, descriptor fingerprint, binding, principal, capacity lease), unique per (run_id, logical_step_id, attempt), running -> terminal CAS lifecycle, JSON summary + append-only history, cancellation fields.
    • backend/alembic/versions/0015_provider_operations.py — guarded creation with idempotent indexes; alembic heads = 0015_provider_operations.
    • backend/src/services/dashboard_testing/execution/provider_operations.py — ScenarioExecution.ProviderOperations.Service: open_provider_operation (before external I/O; duplicate attempt -> PROVIDER_OPERATION_DUPLICATE), complete_provider_operation (CAS terminal, late overwrite -> PROVIDER_OPERATION_TERMINAL), reconcile_provider_operation (only from reconciliation_required; appends audit history), record_late_response (history-only, status unchanged).
    • Browser mutation flow now opens the receipt atomically with the capacity claim (lease id recorded), finalizes it on every path (completed / reconciliation_required / failed+not_started for unsupported), and carries operation_id in typed outcomes. Read-only actions stay receipt-free.
  • Tests: test_provider_operations.py (8) — lifecycle, duplicate, identity validation, reconciliation source gating, late-response history; browser mutation tests now assert persisted receipt states (completed / reconciliation_required).
  • Verification: 33 passed receipts+browser; 388 passed full registry regression; Ruff, compileall, anchors, git diff --check clean; migration chain 0001→0015 applied on fresh postgres:16 with alembic check — no drift (temp DB dropped).
  • Next: reconcile worker wiring (scheduled resolution of reconciliation_required receipts), real DOM-write flow for edit_row/bulk_edit, T042/T042c contract matrix, T042b canaries. Product remains NO-GO.

Checkpoint — 2026-09-01 (provider reconciliation worker)

  • Implemented the scheduled reconciliation sweep per the 044 ProviderOperations contract:
    • provider_operations.py new region ScenarioExecution.ProviderOperations. ReconcileWorker: explicit per-provider reconciler registry (register_provider_reconciler, reviewed registrations only) and reconcile_stale_provider_operations — a bounded, idempotent sweep over stale reconciliation_required receipts (staleness cutoff on updated_at). Verdicts are validated (resolution completed|failed, registered effect_state) and resolved via the write-once CAS with an audit history entry. @REJECTED: auto-resolving receipts without a reconciler — silent completion would forge effect evidence.
    • core/scheduler.py — execute_scheduled_provider_reconciliation scheduled entry (own session, commit, failure-safe rollback+log) registered as APScheduler job provider_operation_reconciliation (IntervalTrigger 30s, max_instances=1, coalesce=true), mirroring the queued-dispatch job pattern.
  • Tests: test_provider_reconcile_worker.py (5) — registered reconciler resolves a stale receipt with history note; missing reconciler leaves the receipt untouched; invalid verdict is skipped and logged; staleness cutoff ignores fresh receipts; the scheduled entrypoint resolves stale receipts in the global DB end-to-end.
  • Verification: 5 passed new + 393 passed full registry regression; Ruff, compileall, anchors, git diff --check clean.
  • Next: real DOM-write flow for edit_row/bulk_edit with a browser reconciler (target-state verification against precondition hash), T042/T042c contract matrix, T042b PREPROD canaries. Product remains NO-GO.

Checkpoint — 2026-09-01 (T041/T042 provider contract matrix)

  • Added the cross-provider contract matrix per the 044 acceptance profile, covering vectors not proven by the per-provider suites (no duplication):
    • test_provider_contract.py (6) — lease safety on transport crash for BOTH providers (lease released, quota zero, no PASS manufactured), storage-ref ownership mismatch never passes (typed SCREENSHOT/BROWSER_EVIDENCE_REF_INVALID, empty artifact_refs), mutation receipt isolation per run (distinct receipts/leases, all released), and health redaction (readiness payload serialized contains no password/secret/token/cookie/credential shapes and carries exactly the five contract keys).
    • Parametrized across screenshot and browser seams; only [EXT:Browser] transports are doubles; capacity, receipts, storage and DB are real.
  • Verification: 6 passed matrix + 399 passed full registry regression; Ruff, anchors (stack check), git diff --check clean.
  • Remaining open: real DOM-write flow for edit_row/bulk_edit (needs a live Superset stand for verification), browser reconciler target-state check, T042c binding revalidation matrix rows, T042b PREPROD canaries. Product remains NO-GO.

Checkpoint — 2026-09-01 (T042b read-only canary GREEN on live stand)

  • Ran the read-only BrowserProvider canary against the live stand https://ss-prod.bebesh.ru (dashboards API verified; stand classified PROD — mutating canaries are forbidden and were not run):
    • specs/044-dashboard-scenario-execution/prototype/browser_readonly_canary.py exercises the REAL production path: build_playwright_browser_transport over the deployment-owned ScreenshotService login flow, shared ProviderEventLoop, ExecutionCapacityManager admission/release, durable draft evidence; no test doubles anywhere in the path.
    • Fixed during the canary: build_playwright_browser_transport returned a bare function while the adapter Protocol expects an object with .execute (AttributeError caught by the provider's typed fail-closed path, exposed by the canary's EXPLORE logs). Wrapped in a transport object; 11 browser unit tests green.
    • Results: 3/3 passed — open_dashboard (5.2s, checkpoint dashboard_open), wait_for_state/networkidle (5.7s), refresh (4.9s); evidence SHA-256 digests + durable refs for every action; page_url confirmed dashboard navigation.
    • Evidence retained: specs/044-dashboard-scenario-execution/evidence/browser-provider/ readonly-canary-20260901T143950Z.json + 3 PNGs (one per action). Credentials are not stored in evidence.
  • Regression: 399 passed full registry suite; anchors, ruff, diff-check clean.
  • Contract target note (T042b): read-only canary + forced timeout/cleanup canary + safe-checkpoint reconstruction still required for GO; mutation canaries require a PREPROD-classified stand. Product status remains NO-GO (sandbox, browser E2E, scheduler soak, live Superset/Screenshot binding composition proof open).

Checkpoint — 2026-09-01 (T042b forced timeout/cleanup canary GREEN; stand reclassified)

  • Stand owner confirmed https://ss-prod.bebesh.ru is a TEST stand — mutation-capable classification is authorized. Canary runs now use SS_STAND_STAGE=PREPROD.
  • Implemented and ran the forced timeout/cleanup canary vector (T042b): a 1s action deadline against the real stand forces the provider to cancel the in-flight Chromium coroutine, release the capacity lease, close the browser context and return typed BROWSER_ACTION_TIMEOUT with zero evidence — GREEN. Canary script gained SS_CANARY_MODE=timeout; evidence JSON retained.
  • Re-ran the actions canary with the PREPROD classification: 3/3 passed again (login form-post fallback → authenticated 302 confirmed in belief logs).
  • Evidence: evidence/browser-provider/readonly-canary-20260901T144732Z.json + PNGs, browser_forced_timeout_cleanup JSON. Regression 399 passed; ruff/anchors clean.
  • Mutation DOM-write status: the stand's charts are standard read-only Superset table/word_cloud visualizations — no editable widget exists, and the repo ships no editable-table plugin. The provider mutation path (contract, receipts, capacity, reconciliation) is complete and unit-proven; the concrete edit_row/bulk_edit DOM mechanics require a design decision (SQL-Lab-mediated data mutation vs a custom editable plugin on the target dashboards) before a mutation canary can be honest. Product status remains NO-GO (that decision + scheduler soak + browser E2E + sandbox open).

Checkpoint — 2026-09-01 (mutation canary GREEN on live stand)

  • Design decision (044 BrowserProvider): edit_row/bulk_edit execute as SQL-Lab-mediated data mutations performed by the authenticated browser session (page-context fetch to the CSRF-protected SQL Lab execute endpoint; DML gated by the database allow_dml policy — enabled on the test stand per owner authorization, database 1). Precondition hash = SHA-256 of the canonical pre-mutation SELECT; mismatch aborts before the UPDATE (not_started); cleanup_policy=restore re-applies pre-mutation values through the same provider path; reconciliation = post-mutation SELECT comparison. Documented as @RATIONALE in contracts/modules.md BrowserProvider. UI-DOM editing rejected (stock Superset tables are read-only; no editable plugin exists).
  • Implemented the mutation transport:
    • providers/browser_mutation.py — mutation contract gate (moved from browser.py for INV_7) + validate_mutation_inputs (table/key_columns/assignments/database_id; fields allowlisted; typed values) + build_mutation_script (page-evaluate JS with identifier allowlisting, typed literals, target-key-bound WHERE, crypto SHA-256 precondition check).
    • providers/browser_transport.py — transport seam + real Playwright transport with the async mutation branch (pre-SELECT → hash compare → UPDATE → post-SELECT → evidence screenshot).
    • providers/browser.py — 390 lines (INV_7); provider merges the contract into action_input, maps BrowserTransportPreconditionMismatch to not_started.
    • Found & fixed live: SQL Lab client_id must be <= 11 chars (query.client_id VARCHAR(11)) — psycopg2 StringDataRightTruncation surfaced through the canary.
  • Live mutation canary GREEN on https://ss-prod.bebesh.ru (test stand, owner-authorized): full provider path per mutation — capacity claim → receipt opened → isolated Chromium context → authenticated session → pre-SELECT hash match → UPDATE → post-SELECT → evidence screenshot → receipt completed → lease released. row_edit: mutate 82.74→82.75 (verified), restore →82.74 (verified, original preserved). 2/2 receipts completed.
  • Evidence: evidence/browser-provider/readonly-canary-20260901T151437Z.json + mutate/restore PNGs. Regression: 399 passed; ruff, anchors, diff-check clean; all provider modules < 400 lines.
  • Remaining for GO: browser reconciler (target-state check) wiring into the sweep, safe-checkpoint reconstruction trace, scheduler soak, browser E2E, sandbox runtime. Product remains NO-GO pending those; the mutation execution path itself is now production-proven.

Checkpoint — 2026-09-01 (browser reconciler + live reconciliation canary GREEN)

  • Closed the mutation loop end-to-end:
    • Receipts now persist the full mutation_context (dashboard_id, table, key_columns, assignments, target_keys, field_allowlist, database_id) in summary at open time.
    • browser_mutation.py adds build_select_only_script — the read-only observation flow (SELECT limited to assignment+key columns; SELECT * was rejected after the live run showed cross-column divergence).
    • browser_reconciler.py — build_browser_reconciler(service): observes the live row state through a fresh authenticated Chromium session (SELECT-only; never UPDATE) and maps it to verdicts: completed (state matches assignments) / failed (diverged or context missing).
    • Bootstrap composition registers the real reconciler for browser, so the 30s scheduler sweep resolves stale reconciliation_required receipts automatically.
  • Live reconciliation canary GREEN: an injected reconciliation_required receipt (real mutation context, backdated) was resolved by the sweep through a live Chromium observation — receipt completed, note "target row state matches the recorded assignments". Evidence: evidence/browser-provider/readonly-canary-20260901T152804Z.json.
  • Verification: 403 passed (4 reconciler unit tests added); ruff, anchors, diff-check clean.
  • Remaining for GO: safe-checkpoint reconstruction trace, scheduler soak, browser E2E, sandbox runtime. The mutation lifecycle (contract → receipt → effect → evidence → reconciliation) is now fully production-proven on the test stand.

Checkpoint — 2026-09-01 (T042b safe-checkpoint reconstruction canary GREEN)

  • Implemented the reconstruction_replay marker (044 data-model contract, previously missing): _archive_recovery_attempt records reconstruction_replay: true in the recovery_history entry, and _advance_run merges the marker into the replayed attempt's outcome when the step's recovery history shows reconstruction.
  • Safe-checkpoint reconstruction canary GREEN on the live stand (SS_CANARY_MODE=recovery, PostgreSQL canary DB): abandoned worker (claim → expired lease) → recover_run from the pinned plan → replay through the real provider against the stand → attempt 2 passed, reconstruction_replay=true in both the recovery history and the replayed outcome, dashboard_open checkpoint, evidence digest + PNG retained (20260901T154336Z-recovery-replay.png). Recovery canary harness seeds a real registry entry/revision and registers live providers into the global composition root (PostgreSQL canary DB removes SQLite cross-connection locking).
  • T042b canary scoreboard (all against the live test stand): read-only actions GREEN · forced timeout/cleanup GREEN · mutation row_edit+restore GREEN (with precondition hash + receipts) · reconciliation observation GREEN · safe-checkpoint reconstruction GREEN.
  • Verification: 403 passed; ruff, anchors, diff-check clean.
  • Remaining for GO: scheduler soak, browser E2E, sandbox runtime (start_exploration real provider). Product remains NO-GO pending those.

Checkpoint — 2026-09-01 (T042b scheduler soak GREEN on live stand)

  • Scheduler soak GREEN (SS_CANARY_MODE=soak, 75s, ~15 real APScheduler ticks against the live test stand): 3 queued browser runs dispatched by the real scenario_queued_dispatch job through the production provider path — 3/3 passed, attempts == 1 each (persisted queued→running CAS prevents duplicate provider I/O across overlapping ticks), active_leases == 0, exactly one durable evidence artifact per run, reconciliation job co-scheduled, scheduler.stop() graceful.
  • Canary harness: idempotent registry seeding helper shared by recovery/soak modes; soak asserts run statuses, attempt counts, lease hygiene and artifact counts in failures.
  • Evidence: evidence/browser-provider/readonly-canary-20260901T160242Z.json. Verification: 403 passed; ruff, anchors, diff-check clean.
  • T042b canary scoreboard: read-only GREEN · timeout/cleanup GREEN · mutation row_edit+restore GREEN · reconciliation GREEN · safe-checkpoint reconstruction GREEN · scheduler soak GREEN. All six canary vectors pass against the live stand.
  • Remaining for GO: browser E2E (full UI journey), sandbox runtime (start_exploration real provider), readiness/health payload hardening. Product remains NO-GO.

Checkpoint — 2026-09-01 (readiness hardening + T042c revalidation + sandbox runtime)

  • Readiness hardening (044 T031/T040 tail): /api/ready now carries the cached provider readiness snapshot (providers section captured at bootstrap; no per-request probe, no credential-shaped values).
  • T042c/T024a dispatcher revalidation tests (test_dispatcher_revalidation.py, 4): PROD reclassification invalidates the mutating gate before provider I/O; principal drift → dispatcher BROWSER_BINDING_MISMATCH; unknown binding ref → fail-closed mismatch; stale registry fingerprint → BROWSER_REGISTRY_MISMATCH before I/O.
  • Sandbox runtime (050 T026-T028 core): exploration_sandbox.py — CAS claim (queued→exploring), read-only observation through the authenticated session (dashboard charts/filters via GETs only, bounded 12/8), server-side graph synthesis (charts → nodes, native filters → applies_to edges), durable evidence screenshot, transitions exploring→exploration_passed/failed. exploration_result JSON column (migration 0016_exploration_result — fixed a 33-char revision id that broke the varchar(32) alembic limit; PG chain 0001→0016 verified, alembic check clean). Scheduler job authoring_exploration_dispatch (5s) sweeps queued requests; absent explorer registration keeps requests queued (sandbox_unavailable semantics).
  • Tests: +3 sandbox, +4 revalidation. Verification: 410 passed; ruff, anchors, diff-check clean. Remaining for GO: browser E2E + MCP parity domains + doc sync.

Checkpoint — 2026-09-01 (spec checkbox sync)

  • Synced stale task checkboxes against implementation evidence: 044 T028–T042c (provider vertical, canaries, revalidation) and 050 T001–T033 (MCP interface, authoring chain, handoff/demolition flags) now reflect done status. Remaining open: 050 T008 CI walkthrough, T012/T013/T014 parity domains, T016 parity tests, T023 promotion E2E, T028 promotion E2E, T040–T043 demolition tails, 044 documentation tails.
  • Session totals (2026-09-01): provider vertical (~1 975 LoC), capacity-aware dispatch, mutation SQL surface, reconciler + sweep, readiness hardening, T042c revalidation, sandbox runtime, 6/6 live canaries GREEN, migrations 0014–0016 PG-verified, 410 passed full regression.
  • Remaining for GO: MCP parity domains (git/deploy/baseline/superset, ~400 LoC), browser/MCP E2E (~600 LoC), promotion E2E wiring, doc tails. Product status: NO-GO.

Checkpoint — 2026-09-01 (050 T012/T013/T014 parity domains)

  • Closed the three open parity domains of the 050 catalog (45 tools now registered — one-to-one with the explicit catalog):
    • backend/src/mcp_server/ops_inputs.py — bounded strict input models for the git/migration/backup/llm and Superset domains; extra=forbid everywhere; the continuation payload re-validates identically under the approved hash.
    • backend/src/mcp_server/ops_tools.py — register_ops_tools: T012 gated bodies (curated schemas, approval-fallback only), Superset read tools (superset_list_databases, superset_explore_database schemas/tables/ table_metadata/select_star, superset_format_sql, superset_audit_permissions), gated Superset mutations (superset_create_dashboard, superset_copy_dashboard, superset_create_dataset), and the baseline 037-parity surface (capture_baseline_candidate, request_baseline_approval, decide_baseline_approval, consume_baseline_approval, create_verification_run) wrapping the exact REST-surface services.
    • backend/src/services/mcp_ops_dispatch.py — reviewed dispatch adapters for every approved continuation: GitService create/commit on the app loop, git-integration/superset-migration/superset-backup/llm_documentation TaskManager enqueue, ValidationTaskService policy creation, SupersetClient dashboard/dataset writes, and one-shot consume_approval for baseline publish. Payload re-validation before any I/O; idempotency keys per invocation.
    • poll_approved_mcp_dispatches chain extended explicitly (no dynamic lookup); server.py catalog +13 entries; RbacFastMCP.call_tool gains the SQL context gate: superset_execute_sql against a PROD-classified environment returns approval_required (MCPX-FR-018 style), non-PROD executes directly under the dedicated plugin:superset_sql permission.
    • rbac_permission_catalog.py: SUPERSET_SQL_PERMISSIONS + default-deny role mapping + is_superset_sql_permission, unioned into discover_declared_permissions so sync seeds the row for admin assignment.
  • SQL risk class: superset_execute_sql is denied to service principals, requires the dedicated permission (not implied by plugin:superset_proxy), bounded projection (50 rows), client-side dangerous-SQL guard intact.
  • Tests: backend/tests/test_mcp_ops_parity.py (13) — catalog risk/permission matrix, 45/45 registration with curated schemas, SQL dedicated-permission visibility, service-principal denial, create_branch gate-only idempotent flow, execute_sql PROD gate / DEV bounded execution / unknown-env fail-closed, baseline principal denial, adapter payload validation + one-shot consume, full poller route for an approved consume_baseline_approval invocation. Existing test_mcp_server.py catalog expectation updated.
  • Verification: MCP slice 70 passed (test_mcp_ops_parity + server + approvals + maintenance + oauth); RBAC catalog tests 41 passed; full backend suite 11167 passed, 240 skipped, 1 xpassed; Ruff and compileall clean; anchors balanced in all touched files (ops_inputs 6/6, ops_tools 7/7, mcp_ops_dispatch 6/6, parity tests 7/7); scoped git diff --check clean.
  • Remaining for GO: T016 parity fixture tests (MCP outputs vs legacy wrappers on shared fixtures), T023/T028 vertical E2E over MCP, browser UI E2E, scheduler soak for the new ops dispatchers, semantic index rebuild (Axiom not attached this session). Product remains NO-GO.

Checkpoint — 2026-09-01 (050 T016 parity tests + T023 MCP vertical E2E)

  • Closed the parity-proof and the first vertical E2E of the 050 gate:
    • backend/tests/api/test_mcp_parity_baseline_037.py (3) — one shared 037 fixture set (api conftest builders) drives BOTH surfaces with one persisted principal: MCP request_baseline_approval gate validates as the exact REST ApprovalGateResponse shape with field-by-field equality (operation, risk, required_permission, status, reason_required, target_paths); decide outputs agree including actor_id; create_verification_run MCP vs REST /verification-runs agree on overall_status/category statuses/evidence/ created_by; superset_format_sql MCP output is byte-identical to the legacy REST format path (only the [EXT:Superset] client is doubled). Discovered domain semantics pinned by the tests: one pending gate per AgentRun.
    • backend/tests/test_mcp_scenario_e2e.py (2) — SC-001 vertical walkthrough over tools/call with a REAL scenario:EDIT/RUN principal (no RBAC stubs): tools/list visibility → authoring session → server-derived graph proposal → owner diff review → promotion → request_save creating one immutable candidate ScenarioRevision in the registry → 038 compile/resolve/validate over MCP → separate-CAS activate_revision advancing current_revision_id; a no-EDIT principal neither sees nor can call save/activate (typed denial proven).
    • Fixed a production defect surfaced by the parity/E2E effort: the 038 resolver selector path crashed on steps with description=None (TypeError: NoneType + str in _apply_selector); the hint is now safely recorded, regression pinned by test_selector_hint_on_step_without_description.
    • Baseline MCP tool outputs normalized to non-colliding envelopes: request_baseline_approval -> {"status","gate"}, decide_baseline_approval -> {"status","decision"}.
  • Verification: parity slice 3 passed; E2E 2 passed; resolver suite 18 passed; combined MCP+037 slice 217 passed (in canonical ordering); full backend suite 11173 passed, 240 skipped, 1 xpassed (was 11167; +6 new tests net); Ruff and compileall clean; anchors balanced (E2E 3/3, parity 5/5).
  • Known collection-order artifact (pre-existing, reproducible with untouched files only): mixing root-level and tests/api package paths in a single ad-hoc argument order can drop tests/api/conftest.py fixtures for the second-collected api module (pytest prepend import mode). The canonical full-suite ordering is unaffected; not chased further.
  • Closed in this session: T028 authoring promotion E2E, ops-dispatcher soak. Remaining 050 items (CORRECTED the same day: the earlier "frontend surface remaining" wording confused 044 T030 numbering with 050 — 050 Phase 3 frontend T030–T033 is already [x] in tasks.md): only Phase 4 removal (T040–T043) and semantic index rebuild (Axiom binaries absent here). Parity + E2E gate criteria on the backend: met.

Checkpoint — 2026-09-02 (050 T028 authoring promotion E2E + exploration wiring)

  • Completed the last backend E2E of the 050 gate and wired the exploration vertical end-to-end:
    • backend/tests/test_mcp_authoring_promotion_e2e.py (8) — full sandbox-to- revision chain over tools/call with real RBAC: queued exploration → real execute_scheduled_exploration_dispatch from scheduler thread context (via worker thread; run_until_complete is not callable inside a live loop) → exploration_passed with observations/proposed_graph/ draft:exploration-* evidence while only the [EXT:Browser] boundary is doubled; bounded get_exploration_result projection (no observation leak); typed proposal derived from the observation (observed_dashboard_title lands in the saved graph); promote/diff/save/activate terminates in a current revision without a pre-activation save. Negative branches: SQL/path/backslash/drop typed ops rejected before proposal/CAS advance; code-token and unregistered-action exploration specs rejected without persistence; all five authoring input models structurally reject caller digest/content_hash.
    • Production wiring closed (was the T026-T028 seam gap): MCP start_exploration now reflects deployment reality (provider_available = get_registered_runner() is not None); bootstrap_live_execution_composition registers default_runner plus the deployment-owned ScreenshotService/draft-storage context for every enabled live binding via the new reviewed register_exploration_context(service, storage) seam. Without an enabled binding the behavior stays typed sandbox-unavailable; requests continue to queue instead of executing.
    • Ops-dispatcher soak (tests/test_mcp_ops_parity.py, now 14): two approved task-backed ops (backup, migration) over three poller cycles — each dispatched exactly once, second/third cycles are no-ops, both carry invocation-scoped idempotency keys, dispatch_status completed.
  • Verification: T028 module 8 passed; parity+soak 14 passed; MCP vertical slice 83 passed (8 files); registry/sandbox/workspace/scenario slices green in canonical ordering; full backend suite 11182 passed, 240 skipped, 1 xpassed (was 11173); Ruff and compileall clean; anchors balanced in all touched files (exploration_sandbox 5/5, live_composition 4/4, server 32/32, T028 E2E 6/6, parity 7/7).
  • Remaining 050 items (CORRECTED: Phase 3 frontend T030–T033 already [x]; the earlier list conflated 044 T030 numbering): Phase 4 removal (T040–T043) and semantic index rebuild (no Axiom binaries in this environment). Product status: parity + E2E criteria met on the backend; overall GO awaits the Phase 4 demolition tail only.

Checkpoint — 2026-09-02 (orthogonal audit + audit-driven hardening)

  • Two independent read-only audits (orthogonal QA + security) over the whole uncommitted 050 slice: QA re-ran all evidence (full suite reproduced exactly: 11182 pre-hardening), spot-verified 7 task claims (all CONFIRMED), audited mocks in the four new test files (no logic mirrors; external-boundary doubles only); security produced 17 CWE-mapped findings. No CRITICAL; no approval bypass; no service-principal leakage; dispatch stays CAS-claimed and static.
  • Fixed in this checkpoint (convergent HIGH findings):
    • backend/src/core/superset_client/safety.py — the SQL risk-class guard hardened from keyword blocklist to coverage of INTO/CALL/SET/REFRESH/ VACUUM/REINDEX/ATTACH/LOAD/PREPARE/COMMENT, server-side primitives (lo_import, lo_export, pg_read_file, pg_write_file, pg_ls_dir, dblink, dblink_exec, set_config, pg_sleep) and multi-statement smuggling (; after string/comment stripping). Tests in tests/test_core/test_superset_safety.py pin bypasses and false-positive guards (order_set/dataset/setval/single trailing ; stay green).
    • backend/src/mcp_server/server.py — PROD-classified SQL execution is now a TERMINAL rejection (production_sql_execution_rejected, recorded as denied provenance, no gate created), using the canonical production criterion resolve_environment_execution_policy (is_production OR stage=PROD) — this removes the approvable-but-never-dispatched black hole and aligns the criterion with scenario starts.
    • backend/src/services/mcp_ops_dispatch.py — baseline consume adapter now refuses deactivated actors (mcp_actor_inactive) matching the direct-tool guard.
  • Evidence after hardening: affected slice 214 passed (MCP vertical + both safety suites + superset extended + resolver); full backend suite 11202 passed, 240 skipped, 1 xpassed; Ruff/compileall clean.
  • Closed in the follow-up hardening round (2026-09-02, same day):
    • QA M2 / T017 gap — response_limit is now ENFORCED: RbacFastMCP.call_tool fails closed with the typed response_too_large envelope and denied provenance when a structured reply exceeds the configured limit (server.config), pinned by test_oversized_reply_is_truncated_to_typed_error.
    • QA M3 — poisoned exploration requests fail closed: default_runner terminates a claimed request whose workspace vanished (EXPLORATION_TARGET_UNRESOLVED / workspace_not_found) instead of raising back into the queue; pinned by test_default_runner_fails_closed_when_workspace_missing.
    • SEC M-10 — exploration evidence digest is now the sha256 content hash of the observation fingerprint; the ref is retrievable and verifiable via DraftStorage.retrieve, pinned by test_evidence_digest_is_sha256_of_observation and an E2E retrieval assertion.
    • SEC M-16 — TaskManager adapters now enqueue under the actor's durable user UUID (_actor_user_id, inactive actors rejected before dispatch), so MCP-originated tasks are visible to their owner via get_task_status; pinned in the migration adapter test.
  • Evidence after hardening: audit slice 161 passed; full backend suite 11205 passed, 240 skipped, 1 xpassed; Ruff/compileall clean; anchors balanced (mcp_ops_dispatch 6/6, ops_tools 7/7, parity tests 8/8).
  • Remaining deferred (product/ops decisions recorded, none block the parity+E2E gate): M-03 self-approval (the gate is a confirmation step; four-eyes needs a product decision on approver separation); M-04 service principals bypass catalog permissions by design of the service branch (superset_audit and table_metadata are the widest service-visible projections); M-06 at-least-once adapter idempotency after lease recovery; M-07 multi-binding exploration context last-writer-wins (single-binding deployments unaffected); M3/M-08 exploration queue TTL/limits; M-05 remnants (expired-gate replay _find_retryable_approval, decide-expiry TOCTOU).
  • Product status unchanged: parity + E2E criteria met on the backend; overall GO awaits the Phase 4 demolition tail (T040–T043) — the frontend Phase 3 (T030–T033) is verified complete in tasks.md.

Checkpoint — 2026-09-02 (frontend status correction + live happy-path demo)

  • Frontend 050 Phase 3 is COMPLETE (T030–T033 [x] in tasks.md): flag-driven decommission via the typed build-time switch MCP_DECOMMISSION (frontend/src/lib/config/mcp.ts, vite define from the MCP_DECOMMISSION env var), HandoffSurface component (frontend/src/lib/components/agent/HandoffSurface.svelte), and flag gates for /agent, the assistant panel and the top-nav entry (routes/+layout.svelte:134, routes/agent/+page.svelte:129/169).
  • Verified live on the local stand: ./run.sh --skip-install (backend 8000 /api/ready green, frontend 5173, agent 7860); demo admin happy-admin created through src/scripts/create_admin.py; Playwright headless-chromium walkthrough performed a real login and captured 8 screenshots (/tmp/kilo/happy-path/00-login.png … 06-migration.png, 07-handoff-agent.png): login → home → /dashboards → /agent (legacy workspace) → /dashboard-testing/runs (Global Run Operations Center) → /dashboard-testing/analytics (investigation queue) → /validation-tasks → /migration; a second dev server with MCP_DECOMMISSION=true on :5174 renders the handoff surface ("Ассистент переведён на MCP" + connection hint + copyable prompt).
  • Note: scenario detail routes are /dashboard-testing/scenarios/[id]/{edit,runs,analytics}; there is no scenario index page — flow entry is the run center and dashboard cards.

Checkpoint — 2026-09-02 (050 Phase 4 removal — the 050 task list is fully closed)

  • T040 — runtime decommission: run.sh now starts exactly backend+frontend (no AGENT_PORT/watchfiles/chat_uploads), docker-compose.yml and docker-compose.enterprise-clean.yml lost the agent service (nginx configs dropped /api/agent/gradio), build.sh lost build:agent/bundle:agent/ bundle:embeddings and the agent block of the generated deploy compose/manifest/env template; AGENTS.md/INSTALL.md rewritten for the two-service world. Live proof: stand boots with only :8000+:5173, /api/ready green, Playwright reaches the handoff surface.
  • T041 — chat code removed: the entire agent/ tree, docker/Dockerfile.agent, docker/agent.entrypoint.sh, the gradio-proxy backend test, and on the frontend: AssistantChatPanel + chat widgets, the ScenarioWorkspace branch (components/agent/dashboard-testing), AgentChat/AgentRun models, the assistant store, the MCP_DECOMMISSION flag infra (vite define, config module, d.ts) and the gradio dev proxy. /agent now renders HandoffSurface unconditionally; TopNavbar button navigates to it. Frontend: 3454 vitest passed, lint 0 errors, build OK. Backend: 11199 passed, 240 skipped, 1 xpassed. MarkdownRenderer retained (used by GitWorkspacePanel); assistant i18n keys retained because the handoff surface renders them; orphan chat DTO module frontend/src/types/agent.ts removed after confirming zero importers (build+tests re-verified).
  • T042 — all twelve 036–047 ## Drift Amendment — MCP Interface sections now carry **Status (2026-09-02): done** with evidence links (050 tasks T012–T028, T030–T033, T040–T041 + workstate checkpoints).
  • T043 — full-suite evidence recorded above; the stand runs without 7860 (SC-005 satisfied).
  • 050 status: all tasks [x]. Product readiness: backend parity + E2E and the removal tail are closed. Residual (documented, product decisions): approval self-approval semantics, service-branch catalog permissions, adapter idempotency after lease recovery, multi-binding exploration context, TTL/expiry polish, semantic index rebuild (no Axiom binaries in this env).

Checkpoint — 2026-09-02 (browser E2E on isolated stack + live admin walkthrough)

  • Browser E2E closed the last evidence gap of T043: isolated compose stack docker compose -p ss-tools-e2e --env-file .env.e2e -f docker-compose.e2e.yml (fresh db volume, initial-admin bootstrap admin/admin123, backend :8103 + frontend :8102 healthy, no 7860 anywhere) + local Playwright run of login.e2e.js + rewritten agent.e2e.js → 6/6 passed, twice (Chromium). Stack torn down with down -v afterwards.
  • FOOTGUN found and worked around: the documented E2E command without -p joins the DEFAULT compose project ss-tools — the same project run.sh uses for docker compose up -d db — so the E2E backend attached to the DEV database (admin had the real password, runner global-setup failed with 401). Isolation requires an explicit project name (-p ss-tools-e2e); dev db volume postgres_data survived and was re-verified after the incident.
  • E2E suite alignment with the decommission (part of T041's frontend tail): agent.e2e.js rewritten from chat-surface assertions to the handoff contract (T01 surface renders / T02 zero chat elements / T03 deep-link stays on handoff); dashboard-scenario-ui.e2e.js reload-recovery test retargeted to the handoff route; agent-scenario-run.e2e.js anchor updated (backend /api/agent-runs lifecycle tests remain valid — AgentRun persistence is kept); stale locator('nav') strict-mode violations fixed in login/smoke (.first()), invalid-credentials regex now matches the backend passthrough detail ("Incorrect username or password").
  • Live admin walkthrough on the native dev stand (real admin credentials provided by the operator): login → /dashboards, /agent renders the handoff (0 textareas, 0 conversation nodes), /dashboard-testing/runs renders; the agentless run.sh restart and split-process start (uvicorn + vite separately) both verified — screenshots /tmp/kilo/happy-path/09–12.
  • Note: the first post-restart Playwright run of agent T02/T03 flaked against a cold vite dev server; warm re-runs are stable (3.3s and 4.8s passes).

Closure Gate Review — 2026-09-03 (orthogonal, categorical; pinned to 731aaaa8, re-verified at ddbfe00f)

Two independent read-only reviewers (spec conformance + protocol/logging/MODEL-FIRST), plus mechanical sweeps. Verdicts by category:

  1. Spec conformance 050: FAIL. Blockers: FR-019 NONCONFORMS — list_checkpoints/ decide_checkpoint do not exist anywhere (T025 [x] is a false claim); FR-013 — /oauth/authorize requires a Bearer header (api/mcp_oauth.py:134-140), no cookie session/consent page, client_credentials grant missing → browser leg uncompletable by a standard client; SC-007 FAIL — zero HTTP-level proofs of /oauth/* (only service-level), T008 "scripted client" claim unsupported; FR-005 PARTIAL — list_pending_mcp_approvals lists only own MCP gates (automation/web gates invisible → Story 3 AC2 violated), mandatory high-risk confirm reason NOT enforced, envelope lacks targets/expires_at/reason_required; SC-005 PARTIAL — /api/assistant
    • agent_* routers still mounted (app.py:595-602), SystemSettings assistant-retention UI live, frontend/src/lib/api/assistant.ts shipped, .env.example retains AGENT/GRADIO 7860, /api/agent/llm-config exemption remains → T032 [x] false at HEAD; FR-010 absent (no catalog version/deprecation mechanism); FR-015 PARTIAL (no Admin DCR surface); FR-007/007a PARTIAL (oversize → rejection not ref+digest; no JSON-depth limit; no per-session rate limit / E6 Retry-After); SC-002 PARTIAL (output-parity fixtures exist only for the 037 domain; legacy wrappers deleted → remaining ~34 tools structurally tested only); SC-004 PARTIAL (no admin-fixture list-filter test); SC-009 unpinned (no mid-flow role-change test). Process: spec.md release-gate rows E2E-AUTH-001..003 still [ ] OPEN while dependent tasks are [x] — by the spec's own rule the gate is NO-GO; Status still "Ready for Implementation".
  2. Happy path A (external MCP client): FAIL — authorize defect blocks the browser leg; no over-HTTP tools/call proof (all tool-level evidence in-process; HTTP only for initialize/tools/list). Discovery→token lifecycle green at service level (68/68 independent rerun).
  3. Happy path B (UI): PASS (conditional) — login→dashboards→nav→/agent handoff proven live + Playwright 6/6; conditional: HandoffSurface does not embed the dashboard context params into the copyable prompt (Story 5 AC2 partial).
  4. GRACE-Poly INV: CONDITIONAL. Clean: INV_3 scripted over all 116 touched files (zero mismatches; TopNavbar 2/1 flag refuted — #region inside @RATIONALE prose); INV_9/TYPE canonical; no legacy markers; secrets clean. Blocker: INV_6 — dangling @RELATION CALLED_BY -> [agent/app.py] survives at HEAD in shared/src/ss_tools/shared/_llm_health.py:7 (round-3 ddbfe00f redirected 42 dead targets but missed this one); zero tombstones for the 571 contracts deleted by 731aaaa8; dangling VERIFIES/BINDS_TO edges in specs 033/035/036/039 manifests. MAJOR INV_4: duplicated @SIDE_EFFECT/env-mutation block in test_migration_routes.py:20-26 (edited in-commit — INV_9 anti-pattern). MINOR: INV_1 naked migrations 0014-0016; exploration_sandbox module region closes early (Registry block outside); INV_7 browser.py=402, server.py grew 1135→1522 (worst; needs DECOMPOSITION GATE plan).
  5. molecular-cot-logging: CONDITIONAL. MAJOR systemic: two-positional logger misuse (logger.X(_SRC, "intent", ...)) drops the real intent — wire emits the contract ID as intent (dynamically proven); ~120 call sites incl. ops_tools (8), mcp_ops_dispatch (4), exploration_sandbox (4), server authoring tools (18+). MAJOR: poll-dispatch adapter failure path is trace-invisible (except-branch in mcp_approvals has no EXPLORE); exploration sandbox fail-closed paths _finish() silently (DB only); 7/9 dispatch adapters have zero REASON/REFLECT. Clean: every existing EXPLORE carries error= (23/23), no secrets in payloads, no legacy markers.
  6. MODEL-FIRST frontend: CONDITIONAL. PASS: HandoffSurface architecture (correct inline-$state judgment per skill table; $lib/ui atoms; semantic tokens verified against tailwind.config; i18n keys ru+en; runes-only; guarded by contract test); +page.svelte thin wrapper; TopNavbar relations accurate; deletion sweep found zero broken imports. MAJOR: duplicated BINDS_TO [EXT:frontend:taskDrawerStore] edge in TaskDrawer.svelte:11-12 (introduced by the decommission edit). MINOR: vestigial assistantOffset=$derived("0px"); hardcoded English fallbacks in HandoffSurface (strict no-hardcoded-strings rule); silent clipboard fallback without EXPLORE; orphaned assistant.ts kept alive only by its test.

OVERALL CLOSURE GATE: NO-GO as-committed. Backend core (hidden-vs-gated RBAC, CAS gates + fenced lease dispatch, provenance, bounded inputs, parity domains, authoring/promotion chain) is solid and independently re-verified (68/68, 74/74). The gate is blocked by: one unimplemented FR (FR-019), the defective OAuth browser leg (FR-013/SC-007), three false [x] claims (T025/T008/T032), spec release-gate rows OPEN, and the protocol debt above (MCL intent drop, INV_6 blocker).

Remediation queue: P0 — downgrade T025/T008/T032; implement checkpoint tools; fix authorize flow (session/consent or documented manual-token path + client_credentials) with one HTTP-level SC-007 test; reconcile spec release-gate rows. P1 — MCL intent repair + wire-format test; EXPLORE/REASON on dispatch-failure and sandbox fail-closed paths; INV_6 edge + tombstone sweep; TaskDrawer/test_migration_routes dedupes. P2 — SC-005 remnants; FR-010 versioning; depth/rate limits; handoff context prompt; migration anchors; server.py decomposition plan.

Checkpoint — 2026-09-03 (closure-gate remediation round 1 / P0 — intermediate save)

  • FR-019 closed for real: MCP list_checkpoints + decide_checkpoint (044 CAS semantics: server-resolved pending checkpoint, decision_version CAS, continue_after_human_decision, typed conflict/not_found envelopes); catalog ("scenario","RUN") + service_allowed=False — automation has no path; tests/test_mcp_checkpoints.py → 3 passed. T025 re-earned its [x].
  • FR-013/SC-007 closed at wire level: client_credentials grant (confidential DCR clients, one-time client_secret, sha256-only persistence via migration 0017_oauth_client_secret, idempotent guard per house style), signed service-principal tokens (principal_type=service, aud=mcp, NO refresh), McpTokenVerifier service short-circuit, AS metadata extended, INSTALL.md §"MCP клиент" documented. tests/test_mcp_client_flow_http.py → 2 passed: FULL scripted-client flows over real HTTP (machine: discovery → DCR → client_credentials → /mcp initialize → tools/list gated-hidden → tools/call; user: DCR → PKCE S256 authorize with Bearer web session (documented SPA-mediated consent) → code exchange → identity-only token → /mcp live-RBAC listing). Browser-less-cookie consent page remains an open product decision (recorded, non-blocking for machine path).
  • PRODUCTION DEFECTS found while closing E2E-AUTH-003 and fixed: both argument-inspection gates in RbacFastMCP.call_tool understood only the FLAT argument shape while FastMCP delivers {"request": {...}} — start_scenario_run was UNCALLABLE over MCP (scenario_start_policy_unavailable denied every wrapped call) and the PROD-SQL terminal denial was BYPASSABLE by wrapped shape. Fixed via _gate_arguments unwrap (both shapes resolve identically); regression pinned by flat×wrapped parametrization of the PROD matrix (test_execute_sql_is_terminally_rejected_in_production, 6 combos).
  • E2E-AUTH release gates CLOSED with evidence (spec.md rows): 001 — T028 chain gained propose_test_plan (single trace: session→plan→exploration→result); 002 — compile/validate/diff already proven in T023; 003 — T023 extended with post-activation start_scenario_run on the canonical runner-shaped fixture graph, asserting the queued run pins scenario_revision_id + scenario_content_hash of the promoted revision.
  • Record honesty: tasks.md downgraded false/unproven [x] (T025/T008/T008b/T032
    • T005a annotation) at review time; T025 and T008 since re-closed with real executable evidence; T032/T008b/T005a remain OPEN (queued below).
  • Verification at save point: targeted slices green (checkpoints 3; client-flow HTTP 2; oauth 10; scenario E2E 2; promotion E2E 8; ops parity 18 incl. wrapped PROD matrix; mcp server/approvals/maintenance green in combined run 68+).
  • REMAINING remediation queue (round 2+):
    • P1 MCL: two-positional logger intent-drop repair across ~120 call sites + wire-format unit test; EXPLORE on poll-dispatch failure path and the four silent exploration fail-closed branches; REASON/REFLECT for the 7 silent dispatch adapters.
    • P1 GRACE: INV_6 — _llm_health.py dangling [agent/app.py] edge + tombstone/dead-edge sweep of specs 033/035/036/039 manifests; TaskDrawer duplicated BINDS_TO; vestigial assistantOffset; test_migration_routes duplicated @SIDE_EFFECT block; anchors for migrations 0014–0016; exploration_sandbox module-region span; server.py decomposition-gate plan.
    • P2: SC-005 remnants (/api/assistant router unmount, SystemSettings retention UI, assistant.ts, .env.example 7860 vars, /api/agent/llm-config exemption decision); FR-010 catalog versioning/deprecation markers; JSON-depth + per-session rate limits (E6 Retry-After); HandoffSurface context-parameterized prompt (Story 5 AC2); admin fixtures for SC-004 exact-catalog assertions; SC-009 mid-flow role-change test; browser cookie-consent decision.

Checkpoint — 2026-09-03 (closure-gate remediation round 2 / P1 — MCL + GRACE, full suites green)

P1 MCL (logging protocol)

  • Intent-drop repaired (182 sites): every two-positional logger.X(_SRC, "intent", ...) misuse converted to the facade convention logger.X("intent", src=_SRC, ...) via an AST transformer (180 sites across 19 files) plus manual repair of the two client_registry.py shutdown sites (real intents; the post-clear REASON became a REFLECT). The facade binds the FIRST positional as intent; the misuse silently dropped the sentence into *args.
  • Single convention repo-wide (211 sites): all direct SSOT log()/cot_log() production call sites in backend/src (58 files, incl. the log as _clog locals in candidate_guards.py and a dead log import in _llm_async_http.py) migrated to the intent-first facade; cot_logger.log() remains the module-internal primitive (belief_scope/cot_span/task-event bridge/shared layer, which has its own identically shaped facade ss_tools.shared.logger).
  • Facade level= support added to reason/reflect/explore (parity with the primitive's DEBUG plumbing lines), emitted through the level-named methods (target.info/warning/... — NOT target.log(numeric)), preserving the established patch/capture contract.
  • EXPLORE gaps closed: Services.McpApprovals.Poll except-branch now emits typed MCP_DISPATCH_<Exc> EXPLORE; exploration fail-closed branches (workspace vanished / target unresolved / sandbox unavailable / provider crash) are wire-visible through one choke-point EXPLORE inside _finish() — never DB-only again.
  • REASON/REFLECT closed for all 9 silent mcp_ops_dispatch adapters (deploy/migration/backup/llm-doc/llm-validation/superset-create+copy+dataset/ baseline-consume) with bounded identity payloads (sql logged as length only).
  • Skill synced with the module (the skill's own "the module wins" rule): canonical .agents/skills/molecular-cot-logging/SKILL.md §I/§II examples now come from real code (mcp_ops_dispatch, capacity.py), §III documents the single facade convention, the module-internal primitive, the forbidden two-positional shape, and the level= kwarg; ./scripts/sync-skills.sh run — .kilo copy byte-identical.
  • Executable pins: backend/tests/test_core/test_logger_wire_format.py (7) — REASON/ EXPLORE/REFLECT wire fields, dict-positional payload promotion, the intent-drop proof for the forbidden shape, level= override, and two repo-wide AST sweeps (two-positional misuse == 0; direct SSOT log imports in backend production code == 0).

P1 GRACE (semantic protocol)

  • INV_6: dead CALLED_BY -> [agent/app.py] edge removed from shared/_llm_health.py (rationale rewritten for the post-decommission single-container reality). Specs 033/035/036/039 dead-edge sweep (resolver over all 10433 live repo anchors): 0 dead edges remain — 5 wrong-id edges retargeted to live contracts ($lib/ui/Button→Ui.Button, $lib/ui/Icon→Ui.Icon, AgentChatSpec→Spec.AgentChatContext.FeatureSpec), 15 genuinely dead edges tombstoned in place (@DEPRECATED + successor: Api.McpOAuth, McpServer, Services.McpApprovals, AgentChat.HandoffSurface; ParsePdf/ParseXlsx — no in-repo successor, recorded).
  • INV_9 dedupes: TaskDrawer duplicated BINDS_TO [EXT:frontend:taskDrawerStore] removed; vestigial assistantOffset=$derived("0px") eliminated (style pinned right: 0); test_migration_routes.py 4×duplicated @SIDE_EFFECT block consolidated into one (and the duplicated DATABASE_URL statement removed).
  • INV_1: migrations 0014–0016 anchored in the house Migration.<Domain> pattern with verified live DEPENDS_ON targets (Models.ScenarioExecution.Capacity, Models.ScenarioExecution.ProviderOperation, Migration.AgentAuthoringWorkspace.ExplorationRequest); exploration_sandbox module region span fixed (Registry block was outside; module now closes at EOF).
  • server.py (worst INV_7 offender, 1571 LOC): missing MODULE region added (McpServer, @defgroup + core invariants) and a DECOMPOSITION GATE invariant recorded in-contract (phases A–D with line counts: scenario_inputs ~245 / auth ~200 / rbac_server ~305 / ProbeTools split ~330+330; frozen contract IDs and import surface; per-phase gates). Binding plan artifact: specs/050-mcp-interface/plans/server-decomposition-gate.md (constraints, risk register, rollback rule). Duplicate environment_policy import removed. Split execution is a follow-up, not done this round.

Full-suite defect root-caused and fixed (first full rerun since round 1)

  • test_mcp_client_flow_http (2 tests) failed ONLY in-suite: tests/test_dependencies_unit.py assigns MagicMocks into six DI singleton globals (config_manager, plugin_loader, task_manager, scheduler_service, resource_service, ...) without restoration; later in the session get_session_idle_timeout_minutes() resolved the MagicMock settings chain WITHOUT raising (so the fallback never engaged) and idle_minutes > 0 raised TypeError: MagicMock > int on every sid-bearing auth request (dependencies.py _enforce_session_policy). Pre-existing since round 1 (round 1 never ran the full suite); unrelated to this round's migration — proven by the exact traceback frame.
  • Fix A (hygiene): module-scoped autouse fixture in test_dependencies_unit.py restores all seven singleton globals after every test.
  • Fix B (production hardening): get_session_idle_timeout_minutes() validates int (excl. bool); a non-integer resolution now emits EXPLORE SESSION_POLICY_CONFIG_INVALID and falls back to the typed session_policy cache — a poisoned config object can no longer crash the auth path.
  • Repro pair pinned: test_dependencies_unit.py + test_mcp_client_flow_http.py → 62 passed (this ordering fails in-suite without the fixes).

Verification evidence (this round)

  • Full backend suite: 11214 passed, 240 skipped, 1 xpassed, 0 failed (was 11205; +7 wire, +2 client-flow now green in-suite).
  • Frontend: 3454 vitest passed (197 files); lint 0 errors / 340 warnings (baseline was 373).
  • Scoped slices during the round: logger/runner/detail-routes/git 73; migrated-modules slice 1393; MCP slice 95 (incl. both HTTP client flows in canonical ordering); wire-format 7.
  • ruff check . clean on the whole backend (src+tests+alembic); compileall clean; anchor-stack balance swept over every file touched this round — PROBLEM: none (server.py 34/34, logger 13/13, scheduler 30/30, wire tests 9/9, TaskDrawer 5/5); scoped git diff --check clean; skills .agents→.kilo byte-identical.

Remaining (round 3+)

  • Execute the server.py decomposition plan (phases A–D, gate-verified); INV_7 tail: browser.py 402 LOC.
  • P2 queue unchanged: SC-005 remnants (/api/assistant unmount, SystemSettings retention UI, assistant.ts, .env.example 7860 vars, llm-config exemption decision); FR-010 catalog versioning; JSON-depth + per-session rate limits (E6 Retry-After); HandoffSurface context-parameterized prompt (Story 5 AC2); SC-004 admin fixtures; SC-009 mid-flow role-change test; browser cookie-consent decision.
  • Deferred product decisions (recorded, non-blocking): M-03 self-approval, service-branch catalog permissions, adapter idempotency after lease recovery, multi-binding exploration context, TTL/expiry polish.
  • Semantic index rebuild — Axiom binaries absent in this environment (zombie-mode sweeps used).
  • Product status: NO-GO pending the closure-gate re-review of rounds 1+2.

Checkpoint — 2026-09-03 (closure-gate remediation round 3 — server.py decomposition EXECUTED)

Executed the binding plan (specs/050-mcp-interface/plans/server-decomposition-gate.md, now marked EXECUTED with a full execution log) plus one addendum:

  • Phase A: mcp_server/scenario_inputs.py (268) — the 13 typed input models moved verbatim; frozen-surface re-export block in server.py.
  • Phase B: mcp_server/auth.py (238) — Configuration/Principal/TokenVerifier/TransportAuth plus the SINGLE _access_token_context definition site. First gate run 42 failed → correctly caught a frozen-surface breach (tests build fake tokens via mcp_server.AccessToken); import retained → green.
  • Phase C: mcp_server/rbac_server.py (393) — GateArguments/ScenarioGate/Catalog/RbacServer moved verbatim. First run 1 failed → monkeypatch seam request_mcp_approval relocated to the owning module.
  • Phase D: mcp_server/tools_authoring.py (367) + mcp_server/tools_scenario.py (373) — the ProbeTools closure split into register seams on the register_ops_tools precedent; the seam call sequence preserves the original tools/list registration order byte-for-byte; ruff F821 caught the single lost import (timedelta). Tool-body seams (get_config_manager for list_environments/ start_scenario_run, get_task_manager maintenance sentinel) relocated to tools_scenario.
  • server.py = 177 LOC (was 1571): @defgroup + frozen-surface re-exports + the composition seam + CreateApp only; its DECOMPOSITION GATE invariant updated to the executed state (re-inlining a moved contract or re-exceeding 400 LOC violates the gate).
  • Addendum E (pre-existing INV_7 offender the closure audit missed): ops_tools.py 420 → 215; the 044 checkpoint + 037 baseline blocks moved verbatim into tools_review.py (253) behind order-preserving register seams; shared helpers (_SRC/_current_user/_resolve_environment) stay single-sited in ops_tools, seam import function-local → import graph acyclic.

Constraint compliance: contract IDs all frozen (verbatim moves of McpServer.* regions incl. the 9 nested authoring tool regions and the ScenarioTools/Checkpoint/Baseline blocks); import surface frozen (server re-exports every moved public name; AccessToken retained); behavior diff zero. Recorded refinement — a module-attribute monkeypatch seam is NOT healed by re-export: it follows the owning module (4 test files updated: test_mcp_server, test_mcp_ops_parity, test_mcp_scenario_e2e; all listed in the plan's execution log).

Evidence:

  • Final full backend suite: 11214 passed, 240 skipped, 1 xpassed, 0 failed (7:02) — identical to the round-2 suite: zero behavior diff.
  • Per-phase gates: anchors stack-balanced after every move; MCP slice 95 green at every phase end (A: 63→95; B: 95 after surface fix; C: 95 after seam relocation; D: 95; E: 95).
  • Package INV_7 census: __init__ 10 / server 177 / ops_inputs 190 / ops_tools 215 / auth 238 / tools_review 253 / scenario_inputs 268 / tools_authoring 367 / tools_scenario 373 / rbac_server 393 — all < 400.
  • ruff clean (whole backend), compileall clean, git diff --check clean, plan doc + module invariant metadata updated (INV_9: tags carry current local facts only).

REMAINING (round 4+):

  • INV_7 tail: browser.py 402 LOC (scenario provider).
  • P2 queue unchanged: SC-005 remnants (/api/assistant unmount, SystemSettings retention UI, assistant.ts, .env.example 7860 vars, llm-config exemption decision); FR-010 catalog versioning; JSON-depth + per-session rate limits (E6 Retry-After); HandoffSurface context-parameterized prompt (Story 5 AC2); SC-004 admin fixtures; SC-009 mid-flow role-change test; browser cookie-consent decision.
  • Deferred product decisions (M-03 self-approval, service-branch catalog permissions, adapter idempotency after lease recovery, multi-binding exploration context, TTL/expiry polish).
  • Semantic index rebuild pending Axiom binary availability.
  • Product status: NO-GO pending the closure-gate re-review of rounds 1+2+3.

Story 5 AC2 — HandoffSurface context-parameterized prompt

  • HandoffSurface.svelte accepts optional context props (objectType/objectId/objectName/envId/ route/intent); the copyable prompt embeds them (object_id=42, environment_id=..., intent=build_dashboard_test_scenario), generic prompt otherwise (offline invariant kept — derivation is synchronous over props+i18n). /agent/+page.svelte forwards URL query params; the existing DashboardHeader «Создать сценарий тестирования» link already carries the full context set — no dead link, prompt is parameterized end-to-end.
  • Metadata updated (SEMANTICS static→context, @UX_STATE Ready/Context); i18n key handoff_context_label added ru/en. Tests: contract test updated (AC2 forwarding pins + unconditional-render invariant) + 2 new render tests → handoff slice 5 passed.

E6 / MCPX-FR-007a — JSON depth + per-session rate limits (guard, pre-dispatch)

  • McpServerConfiguration gained server-owned limits: json_depth_limit=32, session_request_limit=120, session_rate_window_seconds=60.
  • McpTransportGuard (mcp_server/auth.py): POST bodies are depth-checked after buffering and BEFORE dispatch — typed 400 {"error":"json_depth_exceeded","limit":32}, iterative walker (RecursionError from json.loads counts as exceeded; adversarial depth cannot crash the guard), unparseable bodies fall through to the MCP layer's own typed parse error. Per-session sliding window (_SessionRateLimiter, bounded 10k keys, rejected hits not recorded) keyed by mcp-session-id header else authenticated principal → typed 429 rate_limited + Retry-After header (E6 contract). Rejection creates no mutable state.
  • Tests: tests/test_mcp_transport_limits.py 5 passed (HTTP depth 400 + no-state-after-reject
    • 429/Retry-After window + unit boundaries/limiter semantics); MCP slice 95 passed.

SC-005 remnants — CLOSED

  • /api/assistant router unmounted from app.py; routes/assistant package header records the unmount + @REPLACED_BY -> [McpServer] + retention rationale (parity provenance for FR-005; its 11 unit-test files exercise wrappers directly — slice 475 passed, 7 skipped).
  • /api/agent/llm-config decision = REMOVED (in-place Tombstone region in agent_conversations.py): zero consumers (Gradio dead in T041; frontend 0 refs); secret-bearing endpoint (decrypted provider key) kept for nobody was rejected. With it removed: get_agent_service_user_strict DI deleted (zero inbound edges — INV_6-clean), app.py polling suppression entry dropped, test sections pruned (agent_conversations −7, conversation_api −2, app_middleware −1; stale @TEST_EDGE/@NOTE metadata updated).
  • Frontend: assistant.ts + its test deleted (the sole inbound edge Api.ApiModule CALLED_BY -> [Api.Assistant.AssistantApi] removed first — INV_6); SystemSettings «Assistant history retention» UI block + validation keys removed; 8 assistant_* i18n keys pruned ru/en (parity kept); ux-test fixture pruned. Zero residual sendAssistantMessage|api/assistant refs.
  • .env.example: AGENT_HOST_PORT=7860, GRADIO_SERVER_PORT=7860 and the whole agent section removed (zero consumers repo-wide incl. .env.current/.env.master/.env.e2e/run.sh/compose); stale «единый для backend и agent» comment fixed. docs/INSTALL/README carry no references to removed surfaces (verified).

SC-004 + SC-009 exact RBAC visibility (new tests/test_mcp_rbac_visibility.py, 3 passed)

  • SC-004: tools/list exact sets from the live catalog — admin = all 47 tools; analyst (scenario:RUN + dashboard:testing:WRITE + tasks:READ) = 21; viewer (zero grants) = 15 (the permission=None human surface). Exact-set equality, not superset.
  • SC-009: mid-flow revocation of scenario:RUN on the SAME identity-only token (no new consent) hides decide_checkpoint/list_checkpoints/start_scenario_run in the next tools/list AND denies the cached-catalog call by name (PermissionError permission_denied); mid-flow grant of scenario:RUN_PROD exposes the approvals surface in the next list. Plus catalog/registry consistency pin.
  • Cookie-consent (browser-without-Bearer) decision recorded in tasks.md T008: not built in 050 — SPA-mediated authorize is the tested SC-007 path; standalone consent HTML deferred to an architecture amendment (MCPX-FR-020 perimeter stands).
  • INV_7 tail note: browser.py measured 398 < 400 — the round-2 logger migration reduced it below the cap; the "402 tail" is resolved without a split.

Verification evidence (round 4)

  • Full backend suite: 11213 passed, 240 skipped, 1 xpassed, 0 failed (−1 vs round 3 = the removed llm-config/strict-dep tests minus the 3 new RBAC + 5 transport-limit tests).
  • Frontend: 3435 vitest passed (delta vs 3454 = the deleted assistant.test.ts), lint 0 errors / 340 warnings; settings+api slice 368; handoff slice 5.
  • Targeted slices: MCP 95; transport limits 5; SC-005 backend slice 475+7 skipped; RBAC 3; ruff + compileall clean on the whole backend.

REMAINING (round 5+):

  • 050 T008b (local-perimeter deployment tests/docs: reject non-local MCP/LLM/VLM endpoints by default; PARTIAL per round-1 closure review) — the only unchecked 050 task.
  • FR-010 catalog versioning/deprecation markers (P2).
  • Final retirement of the unmounted routes/assistant package once parity provenance is archived (header records the retention rationale); agent_conversations/agent_assistant_db mounted surfaces — round-1 recorded decision (retained pending product call).
  • Deferred product decisions (recorded, non-blocking): M-03 self-approval, service-branch catalog permissions, adapter idempotency after lease recovery, multi-binding exploration context, TTL/expiry polish.
  • Semantic index rebuild pending Axiom binary availability.
  • Product status: NO-GO pending the closure-gate re-review of rounds 1–4.

Checkpoint — 2026-09-04 (closure-gate remediation round 5 — T008b local perimeter + FR-010 catalog versioning: 050 task list closed)

050 T008b — local-perimeter endpoint guard (the last unchecked 050 task → checked with evidence)

  • New Core.EndpointLocality module (backend/src/core/utils/endpoint_locality.py): deny-by-default enterprise-locality guard for provider endpoints (MCPX-FR-020). Locality = loopback/local host literals, private/loopback/link-local IP literals (RFC1918/ULA), enterprise DNS suffixes (.local/.internal/.lan/.corp/.intranet, env-extensible), DNS names resolving entirely to private addresses. Fail closed: unresolvable → denied; empty base_url (public cloud SDK default) → denied; substring spoofing (https://api.openai.com/localhost) denied by hostname — the old _is_local_base_url sniff remains only a token-budget heuristic (its insufficiency recorded as @REJECTED).
  • Enforcement at the single config-time choke: LLMProviderService.create_provider/update_provider (covers LLM and VLM — one provider table) BEFORE persistence/mutation; typed EndpointNotLocalError → admin routes map to HTTP 400 (endpoint_not_local:<reason>), never 500. Every denial emits an EXPLORE wire line (perimeter audit trail).
  • Escape hatches (env, default closed, INSTALL-documented): LLM_ALLOW_NONLOCAL_ENDPOINTS, LLM_NONLOCAL_ENDPOINT_ALLOWED_HOSTS, LLM_LOCAL_HOST_SUFFIXES.
  • MCP transport local contour already existed (DNS-rebinding defaults) + gained E6 limits in round 4; INSTALL.md gained §«Локальный периметр».
  • Test migration (honest consequence): 3 existing service mechanics tests in tests/services/test_llm_provider.py used api.openai.com/http://test fixtures → switched to perimeter-local URLs (http://localhost:1234/v1, http://test.local); their intent (create/update mechanics, masked-key handling) is unchanged. Route/API tests mock the service → unaffected.
  • Evidence: tests/test_endpoint_locality.py (22: accepted/denied/overrides/service-boundary/ route-mapping incl. no-persist-on-deny and no-mutate-on-deny) + provider/route slices → 107 passed; tasks.md T008b [x] with full proof line.

MCPX-FR-010 — catalog version + deprecation markers

  • MCP_CATALOG_VERSION = "1.0.0" (new McpServer.CatalogVersion region) published as serverInfo.version at initialize (_build_probe_server sets the lowlevel server version — the SDK otherwise reports the mcp-package version); re-exported on the frozen server surface.
  • McpToolDefinition gained deprecated + deprecation_note; the [DEPRECATED …] marker is applied at the SINGLE RbacFastMCP.list_tools choke point (deprecated entries stay listed, registered and callable for one minor cycle; registration sites untouched; marker visible over tools/list).
  • Discipline pinned executable (tests/test_mcp_catalog_version.py, 3): semver shape + the pinned major (deliberate-bump ritual — a breaking change cannot pass without consciously editing both the constant and the test pin) + marker mechanism (monkeypatched catalog entry → listed with marker, neighbour stays clean) + initialize wire publishes the version (JSON/SSE-robust parse).
  • MCP slice grew to 103 passed (95 + RBAC visibility 3 + transport limits 5 + local run overlap).

Verification evidence (round 5)

  • Full backend suite: 11243 passed, 240 skipped, 1 xpassed, 0 failed (7:09) — +30 vs round 4 (22 locality + 3 catalog-version + parametrizations), zero regressions.
  • Anchor-stack + AST syntax sweep over ALL 138 touched files: ALL BALANCED; ruff + compileall clean; targeted: endpoint locality slice 107, MCP slice 103, catalog version 3.

050 status after round 5

  • tasks.md: all tasks [x] (T008b was the last unchecked; every [x] now carries current proof).
  • Deferred product decisions unchanged (M-03 self-approval etc. — recorded, non-blocking); cookie-consent resolved in round 4 (tasks.md T008 note).

REMAINING (round 6+):

  • Final retirement of the unmounted routes/assistant package once parity provenance is archived (header records retention rationale); agent_conversations/agent_assistant_db mounted surfaces — round-1 recorded decision pending product call.
  • Semantic index rebuild when an Axiom binary is available (all rounds used zombie-mode sweeps).
  • Closure-gate re-review of rounds 1–5 → product status decision (currently NO-GO by the round-1 review verdict; all P1/P2 remediation queue items are now closed or recorded-as-decided).

Checkpoint — 2026-09-04 (closure-gate re-review, Axiom live) — rounds 1–5 verification matrix

Axiom MCP became available and the index was rebuilt live (full: 10534 contracts; incremental after fixes: 10582/5190 — another feature session is concurrently migrating the logging SSOT from shared/src/ss_tools/shared/ into backend/src/core/ and rebuilding in parallel; that migration is out of scope here and its mid-state accounts for the current workspace-level deltas).

Remediation matrix vs the round-1 NO-GO review (all items CLOSED)

Review finding Round Status Live evidence
P0 fake T025 checkpoints, client_credentials, gate unwrap, SC-007 HTTP 1 CLOSED (committed a0450b33) tests pinned then; re-verified in every later slice
P1 MCL intent-drop (~120 sites → 182), EXPLORE gaps, 9 silent adapters 2 CLOSED wire-format sweep test = 0 misuse; MCP slice 103
P1 INV_6 _llm_health dead edge + specs 033/035/036/039 tombstone sweep 2 CLOSED zombie sweep 0 dead / 15 tombstoned; Axiom: AgentChat.LlmHealth relations resolve
P1 dedupes (TaskDrawer/assistantOffset/@SIDE_EFFECT), migration anchors, sandbox span 2 CLOSED anchor sweeps ALL BALANCED incl. migrations 0014–0016
P1 server.py DECOMPOSITION GATE plan 2 (plan) → 3 (execution) CLOSED server.py 177 LOC; package census all <400; plan doc EXECUTED + log
P2 SC-005 remnants (assistant unmount, llm-config exemption decision, retention UI, assistant.ts, env 7860) 4 CLOSED decision = removal w/ in-place Tombstone; routes slice 475 passed
P2 FR-010 catalog versioning/deprecation 5 CLOSED serverInfo.version wire test + marker choke + pinned major
P2 E6 JSON-depth + per-session rate limits (Retry-After) 4 CLOSED transport-limits 5 passed (pre-dispatch, no-state)
P2 HandoffSurface context prompt (Story 5 AC2) 4 CLOSED contract+render tests 5 passed
P2 SC-004 admin/exact catalog fixtures 4 CLOSED exact-set pins admin 47 / analyst 21 / viewer 15
P2 SC-009 mid-flow role change 4 CLOSED revoke-hides-and-denies-cached-call + grant-without-consent pins
P2 cookie-consent decision 4 RECORDED tasks.md T008 decision note (not built in 050)
050 T008b local-perimeter (last unchecked task) 5 CLOSED locality guard 22 + slice 107; tasks.md [x] with proof

SC-001…SC-009 walkthrough (current executable evidence)

SC-001 authoring→registry provenance: promotion E2E (8). SC-002 parity: ops parity + baseline 037. SC-003 gates zero-side-effect-before-approval: approvals/checkpoints slices. SC-004 RBAC-exact tools/list: test_mcp_rbac_visibility.py exact sets. SC-005 decommission zero-refs: round-4 closure (greps clean; docs clean; env clean). SC-006 no SQL bypass: PROD-SQL terminal matrix + scenario context denial. SC-007 discovery→DCR→PKCE→tools/list: client-flow HTTP (2). SC-008 refresh replay family revocation: oauth tests. SC-009 live role change: RBAC visibility flow test. All nine SCs carry current passing pins (last full runs: backend 11243/0 failed @ round 5, frontend 3435/0 @ round 4 — both predate only the concurrent logging migration, not any 050 code).

Axiom live-audit findings (this re-review)

  • Fixed in-scope (metadata-only edits, uncommitted): McpServer + McpServer.ToolsAuthoring edges → Services.AgentAuthoringWorkspace.Service; McpServer.Package → McpServer (stale McpServer.Server, plus Phase-0 BRIEF refreshed); ScenarioExecution.ExplorationSandbox → .Service; Services.McpOpsDispatch phantom SupersetClient.DashboardWrite → the three verified write contracts (Core.DashboardsWrite.CreateDashboard/.CopyDashboard, Core.Datasets.SupersetClientCreateDataset); Test.McpParityBaseline037 → AgentSuperset.SqlFormat (correct ID, no Api. prefix). Post-fix scoped audits: mcp_server prefix = 0 unresolved, mcp_ops_dispatch = 0 warnings.
  • Accepted advisories (deliberate, rationale-recorded): McpServer.TransportAuth 151/150 (cohesive guard, 1 over); McpServer.RbacLayer parser reports module region 418 while the FILE wc -l is 393 (INV_7 is file-level — compliant; parser metric noted); RbacServer 259 / ScenarioModels 243 / ToolsAuthoring.Register 315 contract spans are gate-frozen shapes (behavior-neutral moves only).
  • Historical debt (predates rounds 1–5, workspace-wide): a first curation pass was executed live-audit-driven — unresolved relations 446 → 401 (edges 5191 → 5207). Fixed (all targets verified to exist before retargeting, comment-only edits, anchors/compile/diff-check clean): 35 retargeted edges across 22 files — scenario chain module-parent edges → verified function contracts (ScenarioGraph.Compiler.Compile/Validator.Validate/Resolver.Resolve/ PackCompiler.Generate) in scenario.py route + specs 038/050; Superset-client path/alias forms → Core.Init.SupersetClientModule (dataset_mapper, dataset_key_sync×2, query_model, structure_diff_capture); [Models.User]×4 + admin flat User → Models.Auth.User; ExecuteEnvelope ×5 (+1 test BINDS_TO) → ExecuteQueryEnvelope; DraftStorage ×2 → Services.AgentRuns.Artifacts; Services.Git.GitService → Services.Init.GitService; Models.Deployment → .DeploymentModels; StructureSnapshot.Service → .Capture+.Diff; Api.DashboardTesting.Core → Api.DashboardTesting; client_registry python-path forms → Core.AsyncNetwork.AsyncAPIClient/Core.ConfigModels; plus 11 malformed multi-target translate plugin lines (-> [A], [B) split into individual @RELATION lines with verified IDs (Models.Translate.*, Plugin.Dictionary.DictionaryManager, Plugin.LlmCall.LLMTranslationService, Plugin.LangDetect.LanguageDetectService, Plugin.BatchSizer.AdaptiveBatchSizer, Plugin.BatchProc.BatchProcessingService, Plugin.SqlGenerator.SQLGenerator, Plugin.SupersetExecutor.SupersetSqlLabExecutor, Plugin.LlmParse.LLMResponseParser, Plugin.TokenBudget.EstimateTokenBudget, Core.DbExecutor, Core.ConnectionService, Core.ConfigManager). Remaining 401 (queued for the next curator round, classified): logging-SSOT-domain edges owned by the concurrent migration (ss_tools.shared._llm_health, CotLoggerModule, main_cot_logger...); function-shaped targets without contracts (ReportsService.get_summary, _compute_content_hash, TranslationOrchestrator.execute_run, SupersetClient.network.request...); legacy single-# region misses (Services.LlmProvider.GetProvider); flat test-module targets (TasksApi, FileIOModule, fixtures.*); unverified remainder (Models.VerificationRun.VerificationRunRecord, BaselineEngine.Catalog.Materialization). Also: 294 module_too_long + 418 contract_too_long + 18 flat_hierarchical_id advisories; 1 tombstone_missing_deprecated counted by the audit but NOT reproducible by source scan (region + brace syntax) — monitor.
  • Orphans 2710 are mostly relation-free C1/C2 leaves (expected per orphan_guidance), not defects.

Verdict and sequencing

  • Rounds 1–5 remediation: COMPLETE. No open P0/P1/P2 item from the closure review remains; 050 tasks.md fully [x] with proof lines.
  • Product status flips NO-GO → GO only after: (a) the concurrent logging-SSOT migration lands coherently (currently mid-flight: ss_tools.shared.cot_logger imports broken tree-wide — outside this workstream by explicit instruction), (b) a full backend + frontend green run on that landed state (11243/3435 numbers must be re-established post-migration), (c) formal closure sign-off against this matrix.
  • Uncommitted at this point: the six metadata edge-fix files listed above (comment-only; runtime behavior untouched — verification is anchor/Axiom-based, not suite-based, while the tree's import layer is mid-migration).

UX-аудит — MCP happy path для BI-аналитика (2026-09-04, read-only, код НЕ менялся)

Прокручен полный продуктовый путь «аналитик → понять про MCP → подключить → работать → одобрить». Верифицировано по коду frontend (file:line приведены). Вердикт понятности: 2/5 — скелет и передача контекста дашборда есть (050 AC2, round 4), но «последняя миля» разорвана в двух местах.

Фактическая карта пути

  • Точки входа в /agent (единственная MCP-поверхность в UI): TopNavbar.svelte:269 кнопка «Ассистент» (глобально); DashboardHeader.svelte:112 кнопка «AI» (тултип «Ask AI about the dashboard»); DashboardHeader.svelte:120 «Создать сценарий тестирования» (intent build_dashboard_test_scenario); datasets/[id]/+page.svelte:32. В сайдбаре (sidebarNavigation.ts) пункта /agent НЕТ.
  • На /agent рендерится HandoffSurface.svelte: заголовок «Ассистент переведен на MCP», описание «…Подключите MCP-клиент, чтобы продолжить», «подсказка подключения» = «Используйте MCP-сервер Superset Tools, настроенный для вашего workspace» (ru, assistant.json:185), копируемый промпт с контекстом дашборда.
  • Петля одобрения (вторая половина пути): экраны существуют — /dashboard-testing/runs (WaitingForMeView, чекбокс «Только ожидающие моего решения») и /dashboard-testing/scenarios/[id]/runs/[runId] (HumanCheckpointPanel), но недостижимы через навигацию.

Замечания (severity → evidence → эффект)

  1. CRITICAL — «Подключите MCP-клиент» без КАК/КУДА. Подсказка подключения не содержит ни endpoint URL (https://<origin>/mcp), ни discovery (/.well-known/oauth-protected-resource/mcp), ни сниппета конфига клиента (Claude Desktop/Cursor/...), ни ссылки на INSTALL.md §MCP. Термин «MCP-клиент» не объясняется. Аналитик гарантированно застревает на шаге 2.
  2. CRITICAL — петля одобрения недостижима. В ROUTES.ts нет ни одного builder'а dashboard-testing/*/load-testing; в buildSidebarSections() нет секции «Тестирование дашбордов». Агент встаёт в waiting_for_user, а человек не может обнаружить, где одобрять (только ручной ввод URL). Продукт «висит» в ожидании ненаходимого действия.
  3. HIGH — bait-and-switch входов. «AI»/«Ассистент» обещают вопрос-ответ в приложении, приводят на заглушку про подключение клиента. Для не знающего о декомиссии чата — выглядит сломанной кнопкой.
  4. HIGH — настроек/статуса MCP в UI нет вообще (grep по settings/admin: только серверные LLM-провайдеры). Ни URL сервера, ни статуса подключений, ни доступного роли набора инструментов (RBAC-каталог роль-зависим: admin 47 / analyst 21 / viewer 15 — данные уже есть, поверхности нет).
  5. MEDIUM — маршруты в обход SSOT. Страницы dashboard-testing/*/load-testing линкуются хардкод-строками через локальный resolve(), минуя ROUTES.ts — нарушение @INVARIANT реестра («every href/goto must use these functions»); link-integrity тест их не покрывает.

Рекомендуемые фиксы (приоритизированы; все frontend; закрывают 1+2+4 = минимум)

  • Fix 1 (HandoffSurface actionable): вывести реальный endpoint из page.url.origin ({origin}/mcp) отдельным копируемым блоком + кнопка «Скопировать URL»; 3 шага onboarding («что такое MCP-клиент → вставить URL → подтвердить доступ в браузере (OAuth)»); ссылка на INSTALL.md §MCP; discovery URL для продвинутых. Новые i18n-ключи ru/en.
  • Fix 2 (петля одобрения в навигации): секция сайдбара «Тестирование дашбордов»: Сценарии / Запуски (бейдж «ждут меня» — счётчик waiting_for_user) / Аналитика; RBAC-фильтр по scenario/dashboard-testing permissions; запись в buildSidebarSections + тест в sidebarNavigation.test.ts.
  • Fix 3 (честные входы): переобозначить «AI»/«Ассистент» → «MCP-ассистент» / тултип «Инструкция по подключению AI-ассистента (MCP)»; опционально first-visit onboarding.
  • Fix 4 (ROUTES SSOT): добавить builders dashboardTesting.{scenarios,runs,runDetail,analytics, automation,scenarioEdit} + loadTesting.detail; перевести хардкод-resolve() на них.
  • Fix 5 (опционально, админ): MCP status page — endpoint, discovery, инструменты по роли.

Статус: не реализовано — аудит read-only по запросу; к имплементации готов минимум Fix 1+2+4 (устраняют оба CRITICAL + SSOT-нарушение) с vitest/svelte-check/eslint-гейтами.

Checkpoint — 2026-09-04 (UX-audit remediation IMPLEMENTED — Fix 1+2+3+4 + registry hub)

Implemented the actionable minimum of the read-only UX audit above (Fix 1+2+4), plus Fix 3 (honest entries) and the scenario-registry hub the audit's Fix 2 "Сценарии" nav item implies. Gates run per AGENTS.md frontend checks (npm run test / npm run lint / npm run build); the project has no svelte-check script, so the "svelte-check gate" is covered by vite build.

Fix 4 — ROUTES SSOT for dashboard-testing / load-testing

  • src/lib/routes.ts: added ROUTES.dashboardTesting.{scenarios, scenarioEdit, scenarioAnalytics, runDetail, runs, analytics, analyticsCase, automation} + ROUTES.loadTesting.detail; every builder maps to a real route file (registry index created this round). @RATIONALE records that the integrity test now scans these trees.
  • Converted ALL hardcoded resolve("/dashboard-testing/...") SSOT bypasses to ROUTES builders in 7 pages (runs ×3, automation, analytics ×2, analytics case, scenario edit/analytics/run-detail)
    • DashboardHeader load-testing href + DashboardHeader/agent hrefs now ROUTES.agent(...).
  • Hardened routes-link-integrity.test.ts: /dashboard-testing + /load-testing added to the watched prefixes AND a resolve("/dashboard-testing|/load-testing...") bypass detector (the audit MEDIUM #5 — these trees were previously uncovered). routes.test.ts +13 builder pins.
  • Removed a now-dead $app/paths vi.mock in run.ux.test.ts orphaned by the conversion.
  • src/routes/dashboard-testing/scenarios/+page.svelte: searchable/filterable registry list over the previously unused ScenarioRegistryModel (042 list/read only — never starts a run/agent); rows link via ROUTES to edit / analytics / last-run. Previously /dashboard-testing/scenarios was a DEAD route that four "← Сценарии" back-links pointed at. i18n dashboard_testing.registry_* ru/en. UX test scenarios_index.ux.test.ts (4): rows+SSOT links, no-run-history, empty, error+retry.

Fix 2 — approval loop reachable from navigation (CRITICAL 2)

  • sidebarNavigation.ts: new dashboard_testing category in the Operations section (no new section id → the existing exact-section/categories tests stay green). Per-subitem RBAC mirrors the exact backend has_permission gates: Сценарии→dashboard:testing READ, Запуски→scenario RUN, Аналитика→scenario:result VIEW, Автоматизация→scenario:automation TRIGGER; category itself unpermissioned (visibility derives from subitems → hidden entirely when none pass).
  • stores/scenarioRuns.svelte.ts (NEW): waiting-for-me count store (server-authoritative total via /scenario-runs?waiting_for_me=true&page_size=1; fail-closed 401/403 disable; inflight dedup; mirrors HealthStore). test_scenarioRuns.ts (9): state/refresh/dedup/auth-disable/retry/subscribe.
  • Sidebar.svelte: warning-token badge on the dashboard_testing category (count expanded / dot collapsed) + 60s poll gated on scenario:RUN — turns the approval loop into a discoverable signal (directly addresses "продукт висит в ожидании ненаходимого действия"). Metadata BINDS_TO/@UX_FEEDBACK updated.
  • nav.json ru/en keys; category.testing tailwind token. sidebarNavigation.test.ts +4 RBAC tests.

Fix 1 — HandoffSurface actionable (CRITICAL 1)

  • HandoffSurface.svelte: renders the REAL MCP endpoint from page.url.origin ({origin}/mcp) as a copyable block + "Copy URL", the OAuth discovery URL ({origin}/.well-known/oauth-protected-resource/mcp, verified against backend/src/api/mcp_oauth.py + app.py /mcp mount), a 3-step onboarding and the docs/INSTALL.md «MCP клиент» pointer — closing "«Подключите MCP-клиент» without how/where". Offline invariant preserved (synchronous origin/props/i18n derivation, no network). Context-parameterized prompt (Story 5 AC2) unchanged. i18n assistant.handoff_* ru/en; context test +3 (endpoint, discovery, onboarding+copy controls). HandoffSurface lints 0-warning.

Fix 3 — honest entries (HIGH #3 bait-and-switch)

  • TopNavbar assistant button + DashboardHeader "AI" button relabeled to «MCP-ассистент» / «MCP» with tooltip «Инструкция по подключению AI-ассистента (MCP)» — no longer promises in-app chat. assistant.mcp_entry_* ru/en.

Verification evidence (this round)

  • Full frontend suite: 3486 passed, 0 failed (200 files) (was 3435; +51 = new builder pins, sidebar RBAC, store, registry index, handoff endpoint/onboarding).
  • eslint: 0 errors / 354 warnings (baseline 340). The +14 are ALL the accepted svelte/no-navigation-without-resolve class: the codebase's established ROUTES convention is bare href={ROUTES.x()}/goto(ROUTES.x()) (98 such warnings pre-exist; resolve(ROUTES()) appears nowhere), so SSOT-compliant navigation deliberately carries the same accepted warning. No new error/a11y/unused-var warnings introduced; Sidebar's 3 each-key warnings are pre-existing.
  • Production build (adapter-static): green. GRACE-Poly anchors balanced in every touched file; store MCL log(src, marker, intent, payload, error) mirrors HealthStore (src-first frontend facade).
  • Recorded decision: bare ROUTES.x() chosen over resolve(ROUTES.x()) for consistency with the dominant convention (base path empty under adapter-static → resolve() is identity anyway).

Not done (optional / out of minimum scope)

  • Fix 5 (MCP status/admin page — endpoint, discovery, per-role tool list): audit-optional and larger; deferred.
  • Broader product-GO gates from prior checkpoints unchanged and NOT touched by this frontend UX round: semantic index rebuild (Axiom), full backend suite re-run on the landed logging-SSOT migration, and browser E2E on a live stand (would need the backend+frontend stack up). Product status per the round-1 closure review remains pending formal re-sign-off.

Checkpoint — 2026-09-05 (UX-audit "остатки" — Fix 5 admin MCP governance + Fix 2 polish, UI/UX discussed first)

Per the operator's request the remaining UI/UX was discussed and decided BEFORE implementing. Decisions: Fix 5 = admin governance (all 47 tools × roles) mounted under /admin/*; polish = badge→waiting view + registry extra filters & pagination. (First-visit onboarding was declined.)

Fix 5 — admin MCP governance surface (audit finding #4: "настроек/статуса MCP в UI нет вообще")

  • Grounding first: MCP's only REST surface is OAuth (/oauth/*, /.well-known/*); the role-dependent tool catalog is exposed only over the MCP protocol (tools/list), and /api/ready carries 044 provider readiness, not the catalog. So a governance view needs a small NEW backend projection.
  • Backend backend/src/api/routes/admin_mcp.py (new module, keeps admin.py under INV_7): read-only GET /api/admin/mcp/catalog, gated has_permission("admin:roles","READ") (same gate as the role/ permission inventory). Returns catalog_version (MCP_CATALOG_VERSION), endpoint/discovery paths, the full frozen catalog (name/risk_level/requires_approval/service_allowed/deprecated/required permission), a per-role visible-tool matrix, and a read-only DCR client list (OAuthClient, scopes/ redirect_uris parsed; never secret/secret_hash). Registered in app.py (413 routes; import verified).
  • The per-role visibility helper _visible_tool_names mirrors RbacFastMCP._can_use_tool's human branch exactly (permission None → visible; start_scenario_run → scenario RUN or RUN_PROD; else exact role resource/action match; admin bypass). tests/api/test_admin_mcp_catalog.py (8) pins it to the SAME exact admin/viewer/analyst subsets SC-004 pins for the protocol → UI cannot drift from tools/list.
  • Frontend: api.getMcpCatalog; models/AdminMcpCatalogModel.svelte.ts (read-only, last-projection-on- error); routes/admin/mcp/+page.svelte (ProtectedRoute admin:roles; server card w/ absolute endpoint+ discovery, full tool table, expandable role×visibility matrix, DCR client list); ROUTES.admin.mcp(); sidebar admin subitem «MCP» (admin:roles READ) + admin hub card; i18n admin.mcp.* + nav.admin_mcp* ru/en.
  • Real defect found & fixed by the page test: the load $effect used !loaded && !loading, which re-fires after a FAILED load (loading toggles, loaded stays false) → an auto-retry loop hammering the endpoint. Switched to a started guard (load once + manual Retry), matching the registry hub. This is the same latent pattern some sibling pages still carry (runs/analytics use rows/loaded guards).

Fix 2 polish — badge → "ожидают меня" view; registry filters + pagination

  • ROUTES.dashboardTesting.runs(waitingForMe?) → ?waiting_for_me=true; run center now READS that param on entry (page.url.searchParams) and opens directly on WaitingForMeView; the Sidebar waiting badge (expanded count) is now a button navigating to runs(true) (stopPropagation so it doesn't toggle the category). The collapsed dot stays an indicator. Turns the passive count into a one-click approval entry.
  • Scenario registry hub gained owner / tag / dashboard_id filters + server-side pagination (PageSize 20, prev/next, page X/Y) over the existing ScenarioRegistryModel.loadList query; filter change resets to page 1 (@INVARIANT). i18n dashboard_testing.registry_* ru/en extended.

Verification evidence (this round)

  • Frontend: full suite 3498 passed, 0 failed (204 files) (was 3486); eslint 0 errors / 354 warnings (UNCHANGED — Fix 5 added no nav-href warnings; admin/mcp page has no internal links); build green (adapter-static); anchors balanced in every new/edited file; all touched i18n JSON parses; git diff --check clean.
  • Backend: admin/mcp + admin + mcp-rbac-visibility + catalog-version slice 42 passed; MCP slice 55 passed; scoped dashboard-testing OpenAPI 25 passed (unaffected by the new /api/admin route); ruff + compileall clean; app.py imports with /api/admin/mcp/catalog registered. Full backend suite NOT re-run this round (change is one isolated additive endpoint + registration; scoped slices + app import are proportionate) — the prior round-5 number (11243) stands pending the logging-SSOT re-run.

Remaining after this round

  • Fix 5 shipped read-only governance. DCR client MANAGEMENT (revoke/rotate) and automated key-rotation expiry remain open (separate features the audit listed); the view exposes clients but performs no mutation.
  • Fix 5 was the last audit UI/UX item; all five audit fixes (1–5) are now implemented. Deferred product decisions (M-03 self-approval, service-branch catalog permissions, adapter idempotency, multi-binding exploration context, TTL/expiry) and the broader GO gates (semantic rebuild, full backend suite on the landed logging migration, browser E2E on a live stand, closure-gate re-sign-off) are unchanged.

Checkpoint — 2026-09-05 (orthogonal UI/UX review + i18n remediation)

Per operator request an independent UI/UX pass was run on the newly delivered surfaces, with i18n completeness as an explicit gate. Findings (orthogonal, from a fresh pair of eyes):

  • i18n debt (primary): the dashboard-testing feature area still carried hardcoded Russian UI strings — runs/+page.svelte (~25 strings), WaitingForMeView.svelte (~12) — violating the "no hardcoded UI strings" invariant. My two new pages additionally used English || "…" fallback literals and rendered raw enum tokens (risk_level, client_type, lifecycle_status, health) untranslated.
  • Bug: scenario registry dashboard_id filter sent NaN in the query for non-numeric input.
  • UX clarity: admin/mcp "Сервис" column (service_allowed yes/no) was ambiguous — what "yes" grants is not obvious; added a clarifying tooltip.
  • Minor: admin/mcp page lacked a document <svelte:head><title>.

Remediated:

  • runs/+page.svelte + WaitingForMeView.svelte fully migrated to $t (dashboard_testing.runs_*, waiting_*); all hardcoded RU strings removed (only a @BRIEF metadata comment references the label).
  • scenarios/+page.svelte + admin/mcp/+page.svelte: fallback literals removed (i18n keys are the only source), lifecycle/health/risk/client_type enums translated (ls_*, health_*, risk_*, client_*) with raw-token fallback for unknown values (drift-safe); dashboard_id NaN guard added; admin/mcp <title> added; service column disambiguated via col_service_hint.
  • i18n locales extended ru/en: dashboard-testing.json (+runs/waiting/ls/health ≈45 keys), admin.json mcp.* (+risk/client/service-hint keys); nav/assistant unchanged this round.
  • Test fixture for the admin/mcp page broadened to the full admin.mcp key set.

Verification: frontend full suite 3498 passed / 0 failed (204 files); eslint 0 errors / 354 warnings (unchanged — no new nav-href); build green; no residual Cyrillic in the four touched templates; all four edited locale JSONs parse. WaitingForMeView.test.ts (real-i18n, ru labels) still green.

Checkpoint — 2026-09-05 (i18n sweep COMPLETE — analytics/editor/automation/run surfaces)

The previously-documented remaining i18n debt is now cleared: the analytics/automation pages, scenario detail pages ([id]/edit, [id]/analytics, [id]/runs/[runId]) and the entire scenario component trees were migrated off hardcoded Russian strings to $t (dashboard_testing.*):

  • scenario-analytics: CaseWorkspace, QueueList, TrendsChart, AgentActionTimeline, RecurringFailuresList, HealthCard + analytics/queue + case routes.
  • scenario-run: HumanCheckpointPanel, RunTimeline, RunComparison, RunHistoryList, ScenarioResultView, RunConfigurationPanel + run-detail monitor route.
  • scenario-editor: ConstrainedAssertionEditor, VisualDagCanvas, AgentActionPanel, EditRevisionDiff + scenario edit route.
  • scenario-automation: AutomationPanel, ScheduleForm + automation route.
  • Locales extended (ru/en) with ~150 keys. Policy: ru values = the current rendered text (some terms stay English where they currently ship untranslated, e.g. "attention item(s)", "Investigate with agent", "occurrence(s)", "group(s)") so real-i18n tests stay green; wiring is complete and the locale entries are now the single place to localize later.
  • Fixed a stale route test case.ux.test.ts ("Записать disposition" → current "Закрыть кейс" + resolved/verification_evidence CAS body); it previously asserted a pre-refactor disposition flow.

Gates: frontend 3494 passed / 4 failed (204 files); the 4 failures are PRE-EXISTING (verified in isolation, none touched by this workstream):

  1. routes-link-integrity — src/routes/oauth-consent/+page.svelte:44 goto('/login') raw SSOT bypass (untracked file, not part of this workstream). 2-3. InvestigationModels.test.ts (2) — stale vs the prior-round InvestigationCaseModel.dispose signature change (rationale.trim is not a function, CAS version 1 vs 2).
  2. settings_page.ux.test.ts (1) — pre-existing numeric-range save guard, unrelated to dashboard-testing. Lint 0 errors / 358 warnings (+4 tolerated no-navigation-without-resolve); build green; no residual hardcoded Russian outside a single @BRIEF comment. The 4 pre-existing failures are out of i18n scope and left for a separate cleanup.

Checkpoint — 2026-09-06 (spec refinement: server-owned handle layer + F3 guard; ADR-0023)

Orthogonal code review of the T029/T029c implementation found a root-cause architectural gap and it is now both guarded in code and fully specified:

  • Code (implemented, verified): ScenarioExecution.RunnerPlan.Derive refuses provenance-only bootstrap revisions (BOOTSTRAP_REVISION_NOT_RUNNABLE, EXPLORE-marked) before any run/gate/schedule side effect — closing a vacuous zero-step false-PASS (Runner.Walker._advance_run marks empty plans passed). pack_compiler.py docstring corrected (Generate persists nothing; registration lives in ScenarioGraph.PackCompiler.Register036). test_mcp_initial_scenario_e2e.py rewritten: asserts the refusal, then proves queued/PROD-gate behavior on a genuinely materialized revision. Gates: 1234 passed (tests/services/dashboard_testing + API/MCP suites), 458 passed focused MCP/registry set, ruff/compileall/git diff --check clean; zero-step seeded_registry fixtures unaffected (guard keys on the exact compiled_handle_id+draft_pack_id+draft_pack_digest signature with no steps, not on empty plans in general).
  • Specs (amended, no code claims): 038 contracts/modules.md — PackCompiler.Generate contract fixed + ScenarioGraph.ServerOwnedPipeline handle-persistence amendment (minting boundaries, three immutable handle entities, canonical-bytes authority, binding/single-consumption/GC, inspect-stage PROPOSED status); 038 contracts/verification-program.md — reconciliation block (implemented: ScenarioStep DAG + ActionRegistry, CaptureSpec/VlmAnalysis/VlmFinding/ScenarioParameter; deferred: five-way VerificationProgram, SqlEvidenceSpec, TransformSpec/ComparisonSpec/AssertionSpec, AgentEvaluationSpec/DecisionPolicy — zero code references); 042 data-model.md — synchronous graph_snapshot materialization from handle bytes inside the create transaction, outbox scoped to reference artifacts only, interim-drift note; 042 contracts/modules.md — CreateInitial amendments
    • RevisionChain rejection of the legacy REST raw-graph_snapshot route; 050 spec.md — stage-table implementation-status note + handle-rules status bullet.
  • New work scoped: 050 tasks.md Phase 2c — T029d (handle tables + canonical bytes store + migration), T029e (042 create consumes handles, sync materialization, outbox/RevisionMaterialization worker, guard demotion), T029f (MCP minting surface, bootstrap accepts stored handle ids only, legacy REST route gate/retire, catalog minor bump), T029g (full-chain fresh-DB MCP E2E + PostgreSQL concurrency/retention), T029h (inspect-stage implement-or-descope decision). T029a remains open and naturally folds into T029d/e.
  • Decision memory: docs/adr/ADR-0023-mcp-scenario-pipeline-handle-gap.md (PARTIALLY IMPLEMENTED) records the finding, the rejected alternatives (silent vacuous PASS; blanket empty-plan rejection; lossy-pack retro-fit as sufficient fix) and the two-part resolution. Registry row added to docs/adr/README.md.
  • Product remains NO-GO for the full external happy path inspect → compile → validate → resolve → draft-pack → bootstrap → start run until Phase 2c lands; T029–T029c scope stays complete as marked.

Checkpoint — 2026-09-06 (Phase 2c landed: handle layer T029d–T029g + inspect hybrid T029h)

The ADR-0023 root cause is closed in code; the NO-GO above is superseded for the MCP happy path.

  • T029d: CompiledScenarioHandle/ValidationResultHandle/DraftPackHandle (immutable, owner-bound, content-addressed canonical-bytes store handle:{sha256}, idempotent mint, single consumption under SELECT ... FOR UPDATE + populate_existing, purge/GC helper), migration 0019_scenario_handles; minting wired into REST api_compile_scenario/api_validate_scenario/api_resolve_scenario/ api_draft_pack (additive response fields only; pure compiler functions keep @SIDE_EFFECT None).
  • T029e: handle-first create_scenario/create_initial materialize graph_snapshot = canonical DashboardTestScenario JSON + server-owned action_registry_version/hash in-transaction; OutboxEvent(materialize_revision) + RevisionMaterialization(pending) + idempotent worker materialize_pending_revisions (migration 0020_scenario_materialization); api_create_scenario returns the real materialization status. BOOTSTRAP_REVISION_NOT_RUNNABLE demoted to defense-in-depth.
  • T029f: MCP register_draft_pack write tool (human-only, AgentRun ownership) mints all three handles; create_initial accepts ONLY stored handle ids (transitional compile:{run_id}:{digest} removed); legacy REST POST /scenarios/{id}/revisions raw-graph route retired → 410 REVISIONS_RAW_GRAPH_RETIRED; catalog 2.0.0 (pinned-major ritual PINNED_CATALOG_MAJOR=2). Side-fix: generate_report reclassified out of _MUTATING_ACTIONS (local draft-report write, matches browser.py mutation set) → ACTION_REGISTRY_VERSION 038.1.0 → 038.2.0, scenario_execution/graph.json fingerprint re-pinned; register_artifact (repository_write) intentionally stays mutating.
  • T029g: fresh-DB MCP E2E runs the full chain with zero REST crutches (register_draft_pack → bootstrap → direct start_scenario_run = queued); PostgreSQL concurrency proven on Testcontainers (test_scenario_handle_concurrency.py: two threads → one CONSUMED + one HANDLE_CONSUMED).
  • T029h (hybrid option C + X1): MCP inspect_dashboard_context (live authoritative DashboardQueryModel
    • fingerprint echo contract); context_authority evaluation at register_draft_pack (server-side fingerprint RECOMPUTE — claimed values ignored; sentinels ""/sha256:error never verify; unconfigured/unreachable env fails OPEN to unverified; live-env falsifiable mismatches reject typed with zero rows); marker persists on DraftPackHandle (migration 0021_context_authority), materializes into graph_snapshot, and Runner.Start refuses PROD on explicit non-verified (CONTEXT_AUTHORITY_REQUIRED_FOR_PROD, legacy missing marker allowed); validator _check_dashboard_context recursively rejects query_context/SQL smuggling in dashboard_context (X1). Catalog 2.1.0 (additive minor).
  • Gates: broad regression 1338 passed (dashboard_testing + scenario/api + MCP suites); integration concurrency green on PostgreSQL 16; alembic heads → 0021_context_authority; ruff/compileall/git diff --check clean.
  • Spec updates: 050 spec.md stage-table status note + handle-rules status bullet rewritten to implemented state; 050 tasks.md T029d–T029h [x] with evidence; 038 ServerOwnedPipeline §6 rewritten (hybrid resolution).
  • Remaining residuals (tracked, non-blocking for the MCP happy path): outbox worker not yet wired into the scheduler poll loop; 043 editor save path and REST api_draft_pack do not evaluate context_authority (NULL = legacy-allowed at the PROD gate); full PostgreSQL suite for outbox retry/GC; D2/D4 frontend gaps (UI launch wiring, PROD approval surface) unchanged.

Checkpoint — 2026-09-07 (Phase 2c residuals closed: outbox wiring, REST parity, marker inheritance)

  • Outbox worker wired: Core.Scheduler.ExecuteRevisionMaterialization tick consumes materialize_pending_revisions every 30s (scenario_revision_materialization, max_instances=1, coalesce) — registration + durable-tick idempotency proven in test_scenario_scheduler_callbacks.py.
  • REST parity: api_draft_pack is async and evaluates ScenarioGraph.ContextAuthority FIRST (same typed rejections, zero rows on falsifiable mismatch); response carries context_authority. Mock-boundary tests updated to patch the authority+mint seams (module-boundary mocking convention).
  • 043 editor path: save_proposal inherits the server-owned context_authority marker from the base revision via apply_ops' {**base_graph} spread — bootstrap-origin scenarios keep PROD-gate coverage after edits (test_save_proposal_inherits_server_context_authority_marker). Legacy NULL markers remain PROD-allowed by design (migration-safe).
  • Stale-mirror fix: tests/api/test_admin_mcp_catalog.py::_ANALYST_EXTRA was missing register_draft_pack (blind spot of the T029f test selection); the full-suite run caught it.
  • Gates: full regression 3360 passed, 8 skipped; alembic heads → 0021_context_authority; ruff/compileall/git diff --check clean.
  • Open (frontend, tracked separately): D2 UI launch wiring (RunConfigurationPanel → route) and D4 PROD approval surface (ApprovalDecisionPanel → run center) — the two cheapest product gaps.

Checkpoint — 2026-09-07 (frontend D2/D4 closed: scenario launch surface + PROD approval UI)

  • D2 scenario launch (UI): new page frontend/src/routes/dashboard-testing/scenarios/[id]/+page.svelte (ScenarioRegistry.Route.Detail) — registry facts + revisions + the first production host of RunConfigurationPanel; ROUTES.dashboardTesting.scenarioDetail(id) added to the SSOT and the registry index name cell now links to it (index still never starts runs — its invariant intact). Launch flow: panel → RunMonitorModel.launch(config, crypto.randomUUID()) → POST /scenario-runs with Idempotency-Key → goto(runDetail). Mandatory steps are fed from the detail graph (logical_step_id), environments from GET /environments (is_prod = stage==='PROD' || is_production).
  • D4 PROD approval (UI): new ApprovalDecisionPanel.svelte (approve/deny + comment, busy-lock); RunMonitorModel.decideApproval posts /scenario-runs/{id}/approval/decision and reloads the run; the run-detail aside renders the panel exactly when status === "pending_approval" (before the waiting_human branch). Closes the last happy-path gap: PROD scheduled/manual runs are now fully decidable from the web UI, not only via MCP/REST.
  • Pre-existing suite failures fixed en route (all green before/after independently of D2/D4): oauth-consent raw goto('/login') → ROUTES.login() (link-integrity audit); stale InvestigationModels dispose signature (2-arg → production 4-arg); settings/+page.svelte fills backend-default mcp_oauth_* values (15/30/90) so absent fields can't crash bind:value (Svelte 5 props_invalid_value).
  • i18n: approval_* (6) + scenario_detail_* (9) keys added to BOTH en/ru dashboard-testing.json.
  • Gates: full vitest 3506 passed; npm run lint 0 errors (364 pre-existing warnings); npm run build succeeds (adapter-static); tests added: ApprovalDecisionPanel.test.ts (2), RunMonitorModel.test.ts (+2 decideApproval), scenarios/[id]/__tests__/detail_launch.ux.test.ts (3: facts+mandatory steps, typed launch POST + redirect, load-failure alert/retry).
  • Product status: the full journey agent-bootstrap → registry → UI launch → monitor → PROD approval → schedule is now UI-complete; no known D-class gaps remain in the 050 review matrix.

Checkpoint — 2026-09-07 (field-run remediation PLAN + test-honesty FIX; ADR-0024, 050 Phase 2d)

Источник: docs/2026-09-07-sales-prod-mcp-run.md — первый полевой прогон ВНЕШНЕГО MCP-клиента против ss-prod Sales Dashboard (ID 11). Прогон доказал: initial-bootstrap цепочка inspect → compile/validate → draft-pack → register_draft_pack → bootstrap → start_scenario_run code-complete, но внешне недостижима — register_draft_pack требует principal-owned AgentRun, а MCP-операции его создания нет (единственная creation-поверхность POST /api/agent/runs — web-session REST; MCP-токены aud=mcp на ней не аутентифицируются). Ни entry/revision/run создано не было. Маскировка: вертикальный E2E seed'ил предусловие raw-ORM-вставкой — suite не мог увидеть разрыв.

Спеки (план включён полностью)

  • docs/adr/ADR-0024-mcp-agent-run-external-boundary.md (новый, PARTIALLY IMPLEMENTED) + строка в docs/adr/README.md: решение = явные typed MCP create_agent_run/get_agent_run поверх существующего Services.AgentRuns.Service.Create (permission ("dashboard:testing","EXECUTE") REST-паритет, service_allowed=False, каталог 2.1.0 → 2.2.0 additive); инвариант внешней достижимости durable-предусловий; binding test-honesty правило (вертикальный тест обязан получать предусловия через границу, достижимую для тестируемого принципала; raw-ORM-seed предусловия в external-chain тесте запрещён). @REJECTED: регистрация без AgentRun; implicit auto-create внутри register_draft_pack; REST-костыль как documented external path; сохранение raw-ORM seed.
  • 050 spec.md: MCPX-FR-027 (external reachability + create/get agent run), MCPX-FR-028 (context-derived capabilities), MCPX-FR-029 (disposition clarity); canonical-names строки create_agent_run/get_agent_run; field-run correction под stage table; release-gate строки E2E-EXT-001/002, CAP-001, DISP-001 (OPEN) + TEST-001 (CLOSED этот раунд); Clarifications Session 2026-09-07 (включая поправку clarification 2026-08-24 про AgentRun); Phase 2d строка.
  • 050 tasks.md: Phase 2d — T029i (MCP AgentRun surface + конвертация E2E в fully external chain), T029j [x] (test-honesty remediation этого раунда), T029k (derive_capabilities от авторитетной DashboardQueryModel/environment policy/provider readiness — ложные human_checkpoint устраняются), T029l (labels «Подтвердить соответствие»/«Проблема не подтверждена»/«Недостаточно данных» + restyle confirm с bg-destructive + tool descriptions; lifecycle mapping неизменен), T029m (live-stand replay полевого сценария). T029g evidence снабжён honesty-amendment (raw-seed вскрыт и заменён).
  • 038 spec.md: Field-run Amendment — context-derived capability authority (SPECIFIED-PENDING): детерминированная серверная derivation, caller-declared capabilities только как сужение, needs_context/ needs_selector/needs_baseline для неразрешённых фактов, human_checkpoint только для реально небезопасного mutation/human judgement; селекторы/метрики/бейзлайны не изобретаются.
  • 044 spec.md: Field-run Amendment — HumanCheckpoint disposition clarity (mapping confirm→passed — канон и неизменен; human-facing формулировки именуются по персистентному исходу; destructive-стилизация confirm запрещена; API-вокабуляр не переименовывается). 045 spec.md: mirror-amendment для монитора (WaitingForMeView/HumanCheckpointPanel + vitest pins).

Test-honesty remediation (исполнено этот раунд, T029j)

  • backend/tests/test_mcp_initial_scenario_e2e.py::_pack(): raw AgentRun(...) ORM-вставка ЗАМЕНЕНА на продуктовую границу create_agent_run(db, CreateAgentRunRequest(context=UIContextV2(...)), user_id=...) (тот же сервис, что wraps POST /api/agent/runs; run существует в точной продуктовой форме — status RUNNING + run_started event). Module metadata честно фиксирует: MCP creation surface отсутствует (T029i open), конвертация в fully external chain — после T029i.
  • Новый backend/tests/test_mcp_agent_run_reachability.py (2 теста): (a) strict-xfail requirement pin MCPX-FR-027 — XPASS и падение suite в момент появления create_agent_run/create_authoring_agent_run в каталоге (форсирует конвертацию E2E + снятие маркера); (b) typed denial pin текущего состояния — внешний принципал без owned AgentRun получает ровно {"status":"blocked","error":"DRAFT_PACK_ACCESS_DENIED"} с НУЛЕВЫМИ строками CompiledScenarioHandle/DraftPackHandle/AgentRun (воспроизведение полевого результата).

Verification (этот раунд)

  • Targeted: 2 passed, 1 xfailed (E2E vertical green через сервис-границу; requirement pin strict-XFAIL; denial pin green) — 2.16s.
  • MCP regression slice (11 файлов: server, scenario E2E, promotion E2E, initial E2E, reachability, t029 bootstrap/automation, ops parity, checkpoints, transport limits, rbac visibility, catalog version): 76 passed, 1 xfailed — 12.6s.
  • ruff + compileall чистые; anchors balanced (reachability 4/4, initial E2E 3/3); scoped git diff --check чистый. Полный backend suite НЕ rerun (изменения — два тестовых файла + specs/ADR docs; production-код не тронут).

Статус и 남은 работы

  • Внешний initial-bootstrap путь остаётся NO-GO до T029i (MCP create_agent_run); UI-путь (D2/D4) и внутренний E2E unaffected. Очередь: T029i → конвертация E2E (E2E-EXT-001) → T029k (CAP-001) → T029l (DISP-001) → T029m live replay (E2E-EXT-002).
  • Все изменения не закоммичены (working-tree note предыдущих чекпоинтов в силе).

Checkpoint — 2026-09-07 (Phase 2d EXECUTED: T029i+T029k+T029l реализованы; первый полный suite после батчей)

T029i — MCP AgentRun surface (E2E-EXT-001 CLOSED)

  • Новый backend/src/mcp_server/tools_agent_run.py (155 LOC, McpServer.ToolsAgentRun): create_agent_run (wraps Services.AgentRuns.Service.Create; server-pinned objectType/route/contextVersion/intent; caller владеет dashboard_id/environment_id/dashboard_name/idempotency_key → conversation-reuse; продуктовая форма run: RUNNING + run_started event; bounded envelope) и get_agent_run (ownership-scoped проекция + draft_count; чужой/неизвестный → typed not_found, ноль existence-oracle).
  • Каталог 2.1.0 → 2.2.0: +2 записи (("dashboard:testing","EXECUTE") / ("dashboard:testing","READ"), service_allowed=False); seam в _build_probe_server (automation → agent-run → scenario); PINNED_CATALOG_MAJOR=2 не тронут; RBAC-зеркала (rbac_visibility/admin_mcp_catalog) правок не потребовали (derived-наборы; analyst без EXECUTE/READ).
  • Тесты: новый test_mcp_agent_run_tools.py (5); catalog-pin test_mcp_server.py расширен; test_mcp_initial_scenario_e2e.py конвертирован в fully external chain — AgentRun минтится вызовом create_agent_run через tools/call (ноль не-MCP seeding, _pack() удалён); strict-xfail pin в test_mcp_agent_run_reachability.py сработал как спроектирован — снят и закалён в hard requirement-тест; denial-pin (typed DRAFT_PACK_ACCESS_DENIED, zero side effects) сохранён.

T029k — context-derived capability authority (CAP-001 CLOSED)

  • Новый backend/src/services/dashboard_testing/scenario/capability_authority.py (236 LOC, ScenarioGraph.CapabilityAuthority): derive_capabilities — facts-only (native_filters; text_filter←STRING; time_rollover←DATE/TIME/TIME_GRAIN; table_filter+pagination←executable table-viz; xlsx_export←capabilities; dataset_field_read+has_dataset_fields←accessible dataset c колонками; browser←T040 readiness-снимок, иначе undetermined); NEVER_DERIVED (row_edit/bulk_edit/persistence_refresh/safe_test_data/safe_clock_fixture/ cross_dashboard) — unsafe-mutation автоматизация метаданными невозможна (@INVARIANT); merge_capabilities — derived-wins в обе стороны + overrides-аудит; legacy-payload → caller-declared passthrough; единый choke point build_capability_authority.
  • Wiring: MCP inspect_scenario (+capability_authority секция), inspect_dashboard_context (+derived_capabilities), REST api_compile_scenario (parity, +секция).
  • Тесты: фикстура query_model_sales.json; unit truth-table + CAP-001 classification-fix через map_all (sales-shape декларации → B01–B04/T01–T03 automated, C04–C06 unsupported, B05–B09/C01–C03/C02/C07 легитимно human_checkpoint) — 7; tool-level 3; REST-parity pin в test_scenario_routes.py. Readiness-seam запинен monkeypatch (детерминизм против глобального composition-root в full-suite).

T029l — disposition clarity (DISP-001 CLOSED)

  • RU/EN: «Подтвердить соответствие»/"Confirm conformance", «Проблема не подтверждена»/"Issue not confirmed", «Недостаточно данных»/"Insufficient data"; confirm-кнопки bg-destructive→bg-primary (HumanCheckpointPanel + WaitingForMeView, @RATIONALE); MCP decide_checkpoint docstring несёт immutable outcome-mapping; API-вокабуляр и lifecycle mapping не менялись; vitest-пины обновлены + новый DISP-001 style/label/dispatch pin (RunMonitorViews), WaitingForMeView.test.ts, run.ux.test.ts.

Полный suite впервые после батчей 4d5ef6be/58c5ae39 — вскрыты и закрыты 2 PRE-EXISTING регрессии HEAD

Первый full-run после этих коммитов дал 15 failed — все НЕ из кода этого раунда (app.py/test_app_lifespan.py/ mcp_oauth.py/test_mcp_client_flow_http.py в working tree не модифицированы; git show доказал самопротиворечие HEAD):

  1. SC-007 resource pin: код и twin-pin test_mcp_server.py:81 канонизировали post-redirect идентификатор /mcp/ (Mount 307 /mcp→/mcp/; 4d5ef6be изменил код, 58c5ae39 выровнял twin), но stale-пин test_mcp_client_flow_http.py:123 endswith("/mcp") не обновлён → обновлён на /mcp/ с комментарием (эмпирическая поддержка: живой внешний клиент 2026-09-07 прошёл discovery против этого значения).
  2. Lifespan re-entry (13 тестов): src.app.mcp_transport_app — модульный синглтон; SDK StreamableHTTPSessionManager.run() once-per-instance → второй вход в lifespan внутри pytest-процесса = RuntimeError. Продакшн стартует lifespan один раз на процесс — синглтон корректен; фикс test-side: autouse fixture _fresh_mcp_transport_app (свежий create_mcp_asgi_app() на тест, оригинал восстанавливается).
  3. Capability-тест: глобальный readiness-снимок (заполняется lifespan-тестами) влиял на derivation browser — seam-patch (выше). Группа test_app_lifespan + client_flow + capability после фиксов: 26 passed (было 15 failed/11 passed).

Verification (этот раунд)

  • Full backend suite: 11357 passed, 243 skipped, 1 xpassed, 0 failed (7:00) — первый зелёный полный прогон после батчей. targeted-срезы: 8-файловый MCP+catalog 70 passed; capability-срез 67 passed.
  • Frontend: vitest 3507 passed (206 файлов), lint 0 errors / 364 warnings (baseline), npm run build OK.
  • ruff + compileall чистые; anchors balanced (все 13 touched backend-файлов + svelte-компоненты); git diff --check чистый.
  • INV_7 хвосты (зафиксированы, не ухудшать дальше; новый код уведён в новые модули 155/236 LOC): mcp_server/tools_scenario.py 508 LOC (пре-existing >400 после T029f/h; +28 этот раунд) и api/routes/dashboard_testing/scenario.py 442 — кандидаты на декомпозицию следующим gate-раундом (например, вынос InspectDashboardContext+register_draft_pack в tools_pipeline seam / разбивка route-модуля).

Статус

  • 050 Phase 2d: T029i [x], T029j [x] (обновлён note о срабатывании xfail-ритуала), T029k [x], T029l [x]; release-gate rows: E2E-EXT-001/CAP-001/DISP-001/TEST-001 CLOSED; единственная OPEN — E2E-EXT-002 (T029m, live-stand replay полевого sales-сценария). Внешняя initial-bootstrap цепочка достижима через MCP alone.
  • Учёт обновлён: 050 spec.md/tasks.md, 038/044/045 amendments → IMPLEMENTED, ADR-0024 → IMPLEMENTED (+closure consequence), README-реестр, follow-up секция docs-отчёта.
  • Commit (2026-09-07): всё вышеперечисленное этого и plan-раунда (ADR-0024 + спеки + T029i/k/l код/тесты + фикс двух pre-existing HEAD-регрессий + docs-отчёт полевого прогона) закоммичено единым коммитом feat(mcp): Phase 2d field-run remediation — ADR-0024 agent-run surface, derived capabilities, disposition clarity (hash — см. git log). Намеренно НЕ включены (не этого workstream'а, остаются в рабочем дереве): translate/migration integration-тесты и _job_routes.py (чужой незакоммиченный workstream), .kilo/agent-manager.json (live UI/recovery state), корневой снапшот specs-036-050-20260907-111314.md. T029m (E2E-EXT-002) — единственная OPEN gate-row.

Checkpoint — 2026-09-16 (recovered re-baseline 2026-09-07 → HEAD 4746af2f)

Восстановленный чекпоинт: WORKSTATE не обновлялся 9 дней, факты ниже собраны из git log (45 коммитов) и handoff-отчётов (docs/reports/agentic-runtime-*.md, ux10-*.md, docs/2026-09-11-sales-prod-mcp-replay.md), а не из памяти. Suite-числа этого чекпоинта — свежие прогоны на HEAD; исторические числа помечены датой источника.

2026-09-08..10 — offline agentic runtime chain + INV_7 + live canaries v1/v2

  • 83727aa7 complete offline agentic runtime chain; INV_7-декомпозиция execution-пакета: runner.py 1207→121 (фасад; 6 модулей), lifecycle.py 611→299, executors.py 508→257; все модули < 400 LOC (a21481ea, 8f057f30, 3e5cc942).
  • D1 published-catalog source (fail-closed, 9 тестов) + D2 server-side live-binding resolution из settings.scenario_live_execution_bindings (12 тестов) — клиент binding не передаёт.
  • Live canary v1 (run 597274d3: browser open_dashboard + 8 durable screenshots) и v2 (REST-only, run 4eebfab3: server-side binding → PROD gate → scheduler dispatch → browser → реальный LLM → AgentEvaluation persisted → DecisionPolicy row 11). Три production-дефекта найдены и закрыты живыми прогонами (preflight deadlock ba2f1f45, artifact_byte_lengths, нормализация provider-ответов 3458343d).
  • Suite (2026-09-10, источник: evening handoff): 11527 passed / 0 failed / 244 skipped.

2026-09-11 — T029m live replay + T045 publish + 046 scheduled happy path

  • E2E-EXT-002 CLOSED: live_mcp_replay.py прогнал полную внешнюю цепочку на ss-prod (run 110a6517…: inspect → create_agent_run → compile B01 → register context_authority=verified → bootstrap → PROD gate pending_approval (идемпотентный retry = один durable gate) → approval → live capture_screenshot passed (8 refs) → typed BROWSER_ACTION_NOT_SUPPORTED → честный inconclusive); human loop живьём через live_mcp_human_loop.py (B05 HumanCheckpoint → waiting_human → list_checkpoints → decide_checkpoint confirm (CAS) → terminal passed). Два fail-closed дефекта найдены и исправлены. Trace: docs/2026-09-11-sales-prod-mcp-replay.md.
  • 050 T045 CLOSED: gated MCP publish_baseline_catalog + 037 publication worker (CAS, receipts, reconcile), REST parity; live pin-from-Gitea proven (canary v4: pin stamped into AgentEvaluation.baseline_pin, strict equality). Каталог 2.2.0 → 2.3.0.
  • 046 T018 CLOSED live: scheduled runs execute as schedule-owner principal; deterministic scheduled-run idempotency key; typed scheduled-reject observability (d0466a2c, b4148cfe).
  • 044 T043/T046: live baseline pin PROVEN; residuals — graph-level deterministic comparison PASS (T043) и graph-level terminal PASS (T046, baseline-semantic policy EVALUATION_UNAVAILABLE для compare без bound evaluation).

2026-09-12..14 — UX-10 Wave C (закоммичено как 8522a2ee 2026-09-15)

  • DG-1 реализован: run-scoped browser session (browser_session.py 501), 9 read-only действий (browser_readonly_actions.py 534), apply_native_filter + safe checkpoint (browser_native_filter.py 275), admission split (browser_admission.py 208).
  • 046 retention: deletions.py (331) + расширение retention.py + миграция 0024_retention_deletions (PG-verified; после инцидента с 33-char revision id добавлен guard test_revision_ids_fit_varchar32; правило: revision id ≤ 32, цепочку проверять на PostgreSQL).
  • Frontend: TerminalReasonBanner + typed terminal-reasons; docs/mcp-client-setup.md (UX-4).
  • 044 T040 [x] (deployment-evidence rule); T042 offline contract suites complete — остаток: live Superset query + явный RESULT_TOO_LARGE bound.
  • Suite (2026-09-14, источник: ux10 handoff, HEAD b4148cfe + dirty tree): backend 11796 passed / 264 skipped / 1 xpassed; frontend 3585 passed / 214 файлов; alembic head 0024_retention_deletions.

2026-09-15..16

  • 8522a2ee закоммитил дерево UX-10 (94 файла, +11014/−589): MCP automation parity tests, terminal reasons, scenario UX flow.
  • 3de0756d удалил legacy LLM dashboard validation (routes/service/schemas/ DashboardValidationPlugin; миграция 0025_drop_legacy_validation; frontend validation models/routes/i18n удалены; settings/automation редиректит на scenario automation).
  • 4746af2f health consolidation review fixes.

Свежие прогоны на HEAD 4746af2f (этот чекпоинт)

  • Backend full suite: 11269 passed, 252 skipped, 1 xpassed, 0 failed (7:42). Дельта против 11796 — удалённые legacy-validation тесты (3de0756d), не регрессии.
  • Frontend: 3359 passed / 208 файлов, 0 failed (49s). Дельта против 3585 — удалённые validation frontend-тесты. npm run lint → 0 errors / 333 warnings (baseline 340–364); npm run build (adapter-static) → green.
  • ruff check . clean; compileall -q src clean; alembic heads → единственная голова 0025_drop_legacy_validation.
  • Axiom rebuild (incremental, live): 10870 contracts / 5365 edges; unresolved relations 393 (было 401 на re-review 2026-09-04); исправлены 3 malformed multi-target @RELATION в backend/src/core/cot_logger.py (split на individual lines с verified targets).
  • MCP slice (T014 re-verification): 84 passed; каталог 61 запись, MCP_CATALOG_VERSION=2.3.0.
  • 037 visual slice (T044 re-verification): 61 passed.

Реконсиляция учёта (этот чекпоинт)

  • 050 tasks.md: T014, T016, T042 → [x] (боксы с evidence-текстом, но не отмеченные; evidence перепроверено свежими прогонами/grep на HEAD).
  • 037 tasks.md: T044 → [x] (visual candidate реализован в candidates.py веткой kind == "visual" + mandatory human disposition гейтом; closure note «All 47 tasks completed» теперь соответствует факту для T001–T047; T082–T084 — отдельные production rows 2026-09-08).
  • 050 quickstart.md: устаревшие OPEN-строки исправлены (T029m/E2E-EXT-002 CLOSED 2026-09-11; T045/T046 CLOSED; каталог 2.3.0/61).
  • 050 CHK003 OPEN — не противоречие: строка покрывает и T046 (CLOSED), и frontend-boundary (T030 OPEN); оставлено OPEN корректно.
  • 044 T042b остаётся OPEN осознанно: канарейки 2026-09-01 предшествуют Wave-C session architecture (8522a2ee) и не являются evidence для текущего кода.

Текущий фронт работ (актуализировано на HEAD)

  1. 044: T042b PREPROD-канарейки под Wave-C архитектуру; T044 live ACL/status/header canary; T045 fault-injection canary; T046 graph-level terminal PASS; T043 graph-level comparison PASS; T022 полный аудит; T042 live Superset query + RESULT_TOO_LARGE bound.
  2. 046: T017 lifecycle notifications (callers отсутствуют); T021 5/15/50-tab canaries; T013e бокс vs landed retention code (сверить); T013 формально открыт при покрытии.
  3. 050: T044 (REST-vs-MCP error-shape parity fixtures — последний OPEN production gate); CHK004 negative UI test; T029a; T030/T032 — продуктовое решение.
  4. 047: T017–T021 (SCAN-FR-015 atomic triage). 042: T024–T032. 043: T015/T023, T021–T022. 045: T020–T022. 037: T082–T084. 038: T041 (VLM), T060–T062.
  5. Skill drift: .agents/skills/semantics-python и semantics-svelte ссылались на удалённый ss_tools.shared.cot_logger (ADR-0022 absorb) — исправлено в этом чекпоинте: facade src.core.logger (intent-first), notify() из $lib/toasts.svelte.ts вместо несуществующего addToast(), i18n dictionary-proxy стиль ($derived($t.migration ?? {})) вместо function-call $t("key"); ./scripts/sync-skills.sh прогнан, .kilo копии byte-identical.

Product status: NO-GO — production-acceptance rows 2026-09-08 открыты во всех спеках; формальный closure re-sign-off не проводился.

Checkpoint — 2026-09-16 (T042b CLOSED: шесть canary-векторов GREEN под Wave-C архитектурой)

  • Все шесть векторов T042b перепроверены живьём на owner-authorized тестовом стенде (https://ss-prod.bebesh.ru, SS_STAND_STAGE=PREPROD, dashboard 11) через реальную продуктовую цепочку Wave-C (run-scoped browser session, capacity admission, receipts, durable evidence). Evidence в specs/044-dashboard-scenario-execution/evidence/browser-provider/ (readonly-canary-20260916T*.json + PNG; креды в evidence не пишутся):
    Вектор Результат Evidence
    read-only actions 3/3 passed (open_dashboard 7.2s, wait_for_state 3.9s, refresh 4.3s) 150158Z
    forced timeout/cleanup typed BROWSER_ACTION_TIMEOUT, 0 refs, lease released 150407Z
    mutation row_edit+restore 82.74→mutate→restore verified, 2/2 receipts completed 150518Z
    reconciliation sweep stale receipt resolved через live SELECT-only observation 150754Z
    safe-checkpoint reconstruction attempt 2 passed, reconstruction_replay=true 150908Z
    scheduler soak 75s 3/3 runs passed, attempts==1 (CAS), 0 active leases, 3 artifacts 155013Z
  • Harness defect найден и исправлен: в bare-script контексте DI-синглтон SchedulerService захватывал мёртвый event loop (asyncio.new_event_loop() без runner), из-за чего async-job'ы (maintenance_auto_end → AsyncJobRunner.run) блокировались на 300s safety cap, а stop() ждал их — первый soak-прогон завис. Фикс в specs/044-dashboard-scenario-execution/prototype/browser_readonly_canary.py: soak-окно выполняется внутри asyncio.run (singleton создаётся при работающем loop), scheduler.stop() вынесен через asyncio.to_thread (loop свободен для drain in-flight coroutines). Продуктовый код не менялся; в production loop FastAPI всегда работает, дефект — артефакт harness'а, но он же демонстрирует задокументированный drain-first компромисс (job занимает слот до cap).
  • PG-канареечная БД canary_044 создана в локальном postgres (recovery/soak/reconcile-режимы); read-only/timeout/mutation используют temp SQLite по умолчанию.
  • 044 tasks.md: T042b → [x] с полным evidence.

Checkpoint — 2026-09-16 (T022 audit: [~] — gates зелёные, открыт INV_7 хвост Wave-C)

  • T022 команды выполнены (см. текст задачи): scoped 044 suite 1590 passed; full backend 11269 / 0 failed; frontend 3359 passed + lint 0 errors / 333 warnings + build green; свежий PostgreSQL 16 0001→0025_drop_legacy_validation + alembic check без drift (проверочная БД alembic_044); Axiom live rebuild — 10870 contracts / 5367 edges, unresolved 393.
  • НЕ [x]: Axiom-аудит execution/ показал 8 structural warnings — Wave-C модули выросли за INV_7: providers/browser_readonly_actions.py 534, providers/browser_session.py 501, providers/browser.py 407 (регрессия против 398 из round 4). Требуется decomposition-gate по прецеденту server.py (specs/050-mcp-interface/plans/ server-decomposition-gate.md): binding plan + behavior-neutral split с frozen contract IDs.
  • T044 (044): offline evidence matrix уже зелёная — test_scenario_artifact_content_api.py 11/11 (GET/HEAD same-headers, 401/403/404-cross-owner, digest/mime 409 без байт, inactive 410, range 416, oversized 413, storage 503); остаток — live ACL/status/header canary на стенде.
  • Skill drift закрыт (semantics-python/svelte → facade src.core.logger, notify() из $lib/toasts.svelte.ts, dictionary-proxy i18n); sync-skills прогнан; malformed multi-target @RELATION в cot_logger.py разбиты (3 линии) — cot_logger.py 0 unresolved.
  • Ruff clean на всём backend; compileall clean; anchors сбалансированы во всех touched-файлах (pre-existing внутри-комментария #region в semantics-svelte line 86 — текст, не якорь).

Checkpoint — 2026-09-17 (Wave 1.6 EXECUTED: provider decomposition gate → T042b+T022 закрыты)

  • Wave 1.6 исполнен по binding-плану specs/044-dashboard-scenario-execution/plans/ provider-decomposition-gate.md (статус PLAN → EXECUTED, полный execution log внутри). Три gated фазы, нулевой behavior diff (per-phase scoped 1590 + финальный полный suite):
    Модуль Было Стало Новые sibling-модули
    browser_readonly_actions.py 534 160 browser_readonly_limits.py 82, flows_nav 199, flows_interact 192
    browser_session.py 501 58 browser_session_handle.py 64, browser_session_checkpoint.py 162, browser_session_managers.py 57, browser_session_registry.py 287
    browser.py 407 398 browser_factory_helpers.py 35
  • Frozen contract IDs + import surface сохранены: фасады ре-экспортируют все перенесённые публичные имена; monkeypatch-сеамы последовали за владеющими модулями (_MAX_EXTRACT_OUTPUT_BYTES → flows_nav, _MAX_DOWNLOAD_BYTES → flows_interact); _register_manager вынесен в browser_session_managers.py (иначе registry⇄facade цикл). Все переносы verbatim.
  • Гейты: per-phase scoped 1590 passed ×3; финальный полный backend suite 11269 passed / 252 skipped / 1 xpassed / 0 failed (7:10) — нулевая дельта против базовой 11269 до декомпозиции; ruff + compileall clean; anchors сбалансированы во всех 9 модулях.
  • Post-decomposition Axiom audit_contracts (execution/providers): 0 module_too_long (было 3). Принятые advisory: 2 × contract_too_long (BrowserProvider.Factory 305, ScreenshotProvider.Factory 200) — typed-лестницы исключений оставлены inline осознанно.
  • 044 tasks.md: T042b [x] (2026-09-16, шесть векторов GREEN) и T022 [x] (2026-09-17, zero P0/P1). Оба production-критичных бокса 044 закрыты с evidence.
  • Skill drift закрыт окончательно: заголовок примера semantics-svelte ссылался на удалённый notificationStore — заменён на $lib/toasts; sync-skills прогнан.

Checkpoint — 2026-09-17 (Wave 1.2/1.3 закрыты; T046 диагностирован с исполняемым proof)

T044 CLOSED [x] — live ACL/status/header canary (11/11 GREEN)

  • Новый harness specs/044-dashboard-scenario-execution/prototype/artifact_content_canary.py (351 LOC, 5 anchors) гоняет РЕАЛЬНОЕ FastAPI-приложение (без dependency overrides) поверх реального PostgreSQL (canary_044) с реальными JWT (реальный create_access_token, реальный поиск user→role→permission; без sid — session-policy пропускается) и РЕАЛЬНЫМИ live-байтами браузерного evidence (soak-прогон soak-canary-c071fb98, PNG 104746 B, sha 22b4b8c1…).
  • Evidence: evidence/artifact-content/artifact-content-canary-20260917T085438Z.json — 11/11: anonymous→401 AUTHENTICATION_REQUIRED (+HEAD пустой); viewer GET→200 byte-exact + согласованные Content-Length/ETag/Disposition/Cache-Control/nosniff/Accept-Ranges; viewer HEAD→200 полный header-parity + пустое тело; no-VIEW→403 PERMISSION_DENIED; unknown run / unknown artifact / foreign-owned → неразличимые 404 NOT_FOUND с идентичным message; expired→410; corrupt sha→409 ARTIFACT_INTEGRITY_FAILED; declared-MIME off-allowlist→409; oversized→413; missing bytes→409 ARTIFACT_MISSING; Range→416 RANGE_NOT_SUPPORTED + Accept-Ranges: none.
  • Harness-инфраструктура: добавлен knob SS_CANARY_STORAGE_ROOT (персистентный evidence-root; ВАЖНО: выделенный root на прогон — soak-ассерт считает файлы в root). Для HTTP-канарейки нужен STORAGE_ROOT_PATH из одобренных корней (/app/storage, проект, ../ss-tools-storage) — использован /home/busya/dev/ss-tools-storage/canary-artifact-acl.
  • Инциденты harness'а (исправлены): (1) повторный прогон брал собственный oversized-clone как «good»-row → фикстуры-клоны теперь именуются canary-* и удаляются в начале прогона; (2) soak с переиспользованным root дал failures: ["artifacts=6"] → выделенный root.

T045 CLOSED [x] — fault-injection canary (16/16 GREEN)

  • Новый harness specs/044-dashboard-scenario-execution/prototype/fault_injection_canary.py (365 LOC, 7 anchors) поверх реального PostgreSQL (canary_044): реальный provider runtime, capacity manager, receipt CAS, cancel lifecycle. Evidence: evidence/fault-injection/fault-injection-canary-20260917T105541Z.json.
  • Векторы: submit до старта/истёкший deadline → типизированные PROVIDER_LOOP_NOT_RUNNING / PROVIDER_SUBMIT_DEADLINE за bounded время; медленная корутина отменяется по caller-deadline; shutdown во время in-flight работы разворачивает вызывающего по его же deadline, loop → not-running (без зависания); crash (lease не освобождён) карантинит слот (второй claim → CAPACITY_UNAVAILABLE) до reconcile_expired_leases, который освобождает ровно эти units; unknown effect → receipt reconciliation_required/unknown → reconcile_provider_operation разрешает; late response добавляет только history (терминальный статус неизменен), reconcile по терминальному receipt отвергнут (PROVIDER_OPERATION_TERMINAL); cancel открывает drain-окно (draining), дедлайн-финализатор терминализирует (0 running steps, 0 unexpired worker leases), capacity-lease отменённого прогона согласуется по TTL, не исчезает молча; немедленный cancel терминализирует одним проходом.
  • Найдено при отладке: cancel_run экспайрит worker-leases (ScenarioStepLease), а не provider capacity-leases (CapacityLease) — последние освобождает provider в finally или TTL-reconcile (совпадает с семантикой «released or reconciled» из SC-008). Первый вариант ассерта это смешивал; исправлено.
  • Offline-референсы (зелёные): test_provider_operations.py (CAS/late/reconcile), test_scenario_cancel_timeout.py (6), test_dispatch_capacity_lifecycle.py, test_provider_capacity.py, test_provider_contract.py, test_provider_preflight.py, test_provider_runtime.py — 41 + 47 passed.

T046 [~] — диагноз с исполняемым proof (binding, не truth table)

  • Остаток T046 — walker-level binding, а не ошибка политики: walker.py считает decide_step_outcome(policy_inputs_from_outcome(...)) ПОШАГОВО, и evaluation берётся только из step_outcome["evaluation_input"], который эмитит единственный шаг agent_evaluation (evaluation_adapter.py:379). Шаг assertion compare_to_baseline всегда видит evaluation=None; derive_runner_plan включает mandatory-режим при наличии agent_evaluation шага → каждый нормативный шаг (включая compare) резолвится в EVALUATION_UNAVAILABLE.
  • Proof на HEAD (pure function, без БД): (A) compare-only + mandatory → inconclusive ["EVALUATION_UNAVAILABLE"]; (B) тот же + disabled → passed ["BASELINE_PASS"]; (C) тот же + bound evaluation_input (succeeded/pass/0.9) + mandatory → passed ["BASELINE_AND_SEMANTIC_PASS"].
  • Канон (contracts/production-chain.md §8) ставит optional evaluation МЕЖДУ comparison и pinned DecisionPolicy → решение обязано видеть оба; сейчас оно считается изолированно по шагу.
  • Опции закрытия (следующий пакет, truth table не меняется): (1) dependency-ordered binding — compare зависит от покрывающего evaluate-visual, walker инжектит persisted AgentEvaluation (agent_evaluation_ids) как evaluation_input для зависимого решения; (2) deferred decision — walker откладывает решение comparison-шагов до персиста покрывающей оценки и считает одно решение на агрегированных входах. Закрытие = live-рерun v4-графа с терминальным passed.

Открытые строки 044 после этого чекпоинта

  • T043 [ ]: residual — graph-level deterministic comparison PASS (050 T045 publish закрыт).
  • T046 [~]: binding (см. выше) + live terminal PASS.
  • Ранее закрыты: T042b [x], T044 [x], T045 [x], T022 [x], T040 [x].

Checkpoint — 2026-09-17 (DESIGN AMENDMENTS: complex scenarios, D/M/E, zero-human, tiering, investigation loop)

Спецификационный пакет (только спеки, implemented=false): пять Design Amendments из дизайн-сессий 2026-09-17 (кастомные Playwright-сценарии → типизированный реестр; R1 «метрики — только baseline»; R2 «zero-human runtime»; hot/cold evidence tiering; аудит интерфейса расследования). Код не менялся; все новые требования закрываются только executable evidence.

Внесённые amendments

Спека Amendment Суть
038 spec.md Design Amendment 2026-09-17 AGSCN-FR-016..025: R1 (числа только через baseline, валидатор-пин embedded literals), R2 (запрет эмиссии human-шагов, LLM не авторизует небезопасное, co-authoring промптов E-класса, DecisionPolicy hardening), таксономия D/M/E, set_step_inputs op, новые read-only действия (assert_dom, inspect_filter_options, navigate_tabs, wait_for_selector; входы search_text/mode/date у фильтров; click→URL-исход), дисциплина registered≠implemented, @REJECTED code-backed provider (не тронут)
038 contracts/browser-actions.md 2026-09-17 блок additive-строки 038.5.0 design; disabled-дисциплина; navigate_dashboard = disabled: pending driver
037 spec.md Design Amendment 2026-09-17 AGBASE-FR-015..018: двухъярусная политика (курируемый gating ≤50 записей + observatory non-gating с запретом→PASS), BaselineSelectionProposal ко-авторинг (агент предлагает facts-only координаты, оператор курирует), scenario-derived кандидаты (kind=scenario_transform, тот же lifecycle/immutability), staleness→кандидат не автосмена; visual-baseline lossless-исключение из tiering
044 spec.md Design Amendment 2026-09-17 SCEX-FR-031..036: evidence tiering (immutable ref ≠ физический тир, chain-of-custody transcode-receipt, hot_runs=2 + holds: открытые кейсы/pending-кандидаты, форматная политика webp:lossless/lossy + XLSX→Parquet Phase 2), LLM-as-judge hardening (evidence-as-data, verdict-схема, zero tool access, LLM-INJ-001), zero-human dispatch (scheduled без касаний, LLM не авторизует PROD-мутации/гейты), registered≠implemented; гейт-строки BSC-FILT/ASSERT/METRIC/ZEROHUMAN/LLMSTAB/EVID-TIER/LLM-INJ-001
045 spec.md Design Amendment 2026-09-17 RUNMON-FR-015..017: waiting-me сужается до launch-гейтов, чекпоинт-панели deprecated для новых ревизий (legacy работают), tiering-совместимость EvidenceViewer, scheduled read-only наблюдение
046 spec.md Design Amendment 2026-09-17 SCAUTO-FR-020..024: hot_runs retention + holds, zero-human schedule eligibility, уведомления wiring (NOTIFY-001), DLQ/quarantine + 72h SOAK (SCHED-SOAK-001), capacity-SLO модель (CAP-SLO-001)
047 spec.md + tasks.md Design Amendment 2026-09-17 + T022–T027 SCAN-FR-016..022: MCP investigation-контур (list_queue/get_case read, record_case_note, propose_disposition pull-only — решение human CAS; service-principal decide denied), case index/findability (in_case видимы, human-readable заголовки), workspace-рендер linked runs/evidence/tier-маркеры, уведомления; гейты INV-MCP/FIND/ATOMIC/NOTIFY-001; tasks T022–T027
050 spec.md Design Amendment 2026-09-17 MCPX-FR-031..034: investigation tools в каталоге, authoring input ops (set_step_inputs/set_step_evaluation + parity), опциональный exploration_step (Phase 2), prompt dangerous-content профиль; spec-impact 047 строка обновлена

Ключевые решения сессии (зафиксированы в amendments)

  1. R1: числовые бизнес-истины не встраиваются в шаги — только transform→baseline→approve→compare (AGSCN-FR-016); таксономия D/M/E (FR-017); TransformSpec — деривация, не проверка (FR-018).
  2. R2: ноль human в рантайме — эмиссия human-шагов запрещена (FR-019), LLM заменяет суждение, не авторизацию (FR-020: unsafe-мутации без fixture = unsupported; PROD-гейты человеческие), промпты E-класса co-authored и версионируются в content_hash (FR-021), DecisionPolicy pinned (FR-022).
  3. Tiering: hot=2 последних прогона оригиналы, дальше WebP-архив с chain-of-custody; baseline-сеты lossless-исключение; holds на открытые кейсы/pending-кандидаты.
  4. Baseline: курируемый набор (агент предлагает facts-only, оператор аппрувит) + observatory-ярус non-gating; scenario-derived кандидаты через transform-выход.
  5. Investigation: агентская петля через MCP (read/note/propose), решение human; findability (case index, in_case-видимость, human-readable заголовки).

Открытые точки для clarify (до имплементации)

  1. Семантика clear для multi-select фильтров (один/все чипы) и формат search_text-входа.
  2. Политика E-вердиктов: single-shot vs best-of-N (порог консистентности).
  3. Пороговые значения element_count/bounding_box как входы D-ассертов (формат pin в ревизии).
  4. Жизненный цикл baseline_set при смене отчётного периода (подтверждение UX staleness→кандидат).
  5. XLSX→Parquet Phase 2: критерий включения (доля XLSX в профиле хранилища).

Связность

  • Все amendments additive к существующим контрактам; ни один upstream @REJECTED не переопределён (code-backed provider остаётся future/unimplemented; per-step изоляция DG-1 остаётся rejected).
  • Гейт-строки добавлены в spec-файлы; WORKSTATE-статус продукта не менялся: NO-GO pending closure re-sign-off — новый пакет добавляет OPEN-требования, не закрывает существующие.
  • Следующий шаг по пакету: speckit-цикл уточнения (clarify по 5 открытым точкам) → plan → tasks per воркстрим (WS-1 реестр/входы; WS-2 фильтры; WS-3 D-ассерты; WS-4 M-путь+037; WS-5 E-класс; WS-6 navigate_dashboard; tiering-воркстрим S-M).

Checkpoint — 2026-09-18 (Wave 1 implementation: D/M/E foundation merged)

Волновая имплементация Design Amendments 2026-09-17. Все четыре ветки независимо верифицированы (scoped-сьюты + ruff + compileall) и слиты в master без NO-GO-статус-изменений продукта: сетевые гейты живых стендов (BSC-FILT-001 live canary, EVID-TIER-001 и др.) остаются OPEN.

Слитые ветки и верификация

Ветка Коммиты Скоуп Тесты
feat/038-step-inputs 534d488f, f123210b → merge e495f9db op set_step_inputs (AGSCN-FR-023), per-action валидация step_inputs.py, валидатор EMBEDDED_METRIC_LITERAL (FR-016), MCP parity; контракт-фикс: ScenarioStep.inputs остался list[Ref], аргументы — в новом optional action_inputs (канонические байты старых ревизий неизменны — пин-тест) 602+12 green
feat/047-investigation-mcp 23414f9a, 5eea72bc → merge c0224486 MCP-инструменты list_investigation_queue / get_investigation_case / record_case_note / propose_case_disposition (SCAN-FR-016..019); queue-проекция queued+in_case с human-readable заголовками; case: linked_run_ids + evidence_refs; каталог 2.4.0; disposition строго resolved accepted; service-principal denied на writes
feat/047-investigation-ui 6fcc8616, e7925d0e → merge 80980cee Queue findability (title, severity/scenario-фильтры, пагинация, in_case-бейдж), case-index маршрут /analytics/cases (SCAN-FR-020), workspace linked runs/evidence + i18n timeline (FR-021), RU/EN i18n 3360 vitest green, lint, build
feat/038-registry-filters d29959f6, 683347be → merge 5d86da54 Реестр 038.5.0: + disabled-строки assert_dom/inspect_filter_options/navigate_tabs/wait_for_selector; pagination/navigate_dashboard → disabled: pending driver; ACTION_DISABLED валидация до I/O (FR-025, негатив-тесты); apply_native_filter: search_text (B02) / mode=set clear / date (DatePicker) + CLEAR_CONFLICT-гард (FR-024); clear фолдится в checkpoint/replay (_FILTER_REPLAY_KEYS); click → page_url/popup_url с закрытием попапа до page-bound

Пост-merge master: 869 passed (registry+scenario+MCP+analytics+catalog+authoring) + ruff green.

Инцидент контаминации (2026-09-18)

Registry-агент волны 1 писал в главный репозиторий (master) вместо своего worktree: templates/__init__.py (реестр 038.5.0) + README.md + untracked docs/dashboard-testing-guide.md. Патч перенесён в worktree ветки и слит штатно через merge; master очищен. Урок: worktree-сессии Agent Manager могут мутировать CWD главного репозитория при ошибке cd — контролировать git status главного репо после каждой остановки агента. Судьба dashboard-testing-guide.md (scope drift, вне ТЗ) — на решение оператора.

Остатки волны 2 (OPEN, не имплементировано)

  1. Драйверы disabled-действий: assert_dom, inspect_filter_options, navigate_tabs, wait_for_selector ЗАКРЫТО 2026-09-18 4274e403: драйверы реализованы в browser_readonly_flows_observe.py (fail-closed, bounded, 8 unit-тестов), реестр включил их (disabled остались только pagination/navigate_dashboard), admission/transport/facade подключены, чекпоинты dom_asserted/filter_options_inspected/tabs_swept/selector_awaited. Живые канарейки на стенде (гейт BSC-ASSERT-001 live) — остаются OPEN до стенда.
  2. set_step_evaluation op + E-класс DecisionPolicy hardening (AGSCN-FR-021/022, WS-5).
  3. 037 scenario-derived baseline candidates + BaselineSelectionProposal (AGBASE-FR-016/017, WS-4).
  4. Evidence hot/cold tiering (SCEX-FR-031..034) + 046 hot_runs retention.
  5. navigate_dashboard драйвер (cross-dashboard сверки).
  6. Zero-human mapper: прекращение эмиссии human_checkpoint ЗАКРЫТО 2026-09-18 a59f5e3c: capability_mapper больше не эмитит human_checkpoint; former-human → unsupported (unsafe mutations — никогда LLM-авторизованы, FR-020); компилятор не создаёт шагов для unsupported-кейсов; 1577 тестов green. Runtime-механика HumanCheckpoint сохранена для legacy-ревизий (deprecated). E-класс (evaluate_declared_spec) остаётся declared-only до пункта 2.
  7. Уведомления wiring (SCAUTO-FR-022), DLQ/SOAK (FR-023/024), MCPX-FR-033 exploration_step.

Checkpoint — 2026-09-18 (LIVE SUPSERSET CANARIES: wave-2 observe + mutation restore GREEN)

Пользователь авторизовал тестовый Superset-контур https://ss-prod.bebesh.ru как PREPROD для read-only и mutation canary (credentials provided out-of-band; секреты не коммитятся). Dashboard 11 (Sales Dashboard) использован как структурный fixture; все evidence-retention файлы сохранены под specs/044-dashboard-scenario-execution/evidence/browser-provider/.

Live Superset readiness

  • /health → 200; /api/v1/security/login (db) → 200; dashboard catalog count=11.
  • Browser login через direct form fallback → authenticated redirect /; dashboard 11 открыт на /superset/dashboard/11/?standalone=true&native_filters_key=....
  • Базовая canary (SS_CANARY_MODE=actions): open_dashboard / wait_for_state / refresh — 3/3 PASS; durable PNG refs/digests, capacity leases released, evidence JSON readonly-canary-20260918T125144Z.json.

Wave-2 observe canary (SS_CANARY_MODE=wave2) — 4/4 PASS

Action Live факт Checkpoint/evidence
wait_for_selector .dashboard-content visible selector_awaited; PNG sha f54ea056…; 104774 B
navigate_tabs visited 2 tabs: 🎯 Sales Overview, 🧭 Exploratory tabs_swept; PNG sha 481032a4…; 49486 B
assert_dom .dashboard-content, min_count=1, actual=1 dom_asserted; passed=true; PNG sha f95e4b30…; 104784 B
inspect_filter_options Region: 19 live options (Australia…USA) filter_options_inspected; PNG sha f842ea03…; 110560 B

Evidence: readonly-canary-20260918T125451Z.json + per-action PNG. failures=[]. BSC-ASSERT-001 live-part CLOSED для четырёх wave-2 драйверов; registry disabled flags сняты в 4274e403 обоснованно (unit 811 + live 4/4).

Mutation canary (SS_CANARY_MODE=mutation, owner-authorized PREPROD) — 2/2 PASS

  • Fixture: video_game_sales, key (Wii Sports, Wii), field global_sales.
  • Original 82.74 → mutate expected 82.75, provider post_rows 82.75; SQL verify value 82.74 (stand SQL path observed a cache/isolation discrepancy after mutation; provider receipt/post_rows are authoritative operation evidence — needs follow-up); restore expected/verified 82.74.
  • Receipts: два durable completed/completed (2c8a5630…, f813cc03…); оба artifact PNG; cleanup checkpoint fixture_restored; final value restored 82.74; failures=[].
  • Evidence: readonly-canary-20260918T125542Z.json + mutate/restore PNG.

Live Superset residuals

  1. Playwright connection tasks emit Task was destroyed but it is pending! after provider loop stop (the result/evidence is complete, but process teardown leaks Connection.run tasks) — follow-up cleanup test/fix required before closing soak quality.
  2. Mutation canary mismatch: mutate provider post_rows=82.75 but immediate stand SQL verify=82.74; restore returns 82.74. Investigate cache/transaction/view path; do not classify false PASS solely from post_rows until independent SQL proof is reconciled.

Checkpoint — 2026-09-18 (P0 RUNTIME SAFETY WAVE: PROVIDER-SHUTDOWN + MUT-RECON CLOSED, BSC-FILT live PASS)

Волна 1 release-последовательности матрицы (PRODUCTION-ACCEPTANCE-MATRIX-043-050.md §3): G-PROVIDER-SHUTDOWN и G-MUTATION-RECON CLOSED, G-BROWSER-FILTERS live-векторы PASS (date-вектор BLOCKED внешне). Full backend suite 11332 passed; ruff/compileall/validate_static green.

SCEX-FR-037 Provider shutdown (PROVIDER-SHUTDOWN-001)

  • browser_session_managers.close_all_sessions(reason): weak-registry fan-out, fail-closed.
  • ProviderEventLoop.stop(): sessions close на живом loop ДО detach (submit() ещё видит self._loop — первичный вариант с detach-рано терял fan-out, поймано unit-тестом), затем loop.stop(); _run_loop finally: bounded drain (cancel + asyncio.wait(timeout=5s)), escapee-лог PROVIDER_LOOP_DRAIN_INCOMPLETE, только затем loop.close().
  • Unit: test_provider_shutdown.py — session-close-on-live-loop / pending-task drain без «Task was destroyed» / idempotent stop (3 passed).
  • Live: actions+mutation canaries readonly-canary-20260918T{174659,184343}Z — task_destroyed_warnings=[], thread_alive=false, leftover_sessions=0.

SCEX-FR-038 Mutation readback (MUT-RECON-001)

  • Root-cause 82.75/82.74: post_rows — честный mid-step SELECT (после UPDATE, ДО in-step restore_fixture cleanup); внешний канареечный SQL читает ПОСЛЕ полного шага (restore уже вернул 82.74). Lifecycle-skew, не cache/transaction-дефект; оба значения корректны в своих точках. Предыдущий residual-диагноз в чекпоинте 2026-09-18 (cache/isolation) снят.
  • browser_readback.py (новый): independent readback = тот же авторизованный page-сеанс, свежий SQL Lab client_id, SELECT-only скрипт; cleanup-aware expectation (restore→pre-image, retain→assignments); каноническая нормализация (float 6dp, key-sorted) — рендеринг float не может сфабриковать mismatch; crash → typed BrowserReadbackError (никакого выдуманного ok).
  • Transport: после mutation-флоу (включая in-step cleanup) — readback; divergence → BrowserTransportReadbackMismatch → provider: inconclusive, effect unknown, receipt reconciliation_required, retry blocked; PASS-путь: checkpoint readback_verified, receipt summary несёт readback_rows/hash/ok.
  • Unit: test_browser_readback.py (expectation/normalization/evaluate ok-mismatch-crash) + provider mismatch→reconciliation; cleanup-тесты обновлены под readback-контракт (3 evaluated scripts на restore-шаге).
  • Live: mutation canary readonly-canary-20260918T184343Z — readback_ok=true ×2 в outcome и receipts, failures=[], shutdown clean.
  • INV_7 (browser.py ≤400): вынесены download-side-artifact, transport_factory (через session_plan_box — late-bound plan), mutation_readback_summary в browser_factory_helpers.py; итог 395 LOC.

BSC-FILT-001 live filters canary

  • Новый prototype/browser_filters_canary.py: page-based discovery по всем 11 дашбордам (REST query-model не экспонирует Superset 4.x data-mode фильтры) + acceptance-векторы в ОДНОЙ run-scoped сессии (эпемерный native_filters_key не переносим между контекстами).
  • PASS на dashboard 11: values-apply (applied_mode=values, chart_data_observed=true, chip Japan наблюдается inspect-ом после apply), clear (mode=clear, chart=true), clear-checkpoint replay capture (1 replayable entry). Evidence filters-canary-20260918T195946Z.json + PNG.
  • Структурные факты стенда: Region dropdown — checkbox-list БЕЗ search input (search-вектор typed-N/A здесь; search-флоу остаётся unit-доказанным); фильтры есть только на dashboards 5, 11.
  • date-вектор BLOCKED внешний: ни на одном дашборде стенда нет date/time-фильтра; harness авто-детектит date_filters и прогонит вектор, когда owner добавит фильтр. Строка матрицы G-BROWSER-FILTERS → PARTIAL (BLOCKED external).
  • Продуктовые robustness-фиксы по live-флейку (оба в browser_native_filter.py): (1) pre-apply Escape — в reused run-сессии dropdown предыдущего шага остаётся открытым и control-click его TOGGLE-закрывает (опции исчезают); (2) bounded dropdown settle 400ms — antd slide-up анимация гонит same-frame visibility snapshot. Оба покрыты существующим сьютом (FakePage-совместимость через getattr).

Live Superset residuals (обновление)

  1. Playwright connection tasks leak — CLOSED (SCEX-FR-037 выше).
  2. Mutation 82.75/82.74 mismatch — CLOSED (lifecycle-skew root-cause + SCEX-FR-038 readback).
  3. NEW (external): date-фильтр на тестовом дашборде стенда — нужен owner; после добавления перезапустить browser_filters_canary.py (date-вектор прогонится автоматически).

Checkpoint — 2026-09-21 (WAVE 2 CORRECTNESS CHAIN: BSL-CATALOG core + BSL-DERIVED 3/4 axes CLOSED)

Волна 2 release-последовательности матрицы §3 (correctness truth chain). Full backend suite 11358 passed, ruff clean, compileall, GRACE-anchors сбалансированы.

G-BSL-CATALOG core (037 T082–T084) — catalog_revision_log.py

  • Append-only JSONL revision log *.revisions.jsonl рядом с каталогом (inter-process per-catalog lock; D11: без новой Alembic-миграции — ORM CatalogRevision отвергнут).
  • T082: duplicate baseline_id / approved coordinate_hash внутри одной ревизии → CatalogDuplicateEntry (первый approval выигрывает); stale If-Match → CatalogCasConflict с указанием текущего head; два конкурентных approval → ровно один head (thread-race тест с барьером, проигравший получает typed conflict — no lost update).
  • T084: append_status_transition (superseded/retired/invalidated) — аппенд, история байт-идентична (тест startswith); record_publish_outcome: publish_failed требует typed error_code, retry-receipt ложится на ту же ревизию (без дублей).
  • Дурэблити: append = полный префикс + новая строка через temp+os.replace — краш оставляет либо старый лог, либо полный новый, но не частичную запись.
  • Статус матрицы: PARTIAL (core proven, wiring open) — интеграция в consume/materialization API-поверхность остаётся (лог — CAS-авторитет рядом с YAML-каталогами).

G-BSL-DERIVED (037 T086–T088) — 3/4 оси закрыты

  • T086 scenario_transform_provenance.py: полная цепь провенанса от server-owned 044 ScenarioArtifact — run существует → artifact owner_type/owner_id/kind/is_active → sha256 == source_response_hash → transform-шаг passed → candidate_value == server-computed actual. Подделки (forged run, foreign/retired artifact, digest/value mismatch) — typed rejection до создания candidate. 8/8 test_scenario_transform_provenance.py; wiring в candidates.create_candidate.
  • T087 baseline_selection.py + schemas/dashboard_testing/baseline_selection.py (AGBASE-FR-016): facts-only ранжирование (cross_check → checklist → big-number/closed period → filter_sentinel → lineage_repr), hard-кап ≤50 с budget_note об исключениях, default+business filter contexts (default обязателен), tolerance rationale; review с CAS (SELECTION_CAS_CONFLICT/SELECTION_CAS_MISMATCH typed), accept/drop/adjust; unknown rank — fail-closed. 8/8 test_baseline_selection.py.
  • T088 observatory tier (AGBASE-FR-015): отдельная секция observatory_entries в catalog-revision.schema.json (non-gating by construction); resolver-guard — observatory entry, провезенная в entry_revisions, → BASELINE_AMBIGUOUS; секция исключена из пина. 11/11 test_baseline_resolver.py (2 новых observatory-edge).
  • Остаток G-BSL-DERIVED: transform execution→artifact capture (038 T068/T070) и период stale→blocked→candidate.

Остатки волны 2 (не в этой волне)

  • G-E2E-EVALUATION (044 T051 walker binding) — следующая итерация.
  • G-LLM-SECURITY (scheduled real-LLM terminal PASS) — отдельный прогон.

Checkpoint — 2026-09-21 (T051 WALKER BINDING CLOSED: E2E evaluation chain decision now observes the covering evaluation)

G-E2E-EVALUATION offline-часть закрыта. Full backend suite 11362 passed, ruff clean, walker.py 398 LOC (INV_7 удержан через вынос capacity-block в capacity_block.py).

T051 dependency-ordered binding — execution/evaluation_binding.py

  • Диагноз 2026-09-17 подтверждён и закрыт: compare-шаг видел evaluation=None, потому что policy_inputs_from_outcome биндит evaluation только из шага-владельца; sibling-топология v4-графов оставляла compare без оценки → EVALUATION_UNAVAILABLE при mandatory-режиме.
  • Решение (вариант 1 из диагноза): bind_covering_evaluation(db, run_id, plan, outcome, step_meta) — если план декларирует покрывающий agent_evaluation-шаг (comparison_refs на step-уровне, server-owned plan facts), walker ищет PERSISTED immutable AgentEvaluation этого run с покрывающим comparison_id и инжектит его как evaluation_input compare-решения. Truth table не изменена (SCEX-FR-028 инвариант сохранён).
  • Deferral: пока покрывающая evaluation не persisted, compare-решение откладывается — шаг → queued, _deferred_this_call-множество в вызове walker; release, когда все покрывающие eval-шаги terminal (persisted/failed/blocked/skipped) — без busy-loop, перенос в следующий dispatch-цикл. Отсутствие покрывающего шага в плане — прежний немедленный путь (EVALUATION_UNAVAILABLE остаётся честным ответом mandatory-режима без объявленной оценки).
  • Bound ids (agent_evaluation_ids, evaluation_input) штампуются на compare-шаг outcome для трассируемости; запись остаётся принадлежащей agent_evaluation-шагу.
  • 4/4 test_evaluation_binding.py: (1) graph-level terminal PASS — compare step BASELINE_AND_SEMANTIC_PASS, run passed (T046 residual closed offline); (2) deferral без ложного PASS; (3) no-covering — прежний путь; (4) disabled mode никогда не биндит.

INV_7 рефактор

  • capacity_block.py (новый): block_run_on_capacity + reconcile_capacity_blocked_runs + CAPACITY_RETRY_CODES вынесены из walker.py (394→398 LOC после добавления binding); импорты в runner.py/dispatch_runs.py переключены, обратная совместимость сохранена.

Остатки G-E2E-EVALUATION

  • Live rerun v4-style графа с terminal passed (живой стенд + провайдер, ближайшее окно).
  • Scheduled zero-human real-LLM PASS и single-shot/best-of-N policy-варианты.
  • G-LLM-SECURITY LLM-INJ-001 live/runtime proof — отдельный live-прогон.

Checkpoint — 2026-09-21 (T068/T069 CLOSED: bounded transform DSL + best-of-N verdict policy)

Обе оставшиеся offline-оси волны 2 закрыты. Full backend suite 11384 passed (+22 новых), ruff clean, compileall, anchors сбалансированы.

T068 (AGSCN-FR-018) — execution/transform_dsl.py + wiring в bounded_transform

  • Structured op-tree DSL (никаких строковых парсеров): sum_column (Σ колонки declared ref, ≤10000 rows), difference (a−b), ratio (a/b, 6dp canonical, деление на ноль typed), scale (k×a). Depth ≤ 4. Decimal-арифметика с canonical 2dp/6dp рендерингом — byte-stable (тест ×25 повторов). Не-числовая ячейка → typed TRANSFORM_CELL_NON_NUMERIC (никакой тихой конверсии). Unknown op/keys → typed. Никакого SQL/Python/shell/network by construction.
  • Wiring: bounded_transform принимает derive_value op-tree на step; outcome несёт derived_value + normalized decimal actual (готово к scenario_transform candidate T086 через server-owned артефакт-цепь); DSL-ошибки → typed inconclusive с dsl_error.
  • 12/12 test_transform_dsl.py.

T069 (AGSCN-FR-022) — execution/evaluation_aggregation.py

  • VerdictPolicy (revision-authored элемент): single_shot | best_of_n (n∈[2..5]), confidence_threshold ∈[0..1] — валидация typed.
  • aggregate_verdicts pure function: strict majority (>N/2) по admissible голосам (succeeded-with-verdict; provider/parser/budget/cancel/timeout не голосуют); tie → EVALUATION_AGGREGATION_TIE; all-error → EVALUATION_AGGREGATION_ALL_ERROR; quorum-miss → EVALUATION_QUORUM_INPUTS_MISSING; пусто → EVALUATION_AGGREGATION_EMPTY; pass-below-threshold → inconclusive EVALUATION_CONFIDENCE_BELOW_THRESHOLD (no quiet PASS). Mean confidence 6dp. Aggregate evaluation_id детерминированно именует когорту.
  • 044-binding: агрегат потребляется существующим T051 evaluation_binding путём.
  • Остаток: revision-schema поверхность (пин VerdictPolicy в AgentEvaluationSpec) — additive authoring-работа.
  • 10/10 test_evaluation_aggregation.py.

T070 — уже закрыт волной 037 (T087 proposal + T088 observatory)

Матрица: G-BSL-DERIVED → PARTIAL (offline оси все закрыты; остатки — period stale→blocked →candidate lifecycle и live capture wiring transform-выхлопов в ScenarioArtifact); G-E2E-EVALUATION residual сузился до live rerun + scheduled real-LLM + revision-schema.

Остатки волны 2 (все live-зависимые)

  • Live rerun v4-style графа с terminal passed (T046 closing evidence).
  • Scheduled zero-human real-LLM PASS (G-LLM-SECURITY LLM-INJ-001 live proof).
  • Period stale→blocked→candidate + live transform capture wiring.
  • Catalog revision log wiring в consume/materialization.

Checkpoint — 2026-09-21 (LIVE WAVE: E2E evaluation chain proven on the stand; terminal PASS external-blocked)

Стенд запускался run.sh --skip-install, затем uvicorn с captured stdout (/tmp/kilo/uvicorn.log). Полный live-цикл T051-цепи выполнен на ss-prod (dashboard 11) + Gitea-published baseline envelope.

Live canary v5 — prototype/live_canary_v5_binding.py (новый)

Sibling-топология: open-dashboard → capture-evidence → {compare-to-baseline, evaluate-visual}; evaluate декларирует inputs.comparison_refs (server-owned plan fact); decision_policy required. Полная цепь по прогону live-canary-v5-binding-20260921T164034Z.json (+eval-record):

  • open-dashboard passed (браузер, PREPROD-стенд);
  • capture-evidence passed BASELINE_PASS (8 durable артефактов) — semantic-фикс этой волны;
  • evaluate-visual: real-LLM evaluation PERSISTED (ce4685e7…, status succeeded, finding с evidence_artifact_ids);
  • compare-to-baseline: walker BINDING сработал — decision несёт evaluation_input (persisted id) + agent_evaluation_ids; deferral-цикл прошёл (compare ждал evaluate);
  • policy: LOW_CONFIDENCE (честный verdict inconclusive 0.0 от text-only модели) — fail-closed.

Продуктовые фиксы, найденные live-прогоном

  1. Semantic overlay scope (decision_policy.py): required-режим больше не требует собственную evaluation от evidence-продюсеров (screenshot/sql без comparison-payload) — они проходят rows 1-8 BASELINE_PASS; semantic-оверлей (rows 9-14) привязан к comparison-шагам (semantic_comparison в StepPolicyInputs; mapper ставит по tool==assertion / comparison-payload). Соответствует production-chain («deterministic-only runs need no LLM»); инвариант truth table сохранён.
  2. Defer deadlock guard (walker.py::_step_feeds_evaluation): deferring шага, от которого зависит покрывающая evaluation, невозможен — живой прогон выявил цикл (capture ждёт evaluate, evaluate ждёт capture).
  3. Test-фикстуры decision-policy обновлены под semantic_comparison (21+1 passed).

Внешние блокеры, обнаруженные live-прогоном (2026-09-21)

  • LLM gateway vision routes ALL 502: omniroute мультимодальные роуты (auto/best-vision, auto/pro-vision, auto/vision, auto/claude-*, auto/gemini, auto/glm) возвращают upstream 502 «Cloudflare Playground browser session failed: browserType.launch: Executable doesn't exist». Text-only auto/best-fast (qwen3.8) работает — провайдер переключён на него с is_multimodal=false (честный text-only manifest по контракту). Terminal PASS (BASELINE_AND_SEMANTIC_PASS) требует визуальный вердикт ≥0.7 — недостижим, пока gateway vision-роут не починен; chain закрывается честно LOW_CONFIDENCE.
  • Операционные правки локального стенда: llm_providers.base_url исправлен https://omniroute.bebesh.ru → https://omniroute.bebesh.ru/v1 (OpenAI SDK строит {base}/chat/completions; без /v1 — 401/404 на каждый вызов, маскировался под EVALUATION_PROVIDER_ERROR). run.sh-стенд не передаёт stdout uvicorn — для live-дебага использован отдельный uvicorn с env из backend/.env (DRAFT_STORAGE_ROOT не нужен — settings.storage.root_path через STORAGE_ROOT_PATH).

Учёт

  • T046: offline+live CLOSED (binding); terminal passed — external-blocked (vision gateway).
  • T051: CLOSED (binding + deferral + deadlock guard; best-of-N через T069 aggregation).
  • G-E2E-EVALUATION → PARTIAL (chain live-proven; terminal PASS blocked external).
  • Evidence: specs/044-dashboard-scenario-execution/evidence/browser-provider/ live-canary-v5-binding-20260921T164034Z.json + -eval-record.json.

Checkpoint — 2026-09-22 (WAVE 3 OPERATIONS integrated: NOTIFICATIONS, ATOMIC INVESTIGATION, POISONED-RUN, MCP PARITY)

Четыре фазы волны 3 выполнены параллельными Agent Manager worktree-сессиями (ds/deepseek-flash, omniroute) и интегрированы в master патч-моделью (база веток была устаревшей: origin/master = 1411c03f, −48 коммитов; rebase не применялся, патчи наложены на master с разрешением пересечений A∩C в analytics/investigation.py и B∩D в api/routes/.../scenario_automation.py).

Коммиты:

  • 93e97228 feat(automation) — NOTIFY-001: fail-closed emit/durable домены, channel/sla, идемпотентные receipts, queue/case/baseline_stale wiring.
  • def2bf40 feat(automation) — SCHED-SOAK offline: poisoned_store.py (durable JSONL идентичных-сбоев, flock+fsync, CAS epoch, без миграций) + poisoned.py N=3 → quarantine, один blocked DLQ-receipt, операторский CAS выхода; REST /quarantine.
  • 223296bc feat(analytics) — INV-ATOMIC-001: locking.py (pg_advisory_xact_lock), savepoint в set_disposition, context.py exact-context comparability; 22 unit + 7 PG-векторов.
  • 65684d6c test(mcp) — 050 T051: единый classify_start_error (REST+MCP), parity-матрица 17 векторов.

Verification: полный backend suite 11472 passed, 290 skipped, 1 xpassed, 0 failed; ruff clean.

Интеграционные адаптации (дрейф устаревшей базы, внесены при слиянии):

  1. test_mcp_rest_error_parity.py — investigation-ACL вектор переписан с устаревшей посылки «нет MCP-инструментов investigation» (на master они есть, G-INVESTIGATION-MCP CLOSED) на реальный no-bypass инвариант: чужой case схлопывается в not_found через MCP без утечки содержимого; каталог-присутствие инструмента проверяется.
  2. test_due_admission_atomicity.py — ожидания обновлены: терминальный прогон эмитит failed + queue_created, replay идемпотентен (2 receipt, не растут).

Матрица: G-NOTIFICATIONS → CLOSED, G-INVESTIGATION-ATOMIC → CLOSED, G-MCP-PARITY → CLOSED, G-DLQ-SOAK → PARTIAL (offline закрыт; 72h soak/restart storms и 5/15/50-tab — live). Инфра-инцидент: недоступность PostgreSQL с хоста из-за VPN-маршрута 172.19.0.0/24 dev throne-tun, перекрывавшего docker-bridge 172.19.0.0/16; маршрут исчез сам, БД восстановилась.

Checkpoint — 2026-09-22 (WAVE 4: offline P0-остатки — catalog revision-log wiring, period stale lifecycle, LLM-INJ-001 offline)

Волна выполнена в одном рабочем дереве (без Agent Manager worktrees); изменения на момент записи НЕ закоммичены (коммиты — только по явному запросу, по одному на фазу). План: .kilo/plans/1790082327000-wave4-offline-p0-residuals.md.

Phase A — G-BSL-CATALOG wiring (037 T082/T084) → CLOSED

  • backend/src/services/dashboard_testing/catalog_revision_wiring.py (новый): server-owned проекция consumed-entry на envelope лога — compute_coordinate_hash (expected value исключён), build_capture_profile_hash (из server-issued capture artifact, caller-хэш отвергнут), build_entry_revisions, settle_revision_outcome (publish outcome только на digest-matched head).
  • materialization._commit_with_catalog_compensation(..., revision=…): append ревизии внутри того же _catalog_lock ДО YAML-записи (duplicate/CAS отказ до мутаций) + компенсация log-bytes вместе с catalog-bytes при отказе DB commit — фантомная ревизия в CAS-логе невозможна. Новые хелперы: _append_revision_locked, _restore_bytes, _restore_revision_log.
  • approvals.consume_approval строит ревизию из server-built entry + capture meta (единственная точка materialize).
  • publication_worker.publish_baseline_catalog(local_catalog_path=…): settle-mirror в трёх точках (verify-reconcile, attempt success, failure) через _record_revision_outcome/_settle_published; _trace_revision_outcome не маскирует исходную типизированную ошибку.
  • Тесты: tests/services/dashboard_testing/test_catalog_revision_wiring.py (10) + execution/test_publication_worker.py (+3).

Решение (reactive micro-ADR): привязка локального лога к publish — через явный server-side параметр local_catalog_path; hard-fail при его отсутствии отвергнут (publish-контракт владеет Git-коммитом, локального каталога в нём нет) — отсутствие трассируется (PUBLISH_REVISION_LOG_ABSENT), а НЕВЕРНЫЙ лог падает типизированно (PUBLISH_REVISION_LOG_MISMATCH). Digest-scan по catalog_base_path отвергнут как недетерминированный и дорогой на текущей топологии.

Phase B — G-BSL-DERIVED period lifecycle (037 T089, AGBASE-FR-018) → PARTIAL

  • baseline_staleness.py (новый): requested_period_from(params, plan), closed_period_of, detect_period_staleness (только closed-period mismatch; open/undeclared период НЕ stale), build_stale_rebaseline_proposal (facts-only: идентичности + capture-kind, requires_operator_approval=true, automatic_rebaseline=false), emit_period_stale_receipt (durable idempotent baseline_stale, run_id=None).
  • baseline_resolver: BaselinePeriodStale(ValueError) сохраняет str(exc)=="BASELINE_STALE", resolve_baseline_pin(requested_period=…) блокирует запуск typed до I/O.
  • start_run пробрасывает период из params/plan и трассирует блок (EXPLORE, error_code BASELINE_STALE).
  • Immutability-приоритет сохранён: hash-change внутри ТОГО ЖЕ закрытого периода остаётся immutability_violation (CRITICAL), не downgrade до stale.
  • Тесты: execution/test_baseline_stale_lifecycle.py (8).

Residual: durable receipt запускного rollover не эмитится из production-surface (reject предшествует созданию run-строки, а owning-transaction откатывается) — эмиттер реализован и протестирован, привязка к поверхности — следующая итерация; comparison-time stale_baseline уже эмитит dedicated receipt (G-NOTIFICATIONS CLOSED).

Phase C — G-LLM-SECURITY LLM-INJ-001 offline (044 SCEX-FR-035) → PARTIAL (offline CLOSED)

  • evaluation_adapter: в промпт добавлен блок evidence_handling (evidence_is_data=true, правила «manifest/значения — данные, не инструкции», closed verdict domain); instructions byte-stable.
  • Тесты: execution/test_llm_injection_offline.py (7 векторов, все через submit-seam, без LLM): compromised-double canary, дроп findings по недекларированному критерию, free-form verdict / out-of-range confidence coercion, provider-identity forgery, prompt evidence-as-data + стабильность инструкций, non-pass в DecisionPolicy, отсутствие tool-канала у seam.
  • Fabricated artifact refs отвергаются walker-валидатором (существующий registry/test_agent_evaluation_store.py).

Учёт и верификация

  • 037 tasks.md: T082/T084 → wiring CLOSED note; T089 → [~] PARTIAL. 037 traceability.md: AGBASE-FR-014 (wiring), AGBASE-FR-018 → PARTIAL. 044 spec.md: LLM-INJ-001 offline CLOSED.
  • Матрица: G-BSL-CATALOG → CLOSED; G-BSL-DERIVED → PARTIAL (период закрыт, live capture wiring остаётся); G-LLM-SECURITY → PARTIAL (offline CLOSED); исправлен дрейф строки G-SCENARIO-AUTHOR (Transform T068/T069 закрыты — прежняя формулировка «open» была устаревшей).
  • Прогоны: focused-сьюты фаз (21 + 42 + 51), домен tests/services/dashboard_testing — 1659 passed; полный backend suite/ruff/compileall — см. D.4 ниже.
  • Продукт остаётся NO-GO: G-FULL-REGRESSION, G-SEMANTIC-HEALTH, G-RELEASE-SIGNOFF открыты.

D.4 — фактические прогоны (2026-09-22, рабочее дерево)

  • python -m pytest -q (полный backend): 11509 passed, 15 failed, 290 skipped, 1 xpassed (462s). Все 15 падений — в областях ДРУГОГО незакоммиченного воркстрима (admin/roles/permissions, maintenance templates), который уже был в дереве до волны 4: src/api/routes/admin.py, src/core/auth/permission_utils.py, src/services/security_badge_service.py, src/api/routes/agent_status.py, tests/api/test_admin.py — все помечены M в git status, и ни один падающий тест не импортирует модули волны 4. Волна 4 эти файлы не трогала.
  • python -m pytest -q tests/services/dashboard_testing (домен): 1675 passed (101s).
  • python -m ruff check . (весь backend): All checks passed.
  • python -m compileall -q src: без ошибок.
  • Axiom verify_after_edit по 12 путям волны: orphans delta 0, unresolved delta 0, parse_warning_count 0, naked_count 0.
  • Найден и исправлен в ходе верификации дефект: правка теста publication worker потеряла закрывающий якорь Test.ScenarioExecution.PublicationWorker.Reconcile (INV_3); пара восстановлена, повторный verify_after_edit → 0 parse warnings, 37 тестов волны собираются.

Ортогональное ревью волны 4 + аудит GRACE INV_1–INV_7 (2026-09-22)

Инструменты: axiom audit detect_missing_contracts (домен dashboard_testing: 176 файлов, naked=0), axiom search read_outline, axiom search verify_after_edit, ruff --select C901, сравнение с git show HEAD:<file> для атрибуции (моё vs унаследованное).

Аудит инвариантов

INV Вердикт Доказательство
INV_1 (region на каждый definition) PASS naked=0; ложно-сработавшие container-coverage по RevisionLogMismatch/BaselinePeriodStale закрыты листовыми регионами, ID = именам символов
INV_2 (NEED_CONTEXT при слепоте) n/a слепых зависимостей не было; все контракты выведены из кода
INV_3 (пары region/endregion) PASS (после фикса) в ходе верификации был потерян #endregion ...Reconcile — восстановлен; сейчас 0 parse-warnings по всем 12 путям
INV_4 (теги до кода, контигуально) PASS теги идут сразу после якоря, код после блока метаданных
INV_5 (локальный workaround не отменяет ADR) n/a ADR не затрагивались; publish-решение задокументировано как micro-ADR
INV_6 (не удалять контракт с входящими рёбрами) PASS контракты только добавлялись; тексты существующих тегов дополнялись
INV_7 (модуль <400 строк; CC ≤10) FAIL размер: approvals 387→401, publication_worker 373→464, baseline_resolver 391→426, evaluation_adapter 390→406, start_run 411→429; CC: publish_baseline_catalog 16 (унаследовано 16), start_run 12→13 (+1)

Находки, исправленные в этом проходе

  • F1 (High, корректность): падение записи каталога после append ревизии оставляло фантомную ревизию (CAS-голова ссылалась на нематериализованную запись). Исправлено: append откатывается внутри того же лока перед пробросом ошибки; +2 регресс-теста.
  • F3 (Medium, точность графа): 5 новых межмодульных зависимостей не были объявлены @RELATION DEPENDS_ON — добавлены (materialization, approvals, baseline_resolver, start_run, publication_worker).
  • F2 (Low, INV_1): листовые регионы для RevisionLogMismatch и BaselinePeriodStale.
  • F4 (Low, наблюдаемость): легитимное отсутствие привязанного лога логировалось как WARNING на каждом publish → alert fatigue; переведено на level="DEBUG".
  • F5 (Low): удалена мёртвая константа _UNAVAILABLE.

Находки, оставленные сознательно (с обоснованием)

  • F6 (Medium, INV_7 размер): 4 файла пересекли 400 строк, start_run уже превышал. Не рефакторил в этом проходе (риск для верифицированного кода); рекомендация: выделить settle-хелперы publication_worker в сиблинг-модуль, вынести период-детекцию из baseline_resolver, а период-блок из start_run — отдельным PR.
  • F7 (Medium, CC): publish_baseline_catalog 16 — унаследовано, не мной; start_run +1. Рекомендуется извлечь ветки settle/period-block в хелперы.
  • F8 (Low, дизайн): digest-mismatch привязанного лога демотирует уже опубликованную операцию в publish_failed (Git-коммит уже есть). Самоизлечимо через reconcile (commit сохранён); рекомендуется записать в runbook.
  • F9 (Medium, интеграционный пробел): у launch-периода нет продюсера: close_period (сторона closure) имеет и продюсера, и guard'ы, но reporting_period в params/plan никто не выставляет — детектор периода остаётся capability без end-to-end прогона. Это подтверждает PARTIAL матрицы. Следующий шаг: пробрасывать объявленный период запуска в params либо читать период из координаты запиненного плана.
  • F10 (Low): Phase C изменил контракт промпта — live zero-human гейт требует повторной валидации (prompt drift), offline-тестов недостаточно.
  • F11 (Low): metric-entry без capture_artifact_ref молча даёт null capture-поля (не пинится) вместо типизированного отказа на consume. Приемлемо (схема требует ref для metric), но consume-time проверка была бы строже.

Проход «правь все» — закрытие находок F6–F11 (2026-09-22)

F6/F7 — INV_7 (размер <400 строк, CC ≤10)

Разбиты 5 модулей; все затронутые теперь <400, а start_run и publish_baseline_catalog ≤10:

Было Стало Как
approvals 401 399 revision-payload собран в catalog_revision_wiring.build_revision_payload
publication_worker 464 356 publication_errors (тип ошибки), publication_request (валидация+target), publication_mirror (settle-зеркало)
baseline_resolver 426 372 execution/baseline_codes (D11-словарь+reject), load_published_catalog → published_catalog_source, BaselinePeriodStale → baseline_staleness
start_run 429 363 execution/start_contract (request-hash+classify), execution/prod_guards (PROD-гварды), request-hash/classify оформлены как Tombstone c @REPLACED_BY (INV_6)
evaluation_adapter 406 368 execution/evaluation_prompt.build_evaluation_prompt

CC: publish_baseline_catalog 16→≤10 (валидация вынесена в publication_request.validate_publish_request); start_run 13→≤10 (PROD-гварды вынесены). Остальные C901-хиты в репозитории (~40 функций, включая decide_step_outcome 27 и validate_evaluation_evidence 16 в 044) — унаследованный долг, волной 4 не затронуты; рекомендация вынесена отдельно.

F8 — зеркало publish не демотирует опубликованную операцию

publication_mirror.settle_published больше не вызывает _mark_failed: Git-коммит уже состоялся, поэтому digest-mismatch привязанного лога теперь громкий (PUBLISH_REVISION_LOG_MISMATCH, EXPLORE) но НЕ переписывает durable publish truth; операция остаётся published с сохранённым commit (reconcile-путь). Тест test_publish_wrong_bound_log_is_typed обновлён соответственно.

F9 — период запуска стал контрактом, а не конвенцией

requested_period_from теперь fail-closed: присутствующая, но не non-blank-string декларация reporting_period реджектится REPORTING_PERIOD_INVALID вместо молчаливого отключения guard'а; контракт зафиксирован в StartRunRequest.params (REST-комментарий) и @INVARIANT start_run. Первопартийный продюсер (UI/автоматизация) остаётся работой следующей итерации; сама capability достижима через свободный params.reporting_period (REST и MCP используют один и тот же путь).

F10 — prompt drift учтён

Изменение промпта Phase C — событие live-гейта: offline-тестов достаточно только для LLM-INJ-001 offline; live zero-human прогон обязан быть повторён на новом промпте (записано в матрице).

F11 — типизированный отказ на gating-ревизию без capture-ref

build_entry_revisions реджектит metric/scenario_transform entry без server capture artifact (CATALOG_REVISION_CAPTURE_REF_REQUIRED) вместо материализации не-пинируемой ревизии; visual по-прежнему допускает null (+тест).

Верификация после прохода

  • python -m pytest -q (полный backend): 11534 passed, 290 skipped, 1 xpassed, 0 failed (438s). Прежние 15 падений чужого воркстрима (admin/roles/permissions) на момент прогона также отсутствуют.
  • python -m pytest -q tests/services/dashboard_testing: 1679 passed.
  • python -m ruff check .: All checks passed; compileall -q src: OK.
  • axiom verify_after_edit (20 путей): parse_warnings 0, naked 0, orphans Δ0; workspace_health по start_run.py: unresolved 0, orphans 0; tombstone-цели ScenarioExecution.StartContract.* существуют в индексе.

Checkpoint — 2026-09-23 (WAVE UI: фронтенд-пробелы матрицы — DLQ recovery, run-monitor reload, disposition E2E-спек)

План: .kilo/plans/1790163777000-wave5-frontend-ui.md (каталог планов gitignored). Изменения на момент записи НЕ закоммичены (коммиты — только по явному запросу).

UI-1 (P0, G-DLQ-SOAK) — операторское восстановление карантина → offline CLOSED

  • Backend-REST уже существовал (GET /quarantine, POST /quarantine/{scenario_id}/release с CAS expected_version; 404 NOT_QUARANTINED / 409 STALE_QUARANTINE_VERSION) — backend не менялся.
  • frontend/src/lib/types/scenario-automation.ts: QuarantineRecord + QuarantineReleaseResult (mirror PoisonedRunStore.Fold: scenario_id, environment_id, error_code, identical_count, version, reason, quarantined_at).
  • frontend/src/lib/api/scenario-automation.ts: listQuarantines + releaseQuarantine.
  • frontend/src/lib/models/AutomationQuarantineModel.svelte.ts (новый, [TYPE Model]): FSM idle/loading/loaded/confirming/releasing/error; release отсылает показанную версию; 409 → reload + notice stale_version (без молчаливого повтора); 404 → reload + notice already_released; неизвестная ошибка → error без reload.
  • frontend/src/lib/components/scenario-automation/QuarantinePanel.svelte (новый) + монтаж на routes/dashboard-testing/automation/+page.svelte; $lib/ui (Button/Badge/ConfirmDialog), семантические токены, i18n (15 ключей ru+en).
  • Тесты: AutomationQuarantineModel.test.ts (7) + QuarantinePanel.test.ts (4) + обновлён мок automation_page.ux.test.ts.

UI-3 (P1, G-MONITOR) — reload/reconnect → offline CLOSED

  • RunMonitorModel.test.ts +2 L1-вектора: (1) свежая модель после «reload» восстанавливает waiting_human с сервера, ровно один checkpoint-шаг (перезагрузка не добавляет второй); повторный loadRun заменяет снапшот; (2) повторный bindEvents закрывает предыдущий EventSource (closed), named-листенеры не стекаются (ровно один на имя).
  • Tiering-маркеры остаются отсутствующими (blocked на backend G-EVIDENCE-TIER).

UI-2 (P1, G-INVESTIGATION-UI) — disposition E2E-спек → написан, исполнение заблокировано

  • frontend/e2e/tests/scenario-disposition.e2e.js (новый): queue→case→disposition через case-workspace (section[aria-label="Рабочее пространство кейса"], select resolved, кнопка «Закрыть кейс») + отдельный CAS-вектор (stale expected_version → 409 STALE_CASE).
  • Residual: на изолированном стенде нет пути засева очереди (backend не отдаёт API создания queue-item; сигнал создаёт падающий прогон) → спек скипается с явной причиной. Это названный остаток строки, не зелёный прогон.

UI-4/UI-5 — только объявленные контракты (BLOCKED)

  • Evidence.TierBadge (tier_hot/tier_cold/untiered) и ScenarioBaseline.SelectionReview (accept/drop/adjust + CAS) объявлены в плане; код не писался, т.к. backend-полей/поверхности нет (иначе UI биндил бы несуществующий контракт данных). Строки матрицы обновлены пометкой «contract-declared, not implemented».

Верификация UI-волны

  • npm run test -- --run: 3425 passed (220 файлов); npm run lint: 0 errors (341 warning svelte/require-each-key); npm run build: ✓.
  • Новые файлы: 2 модели/компонента + 1 e2e-спек + 16 новых frontend-тестов (7+4+2 сверх обновлённого page-мока).
  • Матрица: G-DLQ-SOAK (recovery UI закрыт offline), G-MONITOR (reload/reconnect закрыт), G-INVESTIGATION-UI (спек написан, остаток — seeding), G-EVIDENCE-TIER/G-MCP-BASELINE-SELECTION (blocked-пометки).

Handoff для следующего агента (2026-09-23)

Промежуточное состояние передано: specs/agent-handoffs/wave4-ui/README.md (полный инвентарь коммитов, «мои» vs «чужие» файлы рабочего дерева, статусы строк матрицы, оставшиеся шаги по зависимостям, команды верификации, инфра- и координационные грабли).

  • Закоммичено: 5a5ac4dd (wave 4, 26 файлов). Не закоммичено: wave UI (UI-1/UI-3 + спек UI-2) и учётные правки матрицы/traceability/WORKSTATE — файлы в дереве, страховочный патч specs/agent-handoffs/wave4-ui/wave-ui-and-specs.patch (17 файлов).
  • Планы (каталог .kilo/plans/ gitignored, ключевое продублировано в README): 1790144880000-wave5-offline-p0-residuals.md (backend: DLQ retry-dispatcher, launch-receipt, SECURITY-OPS), 1790163777000-wave5-frontend-ui.md (UI-1..UI-6).
  • Внимание: в дереве параллельно работает вторая frontend-сессия (модальные binding-пробы, runes-contract, admin settings) — её файлы перечислены в README §2 и не должны попадать в коммиты; .kilo/agent-manager.json не коммитить.
  • Остаток для GO (кратко): 10 незакрытых P0-строк (2 внешних блокера, 4 live/offline-микс, 3 релизных гейта, 1 signoff) + 8 P1; продукт остаётся NO-GO.

Wave 5 Phase A — G-DLQ-SOAK retry-dispatcher wiring (2026-09-23)

Мастер-план: .kilo/plans/1790240000000-wave5-6-execution-master.md (порядок фаз и контракты).

  • Phase 0 закоммичена c4d6faa0: wave UI (17 файлов §2 + handoff-каталог) — quarantine recovery UI, run-monitor reload/reconnect, disposition e2e spec, учёт матрицы/traceability/WORKSTATE.
  • Phase A (offline, закрыта в дереве): automation/retry_policy.py (закрытый allowlist retryable infra-кодов, fail-closed; MAX_ATTEMPTS = порог карантина), execution/retry_dispatcher.py (failed→queued CAS re-queue; durable identical-failure counter = attempt-authority — у ScenarioRun нет attempt-колонки, D11 замораживает миграции; карантинные пары и потолок пропускаются), терминальный учёт в execution/terminal_effects.py (один терминальный отказ = одна запись счётчика; automation-источники + allowlist только; passed automation-ран сбрасывает стрик; карантин чистит только операторный CAS), карантин-гейт в execution/dispatch_runs.py (устаревшие queued-строки карантинной пары не диспатчатся до CAS-release), scheduler-тик core/scheduler.py композирует retry→dispatch. Тесты: tests/services/dashboard_testing/automation/test_retry_dispatch.py — 10 векторов (CAS exactly-once, N=3 quarantine + одна нотификация, distinct-code reset, content/manual исключение, гейт+release, потолок, allowlist-membership).
  • Верификация: pytest -q tests/services/dashboard_testing 1689 passed (было 1679); полный pytest -q — 11579 passed, 290 skipped, 1 xpassed, 0 failed; ruff check . All checks passed; compileall -q src OK; axiom verify_after_edit (automation/, execution/, scheduler.py, automation-tests): parse_warnings 0, naked 0, orphans Δ0, unresolved Δ0; make docs-nav — контракты и рёбра обновлены (новые: ScenarioAutomation.RetryPolicy, ScenarioExecution.RetryDispatcher, ScenarioExecution.Runner.QuarantineGate, ScenarioExecution.Runner.PoisonedAccounting, ScenarioExecution.Runner.PoisonedStreakReset).
  • Матрица: G-DLQ-SOAK PARTIAL — offline + recovery UI + retry wiring CLOSED; остаток только live (72h soak, 3 restart storms, 5/15/50-tab canaries). 046 tasks.md T023/T024 и traceability SCAUTO-FR-023/024 обновлены.
  • Следующие фазы (по мастер-плану): B (launch-receipt, после A — теперь разблокирована), C (SECURITY-OPS, ∥ B), S (UI-2 queue-seed, ∥ B/C), затем P1-волна и релизные гейты.
  • Коммит Phase A — по явному запросу (файлы в дереве, чужие файлы второй сессии не тронуты).

Wave 5 Phase B + C — G-BSL-DERIVED launch-receipt + G-SECURITY-OPS offline (2026-09-23)

  • Phase B (launch-receipt, закрыта в дереве): BaselineEngine.BaselineStaleness.LaunchReceipt (emit_launch_period_stale_receipt) — durable baseline_stale receipt на отдельной закоммиченной сессии после отката транзакции запуска; fail-safe (провал эмита — EXPLORE-трейс, транспортный ответ не маскируется). Вызов ровно один на транспорт: REST scenario_runs.py (BaselinePeriodStale до generic ValueError → 422 BASELINE_STALE, ответ байт-идентичен прежнему контракту) и MCP tools_scenario.py (blocked + classify_start_error, 050 T044 parity). Idempotency-key стора схлопывает двойной вызов REST+MCP в одну строку. Тесты: tests/.../test_start_run_period_receipt.py — 6 векторов (REST/MCP receipt, exactly-once, digest-mismatch без receipt, отсутствие run-строки, провал эмита не маскирует блок); parity-матрица 17 векторов зелёная.
  • Phase C (SECURITY-OPS offline, закрыта в дереве):
    • C.1: Core.EncryptionKey.KeyAge — fail-closed вердикт возраста ключа по mtime .env (неизвестный возраст = warning, не ok; превышение интервала = age_exceeded); тест ротации (_reencrypt_value+Fernet): старый шифротекст нечитаем новым ключом, после re-encrypt читаем новым и нечитаем старым; legacy-значения репортятся, не конвертируются.
    • C.2: redaction на границе записи evidence — ScenarioExecution.EvaluationAdapter.Raw + _redact_tree (per-string-leaf через plugin RedactionService; сериализованную строку редактить нельзя — регексы не JSON-aware и ломают структуру, @REJECTED); findings и структура JSON сохраняются. Второй реализации redaction нет (reuse единственного модуля).
    • C.3: retention-политика evidence объявлена в GlobalSettings (scenario_evidence_retention_days default 90 + scenario_evidence_archive_target), falsifiable config-check в тесте; enforcement-sweep не реализован (scope-забор → G-STORAGE-BUDGET, P1.e).
  • Верификация (волна 5 целиком): полный pytest -q — 11589 passed, 290 skipped, 1 xpassed, 0 failed; dashboard_testing-сьют 1699 passed; ruff check . All checks passed; compileall OK; axiom verify_after_edit: parse_warnings 0, orphans Δ0, unresolved Δ0; naked=0 на моих файлах (11 унаследованных naked в чужих тест-файлах test_comparison/test_normalization/ test_visual_baseline_staleness — вне скоупа волны, INV_1-долг на P1/релизную гигиену).
  • Матрица: G-BSL-DERIVED PARTIAL (launch-receipt CLOSED; остаток live capture wiring); G-SECURITY-OPS PARTIAL (offline CLOSED: rotation/age/redaction/retention-policy; остаток live drill). G-DLQ-SOAK — см. checkpoint Phase A выше.
  • Следующие фазы (мастер-план): S (UI-2 queue-seed для e2e), P1-волна (P1.a G-EVIDENCE-TIER backend → UI-4; P1.b MCP-surface → UI-5; P1.c/d/e/f), релизные гейты R.1–R.3.
  • Коммит волны 5 (A+B+C одним или тремя коммитами) — по явному запросу; файлы в дереве, чужие файлы второй сессии не тронуты.

Post-wave-5 P0 contract closure — transform evidence + retention enforcement (2026-09-23)

  • G-BSL-DERIVED: успешный bounded transform теперь сохраняет канонические JSON-байты результата через существующий DraftStorage и передаёт ref/digest/MIME/length в outcome. Walker регистрирует их как ScenarioArtifact с владельцем scenario_run, logical_step_id и attempt; provenance-проверка кандидата связывает ref, digest и вычисленное значение. Подмена ref/digest/value отвергается. Отдельный bridge-тест проверяет bytes → artifact → candidate. Offline-векторы зелёные; live canary остаётся открытым.
  • G-SECURITY-OPS: регистрация evidence назначает raw VLM класс raw_vlm (7 дней), изображениям screenshots (30 дней), прочим step-артефактам artifacts (30 дней). Плановый retention job находит истёкшие ScenarioArtifact, идемпотентно создаёт deletion receipts и коммитит их до удаления байтов; существующий sweep соблюдает baseline/active-run/open-case holds, проверяет отсутствие байтов перед tombstone и на момент sweep защищает новый артефакт с тем же content_ref. Проверены границы 8/31/91 дней, прерывание удаления после mark и повторный sweep. scenario_evidence_retention_days управляет только явно назначенным классом scenario_evidence (текущие producer-ы его не назначают); scenario_evidence_archive_target пока не архивирует байты. Live rotation/PII/retention drill остаётся открытым; storage tiering/budget — P1.
  • Верификация: независимый focused pytest — 73 passed; Ruff для затронутых backend-файлов чистый; git diff --check чистый. Полный release regression на замороженном коммите не выполнялся.
  • Live-стенд: ./run.sh --skip-install не дошёл до backend: PostgreSQL отвечает внутри Docker (SELECT 1), но host TCP на опубликованном порту 5432 обрывается до PostgreSQL; /api/ready недоступен. В backend/.env также обнаружены legacy multiline PEM-секции, из-за которых Docker Compose не может разобрать env-файл. Секреты и данные стенда не изменялись.
  • Матрица сохраняет G-BSL-DERIVED и G-SECURITY-OPS как PARTIAL до retained live evidence.

Business Analyst UX acceptance target — 2026-09-23

  • Новая обязательная release-планка G-BA-UX-READINESS зафиксирована в unified matrix (P0): каждый из семи критических сценариев аналитика должен получить отдельную оценку ≥4/5; среднее не маскирует провал отдельного сценария. Приёмка — 5 бизнес-аналитиков, ≥80% шагов без модератора, без false PASS/approval/case closure, с keyboard/a11y и retained browser evidence на одном release candidate.
  • Контракты и backlog связаны из 039 (authoring + scoring), 045 (launch/result/monitor), 046 (schedule operations), 047 (queue→case/disposition). Текущий аудит оставляет gate OPEN: baseline review UI отсутствует; editor неполон; evidence/comparison не подключены; queue→case ID/pagination расходятся с API; resolution не требует ссылку на evidence; cron/ID — в форме расписания. Реальный browser/user study не проводился.
  • Подробная шкала и протокол находятся в specs/039-dashboard-scenario-ui/spec.md → Business Analyst Readiness Amendment; тесты/build являются supporting evidence, не UX sign-off.

047 investigation seeded browser proof — 2026-09-24

  • В изолированном Compose-проекте ss-tools-e2e добавлен одноразовый seed после миграций: фиксированные scenario/run/artifact IDs, failed terminal run, active JSON evidence bytes в общем storage и сигнал через emit_terminal_run_signal. Повторный seed сбрасывает только собственную фикстуру.
  • Playwright scenario-disposition.e2e.js: queue→правильный case через UI→linked run/evidence viewer→resolved с проверяемым ref→закрытие queue→старый decision_version получает 409 STALE_CASE. Chromium: 1 passed, 0 skipped; trace/screenshot сохранены в frontend/test-results/scenario-disposition.e2e.j-cc566-dence-and-rejects-stale-CAS-chromium/.
  • Прогон выявил реальный HTTP 405: модель case disposition использовала GET-обёртку fetchApi. Исправлено на postApi; focused Vitest 10/10, frontend lint/build, backend Ruff/compileall, синтаксис JS, Compose config и git diff --check прошли. Playwright runner обновлён до 1.60.0; ENCRYPTION_KEY берётся из игнорируемого .env.e2e, а backend/.env с multiline PEM больше не передаётся Compose.
  • T027/T032 закрыты на уровне изолированного browser proof. Axiom verify_after_edit по четырём затронутым source/test путям: success, parse warnings 0, naked 0, orphan/unresolved Δ0. Контейнеры удалены с volumes; коммита нет. G-INVESTIGATION-UI имеет scoped implementation proof, но G-BA-UX-READINESS, полный regression на release commit и остальные release gates остаются открыты.

Correction — live browser filters and multimodal provider evidence (2026-09-25)

  • В browser_filters_canary.py добавлена поддержка SS_STAND_NATIVE_FILTERS_KEY, scope discovery при этом ограничен dashboard-владельцем ключа. Browser session сохраняет native_filters_key в URL. Удалены вручную выставлявшиеся глобальные Sec-Fetch-* заголовки: они применялись к JS/CSS subresources и мешали dashboard React-приложению загрузиться в Playwright, хотя обычный браузер открывал страницу.
  • browser_native_filter.py теперь связывает Ant Design label с .select-container фильтра. Для label Country это устранило miss: live canary применил Canada, дождался 5 charts, затем очистил фильтр и дождался 6 charts.
  • Evidence specs/044-dashboard-scenario-execution/evidence/browser-provider/filters-canary-20260925T191720Z.json (+ PNG): inspect/apply/inspect/clear/inspect и replay checkpoint прошли; shutdown leftover_sessions=0, task_destroyed_warnings=[]. G-BROWSER-FILTERS остаётся PARTIAL только из-за отсутствующего date-picker fixture; search structurally N/A для выбранного dropdown.
  • Проверено, что Omniroute deepseek-flash принимает screenshot: OpenAI-compatible image_url request с живым PNG завершился HTTP 200 и IMAGE_CANARY_OK; фактически gateway сообщил route model gpt-5.6-sol. Это опровергает прежний blanket вывод о text-only fallback/неработающих всех vision routes. Это отдельный provider-capability probe, не полный VLM scenario PASS.
  • Для G-E2E-EVALUATION остаётся полный scenario rerun с реальным evaluation spec, persisted AgentEvaluation → walker binding → terminal policy result, а также scheduled zero-human и injection proof. Общий продуктовый статус остаётся NO-GO; релизные регрессия, semantic health, BA readiness и signoff не закрыты.

Evidence для browser vectors: filters-canary-20260925T191720Z.json. Matrix и 044 quickstart/traceability отражают correction; предыдущие исторические checkpoints выше сохраняются как датированные наблюдения, но не как актуальный вывод о мультимодальной способности.

T029a working-tree continuation — 2026-09-30

  • Два ранее падавших MCP-вектора (external_client_creates..., propose_test_pack_profile...) прошли: 2 passed. Синтетический source для baseline-пина теперь совпадает с dashboard, а обычные filter cases не требуют произвольного выбора metric coordinate.
  • Связанный набор MCP (scenario E2E, initial E2E, bootstrap automation, catalog) без зависшего HTTP initialize теста: 8 passed, 1 deselected. Отдельный focused scenario/registry набор: 208 passed. Фикстуры bootstrap теперь используют общую привязку обоих хэндлов к run; initial E2E использует канонический опубликованный source. Это тестовые поправки к новому обязательному receipt gate.
  • HTTP initialize тест test_initialize_publishes_catalog_version остановлен по timeout 15 секунд после восьми успешных случаев. Promotion E2E остановлен по timeout 20 секунд после sandbox exploration, до завершения первого теста. Их результат не засчитан как PASS. Ruff по затронутым backend путям, compileall и git diff --check прошли.
  • G-MCP-TEST-PACK-PROFILE остаётся PARTIAL / NO-GO: ещё не доказаны fresh profile→resolution→eligible register→bootstrap, typed authoring context, точная опубликованная baseline entry/filter binding и runnable сравнение. Полный release regression, PostgreSQL migration gate и внешние live gates не выполнялись. Текущие изменения не закоммичены.
  • Затем добавлен изолированный SQLite MCP-tool тест для selector-only профиля: один principal создаёт AgentRun, получает профиль, разрешает два selector вопроса через CAS, регистрирует save-eligible pack только по profile_handle_id и создаёт current revision. Проверены привязки receipt к сессии/digest/CAS и отказ при изменённом query fingerprint до записи pack/registry. Объединённый целевой набор: 217 passed, 1 deselected. Superset inspection в тесте мокирован, а вызовы MCP выполняются in-process; внешний wire-client и baseline/context ветви остаются открытыми.

T029a published-entry selector — 2026-09-30

  • Добавлен отдельный fail-closed selector 037 metric entry из receipt-bound опубликованного каталога: exact chart/dataset/result key/effective filters/ reference source, пересчёт entry_digest и coordinate_hash существующими 037 функциями, сравнение ожидаемого release ID/version/commit. Прямого raw snapshot аргумента нет; неподходящий релиз даёт BASELINE_STALE. Ранний 037 resolver отклоняет malformed filter/provenance payloads как BASELINE_EVIDENCE_UNAVAILABLE до разыменования. Источник ожидаемого релиза должен быть server-derived.
  • Семантическая проверка INV1–7: сбалансированные anchors, metadata перед кодом, 0 naked/parse warnings, валидные relations, 142 строки selector и 399 строк resolver, C901 clean. Ruff, compileall и diff-check прошли. Behavioral tests для этого среза не добавлялись и не запускались.
  • Selector ещё не вызывается из профиля/MCP. Профиль не хранит server-owned effective filter hash, а metric_name → result_key соглашение ещё не привязано к durable profile. Поэтому needs_baseline остаётся preview_only, T029a — PARTIAL / NO-GO, общий релиз — NO-GO. План data flow, versioned binding и recheck boundaries: specs/050-mcp-interface/plans/T029a-baseline-binding-slice-2026-09-30.md.