273 KiB
WORKSTATE 043–047 (intermediate 2026-09-24 06:52 +03:00; supersedes the 2026-08-20 header)
Статический + targeted-runtime checkpoint. Не является feature-completion. Правило:
[x]только с текущим доказательством;[~]частичная реализация;[ ]отсутствует.
T029a MCP profile update — 2026-09-29
External MCP scenario authoring now has bounded profile-preview handlers in
runtime catalog 2.6.0 with dashboard:testing:READ. The server re-inspects the selected dashboard and
environment, returns complete selected-case coverage, stable chart→dataset→metric
coordinates and native-filter applicability, and keeps unresolved questions
typed. Selector answers are restricted to one exact compiler step and guarded by
profile-digest CAS. Metric-coordinate selection changes needs_metric to
needs_baseline; it never supplies expected numeric truth or a baseline pin.
Bootstrap now checks selected environment and case coverage against digest-bound server handles before any registry/revision/workspace write. Targeted evidence: 32 focused profile/compiler/handle/bootstrap/MCP tests passed; Ruff, compileall and diff checks passed. This is PARTIAL, not production closure. Domain context and published-baseline resolution, save-eligible handle minting, full fresh external-client profile→bootstrap E2E and catalog/schema parity remain open. Overall production status remains NO-GO.
Current UX handoff
2026-09-24 architecture amendment: every new baseline is defined by an explicit same-dashboard Superset URL and its server-parsed immutable native-filter snapshot (037 ReferenceUrl). This is a P0 production invariant, tracked as G-BSL-REFERENCE-URL in the unified matrix, 037 T091 and 050 T052. Both supplied ss-dev dashboard 10 URLs have live read-only parser/query evidence; the native_filters_key=A4_rwhfQta0 link was rerun on 2026-09-25 with platform=3DS, 10/10 metrics executed and 6 distinct metric formulas suggested. A separate, removed-after-test Compose project with a dashboard 10 scenario and test-only synthetic release then proved browser URL entry → filter/metric preview → six draft candidates (Chromium 2/2, including a different-dashboard field error). Its isolated DB retained the exact URL hash and platform IN ["3DS"] snapshot across six captures. The ordinary local database still has no dashboard 10 scenario/release fixture; a later isolated Chromium pass proved source-byte verification, reasoned approval and local catalog materialization for one URL candidate, with the exact URL/filter snapshot retained in its revision. Git publication has no receipt; real release→Git publication→run pin→scheduled replay remains unproven. Do not infer the full invariant or analyst acceptance from the seeded 4/5 review score.
The dashboard-testing candidate is now 9acd9bad (UX slice on top of fa531225) and has live smoke proof for route loading, queue→case navigation, protected evidence, schedule preview and run-center rendering. The UX slice adds grouped queue states, business-language case summary, evidence availability cards and collapsed technical disclosures; focused frontend proof is 13 tests, with lint/build green and Axiom parse warnings at zero. A business-analyst expert review nevertheless scores the seven critical tasks 2, 1, 2, 2, 1, 3, 2 respectively, mean 1.9/5; this is not moderated five-analyst acceptance. See specs/HANDOFF-DASHBOARD-ANALYST-UX-2026-09-24.md. G-BA-UX-READINESS remains OPEN.
The most important remaining product work is live analyst validation of investigation/case and baseline workflows, period-change/curated baseline selection UX, business-language explanations for launch/result states, and the five-analyst study with retained browser/network evidence. The uncommitted seeded baseline review slice reaches a provisional expert 4/5 on its scoped happy path; it does not imply readiness acceptance.
Overall
| Spec | Status | Notes |
|---|---|---|
| 042 Registry | [~] |
CRUD/lifecycle есть. Upstream 037/041 ingestion + queue signal wired. Full quickstart/T024 unverified. |
| 043 Editor | [~] |
Code/route present. No current E2E/a11y proof. Depends on 042/044. |
| 044 Execution | [~] not production-complete |
Walker/lifecycle/API работают. Executors fail-safe. Live Playwright/Superset session still missing. |
| 045 Run Monitor | [~] |
Typed launch contract aligned (release/baseline/toggles). Full route E2E open. |
| 046 Automation | [~] |
Scheduler reload + trigger dispatch + notifications persist. Live due-job E2E open. |
| 047 Analytics | [~] not production-complete |
Signal ingest, CAS closure, recurrence, owner ACL. Chat/AgentRun workspace open. |
Verified facts (this checkpoint)
044 Execution
- DAG walker
_advance_runwalkstopological_order, claims steps, registers artifacts, suspends on human. - Failed/blocked producers block descendants including human; walker does not suspend a terminal failed run.
- API
POST /scenario-runscreates durable queued run (auto_advance=False); worker/tests advance separately. - Launch contract: typed
dashboard_release_id,baseline_set,execution_togglesonStartRunRequest/start_run; frontend no longer hides them inparams.launch_config. - Executors (
execution/executors.py):assertion→ 037compare_values(can fail; never hardcoded PASS).xlsx→ openpyxl parse of real workbook bytes + artifact ref.screenshot/report/artifactbind only with real bytes/refs/digest.browser/superset_apiwithout session/query envelope → typedinconclusive.
- Cancel: queued steps →
skipped, run →cancelled(viacancel_requested/draining). - Timeout:
apply_step_timeout→ step/runinconclusive+STEP_TIMEOUT. - Terminal failed/blocked/inconclusive →
auto_queue_failed_run+ persisted notification.
047 Analytics
ingest_investigation_signal()idempotent queue projection; no auto agent start.- Queue/case persist
evidence_snapshot+linked_run_ids. - Closure CAS:
resolvedneeds reconciled verification evidence;acceptedneeds rationale. - Recurrence after resolved episode opens new episode + queue item.
- Object ACL:
InvestigationCase.owner_id;can_access_case; GET/cases/{id}403 for non-owner view. - Migration
b9c0d1e2f3a4adds evidence snapshot, linked_run_ids, owner_id.
046 Automation
load_schedules()reloads enabledScenarioScheduleand keepsscenario_*jobs in desired set.add_scenario_jobmaps persistedmisfire_grace_time,max_instances, missed-execution/coalesce.dispatch_trigger_eventresolvescurrentrevision and callsstart_run.POST /api/scenario-automation/events/dispatchis the HTTP boundary.persist_notificationwrites lifecycle events (completed/failed/blocked).
042 / 045
ingest_upstream_staleness_event()for normalizedstructure_diff/lineage_blast_radius.- Staleness apply emits investigation queue signal.
- RunMonitorModel posts typed 044 start fields (Vitest: launch body contains
dashboard_release_id, notlaunch_config).
Targeted evidence (do not treat as suite-close)
Recorded in this session with AUTH_SECRET_KEY / SECRET_KEY set:
- Investigation + staleness + executors: 12 passed (later expanded).
- Automation services/API: 32 passed; trigger dispatch pinning: 9 passed.
- Runner/walker/dispatch after fail-safe executors: 11 passed.
- Run API after queued-start restore: 30 passed.
- RunMonitorModel Vitest: 2 passed.
- Core scheduler regression: 17 passed.
- Consolidated 042/044/046/047 slice: 81 passed.
- Recurrence + walker fail-path + executors: 47 passed, ruff clean on touched modules.
- Lifecycle cancel/timeout + investigation ACL: 20 passed.
- Analytics API after resolved-disposition contract: 21 passed.
Full backend/frontend suites, Playwright, semantic rebuild: not rerun.
Remaining (priority)
- 044 live adapters — Playwright session replay and catalog-backed Superset query when a real session/query envelope exists. Current
inconclusiveis fail-safe, not SCEX-FR-002/009 complete. - 044 T025 remainder — approval-to-live-dispatch, infra resume continuation of remaining DAG, timeout during in-flight executor I/O.
- 047 T016 remainder — durable chat / linked AgentRun workspace on InvestigationCase.
- 047 T018 / 044 T022 — independent full scoped verification; do not reuse old aggregate counts.
- 045/043 UI — mount remaining production panels; browser/a11y E2E.
- 039 T058 — wire
_trigger_release_verificationinto preprod deploy (out of 042–047 core, still listed).
Key files (this checkpoint)
backend/src/services/dashboard_testing/execution/{runner,executors,lifecycle,dispatch}.pybackend/src/services/dashboard_testing/analytics/{investigation,recurring}.pybackend/src/services/dashboard_testing/automation/{trigger,notify}.pybackend/src/services/dashboard_testing/registry/staleness.pybackend/src/core/scheduler.pybackend/src/api/routes/dashboard_testing/{scenario_runs,scenario_analytics,scenario_automation}.pybackend/src/models/scenario_investigation.pybackend/alembic/versions/b9c0d1e2f3a4_add_investigation_evidence_snapshot.pyfrontend/src/lib/models/RunMonitorModel.svelte.ts
MCP Handoff (2026-08-28 08:00 +03:00)
Это промежуточный handoff для следующего агента. MCP infrastructure реализована частично. Maintenance adapters ниже являются только техническим доказательством approval/scheduler/task boundary и не являются продуктовым happy path dashboard testing.
Product priority
Основной продуктовый vertical path:
MCP client
-> 042 current ScenarioRevision / WorkingDraft
-> 038 compile + deterministic validation
-> 037 baseline/lineage resolution
-> 044 ScenarioRun creation
-> 045 approval/monitoring
-> 044 ExecutionCapacityManager + live providers
-> 037/044 evidence and StepOutcome
-> 047 InvestigationSignal / Queue / Analytics
Следующий агент НЕ должен продолжать расширять maintenance как самостоятельную feature. Приоритет:
038 -> 042 -> 044 -> 045 -> 047 MCP vertical integration.
Implemented MCP infrastructure
- Official Python SDK
mcp==1.29.1; Streamable HTTP mounted at/mcp. - FastMCP lifespan bridge and local host/origin allowlists (
MCP_ALLOWED_HOSTS,MCP_ALLOWED_ORIGINS). - Request byte limit for both
Content-Lengthand chunked bodies; buffered ASGI messages are replayed. - Explicit catalog with safe, guarded, approval, and read-only tools; no generic
api_call. - Live RBAC catalog filtering and invocation guard; service principal is read-only.
- RS256 MCP JWT,
kid, JWKS/oauth/jwks.json, current/previous public-key overlap. - PKCE S256 authorization code flow, refresh rotation/reuse revocation, DCR rate limit, client revocation.
McpToolInvocationRecord: argument hash, approval/dispatch state, continuation payload, lease/fencing, result hash; raw bearer tokens, arguments, results, and credentials are not persisted.- MCP-owned
ActionApprovalGate, human-onlylist_pending_approvals/decide_approval, CAS decision. - Payload hash and operation binding revalidation before dispatch.
- Scheduler poller with explicit dispatcher registry, lease renewal, stale claim recovery and fencing.
- Current reviewed adapters:
start_maintenanceandend_maintenanceonly; they enqueue existing TaskManager/MaintenanceBannerPlugin work and do not call Superset directly.
Alembic head
0001 baseline
-> 0002 MCP OAuth
-> 0003 MCP provenance
-> 0004 continuation payload
-> 0005 dispatch receipt
-> 0006 dispatch lease
-> 0007 DCR rate limit
-> 0008_mcp_reconcile maintenance ownership/recovery
Evidence
- Focused MCP/OAuth/approval/maintenance unit suite: 41 passed at the latest reconciliation checkpoint.
- PostgreSQL Testcontainers concurrency suite: 4 passed for claim, stale-worker fencing, renewal, and MCP-owned orphan reconciliation.
- Real MCP protocol test covers
initialize -> mcp-session-id -> notifications/initialized -> tools/list -> tools/call. ruff,compileall,git diff --check, and clean Alembic upgrade through head passed at the latest checkpoints.- Full backend suite, full frontend suite, Playwright product E2E, and semantic index rebuild were not rerun after all MCP changes.
Known MCP/product gaps
- OAuth browser consent/session flow and native FastMCP OAuth middleware are not production-complete.
- DCR PostgreSQL concurrent rate-limit evidence, admin client management, and automated key-rotation expiry remain open.
- MCP catalog is not parity-complete with the intended 37 tool surface; current explicit catalog is a staged subset.
- Approved MCP dispatch is wired only to maintenance adapters; no real 044
ScenarioRundispatcher exists yet. - 044 live Browser/Superset/Screenshot provider composition and
ExecutionCapacityManagerare not proven. - No MCP vertical E2E currently proves ScenarioRevision -> compile -> validate -> ScenarioRun -> evidence -> analytics.
- Multi-process scheduler restart/soak, retry/dead-letter policy, and provider I/O failure injection remain open.
- PII is accepted for enterprise-local MCP/LLM/VLM providers. Credentials, cookies, access/refresh/service tokens, secrets, and raw storage paths remain prohibited in payloads, telemetry, provenance, and artifact refs.
Next implementation packet
Start with a dashboard-testing adapter, not maintenance:
- Find the existing 042 revision/WorkingDraft load and 038 compile/validate service APIs.
- Add a typed MCP
inspect_scenario/validate_scenarioread path using those services. - Add a typed gated
start_scenario_runrequest that creates the existing 044 queuedScenarioRunonly after validation and approval; never accept a client-supplied compiled graph as canonical. - Bind approval continuation to
scenario_revision_id,revision_digest, validation digest, environment policy fingerprint, and execution toggles. - Use existing 044 queued dispatcher and capacity boundary; do not create a second runner.
- Add PostgreSQL and MCP protocol tests for stale revision, digest drift, unknown lineage, duplicate run, approval CAS, cancellation, timeout and evidence receipt.
Working-tree note
All current changes are uncommitted. Existing unrelated modifications/deletions in .agents/, .kilo/,
AGENTS.md, and earlier spec files were present during the work and must not be reverted. Review
git status and git diff before any commit.
MCP Checkpoint (2026-08-28)
Dated implementation checkpoint. Claims below are limited to evidence collected in this round.
Verified changes
- MCP catalog removed
show_capabilitiesand unbound approval stubs;list_maintenance_eventsis read-only and non-gated. - Repeated direct
RbacFastMCPgated calls reuse the invocation hash, gate hash, and request hash. - Fixed the gate-id handling that caused
DetachedInstanceError. - Frontend exposes the
MCP_DECOMMISSIONhandoff surface. Default and flag-on builds passed; the focused contract test passed.
Verification evidence
- Backend MCP focused suite: 44 passed.
ruffandcompileall: passed.- Frontend lint: 0 errors / 373 warnings.
- Semantic anchor repair for
mcp_approvals: 13/13. - Browser E2E was unavailable because the stack at ports
8101/8102was unreachable.
Product decision
Product remains NO-GO. Open proof and delivery areas are:
- live providers and execution capacity;
- the full
038 -> 042 -> 044MCP vertical; - parity with the intended 37-tool surface;
- full backend and frontend suites;
- semantic rebuild;
- browser E2E.
MCP Checkpoint (2026-08-28, this round)
Dated implementation checkpoint. Claims below are limited to evidence collected in this round.
Verified changes
- Added typed MCP
inspect_scenario,validate_scenario, andstart_scenario_runusing existing 038/042/044 boundaries. start_scenario_runhas strict, server-owned revision input and does not accept a client-supplied graph.- Permission classification is
RUN_PRODfor PROD andRUNfor non-PROD. - Removed duplicate scenario MCP approval ownership and the stale dynamic dispatcher/allowlist; maintenance-only explicit adapters remain.
Verification evidence
- Focused MCP suites after the final fix: 43 passed.
ruff,compileall, and scopedgit diff --checkpassed; anchors are balanced.
Product decision
Scenario MCP start remains intentionally queued/not-ready because no reviewed scenario dispatcher or live provider adapter exists.
Product remains NO-GO. The following remain open:
- full scenario authoring persistence;
- MCP 37-tool parity;
- live providers and capacity;
- PostgreSQL integration;
- browser E2E;
- full suites;
- semantic rebuild.
Architectural Amendment Checkpoint (2026-08-28)
User-approved amendment to the 043–047 authoring and execution architecture. This checkpoint records the governing boundary; it does not change the completion status or convert unverified implementation into evidence.
Approved architecture
-
The agent is a full co-author of intent and diagnostics through the persistent, server-owned
AgentAuthoringWorkspaceand governed MCP operations. It is not a production authority. -
Exploratory Playwright/code is allowed only inside an isolated authoring sandbox, with explicit time, size and network limits, cancellation, durable receipts and bounded artifact ownership. Credential, filesystem and shell escape is forbidden, as are network-policy escapes and production side effects.
-
AuthoringArtifactis distinct fromExecutionProgram: source, patch, trace, screenshot, diagnostic and operation-receipt artifacts are review material only;ExecutionProgramis the typed, deterministically validated 038 graph/program. -
Promotion is strictly:
sandbox -> typed action candidates / graph proposal -> deterministic 038 compile + validate -> user diff review -> 042 handle-based save (candidate) -> separate activate_revision CAS -> 044 execution -
Raw code is never direct
ScenarioRunauthority. A code-backed production provider remains a future, separate and unimplemented contract; exploratory code support does not imply production code execution.
Canonical MCP operation names
The exact canonical MCP names are the following 15 snake_case operations. REST
operationId values are transport aliases only and are distinct from MCP names; REST
aliases MUST NOT be exposed as canonical MCP names in tools/list.
inspect_dashboard_contextcreate_authoring_sessionpropose_test_planstart_explorationget_exploration_resultpropose_graph_revisionget_graph_diffpromote_to_scenarioscenario_compilescenario_validatescenario_resolvegenerate_draft_packrequest_saveactivate_revisionstart_scenario_run
The corrected 042 rule is explicit: ordinary save creates an immutable candidate;
only the separate activate_revision operation may advance current_revision, subject
to eligibility, materialization, policy, approval and CAS checks. Initial creation may
be the sole explicit exception when its eligible initial revision is created as current.
Verification and release decision
- Scoped
git diff --checkpassed. - Required semantic anchors remained balanced after the append.
- Product remains NO-GO. Open proof/delivery areas are authoring promotion E2E, sandbox runtime isolation and limits, MCP implementation and REST/MCP parity, live 044 providers and execution capacity, PostgreSQL integration evidence, and complete end-to-end/full-suite evidence.
Checkpoint — 2026-08-28: AgentAuthoringWorkspace foundation implementation
- Added the
AgentAuthoringWorkspacemodel and migration chain0009 -> 0010 -> 0011. - The exact normative workspace states are:
draft,exploring,exploration_failed,exploration_passed,proposal_ready,validation_blocked,awaiting_user_review,save_eligible,pending_approval,candidate, andcurrent. - Create, load, transition, attach, and expire operations are owner-bound and use CAS; reads of expired workspaces fail read-only without a hidden mutation.
- Durable per-operation idempotency and provenance receipts support replay after a later mutation and reject a changed-hash conflict.
- The registry remains unmodified by the workspace foundation.
- Twelve tests passed, including migration smoke; Ruff and
compileallpassed; one Alembic head is present at0011. - PostgreSQL upgrade was not executed or verified; the scoped diff operation is absent.
- Sandbox runtime, MCP authoring tools, graph promotion/activation, authoring E2E, provider/capacity evidence, and full production evidence remain open.
- Product remains NO-GO.
Checkpoint — 2026-08-28: MCP authoring session implementation
create_authoring_sessionMCP tool is implemented with strict bounded input, authenticated user-derived ownership, service-principal denial, transactional workspace creation, durable idempotency replay/conflict handling, and no registry mutation.propose_test_planis intentionally deferred because no safe canonical persistence shape exists.- Twenty-five MCP/workspace-scope tests passed; Ruff and
compileallpassed; the Alembic head is0011. - Canonical authoring operations and the scenario chain remain incomplete. Sandbox, proposal/diff/promotion/activation, MCP parity, response bounds, live 044 providers/capacity/PostgreSQL/full E2E remain open.
- Product remains NO-GO.
Checkpoint — 2026-08-28: bounded test-plan intent artifact
- Implemented the
propose_test_planMCP operation as a bounded, server-owned intent artifact. - Added immutable
AgentAuthoringWorkspaceTestPlan, migration0012after0011, and a canonical server-computed digest. - Enforced owner binding, CAS, expiry, idempotency, provenance, and dangerous content bounds. The operation accepts no raw executable graph or code and does not mutate the registry, revision, or run.
- Fixed the fresh SQLite migration-chain issue by making guards in
0010–0012idempotent;0001was not rewritten. - Fixed the owner-before-replay disclosure vulnerability.
- Verification completed: fresh SQLite upgrade twice; 28 targeted tests;
Ruff and
compileall; one Alembic head at0012; semantic anchors andgit diff --checkpassed. - PostgreSQL concurrency and partial-schema migration branches were not verified. Sandbox, exploration, proposal/diff/promotion/activation, MCP parity, and live 044 evidence remain open.
- Product remains NO-GO.
Checkpoint — 2026-08-28: bounded exploration request/result boundaries
- Implemented
start_explorationas a bounded, server-owned, persistence-only request boundary with the exact 038REGISTERED_ACTIONSallowlist, dangerous-content rejection, CAS/idempotency/provenance, and a typedsandbox_unavailableoutcome. - The workspace remains
draft. The operation performs no provider or subprocess invocation and makes no registry, revision, run, or artifact mutation. - Implemented read-only
get_exploration_resultwith owner-before-disclosure authorization, service-principal denial, a typed not-found outcome, and a bounded projection only. - Fixed the SQLite timezone CAS defect and closed the affected semantic anchor. A fresh migration chain now upgrades twice successfully.
- Verification completed: 34 targeted tests passed; fresh SQLite upgrade
twice; Ruff and
compileall; Alembic head0013; semantic anchors andgit diff --checkpassed. - Remaining: no real sandbox/provider; no exploration result or artifact persistence from a provider; no graph proposal/diff/promotion/activation; no PostgreSQL concurrency, browser E2E, MCP parity, or full production verification.
- Product remains NO-GO.
Checkpoint — 2026-09-01 (graph revision boundary)
- Implemented
propose_graph_revisionMCP operation: persists one server-derivedScenarioEditProposalfrom the closed typedEditOperationunion (via existingagent_propose+apply_ops) and attaches it to the workspace. Enforces owner-before-replay disclosure, draft-only state, scenario/revision binding, CAS, and durable idempotency (same key+ops replays the same proposal; changed ops raise a typed conflict). No candidate orcurrentrevision is created or activated; the client never supplies a graph. - Implemented read-only
get_graph_diffMCP operation with owner-before- disclosure and bounded projection (proposal_id,base_revision_id,digest,diff,validation,proposal_status,cas_version). No mutation, provider I/O, or graph-snapshot disclosure. - Renamed the editor's graph-diff helper
_graph_diffto a publicgraph_diffand reused it as the single diff authority acrossagent_proposeand the read projection. - Verification: 47 targeted tests passed (workspace service, editor agent,
MCP server); Ruff and
compileallpass; Alembic head remains0013; anchor pairing balanced (fixed a previously unclosedGetExplorationResultregion); scopedgit diff --checkpasses. - Remaining: exploration sandbox runtime;
promote_to_scenario/request_save/activate_revision; PostgreSQL concurrency; browser E2E; full MCP authoring E2E; 044 live providers/capacity. Product remains NO-GO.
Checkpoint — 2026-09-01 (promotion boundary promote_to_scenario)
propose_graph_revisionnow advances the workspacedraft -> proposal_ready(in addition to attaching the server-derived proposal), aligning with the normative promotion machine.- Implemented
promote_to_scenario: deterministic promotion boundary that reachesproposal_ready -> awaiting_user_review(orvalidation_blocked), recomputes the proposal digest server-side, returns the computed diff and validation, and records a durable idempotent operation receipt. It never creates or activates a revision and never accepts a caller digest. - Verification: 50 targeted tests passed (workspace service + editor agent + MCP
server); Ruff and
compileallpass; Alembic head0013; anchor pairing balanced (10/10service,26/26MCP); scopedgit diff --checkpasses. - Remaining:
request_save(candidate save viasave_proposalwith policy authorization),activate_revisionservice + MCP, isolated sandbox runtime, 038 compile/validate parity names, PostgreSQL concurrency, browser E2E, full MCP authoring E2E, 044 live providers/capacity. Product remains NO-GO.
Checkpoint — 2026-09-01 (candidate save request_save)
- Implemented
request_saveMCP operation: saves the owner-reviewed proposal into one immutablecandidateScenarioRevisionthrough the existing guardedsave_proposal->save_revision->create_revisionpath. The digest is recomputed server-side; the caller never supplies one. The workspace advancesawaiting_user_review -> candidate;current_revisionis never advanced. request_saveis the first authoring operation with a real registry mutation, so its catalog entry requires("scenario", "EDIT")and denies service principals (risk_level="guarded").- Verification: 53 targeted tests passed (workspace service + editor agent +
MCP server); Ruff and
compileallpass; Alembic head0013; anchor pairing balanced (11/11service,28/28MCP); scopedgit diff --checkpasses. - Remaining:
activate_revision(separate CAS;activate_current_revisionis still spec-only), isolated sandbox runtime, 038 compile/validate parity names, PostgreSQL concurrency, browser E2E, full MCP authoring E2E, 044 live providers/capacity. Product remains NO-GO.
Checkpoint — 2026-09-01 (activation activate_revision)
- Implemented the spec-only
activate_current_revisioninregistry/revisions.py: atomically promotes an eligible candidate to current, demotes the incumbent, advancesentry.current_revision_id, enforces a compatibility-family check, and writes aScenarioLifecycleAuditrow. - Implemented the workspace
activate_revision(and MCP tool): requires an owner workspace incandidate, an explicitrevision_id, CAS, and idempotent replay. It callsactivate_current_revisionand advances the workspacecandidate -> current. AScenarioRuncan now be pinned to a genuinely promoted/currentrevision_id. activate_revisionrequires("scenario", "EDIT")and denies service principals (risk_level="guarded"), matchingrequest_save.- Verification: 286 targeted tests passed (workspace service + MCP server +
the full registry service suite); Ruff and
compileallpass; Alembic head0013; anchor pairing balanced (12/12service,30/30MCP,6/6revisions); scopedgit diff --checkpasses. - The authoring promotion chain
session -> plan -> proposal -> promote -> diff -> request_save -> activate_revisionis now implemented end-to-end (exploration/sandbox aside). - Remaining: isolated sandbox runtime, 038 compile/validate parity names, PostgreSQL concurrency, browser E2E, full external MCP authoring E2E, 044 live providers/capacity. Product remains NO-GO.
Checkpoint — 2026-09-01 (038 parity: scenario_resolve + generate_draft_pack)
- Added read-only MCP tools
scenario_resolveandgenerate_draft_packwrapping the existing 038resolve_scenarioandgenerate_draft_pack.scenario_resolveapplies the typedResolveChangeunion (parameter|selector|manual_conversion| remove_step) with a boundedvalue(<=2048 canonical chars), re-validates the resolved graph, and returns the newrevision_hash+ validation.generate_draft_packreturns the server-owned manifest only (no artifact bytes),save_eligibleorpreview_only. Both are read paths and persist nothing. - Catalog now:
inspect_scenario(compile),validate_scenario,scenario_resolve,generate_draft_pack,start_scenario_run(plus the authoring chain). - Verification: 288 targeted tests passed (MCP server + workspace service +
registry service suite); Ruff and
compileallpass; Alembic head0013; anchor pairing balanced (32/32server); scopedgit diff --checkpasses. - The 038 authoring pipeline
compile -> validate -> resolve -> generate_draft_packis now reachable over MCP, complementing the implemented promotion chainsession -> plan -> proposal -> promote -> diff -> request_save -> activate. - Remaining: isolated sandbox runtime (
start_explorationreal provider), PostgreSQL concurrency, browser E2E, full external MCP authoring E2E, 044 live providers/ capacity. Product remains NO-GO.
Checkpoint — 2026-09-01 (PostgreSQL migration + concurrency evidence)
- Fixed a PostgreSQL-only migration defect: revision IDs
0010..0013exceeded thealembic_version.version_num VARCHAR(32)width and failed withStringDataRightTruncation. Shortened the four authoring revision IDs (0010_authoring_contract,0011_authoring_operations,0012_authoring_test_plan,0013_authoring_exploration) and renamed the files; the linear chain is intact. - Fixed schema drift surfaced by
alembic check:AgentAuthoringWorkspaceOperationusedindex=Trueonworkspace_id, auto-generating an index name that differed from the0011migration. The model now declares the explicitix_authoring_workspace_operations_workspaceindex, matching the migration. - Evidence captured against a fresh real PostgreSQL (postgres:16):
alembic upgrade headapplies the full0001 -> 0013chain;alembic checkreports no new upgrade operations (no schema drift);- MCP PostgreSQL concurrency suite passes
4/4(--run-integration): one-winner CAS claim, stale-worker fencing, lease renewal fencing, scoped/idempotent orphan reconciliation.
- Verification:
54 targeted tests passed(migration smoke + workspace + MCP); Ruff andcompileallpass; Alembic head0013_authoring_exploration. - Remaining: isolated sandbox runtime, browser E2E, full external MCP authoring E2E, 044 live providers/capacity. Product remains NO-GO.
Checkpoint — 2026-09-01 (044 ProviderRuntime spec amendment)
- Amended spec 044 to close the three provider-spec gaps identified during the live-provider
effort assessment (
contracts/modules.md,data-model.md,tasks.md; +43/-4 lines):- Async/sync bridge pinned — new
ScenarioExecution.ProviderRuntimeADR contract: exactly one application-owned long-lived provider event-loop thread; sync executors submit viarun_coroutine_threadsafeand block under the step deadline; context/session objects live only on that loop and remain usable across all steps of one run.@REJECTEDper-stepasyncio.run(cannot hold live contexts), warm context pools for Phase 1; dedicated browser worker process deferred with adapter-Protocol survival requirement. - Concurrency defaults pinned — 2 concurrent browser contexts for DEV/PREPROD, 1 for PROD
per environment; screenshot capture admitted through the same CapacityManager lease
accounting; exhausted capacity keeps the run queued with
CAPACITY_BLOCKED; bounded submit queue (overflow waits under the step deadline, never starts I/O). - Trace/video policy pinned —
@REJECTEDPlaywright trace/video as durable evidence (non-deterministic, outside descriptor-declared evidence); operator debug flag keeps them local only, never registered as Artifact/EvidenceReceipt.
- Async/sync bridge pinned — new
- Implementation tasks amended accordingly: T032 (provider_runtime.py: loop thread + bounded submissions + defaults), T034 (loop-bound contexts, no per-step loop, no warm pool), T035 (wrap existing llm_analysis Playwright capture stack, shared loop + admission).
- Anchor balance after edit: modules.md 19/19 region pairs, 2/2 brace contracts.
Checkpoint — 2026-09-01 (T028 provider protocol + ProviderRuntime loop bridge)
- Implemented the first provider-production slice per the amended 044 spec:
backend/src/services/dashboard_testing/execution/provider_runtime.py—ScenarioExecution.ProviderRuntime.Engine: one long-lived application-owned event-loop thread;submit()is deadline-aware with bounded slots (semaphore); overflow raisesProviderSubmissionOverflowbefore the coroutine is created (no I/O started); deadline expiry cancels the in-flight coroutine and keeps the loop reusable; idempotentstart(), drainingstop(); process-wide accessor with test/shutdown reset.backend/src/services/dashboard_testing/execution/provider_protocol.py—ScenarioExecution.ProviderProtocol.Schemas(T028): frozenProviderExecutionContext/ProviderExecutionResultwith closed enums (passed|failed|inconclusive|blocked, effect states, retry dispositions), stablePROVIDER_CONTEXT_INVALID/PROVIDER_RESULT_INVALIDtaxonomy, no-manufactured-PASS invariant (passed requires operation_id; non-pass requires reason_code), and pre-I/Ocreate_operation_receiptwith terminal-status predicate.
- Tests:
backend/tests/services/dashboard_testing/registry/test_provider_runtime.py(11),test_provider_protocol.py(19) — deadline cancellation, overflow-never-starts-IO, precondition rejections, lifecycle join, schema/receipt shapes. - Verification:
30 passed(new) +63 passed(live-binding + MCP + workspace regression); Ruff clean;compileallclean; anchor nesting verified in all 4 new files. - Next: T032 full
capacity.pylease table (migration 0014), then T035 ScreenshotProvider wrapping the llm_analysis Playwright stack on this loop. Product remains NO-GO.
Checkpoint — 2026-09-01 (T032 ExecutionCapacityManager)
- Implemented the durable capacity slice per the amended 044 spec:
backend/src/models/provider_capacity.py—ScenarioExecution.Capacity:provider_capacity_quotas(one server-owned counter per environment+workload pair, unique index) andprovider_capacity_leases(claimed|released|expired|reconciled).backend/alembic/versions/0014_provider_capacity.py— guarded table creation, idempotent index creation, clean downgrade; revision id 22 chars (within VARCHAR(32)).backend/src/services/dashboard_testing/execution/capacity.py—ScenarioExecution.CapacityManager.Service: atomic admission via one CAS increment (active_units + units <= limit_units), typedCapacityUnavailable(CAPACITY_UNAVAILABLE | CAPACITY_LEASE_NOT_ACTIVE | CAPACITY_LEASE_UNKNOWN), pinned defaults (browser/screenshot 1 PROD / 2 otherwise; other classes server-owned), idempotent release with zero floor, bounded expiry reconciliation (reconcile_expired_leases), heartbeat, read-only snapshot; naive/UTC-aware datetime normalization for SQLite and PostgreSQL TIMESTAMP columns; belief-runtime reason/reflect/explore on every branch.- Fixed during slice: heartbeat discovering an expired lease now frees quota units via
the shared CAS
_expire_lease_and_free(no double-free on concurrent paths).
- Tests:
backend/tests/services/dashboard_testing/registry/test_provider_capacity.py(17) — pinned defaults, exhaustion keeps counter unchanged, environment independence, heartbeat extension/rejection, idempotent release, expiry-once reconciliation, precondition taxonomy. - Verification:
17 passed(capacity) +98 passed(provider runtime/protocol + MCP + workspace regression); Ruff clean;compileallclean; anchor nesting OK (4 files); migration smoke3 passed;alembic heads=0014_provider_capacity; full chain0001→0014applied on fresh postgres:16 withalembic check— no drift. - Next: T035 ScreenshotProvider wrapping the llm_analysis Playwright stack on the shared loop with capacity admission; then dispatcher integration (lease into ProviderExecutionContext) and T034 BrowserProvider. Product remains NO-GO.
Checkpoint — 2026-09-01 (T035 ScreenshotProvider)
- Implemented the first real live provider per the amended 044 spec:
backend/src/services/dashboard_testing/execution/providers/screenshot.py—ScenarioExecution.ScreenshotProvider: registered LiveProvider wrapping the existing llm_analysisScreenshotServicetransport; capture submitted to the single shared provider event loop (never per-call run_async); capacity admitted per attempt through the ExecutionCapacityManager (workload_class=screenshot, PROD limit 1 / otherwise 2) and always released infinally; evidence re-digested after capture, size-capped at 10 MiB, stored only as opaquedraft:{run_id}:{digest}refs; typed non-pass taxonomy (SCREENSHOT_BINDING_INVALID/MISMATCH, RUN_UNRESOLVED, TARGET_MISMATCH, CAPACITY_UNAVAILABLE, CAPTURE_TIMEOUT/EMPTY/FAILED, TOO_LARGE, EVIDENCE_REF_INVALID, LOOP_UNAVAILABLE/OVERFLOW).live_composition.pybootstrap now registers the deployment-owned screenshot provider per enabled binding (falls back to typed-unavailable); browser-safe providers remain explicitly unavailable. Amended the bootstrap @INVARIANT.- Fail-closed ordering fixed after the pre-existing lifespan test caught it: the
provider checks loop availability BEFORE any DB/capacity I/O, so a not-running
runtime performs zero side effects (
SCREENSHOT_LOOP_UNAVAILABLE).
- Tests:
test_provider_screenshot.py(9) — happy path with lease release evidence, binding/identity rejections before I/O, capacity exhaustion refuses capture, PROD single-context limit, empty/oversize/multi-image evidence shapes, and a loop-binding probe proving capture runs on the shared loop. Transport ([EXT:Browser]) is the only mocked boundary; capacity, loop bridge, storage and DB are real. - Verification:
122 passedscoped regression (providers + live binding + MCP + workspace + migration smoke); Ruff clean;compileallclean; anchor nesting OK;git diff --checkclean. - Next: dispatcher integration (capacity lease into ProviderExecutionContext) and T034 BrowserProvider on the same loop; then T040 startup readiness preflight. Product remains NO-GO (browser provider, sandbox, E2E, capacity soak open).
Checkpoint — 2026-09-01 (T034 read-only BrowserProvider slice)
- Implemented the browser provider read-only core per the amended 044 spec:
backend/src/services/dashboard_testing/execution/providers/browser.py—ScenarioExecution.BrowserProvider: read-only action setopen_dashboard,wait_for_state,refresh; descriptor validation before any I/O (presence, tool match, mutating rejection, registry fingerprint when present); capture via aBrowserActionTransportseam whose real implementation reuses the deployment-owned ScreenshotService_launch_and_loginflow (server-owned auth, isolated headless Chromium context per attempt, context/browser closed on every path); evidence screenshot re-digested, size-capped at 10 MiB, stored as opaquedraft:{run_id}:{digest}; capacity admitted per attempt through the ExecutionCapacityManager (workload_class=browser, PROD 1 / otherwise 2) with guaranteed release; typed taxonomy (BROWSER_BINDING_*, RUN_UNRESOLVED, TARGET_MISMATCH, ACTION_DESCRIPTOR_REQUIRED, ACTION_TOOL_MISMATCH, ACTION_NOT_SUPPORTED, REGISTRY_MISMATCH, CAPACITY_UNAVAILABLE, ACTION_TIMEOUT, LOOP_UNAVAILABLE/OVERFLOW, EVIDENCE_REQUIRED/TOO_LARGE/REF_INVALID, ACTION_FAILED).live_composition.pybootstrap registers both providers (screenshot + read-only browser) from one deployment-owned ScreenshotService per enabled binding, falling back to typed-unavailable; mutating browser capability stays unregistered.
- Tests:
test_provider_browser.py(11) — happy path with checkpoints/lease release, descriptor rejections before I/O (missing/tool-mismatch/mutating), binding mismatch, capacity exhaustion refuses transport, PROD single-context limit, evidence missing/oversize, shared-loop probe. Only [EXT:Browser] is a test double. - Verification:
133 passedscoped regression (browser + screenshot + capacity + runtime + protocol + live binding + MCP + workspace + migration smoke); Ruff,compileall, anchors,git diff --checkclean. - Next:
wait_for_state/refreshreal-stand canary, mutating-action flow (fixture lease + reconciliation), dispatcher lease integration, T040 readiness preflight, T042b PREPROD canaries. Product remains NO-GO.
Checkpoint — 2026-09-01 (capacity-aware dispatch)
- Wired the capacity policy into the dispatch lifecycle per the amended 044 spec:
runner.pynew regionScenarioExecution.Runner.CapacityBlock: a provider step refused by the shared allocator (BROWSER_CAPACITY_UNAVAILABLE/SCREENSHOT_CAPACITY_UNAVAILABLE) no longer terminalizes the run. The step returns to queued without consuming its attempt (outcome flaggedcapacity_retry), and the run parks asqueuedwithCAPACITY_BLOCKED.reconcile_capacity_blocked_runsclears the parking code (bounded, idempotent) anddispatch_queued_runscalls it at cycle start, so parked runs re-attempt with a fresh capacity claim on the next cycle.@REJECTED: terminalizing capacity-refused runs as inconclusive.
- Tests:
test_scenario_capacity_block.py(3) — requeue semantics (attempt preserved, no terminal side effects), reconcile-then-retry reaching a real terminalpassedwith the same attempt, and scoping (non-capacity error codes untouched). - Verification:
3 passednew +361 passedfull registry suite regression (queued dispatch, crash recovery, providers, live binding, approvals, automation); Ruff,compileall, anchors,git diff --checkclean. - Next: mutating browser-action flow (fixture lease + precondition hash + reconciliation), T040 startup readiness preflight, T042b PREPROD canaries. Product remains NO-GO.
Checkpoint — 2026-09-01 (T040 provider readiness preflight)
- Implemented the startup readiness surface per the amended 044 spec:
backend/src/services/dashboard_testing/execution/providers/preflight.py—ScenarioExecution.ProviderPreflight: JSON-safe readiness snapshot per capability (provider loop, evidence storage, screenshot, browser, bindings) built after bootstrap registration; bounded Chromium-executable probe (resolves the executable without launching a browser) via the caller's run_async; evidence storage construction check (no bytes written); every check failure/crash becomes a typed degraded reason (BROWSER_EXECUTABLE_MISSING,BROWSER_PROBE_TIMEOUT,BROWSER_PROBE_FAILED,EVIDENCE_STORAGE_UNAVAILABLE,PROVIDER_CHECK_CRASHED,*_PROVIDER_UNREGISTERED), never an exception and never a startup blocker.@REJECTED: blocking startup on readiness — typed-unavailable registration already fails closed at dispatch.live_composition.pybootstrap now logs the readiness snapshot at the end of composition (Provider readiness snapshot).
- Tests:
test_provider_preflight.py(5) — ready states with bookkeeping of unavailable bindings, probe failure degradation, check-crash safety, unregistered capability reasons, and a real-checks smoke (returns None or stable codes). - Verification:
5 passednew +366 passedfull registry regression; Ruff,compileall, anchors,git diff --checkclean. - Next: mutating browser-action flow (fixture lease + precondition hash + reconciliation), T042/T042c contract test matrix, T042b PREPROD canaries. Product remains NO-GO.
Checkpoint — 2026-09-01 (T034 mutation admission + reconciliation semantics)
- Extended the BrowserProvider with the mutation contract gate and unknown-effect
semantics per the 044 BrowserProvider contract:
- Admission split into
ScenarioExecution.BrowserProvider.Admission(loop, binding, run identity, target identity, descriptor gating) and.MutationContract(per-key validation: fixture_lease_id, non-empty bounded target_keys <= 100, field_allowlist, 64-hex precondition_hash, cleanup_policy, retry_safe=false — typed codes BROWSER_MUTATION_CONTRACT_INVALID / PRECONDITION_INVALID / TARGETS_INVALID, all before any browser I/O). - Outcome semantics: mutating PASS carries operation_id + effect_state=completed; mutating timeout/crash -> BROWSER_MUTATION_RECONCILE_REQUIRED with effect_state unknown, reconciliation_required=true, retry after_reconciliation (never a plain timeout or PASS); transport-unsupported -> effect_state not_started (nothing started, fail-closed); read-only PASS carries effect_state=none without operation_id.
- Real transport explicitly raises BrowserTransportUnsupported for mutating actions before any I/O until the DOM-write flow lands; real read-only flow unchanged.
- Admission split into
- Tests:
test_provider_browser.pygrown to 24 — valid-mutation receipt, 8 contract rejection shapes, timeout/crash reconciliation semantics, unsupported not_started, read-only effect_state none. - Verification:
24 passedbrowser +379 passedfull registry regression; Ruff,compileall, anchors,git diff --checkclean. - Next: real DOM-write flow for edit_row/bulk_edit (fixture lease claim in transport, precondition verification, cleanup execution), durable provider operation receipts (T030), T042/T042c contract matrix, T042b canaries. Product remains NO-GO.
Checkpoint — 2026-09-01 (T030 durable provider operation receipts)
- Implemented durable write-once operation receipts and wired them into the browser
mutation flow:
backend/src/models/provider_operation.py—ScenarioExecution.ProviderOperation:provider_operation_receiptswith pinned identity (run/step/attempt, provider, descriptor fingerprint, binding, principal, capacity lease), unique per (run_id, logical_step_id, attempt), running -> terminal CAS lifecycle, JSON summary + append-only history, cancellation fields.backend/alembic/versions/0015_provider_operations.py— guarded creation with idempotent indexes;alembic heads=0015_provider_operations.backend/src/services/dashboard_testing/execution/provider_operations.py—ScenarioExecution.ProviderOperations.Service:open_provider_operation(before external I/O; duplicate attempt -> PROVIDER_OPERATION_DUPLICATE),complete_provider_operation(CAS terminal, late overwrite -> PROVIDER_OPERATION_TERMINAL),reconcile_provider_operation(only from reconciliation_required; appends audit history),record_late_response(history-only, status unchanged).- Browser mutation flow now opens the receipt atomically with the capacity claim (lease id recorded), finalizes it on every path (completed / reconciliation_required / failed+not_started for unsupported), and carries operation_id in typed outcomes. Read-only actions stay receipt-free.
- Tests:
test_provider_operations.py(8) — lifecycle, duplicate, identity validation, reconciliation source gating, late-response history; browser mutation tests now assert persisted receipt states (completed / reconciliation_required). - Verification:
33 passedreceipts+browser;388 passedfull registry regression; Ruff,compileall, anchors,git diff --checkclean; migration chain0001→0015applied on fresh postgres:16 withalembic check— no drift (temp DB dropped). - Next: reconcile worker wiring (scheduled resolution of reconciliation_required receipts), real DOM-write flow for edit_row/bulk_edit, T042/T042c contract matrix, T042b canaries. Product remains NO-GO.
Checkpoint — 2026-09-01 (provider reconciliation worker)
- Implemented the scheduled reconciliation sweep per the 044 ProviderOperations
contract:
provider_operations.pynew regionScenarioExecution.ProviderOperations. ReconcileWorker: explicit per-provider reconciler registry (register_provider_reconciler, reviewed registrations only) andreconcile_stale_provider_operations— a bounded, idempotent sweep over stalereconciliation_requiredreceipts (staleness cutoff on updated_at). Verdicts are validated (resolution completed|failed, registered effect_state) and resolved via the write-once CAS with an audit history entry.@REJECTED: auto-resolving receipts without a reconciler — silent completion would forge effect evidence.core/scheduler.py—execute_scheduled_provider_reconciliationscheduled entry (own session, commit, failure-safe rollback+log) registered as APScheduler jobprovider_operation_reconciliation(IntervalTrigger 30s, max_instances=1, coalesce=true), mirroring the queued-dispatch job pattern.
- Tests:
test_provider_reconcile_worker.py(5) — registered reconciler resolves a stale receipt with history note; missing reconciler leaves the receipt untouched; invalid verdict is skipped and logged; staleness cutoff ignores fresh receipts; the scheduled entrypoint resolves stale receipts in the global DB end-to-end. - Verification:
5 passednew +393 passedfull registry regression; Ruff,compileall, anchors,git diff --checkclean. - Next: real DOM-write flow for edit_row/bulk_edit with a browser reconciler (target-state verification against precondition hash), T042/T042c contract matrix, T042b PREPROD canaries. Product remains NO-GO.
Checkpoint — 2026-09-01 (T041/T042 provider contract matrix)
- Added the cross-provider contract matrix per the 044 acceptance profile, covering
vectors not proven by the per-provider suites (no duplication):
test_provider_contract.py(6) — lease safety on transport crash for BOTH providers (lease released, quota zero, no PASS manufactured), storage-ref ownership mismatch never passes (typed SCREENSHOT/BROWSER_EVIDENCE_REF_INVALID, empty artifact_refs), mutation receipt isolation per run (distinct receipts/leases, all released), and health redaction (readiness payload serialized contains no password/secret/token/cookie/credential shapes and carries exactly the five contract keys).- Parametrized across screenshot and browser seams; only [EXT:Browser] transports are doubles; capacity, receipts, storage and DB are real.
- Verification:
6 passedmatrix +399 passedfull registry regression; Ruff, anchors (stack check),git diff --checkclean. - Remaining open: real DOM-write flow for edit_row/bulk_edit (needs a live Superset stand for verification), browser reconciler target-state check, T042c binding revalidation matrix rows, T042b PREPROD canaries. Product remains NO-GO.
Checkpoint — 2026-09-01 (T042b read-only canary GREEN on live stand)
- Ran the read-only BrowserProvider canary against the live stand
https://ss-prod.bebesh.ru(dashboards API verified; stand classified PROD — mutating canaries are forbidden and were not run):specs/044-dashboard-scenario-execution/prototype/browser_readonly_canary.pyexercises the REAL production path:build_playwright_browser_transportover the deployment-owned ScreenshotService login flow, sharedProviderEventLoop,ExecutionCapacityManageradmission/release, durable draft evidence; no test doubles anywhere in the path.- Fixed during the canary:
build_playwright_browser_transportreturned a bare function while the adapter Protocol expects an object with.execute(AttributeError caught by the provider's typed fail-closed path, exposed by the canary's EXPLORE logs). Wrapped in a transport object; 11 browser unit tests green. - Results: 3/3 passed — open_dashboard (5.2s, checkpoint
dashboard_open), wait_for_state/networkidle (5.7s), refresh (4.9s); evidence SHA-256 digests + durable refs for every action; page_url confirmed dashboard navigation. - Evidence retained:
specs/044-dashboard-scenario-execution/evidence/browser-provider/ readonly-canary-20260901T143950Z.json+ 3 PNGs (one per action). Credentials are not stored in evidence.
- Regression:
399 passedfull registry suite; anchors, ruff, diff-check clean. - Contract target note (T042b): read-only canary + forced timeout/cleanup canary + safe-checkpoint reconstruction still required for GO; mutation canaries require a PREPROD-classified stand. Product status remains NO-GO (sandbox, browser E2E, scheduler soak, live Superset/Screenshot binding composition proof open).
Checkpoint — 2026-09-01 (T042b forced timeout/cleanup canary GREEN; stand reclassified)
- Stand owner confirmed
https://ss-prod.bebesh.ruis a TEST stand — mutation-capable classification is authorized. Canary runs now useSS_STAND_STAGE=PREPROD. - Implemented and ran the forced timeout/cleanup canary vector (T042b): a 1s
action deadline against the real stand forces the provider to cancel the in-flight
Chromium coroutine, release the capacity lease, close the browser context and return
typed
BROWSER_ACTION_TIMEOUTwith zero evidence — GREEN. Canary script gainedSS_CANARY_MODE=timeout; evidence JSON retained. - Re-ran the actions canary with the PREPROD classification: 3/3 passed again (login form-post fallback → authenticated 302 confirmed in belief logs).
- Evidence:
evidence/browser-provider/readonly-canary-20260901T144732Z.json+ PNGs,browser_forced_timeout_cleanupJSON. Regression399 passed; ruff/anchors clean. - Mutation DOM-write status: the stand's charts are standard read-only Superset
table/word_cloudvisualizations — no editable widget exists, and the repo ships no editable-table plugin. The provider mutation path (contract, receipts, capacity, reconciliation) is complete and unit-proven; the concreteedit_row/bulk_editDOM mechanics require a design decision (SQL-Lab-mediated data mutation vs a custom editable plugin on the target dashboards) before a mutation canary can be honest. Product status remains NO-GO (that decision + scheduler soak + browser E2E + sandbox open).
Checkpoint — 2026-09-01 (mutation canary GREEN on live stand)
- Design decision (044 BrowserProvider):
edit_row/bulk_editexecute as SQL-Lab-mediated data mutations performed by the authenticated browser session (page-context fetch to the CSRF-protected SQL Lab execute endpoint; DML gated by the databaseallow_dmlpolicy — enabled on the test stand per owner authorization, database 1). Precondition hash = SHA-256 of the canonical pre-mutation SELECT; mismatch aborts before the UPDATE (not_started); cleanup_policy=restore re-applies pre-mutation values through the same provider path; reconciliation = post-mutation SELECT comparison. Documented as@RATIONALEincontracts/modules.mdBrowserProvider. UI-DOM editing rejected (stock Superset tables are read-only; no editable plugin exists). - Implemented the mutation transport:
providers/browser_mutation.py— mutation contract gate (moved from browser.py for INV_7) +validate_mutation_inputs(table/key_columns/assignments/database_id; fields allowlisted; typed values) +build_mutation_script(page-evaluate JS with identifier allowlisting, typed literals, target-key-bound WHERE, crypto SHA-256 precondition check).providers/browser_transport.py— transport seam + real Playwright transport with the async mutation branch (pre-SELECT → hash compare → UPDATE → post-SELECT → evidence screenshot).providers/browser.py— 390 lines (INV_7); provider merges the contract into action_input, mapsBrowserTransportPreconditionMismatchto not_started.- Found & fixed live: SQL Lab
client_idmust be <= 11 chars (query.client_idVARCHAR(11)) — psycopg2 StringDataRightTruncation surfaced through the canary.
- Live mutation canary GREEN on
https://ss-prod.bebesh.ru(test stand, owner-authorized): full provider path per mutation — capacity claim → receipt opened → isolated Chromium context → authenticated session → pre-SELECT hash match → UPDATE → post-SELECT → evidence screenshot → receipt completed → lease released.row_edit: mutate 82.74→82.75 (verified), restore →82.74 (verified, original preserved). 2/2 receiptscompleted. - Evidence:
evidence/browser-provider/readonly-canary-20260901T151437Z.json+ mutate/restore PNGs. Regression:399 passed; ruff, anchors, diff-check clean; all provider modules < 400 lines. - Remaining for GO: browser reconciler (target-state check) wiring into the sweep, safe-checkpoint reconstruction trace, scheduler soak, browser E2E, sandbox runtime. Product remains NO-GO pending those; the mutation execution path itself is now production-proven.
Checkpoint — 2026-09-01 (browser reconciler + live reconciliation canary GREEN)
- Closed the mutation loop end-to-end:
- Receipts now persist the full
mutation_context(dashboard_id, table, key_columns, assignments, target_keys, field_allowlist, database_id) in summary at open time. browser_mutation.pyaddsbuild_select_only_script— the read-only observation flow (SELECT limited to assignment+key columns;SELECT *was rejected after the live run showed cross-column divergence).browser_reconciler.py—build_browser_reconciler(service): observes the live row state through a fresh authenticated Chromium session (SELECT-only; never UPDATE) and maps it to verdicts: completed (state matches assignments) / failed (diverged or context missing).- Bootstrap composition registers the real reconciler for
browser, so the 30s scheduler sweep resolves stalereconciliation_requiredreceipts automatically.
- Receipts now persist the full
- Live reconciliation canary GREEN: an injected
reconciliation_requiredreceipt (real mutation context, backdated) was resolved by the sweep through a live Chromium observation — receiptcompleted, note "target row state matches the recorded assignments". Evidence:evidence/browser-provider/readonly-canary-20260901T152804Z.json. - Verification:
403 passed(4 reconciler unit tests added); ruff, anchors, diff-check clean. - Remaining for GO: safe-checkpoint reconstruction trace, scheduler soak, browser E2E, sandbox runtime. The mutation lifecycle (contract → receipt → effect → evidence → reconciliation) is now fully production-proven on the test stand.
Checkpoint — 2026-09-01 (T042b safe-checkpoint reconstruction canary GREEN)
- Implemented the
reconstruction_replaymarker (044 data-model contract, previously missing):_archive_recovery_attemptrecordsreconstruction_replay: truein the recovery_history entry, and_advance_runmerges the marker into the replayed attempt's outcome when the step's recovery history shows reconstruction. - Safe-checkpoint reconstruction canary GREEN on the live stand
(
SS_CANARY_MODE=recovery, PostgreSQL canary DB): abandoned worker (claim → expired lease) →recover_runfrom the pinned plan → replay through the real provider against the stand → attempt 2passed,reconstruction_replay=truein both the recovery history and the replayed outcome,dashboard_opencheckpoint, evidence digest + PNG retained (20260901T154336Z-recovery-replay.png). Recovery canary harness seeds a real registry entry/revision and registers live providers into the global composition root (PostgreSQL canary DB removes SQLite cross-connection locking). - T042b canary scoreboard (all against the live test stand):
read-only actions
GREEN· forced timeout/cleanupGREEN· mutation row_edit+restoreGREEN(with precondition hash + receipts) · reconciliation observationGREEN· safe-checkpoint reconstructionGREEN. - Verification:
403 passed; ruff, anchors, diff-check clean. - Remaining for GO: scheduler soak, browser E2E, sandbox runtime (start_exploration real provider). Product remains NO-GO pending those.
Checkpoint — 2026-09-01 (T042b scheduler soak GREEN on live stand)
- Scheduler soak GREEN (
SS_CANARY_MODE=soak, 75s, ~15 real APScheduler ticks against the live test stand): 3 queued browser runs dispatched by the realscenario_queued_dispatchjob through the production provider path — 3/3passed,attempts == 1each (persisted queued→running CAS prevents duplicate provider I/O across overlapping ticks),active_leases == 0, exactly one durable evidence artifact per run, reconciliation job co-scheduled,scheduler.stop()graceful. - Canary harness: idempotent registry seeding helper shared by recovery/soak modes;
soak asserts run statuses, attempt counts, lease hygiene and artifact counts in
failures. - Evidence:
evidence/browser-provider/readonly-canary-20260901T160242Z.json. Verification:403 passed; ruff, anchors, diff-check clean. - T042b canary scoreboard: read-only
GREEN· timeout/cleanupGREEN· mutation row_edit+restoreGREEN· reconciliationGREEN· safe-checkpoint reconstructionGREEN· scheduler soakGREEN. All six canary vectors pass against the live stand. - Remaining for GO: browser E2E (full UI journey), sandbox runtime (
start_explorationreal provider), readiness/health payload hardening. Product remains NO-GO.
Checkpoint — 2026-09-01 (readiness hardening + T042c revalidation + sandbox runtime)
- Readiness hardening (044 T031/T040 tail):
/api/readynow carries the cached provider readiness snapshot (providerssection captured at bootstrap; no per-request probe, no credential-shaped values). - T042c/T024a dispatcher revalidation tests (
test_dispatcher_revalidation.py, 4): PROD reclassification invalidates the mutating gate before provider I/O; principal drift → dispatcherBROWSER_BINDING_MISMATCH; unknown binding ref → fail-closed mismatch; stale registry fingerprint →BROWSER_REGISTRY_MISMATCHbefore I/O. - Sandbox runtime (050 T026-T028 core):
exploration_sandbox.py— CAS claim (queued→exploring), read-only observation through the authenticated session (dashboard charts/filters via GETs only, bounded 12/8), server-side graph synthesis (charts → nodes, native filters →applies_toedges), durable evidence screenshot, transitions exploring→exploration_passed/failed.exploration_resultJSON column (migration0016_exploration_result— fixed a 33-char revision id that broke the varchar(32) alembic limit; PG chain 0001→0016 verified,alembic checkclean). Scheduler jobauthoring_exploration_dispatch(5s) sweeps queued requests; absent explorer registration keeps requests queued (sandbox_unavailable semantics). - Tests: +3 sandbox, +4 revalidation. Verification:
410 passed; ruff, anchors, diff-check clean. Remaining for GO: browser E2E + MCP parity domains + doc sync.
Checkpoint — 2026-09-01 (spec checkbox sync)
- Synced stale task checkboxes against implementation evidence: 044 T028–T042c (provider vertical, canaries, revalidation) and 050 T001–T033 (MCP interface, authoring chain, handoff/demolition flags) now reflect done status. Remaining open: 050 T008 CI walkthrough, T012/T013/T014 parity domains, T016 parity tests, T023 promotion E2E, T028 promotion E2E, T040–T043 demolition tails, 044 documentation tails.
- Session totals (2026-09-01): provider vertical (~1 975 LoC), capacity-aware dispatch,
mutation SQL surface, reconciler + sweep, readiness hardening, T042c revalidation,
sandbox runtime, 6/6 live canaries GREEN, migrations 0014–0016 PG-verified,
410 passedfull regression. - Remaining for GO: MCP parity domains (git/deploy/baseline/superset, ~400 LoC), browser/MCP E2E (~600 LoC), promotion E2E wiring, doc tails. Product status: NO-GO.
Checkpoint — 2026-09-01 (050 T012/T013/T014 parity domains)
- Closed the three open parity domains of the 050 catalog (45 tools now registered
— one-to-one with the explicit catalog):
backend/src/mcp_server/ops_inputs.py— bounded strict input models for the git/migration/backup/llm and Superset domains;extra=forbideverywhere; the continuation payload re-validates identically under the approved hash.backend/src/mcp_server/ops_tools.py—register_ops_tools: T012 gated bodies (curated schemas, approval-fallback only), Superset read tools (superset_list_databases,superset_explore_databaseschemas/tables/ table_metadata/select_star,superset_format_sql,superset_audit_permissions), gated Superset mutations (superset_create_dashboard,superset_copy_dashboard,superset_create_dataset), and the baseline 037-parity surface (capture_baseline_candidate,request_baseline_approval,decide_baseline_approval,consume_baseline_approval,create_verification_run) wrapping the exact REST-surface services.backend/src/services/mcp_ops_dispatch.py— reviewed dispatch adapters for every approved continuation: GitService create/commit on the app loop,git-integration/superset-migration/superset-backup/llm_documentationTaskManager enqueue,ValidationTaskServicepolicy creation, SupersetClient dashboard/dataset writes, and one-shotconsume_approvalfor baseline publish. Payload re-validation before any I/O; idempotency keys per invocation.poll_approved_mcp_dispatcheschain extended explicitly (no dynamic lookup);server.pycatalog +13 entries;RbacFastMCP.call_toolgains the SQL context gate:superset_execute_sqlagainst a PROD-classified environment returnsapproval_required(MCPX-FR-018 style), non-PROD executes directly under the dedicatedplugin:superset_sqlpermission.rbac_permission_catalog.py:SUPERSET_SQL_PERMISSIONS+ default-deny role mapping +is_superset_sql_permission, unioned intodiscover_declared_permissionsso sync seeds the row for admin assignment.
- SQL risk class:
superset_execute_sqlis denied to service principals, requires the dedicated permission (not implied byplugin:superset_proxy), bounded projection (50 rows), client-side dangerous-SQL guard intact. - Tests:
backend/tests/test_mcp_ops_parity.py(13) — catalog risk/permission matrix, 45/45 registration with curated schemas, SQL dedicated-permission visibility, service-principal denial, create_branch gate-only idempotent flow, execute_sql PROD gate / DEV bounded execution / unknown-env fail-closed, baseline principal denial, adapter payload validation + one-shot consume, full poller route for an approvedconsume_baseline_approvalinvocation. Existingtest_mcp_server.pycatalog expectation updated. - Verification: MCP slice 70 passed (
test_mcp_ops_parity+ server + approvals + maintenance + oauth); RBAC catalog tests 41 passed; full backend suite 11167 passed, 240 skipped, 1 xpassed; Ruff andcompileallclean; anchors balanced in all touched files (ops_inputs 6/6, ops_tools 7/7, mcp_ops_dispatch 6/6, parity tests 7/7); scopedgit diff --checkclean. - Remaining for GO: T016 parity fixture tests (MCP outputs vs legacy wrappers on shared fixtures), T023/T028 vertical E2E over MCP, browser UI E2E, scheduler soak for the new ops dispatchers, semantic index rebuild (Axiom not attached this session). Product remains NO-GO.
Checkpoint — 2026-09-01 (050 T016 parity tests + T023 MCP vertical E2E)
- Closed the parity-proof and the first vertical E2E of the 050 gate:
backend/tests/api/test_mcp_parity_baseline_037.py(3) — one shared 037 fixture set (api conftest builders) drives BOTH surfaces with one persisted principal: MCPrequest_baseline_approvalgate validates as the exact RESTApprovalGateResponseshape with field-by-field equality (operation, risk, required_permission, status, reason_required, target_paths); decide outputs agree includingactor_id;create_verification_runMCP vs REST/verification-runsagree on overall_status/category statuses/evidence/ created_by;superset_format_sqlMCP output is byte-identical to the legacy REST format path (only the [EXT:Superset] client is doubled). Discovered domain semantics pinned by the tests: one pending gate per AgentRun.backend/tests/test_mcp_scenario_e2e.py(2) — SC-001 vertical walkthrough overtools/callwith a REALscenario:EDIT/RUNprincipal (no RBAC stubs): tools/list visibility → authoring session → server-derived graph proposal → owner diff review → promotion →request_savecreating one immutablecandidateScenarioRevisionin the registry → 038 compile/resolve/validate over MCP → separate-CASactivate_revisionadvancingcurrent_revision_id; a no-EDIT principal neither sees nor can call save/activate (typed denial proven).- Fixed a production defect surfaced by the parity/E2E effort: the 038
resolver selector path crashed on steps with
description=None(TypeError: NoneType + strin_apply_selector); the hint is now safely recorded, regression pinned bytest_selector_hint_on_step_without_description. - Baseline MCP tool outputs normalized to non-colliding envelopes:
request_baseline_approval -> {"status","gate"},decide_baseline_approval -> {"status","decision"}.
- Verification: parity slice
3 passed; E2E2 passed; resolver suite18 passed; combined MCP+037 slice217 passed(in canonical ordering); full backend suite 11173 passed, 240 skipped, 1 xpassed (was 11167; +6 new tests net); Ruff andcompileallclean; anchors balanced (E2E 3/3, parity 5/5). - Known collection-order artifact (pre-existing, reproducible with untouched
files only): mixing root-level and
tests/apipackage paths in a single ad-hoc argument order can droptests/api/conftest.pyfixtures for the second-collected api module (pytest prepend import mode). The canonical full-suite ordering is unaffected; not chased further. - Closed in this session: T028 authoring promotion E2E, ops-dispatcher soak. Remaining 050 items (CORRECTED the same day: the earlier "frontend surface remaining" wording confused 044 T030 numbering with 050 — 050 Phase 3 frontend T030–T033 is already [x] in tasks.md): only Phase 4 removal (T040–T043) and semantic index rebuild (Axiom binaries absent here). Parity + E2E gate criteria on the backend: met.
Checkpoint — 2026-09-02 (050 T028 authoring promotion E2E + exploration wiring)
- Completed the last backend E2E of the 050 gate and wired the exploration
vertical end-to-end:
backend/tests/test_mcp_authoring_promotion_e2e.py(8) — full sandbox-to- revision chain overtools/callwith real RBAC: queued exploration → realexecute_scheduled_exploration_dispatchfrom scheduler thread context (via worker thread;run_until_completeis not callable inside a live loop) →exploration_passedwith observations/proposed_graph/draft:exploration-*evidence while only the [EXT:Browser] boundary is doubled; boundedget_exploration_resultprojection (no observation leak); typed proposal derived from the observation (observed_dashboard_titlelands in the saved graph); promote/diff/save/activate terminates in acurrentrevision without a pre-activation save. Negative branches: SQL/path/backslash/drop typed ops rejected before proposal/CAS advance; code-token and unregistered-action exploration specs rejected without persistence; all five authoring input models structurally reject callerdigest/content_hash.- Production wiring closed (was the T026-T028 seam gap): MCP
start_explorationnow reflects deployment reality (provider_available = get_registered_runner() is not None);bootstrap_live_execution_compositionregistersdefault_runnerplus the deployment-ownedScreenshotService/draft-storage context for every enabled live binding via the new reviewedregister_exploration_context(service, storage)seam. Without an enabled binding the behavior stays typed sandbox-unavailable; requests continue to queue instead of executing. - Ops-dispatcher soak (
tests/test_mcp_ops_parity.py, now 14): two approved task-backed ops (backup, migration) over three poller cycles — each dispatched exactly once, second/third cycles are no-ops, both carry invocation-scoped idempotency keys, dispatch_statuscompleted.
- Verification: T028 module
8 passed; parity+soak14 passed; MCP vertical slice 83 passed (8 files); registry/sandbox/workspace/scenario slices green in canonical ordering; full backend suite 11182 passed, 240 skipped, 1 xpassed (was 11173); Ruff andcompileallclean; anchors balanced in all touched files (exploration_sandbox 5/5, live_composition 4/4, server 32/32, T028 E2E 6/6, parity 7/7). - Remaining 050 items (CORRECTED: Phase 3 frontend T030–T033 already [x]; the earlier list conflated 044 T030 numbering): Phase 4 removal (T040–T043) and semantic index rebuild (no Axiom binaries in this environment). Product status: parity + E2E criteria met on the backend; overall GO awaits the Phase 4 demolition tail only.
Checkpoint — 2026-09-02 (orthogonal audit + audit-driven hardening)
- Two independent read-only audits (orthogonal QA + security) over the whole uncommitted 050 slice: QA re-ran all evidence (full suite reproduced exactly: 11182 pre-hardening), spot-verified 7 task claims (all CONFIRMED), audited mocks in the four new test files (no logic mirrors; external-boundary doubles only); security produced 17 CWE-mapped findings. No CRITICAL; no approval bypass; no service-principal leakage; dispatch stays CAS-claimed and static.
- Fixed in this checkpoint (convergent HIGH findings):
backend/src/core/superset_client/safety.py— the SQL risk-class guard hardened from keyword blocklist to coverage ofINTO/CALL/SET/REFRESH/ VACUUM/REINDEX/ATTACH/LOAD/PREPARE/COMMENT, server-side primitives (lo_import,lo_export,pg_read_file,pg_write_file,pg_ls_dir,dblink,dblink_exec,set_config,pg_sleep) and multi-statement smuggling (;after string/comment stripping). Tests intests/test_core/test_superset_safety.pypin bypasses and false-positive guards (order_set/dataset/setval/single trailing;stay green).backend/src/mcp_server/server.py— PROD-classified SQL execution is now a TERMINAL rejection (production_sql_execution_rejected, recorded as denied provenance, no gate created), using the canonical production criterionresolve_environment_execution_policy(is_production OR stage=PROD) — this removes the approvable-but-never-dispatched black hole and aligns the criterion with scenario starts.backend/src/services/mcp_ops_dispatch.py— baseline consume adapter now refuses deactivated actors (mcp_actor_inactive) matching the direct-tool guard.
- Evidence after hardening: affected slice 214 passed (MCP vertical + both safety suites + superset extended + resolver); full backend suite 11202 passed, 240 skipped, 1 xpassed; Ruff/compileall clean.
- Closed in the follow-up hardening round (2026-09-02, same day):
- QA M2 / T017 gap —
response_limitis now ENFORCED:RbacFastMCP.call_toolfails closed with the typedresponse_too_largeenvelope and denied provenance when a structured reply exceeds the configured limit (server.config), pinned bytest_oversized_reply_is_truncated_to_typed_error. - QA M3 — poisoned exploration requests fail closed:
default_runnerterminates a claimed request whose workspace vanished (EXPLORATION_TARGET_UNRESOLVED/workspace_not_found) instead of raising back into the queue; pinned bytest_default_runner_fails_closed_when_workspace_missing. - SEC M-10 — exploration evidence digest is now the sha256 content hash of the
observation fingerprint; the ref is retrievable and verifiable via
DraftStorage.retrieve, pinned bytest_evidence_digest_is_sha256_of_observationand an E2E retrieval assertion. - SEC M-16 — TaskManager adapters now enqueue under the actor's durable user
UUID (
_actor_user_id, inactive actors rejected before dispatch), so MCP-originated tasks are visible to their owner viaget_task_status; pinned in the migration adapter test.
- QA M2 / T017 gap —
- Evidence after hardening: audit slice 161 passed; full backend suite 11205 passed, 240 skipped, 1 xpassed; Ruff/compileall clean; anchors balanced (mcp_ops_dispatch 6/6, ops_tools 7/7, parity tests 8/8).
- Remaining deferred (product/ops decisions recorded, none block the
parity+E2E gate): M-03 self-approval (the gate is a confirmation step;
four-eyes needs a product decision on approver separation); M-04 service
principals bypass catalog permissions by design of the service branch
(superset_audit and table_metadata are the widest service-visible
projections); M-06 at-least-once adapter idempotency after lease recovery;
M-07 multi-binding exploration context last-writer-wins (single-binding
deployments unaffected); M3/M-08 exploration queue TTL/limits; M-05 remnants
(expired-gate replay
_find_retryable_approval, decide-expiry TOCTOU). - Product status unchanged: parity + E2E criteria met on the backend; overall GO awaits the Phase 4 demolition tail (T040–T043) — the frontend Phase 3 (T030–T033) is verified complete in tasks.md.
Checkpoint — 2026-09-02 (frontend status correction + live happy-path demo)
- Frontend 050 Phase 3 is COMPLETE (T030–T033
[x]in tasks.md): flag-driven decommission via the typed build-time switchMCP_DECOMMISSION(frontend/src/lib/config/mcp.ts, vite define from theMCP_DECOMMISSIONenv var),HandoffSurfacecomponent (frontend/src/lib/components/agent/HandoffSurface.svelte), and flag gates for/agent, the assistant panel and the top-nav entry (routes/+layout.svelte:134,routes/agent/+page.svelte:129/169). - Verified live on the local stand:
./run.sh --skip-install(backend 8000/api/readygreen, frontend 5173, agent 7860); demo adminhappy-admincreated throughsrc/scripts/create_admin.py; Playwright headless-chromium walkthrough performed a real login and captured 8 screenshots (/tmp/kilo/happy-path/00-login.png…06-migration.png,07-handoff-agent.png): login → home →/dashboards→/agent(legacy workspace) →/dashboard-testing/runs(Global Run Operations Center) →/dashboard-testing/analytics(investigation queue) →/validation-tasks→/migration; a second dev server withMCP_DECOMMISSION=trueon :5174 renders the handoff surface ("Ассистент переведён на MCP" + connection hint + copyable prompt). - Note: scenario detail routes are
/dashboard-testing/scenarios/[id]/{edit,runs,analytics}; there is no scenario index page — flow entry is the run center and dashboard cards.
Checkpoint — 2026-09-02 (050 Phase 4 removal — the 050 task list is fully closed)
- T040 — runtime decommission:
run.shnow starts exactly backend+frontend (no AGENT_PORT/watchfiles/chat_uploads),docker-compose.ymlanddocker-compose.enterprise-clean.ymllost theagentservice (nginx configs dropped/api/agent/gradio),build.shlostbuild:agent/bundle:agent/bundle:embeddingsand the agent block of the generated deploy compose/manifest/env template;AGENTS.md/INSTALL.mdrewritten for the two-service world. Live proof: stand boots with only :8000+:5173,/api/readygreen, Playwright reaches the handoff surface. - T041 — chat code removed: the entire
agent/tree,docker/Dockerfile.agent,docker/agent.entrypoint.sh, the gradio-proxy backend test, and on the frontend: AssistantChatPanel + chat widgets, the ScenarioWorkspace branch (components/agent/dashboard-testing), AgentChat/AgentRun models, the assistant store, theMCP_DECOMMISSIONflag infra (vite define, config module, d.ts) and the gradio dev proxy./agentnow renders HandoffSurface unconditionally; TopNavbar button navigates to it. Frontend: 3454 vitest passed, lint 0 errors, build OK. Backend: 11199 passed, 240 skipped, 1 xpassed. MarkdownRenderer retained (used by GitWorkspacePanel); assistant i18n keys retained because the handoff surface renders them; orphan chat DTO modulefrontend/src/types/agent.tsremoved after confirming zero importers (build+tests re-verified). - T042 — all twelve 036–047
## Drift Amendment — MCP Interfacesections now carry**Status (2026-09-02): done**with evidence links (050 tasks T012–T028, T030–T033, T040–T041 + workstate checkpoints). - T043 — full-suite evidence recorded above; the stand runs without 7860 (SC-005 satisfied).
- 050 status: all tasks [x]. Product readiness: backend parity + E2E and the removal tail are closed. Residual (documented, product decisions): approval self-approval semantics, service-branch catalog permissions, adapter idempotency after lease recovery, multi-binding exploration context, TTL/expiry polish, semantic index rebuild (no Axiom binaries in this env).
Checkpoint — 2026-09-02 (browser E2E on isolated stack + live admin walkthrough)
- Browser E2E closed the last evidence gap of T043: isolated compose stack
docker compose -p ss-tools-e2e --env-file .env.e2e -f docker-compose.e2e.yml(fresh db volume, initial-admin bootstrap admin/admin123, backend :8103 + frontend :8102 healthy, no 7860 anywhere) + local Playwright run oflogin.e2e.js+ rewrittenagent.e2e.js→ 6/6 passed, twice (Chromium). Stack torn down withdown -vafterwards. - FOOTGUN found and worked around: the documented E2E command without
-pjoins the DEFAULT compose projectss-tools— the same projectrun.shuses fordocker compose up -d db— so the E2E backend attached to the DEV database (admin had the real password, runner global-setup failed with 401). Isolation requires an explicit project name (-p ss-tools-e2e); dev db volumepostgres_datasurvived and was re-verified after the incident. - E2E suite alignment with the decommission (part of T041's frontend tail):
agent.e2e.jsrewritten from chat-surface assertions to the handoff contract (T01 surface renders / T02 zero chat elements / T03 deep-link stays on handoff);dashboard-scenario-ui.e2e.jsreload-recovery test retargeted to the handoff route;agent-scenario-run.e2e.jsanchor updated (backend /api/agent-runs lifecycle tests remain valid — AgentRun persistence is kept); stalelocator('nav')strict-mode violations fixed in login/smoke (.first()), invalid-credentials regex now matches the backend passthrough detail ("Incorrect username or password"). - Live admin walkthrough on the native dev stand (real admin credentials provided by the operator): login → /dashboards, /agent renders the handoff (0 textareas, 0 conversation nodes), /dashboard-testing/runs renders; the agentless run.sh restart and split-process start (uvicorn + vite separately) both verified — screenshots /tmp/kilo/happy-path/09–12.
- Note: the first post-restart Playwright run of agent T02/T03 flaked against a cold vite dev server; warm re-runs are stable (3.3s and 4.8s passes).
Closure Gate Review — 2026-09-03 (orthogonal, categorical; pinned to 731aaaa8, re-verified at ddbfe00f)
Two independent read-only reviewers (spec conformance + protocol/logging/MODEL-FIRST), plus mechanical sweeps. Verdicts by category:
- Spec conformance 050: FAIL. Blockers: FR-019 NONCONFORMS —
list_checkpoints/decide_checkpointdo not exist anywhere (T025[x]is a false claim); FR-013 —/oauth/authorizerequires a Bearer header (api/mcp_oauth.py:134-140), no cookie session/consent page,client_credentialsgrant missing → browser leg uncompletable by a standard client; SC-007 FAIL — zero HTTP-level proofs of /oauth/* (only service-level), T008 "scripted client" claim unsupported; FR-005 PARTIAL —list_pending_mcp_approvalslists only own MCP gates (automation/web gates invisible → Story 3 AC2 violated), mandatory high-risk confirm reason NOT enforced, envelope lacks targets/expires_at/reason_required; SC-005 PARTIAL —/api/assistant- agent_* routers still mounted (app.py:595-602), SystemSettings assistant-retention
UI live,
frontend/src/lib/api/assistant.tsshipped,.env.exampleretains AGENT/GRADIO 7860,/api/agent/llm-configexemption remains → T032[x]false at HEAD; FR-010 absent (no catalog version/deprecation mechanism); FR-015 PARTIAL (no Admin DCR surface); FR-007/007a PARTIAL (oversize → rejection not ref+digest; no JSON-depth limit; no per-session rate limit / E6 Retry-After); SC-002 PARTIAL (output-parity fixtures exist only for the 037 domain; legacy wrappers deleted → remaining ~34 tools structurally tested only); SC-004 PARTIAL (no admin-fixture list-filter test); SC-009 unpinned (no mid-flow role-change test). Process: spec.md release-gate rows E2E-AUTH-001..003 still[ ] OPENwhile dependent tasks are[x]— by the spec's own rule the gate is NO-GO; Status still "Ready for Implementation".
- agent_* routers still mounted (app.py:595-602), SystemSettings assistant-retention
UI live,
- Happy path A (external MCP client): FAIL — authorize defect blocks the browser
leg; no over-HTTP
tools/callproof (all tool-level evidence in-process; HTTP only for initialize/tools/list). Discovery→token lifecycle green at service level (68/68 independent rerun). - Happy path B (UI): PASS (conditional) — login→dashboards→nav→/agent handoff proven live + Playwright 6/6; conditional: HandoffSurface does not embed the dashboard context params into the copyable prompt (Story 5 AC2 partial).
- GRACE-Poly INV: CONDITIONAL. Clean: INV_3 scripted over all 116 touched files
(zero mismatches; TopNavbar 2/1 flag refuted —
#regioninside @RATIONALE prose); INV_9/TYPE canonical; no legacy markers; secrets clean. Blocker: INV_6 — dangling@RELATION CALLED_BY -> [agent/app.py]survives at HEAD inshared/src/ss_tools/shared/_llm_health.py:7(round-3ddbfe00fredirected 42 dead targets but missed this one); zero tombstones for the 571 contracts deleted by 731aaaa8; dangling VERIFIES/BINDS_TO edges in specs 033/035/036/039 manifests. MAJOR INV_4: duplicated @SIDE_EFFECT/env-mutation block in test_migration_routes.py:20-26 (edited in-commit — INV_9 anti-pattern). MINOR: INV_1 naked migrations 0014-0016; exploration_sandbox module region closes early (Registry block outside); INV_7 browser.py=402, server.py grew 1135→1522 (worst; needs DECOMPOSITION GATE plan). - molecular-cot-logging: CONDITIONAL. MAJOR systemic: two-positional logger
misuse (
logger.X(_SRC, "intent", ...)) drops the real intent — wire emits the contract ID asintent(dynamically proven); ~120 call sites incl. ops_tools (8), mcp_ops_dispatch (4), exploration_sandbox (4), server authoring tools (18+). MAJOR: poll-dispatch adapter failure path is trace-invisible (except-branch in mcp_approvals has no EXPLORE); exploration sandbox fail-closed paths_finish()silently (DB only); 7/9 dispatch adapters have zero REASON/REFLECT. Clean: every existing EXPLORE carries error= (23/23), no secrets in payloads, no legacy markers. - MODEL-FIRST frontend: CONDITIONAL. PASS: HandoffSurface architecture (correct
inline-$state judgment per skill table; $lib/ui atoms; semantic tokens verified
against tailwind.config; i18n keys ru+en; runes-only; guarded by contract test);
+page.svelte thin wrapper; TopNavbar relations accurate; deletion sweep found zero
broken imports. MAJOR: duplicated
BINDS_TO [EXT:frontend:taskDrawerStore]edge in TaskDrawer.svelte:11-12 (introduced by the decommission edit). MINOR: vestigialassistantOffset=$derived("0px"); hardcoded English fallbacks in HandoffSurface (strict no-hardcoded-strings rule); silent clipboard fallback without EXPLORE; orphanedassistant.tskept alive only by its test.
OVERALL CLOSURE GATE: NO-GO as-committed. Backend core (hidden-vs-gated RBAC,
CAS gates + fenced lease dispatch, provenance, bounded inputs, parity domains,
authoring/promotion chain) is solid and independently re-verified (68/68, 74/74).
The gate is blocked by: one unimplemented FR (FR-019), the defective OAuth browser
leg (FR-013/SC-007), three false [x] claims (T025/T008/T032), spec release-gate
rows OPEN, and the protocol debt above (MCL intent drop, INV_6 blocker).
Remediation queue: P0 — downgrade T025/T008/T032; implement checkpoint tools; fix authorize flow (session/consent or documented manual-token path + client_credentials) with one HTTP-level SC-007 test; reconcile spec release-gate rows. P1 — MCL intent repair + wire-format test; EXPLORE/REASON on dispatch-failure and sandbox fail-closed paths; INV_6 edge + tombstone sweep; TaskDrawer/test_migration_routes dedupes. P2 — SC-005 remnants; FR-010 versioning; depth/rate limits; handoff context prompt; migration anchors; server.py decomposition plan.
Checkpoint — 2026-09-03 (closure-gate remediation round 1 / P0 — intermediate save)
- FR-019 closed for real: MCP
list_checkpoints+decide_checkpoint(044 CAS semantics: server-resolved pending checkpoint,decision_versionCAS,continue_after_human_decision, typed conflict/not_found envelopes); catalog("scenario","RUN")+service_allowed=False— automation has no path;tests/test_mcp_checkpoints.py→ 3 passed. T025 re-earned its [x]. - FR-013/SC-007 closed at wire level:
client_credentialsgrant (confidential DCR clients, one-timeclient_secret, sha256-only persistence via migration0017_oauth_client_secret, idempotent guard per house style), signed service-principal tokens (principal_type=service, aud=mcp, NO refresh),McpTokenVerifierservice short-circuit, AS metadata extended, INSTALL.md §"MCP клиент" documented.tests/test_mcp_client_flow_http.py→ 2 passed: FULL scripted-client flows over real HTTP (machine: discovery → DCR → client_credentials → /mcp initialize → tools/list gated-hidden → tools/call; user: DCR → PKCE S256 authorize with Bearer web session (documented SPA-mediated consent) → code exchange → identity-only token → /mcp live-RBAC listing). Browser-less-cookie consent page remains an open product decision (recorded, non-blocking for machine path). - PRODUCTION DEFECTS found while closing E2E-AUTH-003 and fixed: both
argument-inspection gates in
RbacFastMCP.call_toolunderstood only the FLAT argument shape while FastMCP delivers{"request": {...}}—start_scenario_runwas UNCALLABLE over MCP (scenario_start_policy_unavailabledenied every wrapped call) and the PROD-SQL terminal denial was BYPASSABLE by wrapped shape. Fixed via_gate_argumentsunwrap (both shapes resolve identically); regression pinned by flat×wrapped parametrization of the PROD matrix (test_execute_sql_is_terminally_rejected_in_production, 6 combos). - E2E-AUTH release gates CLOSED with evidence (spec.md rows): 001 — T028 chain
gained
propose_test_plan(single trace: session→plan→exploration→result); 002 — compile/validate/diff already proven in T023; 003 — T023 extended with post-activationstart_scenario_runon the canonical runner-shaped fixture graph, asserting the queued run pinsscenario_revision_id+scenario_content_hashof the promoted revision. - Record honesty: tasks.md downgraded false/unproven [x] (T025/T008/T008b/T032
- T005a annotation) at review time; T025 and T008 since re-closed with real executable evidence; T032/T008b/T005a remain OPEN (queued below).
- Verification at save point: targeted slices green (checkpoints 3; client-flow HTTP 2; oauth 10; scenario E2E 2; promotion E2E 8; ops parity 18 incl. wrapped PROD matrix; mcp server/approvals/maintenance green in combined run 68+).
- REMAINING remediation queue (round 2+):
- P1 MCL: two-positional logger intent-drop repair across ~120 call sites + wire-format unit test; EXPLORE on poll-dispatch failure path and the four silent exploration fail-closed branches; REASON/REFLECT for the 7 silent dispatch adapters.
- P1 GRACE: INV_6 —
_llm_health.pydangling[agent/app.py]edge + tombstone/dead-edge sweep of specs 033/035/036/039 manifests; TaskDrawer duplicated BINDS_TO; vestigialassistantOffset; test_migration_routes duplicated @SIDE_EFFECT block; anchors for migrations 0014–0016; exploration_sandbox module-region span; server.py decomposition-gate plan. - P2: SC-005 remnants (
/api/assistantrouter unmount, SystemSettings retention UI,assistant.ts,.env.example7860 vars,/api/agent/llm-configexemption decision); FR-010 catalog versioning/deprecation markers; JSON-depth + per-session rate limits (E6 Retry-After); HandoffSurface context-parameterized prompt (Story 5 AC2); admin fixtures for SC-004 exact-catalog assertions; SC-009 mid-flow role-change test; browser cookie-consent decision.
Checkpoint — 2026-09-03 (closure-gate remediation round 2 / P1 — MCL + GRACE, full suites green)
P1 MCL (logging protocol)
- Intent-drop repaired (182 sites): every two-positional
logger.X(_SRC, "intent", ...)misuse converted to the facade conventionlogger.X("intent", src=_SRC, ...)via an AST transformer (180 sites across 19 files) plus manual repair of the twoclient_registry.py shutdownsites (real intents; the post-clear REASON became a REFLECT). The facade binds the FIRST positional asintent; the misuse silently dropped the sentence into*args. - Single convention repo-wide (211 sites): all direct SSOT
log()/cot_log()production call sites inbackend/src(58 files, incl. thelog as _cloglocals incandidate_guards.pyand a deadlogimport in_llm_async_http.py) migrated to the intent-first facade;cot_logger.log()remains the module-internal primitive (belief_scope/cot_span/task-event bridge/shared layer, which has its own identically shaped facadess_tools.shared.logger). - Facade
level=support added toreason/reflect/explore(parity with the primitive's DEBUG plumbing lines), emitted through the level-named methods (target.info/warning/...— NOTtarget.log(numeric)), preserving the established patch/capture contract. - EXPLORE gaps closed:
Services.McpApprovals.Pollexcept-branch now emits typedMCP_DISPATCH_<Exc>EXPLORE; exploration fail-closed branches (workspace vanished / target unresolved / sandbox unavailable / provider crash) are wire-visible through one choke-point EXPLORE inside_finish()— never DB-only again. - REASON/REFLECT closed for all 9 silent
mcp_ops_dispatchadapters (deploy/migration/backup/llm-doc/llm-validation/superset-create+copy+dataset/ baseline-consume) with bounded identity payloads (sql logged as length only). - Skill synced with the module (the skill's own "the module wins" rule): canonical
.agents/skills/molecular-cot-logging/SKILL.md§I/§II examples now come from real code (mcp_ops_dispatch, capacity.py), §III documents the single facade convention, the module-internal primitive, the forbidden two-positional shape, and thelevel=kwarg;./scripts/sync-skills.shrun —.kilocopy byte-identical. - Executable pins:
backend/tests/test_core/test_logger_wire_format.py(7) — REASON/ EXPLORE/REFLECT wire fields, dict-positional payload promotion, the intent-drop proof for the forbidden shape,level=override, and two repo-wide AST sweeps (two-positional misuse == 0; direct SSOTlogimports in backend production code == 0).
P1 GRACE (semantic protocol)
- INV_6: dead
CALLED_BY -> [agent/app.py]edge removed fromshared/_llm_health.py(rationale rewritten for the post-decommission single-container reality). Specs 033/035/036/039 dead-edge sweep (resolver over all 10433 live repo anchors): 0 dead edges remain — 5 wrong-id edges retargeted to live contracts ($lib/ui/Button→Ui.Button,$lib/ui/Icon→Ui.Icon,AgentChatSpec→Spec.AgentChatContext.FeatureSpec), 15 genuinely dead edges tombstoned in place (@DEPRECATED+ successor:Api.McpOAuth,McpServer,Services.McpApprovals,AgentChat.HandoffSurface; ParsePdf/ParseXlsx — no in-repo successor, recorded). - INV_9 dedupes: TaskDrawer duplicated
BINDS_TO [EXT:frontend:taskDrawerStore]removed; vestigialassistantOffset=$derived("0px")eliminated (style pinnedright: 0);test_migration_routes.py4×duplicated @SIDE_EFFECT block consolidated into one (and the duplicated DATABASE_URL statement removed). - INV_1: migrations
0014–0016anchored in the houseMigration.<Domain>pattern with verified liveDEPENDS_ONtargets (Models.ScenarioExecution.Capacity,Models.ScenarioExecution.ProviderOperation,Migration.AgentAuthoringWorkspace.ExplorationRequest);exploration_sandboxmodule region span fixed (Registry block was outside; module now closes at EOF). - server.py (worst INV_7 offender, 1571 LOC): missing MODULE region added (
McpServer, @defgroup + core invariants) and a DECOMPOSITION GATE invariant recorded in-contract (phases A–D with line counts: scenario_inputs ~245 / auth ~200 / rbac_server ~305 / ProbeTools split ~330+330; frozen contract IDs and import surface; per-phase gates). Binding plan artifact:specs/050-mcp-interface/plans/server-decomposition-gate.md(constraints, risk register, rollback rule). Duplicateenvironment_policyimport removed. Split execution is a follow-up, not done this round.
Full-suite defect root-caused and fixed (first full rerun since round 1)
test_mcp_client_flow_http(2 tests) failed ONLY in-suite:tests/test_dependencies_unit.pyassigns MagicMocks into six DI singleton globals (config_manager,plugin_loader,task_manager,scheduler_service,resource_service, ...) without restoration; later in the sessionget_session_idle_timeout_minutes()resolved the MagicMock settings chain WITHOUT raising (so the fallback never engaged) andidle_minutes > 0raisedTypeError: MagicMock > inton every sid-bearing auth request (dependencies.py _enforce_session_policy). Pre-existing since round 1 (round 1 never ran the full suite); unrelated to this round's migration — proven by the exact traceback frame.- Fix A (hygiene): module-scoped autouse fixture in
test_dependencies_unit.pyrestores all seven singleton globals after every test. - Fix B (production hardening):
get_session_idle_timeout_minutes()validatesint(excl.bool); a non-integer resolution now emits EXPLORESESSION_POLICY_CONFIG_INVALIDand falls back to the typedsession_policycache — a poisoned config object can no longer crash the auth path. - Repro pair pinned:
test_dependencies_unit.py + test_mcp_client_flow_http.py→ 62 passed (this ordering fails in-suite without the fixes).
Verification evidence (this round)
- Full backend suite: 11214 passed, 240 skipped, 1 xpassed, 0 failed (was 11205; +7 wire, +2 client-flow now green in-suite).
- Frontend: 3454 vitest passed (197 files); lint 0 errors / 340 warnings (baseline was 373).
- Scoped slices during the round: logger/runner/detail-routes/git 73; migrated-modules slice 1393; MCP slice 95 (incl. both HTTP client flows in canonical ordering); wire-format 7.
ruff check .clean on the whole backend (src+tests+alembic);compileallclean; anchor-stack balance swept over every file touched this round — PROBLEM: none (server.py 34/34, logger 13/13, scheduler 30/30, wire tests 9/9, TaskDrawer 5/5); scopedgit diff --checkclean; skills.agents→.kilobyte-identical.
Remaining (round 3+)
- Execute the server.py decomposition plan (phases A–D, gate-verified); INV_7 tail:
browser.py402 LOC. - P2 queue unchanged: SC-005 remnants (
/api/assistantunmount, SystemSettings retention UI,assistant.ts,.env.example7860 vars, llm-config exemption decision); FR-010 catalog versioning; JSON-depth + per-session rate limits (E6 Retry-After); HandoffSurface context-parameterized prompt (Story 5 AC2); SC-004 admin fixtures; SC-009 mid-flow role-change test; browser cookie-consent decision. - Deferred product decisions (recorded, non-blocking): M-03 self-approval, service-branch catalog permissions, adapter idempotency after lease recovery, multi-binding exploration context, TTL/expiry polish.
- Semantic index rebuild — Axiom binaries absent in this environment (zombie-mode sweeps used).
- Product status: NO-GO pending the closure-gate re-review of rounds 1+2.
Checkpoint — 2026-09-03 (closure-gate remediation round 3 — server.py decomposition EXECUTED)
Executed the binding plan (specs/050-mcp-interface/plans/server-decomposition-gate.md, now marked
EXECUTED with a full execution log) plus one addendum:
- Phase A:
mcp_server/scenario_inputs.py(268) — the 13 typed input models moved verbatim; frozen-surface re-export block in server.py. - Phase B:
mcp_server/auth.py(238) — Configuration/Principal/TokenVerifier/TransportAuth plus the SINGLE_access_token_contextdefinition site. First gate run 42 failed → correctly caught a frozen-surface breach (tests build fake tokens viamcp_server.AccessToken); import retained → green. - Phase C:
mcp_server/rbac_server.py(393) — GateArguments/ScenarioGate/Catalog/RbacServer moved verbatim. First run 1 failed → monkeypatch seamrequest_mcp_approvalrelocated to the owning module. - Phase D:
mcp_server/tools_authoring.py(367) +mcp_server/tools_scenario.py(373) — the ProbeTools closure split into register seams on theregister_ops_toolsprecedent; the seam call sequence preserves the original tools/list registration order byte-for-byte; ruff F821 caught the single lost import (timedelta). Tool-body seams (get_config_managerfor list_environments/ start_scenario_run,get_task_managermaintenance sentinel) relocated to tools_scenario. - server.py = 177 LOC (was 1571): @defgroup + frozen-surface re-exports + the composition seam + CreateApp only; its DECOMPOSITION GATE invariant updated to the executed state (re-inlining a moved contract or re-exceeding 400 LOC violates the gate).
- Addendum E (pre-existing INV_7 offender the closure audit missed):
ops_tools.py420 → 215; the 044 checkpoint + 037 baseline blocks moved verbatim intotools_review.py(253) behind order-preserving register seams; shared helpers (_SRC/_current_user/_resolve_environment) stay single-sited in ops_tools, seam import function-local → import graph acyclic.
Constraint compliance: contract IDs all frozen (verbatim moves of McpServer.* regions incl. the
9 nested authoring tool regions and the ScenarioTools/Checkpoint/Baseline blocks); import surface
frozen (server re-exports every moved public name; AccessToken retained); behavior diff zero.
Recorded refinement — a module-attribute monkeypatch seam is NOT healed by re-export: it follows the
owning module (4 test files updated: test_mcp_server, test_mcp_ops_parity, test_mcp_scenario_e2e;
all listed in the plan's execution log).
Evidence:
- Final full backend suite: 11214 passed, 240 skipped, 1 xpassed, 0 failed (7:02) — identical to the round-2 suite: zero behavior diff.
- Per-phase gates: anchors stack-balanced after every move; MCP slice 95 green at every phase end (A: 63→95; B: 95 after surface fix; C: 95 after seam relocation; D: 95; E: 95).
- Package INV_7 census:
__init__ 10 / server 177 / ops_inputs 190 / ops_tools 215 / auth 238 / tools_review 253 / scenario_inputs 268 / tools_authoring 367 / tools_scenario 373 / rbac_server 393— all < 400. - ruff clean (whole backend), compileall clean,
git diff --checkclean, plan doc + module invariant metadata updated (INV_9: tags carry current local facts only).
REMAINING (round 4+):
- INV_7 tail:
browser.py402 LOC (scenario provider). - P2 queue unchanged: SC-005 remnants (
/api/assistantunmount, SystemSettings retention UI,assistant.ts,.env.example7860 vars, llm-config exemption decision); FR-010 catalog versioning; JSON-depth + per-session rate limits (E6 Retry-After); HandoffSurface context-parameterized prompt (Story 5 AC2); SC-004 admin fixtures; SC-009 mid-flow role-change test; browser cookie-consent decision. - Deferred product decisions (M-03 self-approval, service-branch catalog permissions, adapter idempotency after lease recovery, multi-binding exploration context, TTL/expiry polish).
- Semantic index rebuild pending Axiom binary availability.
- Product status: NO-GO pending the closure-gate re-review of rounds 1+2+3.
Checkpoint — 2026-09-04 (closure-gate remediation round 4 — P2 queue: AC2, E6 limits, SC-005 closure, SC-004/SC-009, cookie-consent decision)
Story 5 AC2 — HandoffSurface context-parameterized prompt
HandoffSurface.svelteaccepts optional context props (objectType/objectId/objectName/envId/ route/intent); the copyable prompt embeds them (object_id=42,environment_id=...,intent=build_dashboard_test_scenario), generic prompt otherwise (offline invariant kept — derivation is synchronous over props+i18n)./agent/+page.svelteforwards URL query params; the existing DashboardHeader «Создать сценарий тестирования» link already carries the full context set — no dead link, prompt is parameterized end-to-end.- Metadata updated (SEMANTICS
static→context,@UX_STATE Ready/Context); i18n keyhandoff_context_labeladded ru/en. Tests: contract test updated (AC2 forwarding pins + unconditional-render invariant) + 2 new render tests → handoff slice 5 passed.
E6 / MCPX-FR-007a — JSON depth + per-session rate limits (guard, pre-dispatch)
McpServerConfigurationgained server-owned limits:json_depth_limit=32,session_request_limit=120,session_rate_window_seconds=60.McpTransportGuard(mcp_server/auth.py): POST bodies are depth-checked after buffering and BEFORE dispatch — typed400 {"error":"json_depth_exceeded","limit":32}, iterative walker (RecursionError from json.loads counts as exceeded; adversarial depth cannot crash the guard), unparseable bodies fall through to the MCP layer's own typed parse error. Per-session sliding window (_SessionRateLimiter, bounded 10k keys, rejected hits not recorded) keyed bymcp-session-idheader else authenticated principal → typed429 rate_limited+Retry-Afterheader (E6 contract). Rejection creates no mutable state.- Tests:
tests/test_mcp_transport_limits.py5 passed (HTTP depth 400 + no-state-after-reject- 429/Retry-After window + unit boundaries/limiter semantics); MCP slice 95 passed.
SC-005 remnants — CLOSED
/api/assistantrouter unmounted from app.py;routes/assistantpackage header records the unmount +@REPLACED_BY -> [McpServer]+ retention rationale (parity provenance for FR-005; its 11 unit-test files exercise wrappers directly — slice 475 passed, 7 skipped)./api/agent/llm-configdecision = REMOVED (in-place Tombstone region in agent_conversations.py): zero consumers (Gradio dead in T041; frontend 0 refs); secret-bearing endpoint (decrypted provider key) kept for nobody was rejected. With it removed:get_agent_service_user_strictDI deleted (zero inbound edges — INV_6-clean), app.py polling suppression entry dropped, test sections pruned (agent_conversations −7, conversation_api −2, app_middleware −1; stale @TEST_EDGE/@NOTE metadata updated).- Frontend:
assistant.ts+ its test deleted (the sole inbound edgeApi.ApiModule CALLED_BY -> [Api.Assistant.AssistantApi]removed first — INV_6); SystemSettings «Assistant history retention» UI block + validation keys removed; 8assistant_*i18n keys pruned ru/en (parity kept); ux-test fixture pruned. Zero residualsendAssistantMessage|api/assistantrefs. .env.example:AGENT_HOST_PORT=7860,GRADIO_SERVER_PORT=7860and the whole agent section removed (zero consumers repo-wide incl..env.current/.env.master/.env.e2e/run.sh/compose); stale «единый для backend и agent» comment fixed. docs/INSTALL/README carry no references to removed surfaces (verified).
SC-004 + SC-009 exact RBAC visibility (new tests/test_mcp_rbac_visibility.py, 3 passed)
- SC-004:
tools/listexact sets from the live catalog — admin = all 47 tools; analyst (scenario:RUN + dashboard:testing:WRITE + tasks:READ) = 21; viewer (zero grants) = 15 (the permission=None human surface). Exact-set equality, not superset. - SC-009: mid-flow revocation of scenario:RUN on the SAME identity-only token (no new consent)
hides decide_checkpoint/list_checkpoints/start_scenario_run in the next
tools/listAND denies the cached-catalog call by name (PermissionError permission_denied); mid-flow grant of scenario:RUN_PROD exposes the approvals surface in the next list. Plus catalog/registry consistency pin. - Cookie-consent (browser-without-Bearer) decision recorded in tasks.md T008: not built in 050 — SPA-mediated authorize is the tested SC-007 path; standalone consent HTML deferred to an architecture amendment (MCPX-FR-020 perimeter stands).
- INV_7 tail note:
browser.pymeasured 398 < 400 — the round-2 logger migration reduced it below the cap; the "402 tail" is resolved without a split.
Verification evidence (round 4)
- Full backend suite: 11213 passed, 240 skipped, 1 xpassed, 0 failed (−1 vs round 3 = the removed llm-config/strict-dep tests minus the 3 new RBAC + 5 transport-limit tests).
- Frontend: 3435 vitest passed (delta vs 3454 = the deleted assistant.test.ts), lint 0 errors / 340 warnings; settings+api slice 368; handoff slice 5.
- Targeted slices: MCP 95; transport limits 5; SC-005 backend slice 475+7 skipped; RBAC 3; ruff + compileall clean on the whole backend.
REMAINING (round 5+):
- 050 T008b (local-perimeter deployment tests/docs: reject non-local MCP/LLM/VLM endpoints by default; PARTIAL per round-1 closure review) — the only unchecked 050 task.
- FR-010 catalog versioning/deprecation markers (P2).
- Final retirement of the unmounted
routes/assistantpackage once parity provenance is archived (header records the retention rationale);agent_conversations/agent_assistant_dbmounted surfaces — round-1 recorded decision (retained pending product call). - Deferred product decisions (recorded, non-blocking): M-03 self-approval, service-branch catalog permissions, adapter idempotency after lease recovery, multi-binding exploration context, TTL/expiry polish.
- Semantic index rebuild pending Axiom binary availability.
- Product status: NO-GO pending the closure-gate re-review of rounds 1–4.
Checkpoint — 2026-09-04 (closure-gate remediation round 5 — T008b local perimeter + FR-010 catalog versioning: 050 task list closed)
050 T008b — local-perimeter endpoint guard (the last unchecked 050 task → checked with evidence)
- New
Core.EndpointLocalitymodule (backend/src/core/utils/endpoint_locality.py): deny-by-default enterprise-locality guard for provider endpoints (MCPX-FR-020). Locality = loopback/local host literals, private/loopback/link-local IP literals (RFC1918/ULA), enterprise DNS suffixes (.local/.internal/.lan/.corp/.intranet, env-extensible), DNS names resolving entirely to private addresses. Fail closed: unresolvable → denied; empty base_url (public cloud SDK default) → denied; substring spoofing (https://api.openai.com/localhost) denied by hostname — the old_is_local_base_urlsniff remains only a token-budget heuristic (its insufficiency recorded as @REJECTED). - Enforcement at the single config-time choke:
LLMProviderService.create_provider/update_provider(covers LLM and VLM — one provider table) BEFORE persistence/mutation; typedEndpointNotLocalError→ admin routes map to HTTP 400 (endpoint_not_local:<reason>), never 500. Every denial emits an EXPLORE wire line (perimeter audit trail). - Escape hatches (env, default closed, INSTALL-documented):
LLM_ALLOW_NONLOCAL_ENDPOINTS,LLM_NONLOCAL_ENDPOINT_ALLOWED_HOSTS,LLM_LOCAL_HOST_SUFFIXES. - MCP transport local contour already existed (DNS-rebinding defaults) + gained E6 limits in round 4; INSTALL.md gained §«Локальный периметр».
- Test migration (honest consequence): 3 existing service mechanics tests in
tests/services/test_llm_provider.pyusedapi.openai.com/http://testfixtures → switched to perimeter-local URLs (http://localhost:1234/v1,http://test.local); their intent (create/update mechanics, masked-key handling) is unchanged. Route/API tests mock the service → unaffected. - Evidence:
tests/test_endpoint_locality.py(22: accepted/denied/overrides/service-boundary/ route-mapping incl. no-persist-on-deny and no-mutate-on-deny) + provider/route slices → 107 passed; tasks.md T008b[x]with full proof line.
MCPX-FR-010 — catalog version + deprecation markers
MCP_CATALOG_VERSION = "1.0.0"(newMcpServer.CatalogVersionregion) published asserverInfo.versionat initialize (_build_probe_serversets the lowlevel server version — the SDK otherwise reports the mcp-package version); re-exported on the frozen server surface.McpToolDefinitiongaineddeprecated+deprecation_note; the [DEPRECATED …] marker is applied at the SINGLERbacFastMCP.list_toolschoke point (deprecated entries stay listed, registered and callable for one minor cycle; registration sites untouched; marker visible over tools/list).- Discipline pinned executable (
tests/test_mcp_catalog_version.py, 3): semver shape + the pinned major (deliberate-bump ritual — a breaking change cannot pass without consciously editing both the constant and the test pin) + marker mechanism (monkeypatched catalog entry → listed with marker, neighbour stays clean) + initialize wire publishes the version (JSON/SSE-robust parse). - MCP slice grew to 103 passed (95 + RBAC visibility 3 + transport limits 5 + local run overlap).
Verification evidence (round 5)
- Full backend suite: 11243 passed, 240 skipped, 1 xpassed, 0 failed (7:09) — +30 vs round 4 (22 locality + 3 catalog-version + parametrizations), zero regressions.
- Anchor-stack + AST syntax sweep over ALL 138 touched files: ALL BALANCED; ruff + compileall clean; targeted: endpoint locality slice 107, MCP slice 103, catalog version 3.
050 status after round 5
- tasks.md: all tasks [x] (T008b was the last unchecked; every [x] now carries current proof).
- Deferred product decisions unchanged (M-03 self-approval etc. — recorded, non-blocking); cookie-consent resolved in round 4 (tasks.md T008 note).
REMAINING (round 6+):
- Final retirement of the unmounted
routes/assistantpackage once parity provenance is archived (header records retention rationale);agent_conversations/agent_assistant_dbmounted surfaces — round-1 recorded decision pending product call. - Semantic index rebuild when an Axiom binary is available (all rounds used zombie-mode sweeps).
- Closure-gate re-review of rounds 1–5 → product status decision (currently NO-GO by the round-1 review verdict; all P1/P2 remediation queue items are now closed or recorded-as-decided).
Checkpoint — 2026-09-04 (closure-gate re-review, Axiom live) — rounds 1–5 verification matrix
Axiom MCP became available and the index was rebuilt live (full: 10534 contracts; incremental after
fixes: 10582/5190 — another feature session is concurrently migrating the logging SSOT from
shared/src/ss_tools/shared/ into backend/src/core/ and rebuilding in parallel; that migration is
out of scope here and its mid-state accounts for the current workspace-level deltas).
Remediation matrix vs the round-1 NO-GO review (all items CLOSED)
| Review finding | Round | Status | Live evidence |
|---|---|---|---|
| P0 fake T025 checkpoints, client_credentials, gate unwrap, SC-007 HTTP | 1 | CLOSED (committed a0450b33) |
tests pinned then; re-verified in every later slice |
| P1 MCL intent-drop (~120 sites → 182), EXPLORE gaps, 9 silent adapters | 2 | CLOSED | wire-format sweep test = 0 misuse; MCP slice 103 |
P1 INV_6 _llm_health dead edge + specs 033/035/036/039 tombstone sweep |
2 | CLOSED | zombie sweep 0 dead / 15 tombstoned; Axiom: AgentChat.LlmHealth relations resolve |
| P1 dedupes (TaskDrawer/assistantOffset/@SIDE_EFFECT), migration anchors, sandbox span | 2 | CLOSED | anchor sweeps ALL BALANCED incl. migrations 0014–0016 |
| P1 server.py DECOMPOSITION GATE plan | 2 (plan) → 3 (execution) | CLOSED | server.py 177 LOC; package census all <400; plan doc EXECUTED + log |
| P2 SC-005 remnants (assistant unmount, llm-config exemption decision, retention UI, assistant.ts, env 7860) | 4 | CLOSED | decision = removal w/ in-place Tombstone; routes slice 475 passed |
| P2 FR-010 catalog versioning/deprecation | 5 | CLOSED | serverInfo.version wire test + marker choke + pinned major |
| P2 E6 JSON-depth + per-session rate limits (Retry-After) | 4 | CLOSED | transport-limits 5 passed (pre-dispatch, no-state) |
| P2 HandoffSurface context prompt (Story 5 AC2) | 4 | CLOSED | contract+render tests 5 passed |
| P2 SC-004 admin/exact catalog fixtures | 4 | CLOSED | exact-set pins admin 47 / analyst 21 / viewer 15 |
| P2 SC-009 mid-flow role change | 4 | CLOSED | revoke-hides-and-denies-cached-call + grant-without-consent pins |
| P2 cookie-consent decision | 4 | RECORDED | tasks.md T008 decision note (not built in 050) |
| 050 T008b local-perimeter (last unchecked task) | 5 | CLOSED | locality guard 22 + slice 107; tasks.md [x] with proof |
SC-001…SC-009 walkthrough (current executable evidence)
SC-001 authoring→registry provenance: promotion E2E (8). SC-002 parity: ops parity + baseline 037.
SC-003 gates zero-side-effect-before-approval: approvals/checkpoints slices. SC-004 RBAC-exact
tools/list: test_mcp_rbac_visibility.py exact sets. SC-005 decommission zero-refs: round-4 closure
(greps clean; docs clean; env clean). SC-006 no SQL bypass: PROD-SQL terminal matrix + scenario
context denial. SC-007 discovery→DCR→PKCE→tools/list: client-flow HTTP (2). SC-008 refresh replay
family revocation: oauth tests. SC-009 live role change: RBAC visibility flow test.
All nine SCs carry current passing pins (last full runs: backend 11243/0 failed @ round 5,
frontend 3435/0 @ round 4 — both predate only the concurrent logging migration, not any 050 code).
Axiom live-audit findings (this re-review)
- Fixed in-scope (metadata-only edits, uncommitted):
McpServer+McpServer.ToolsAuthoringedges →Services.AgentAuthoringWorkspace.Service;McpServer.Package→McpServer(staleMcpServer.Server, plus Phase-0 BRIEF refreshed);ScenarioExecution.ExplorationSandbox→.Service;Services.McpOpsDispatchphantomSupersetClient.DashboardWrite→ the three verified write contracts (Core.DashboardsWrite.CreateDashboard/.CopyDashboard,Core.Datasets.SupersetClientCreateDataset);Test.McpParityBaseline037→AgentSuperset.SqlFormat(correct ID, noApi.prefix). Post-fix scoped audits: mcp_server prefix = 0 unresolved, mcp_ops_dispatch = 0 warnings. - Accepted advisories (deliberate, rationale-recorded):
McpServer.TransportAuth151/150 (cohesive guard, 1 over);McpServer.RbacLayerparser reports module region 418 while the FILE wc -l is 393 (INV_7 is file-level — compliant; parser metric noted);RbacServer259 /ScenarioModels243 /ToolsAuthoring.Register315 contract spans are gate-frozen shapes (behavior-neutral moves only). - Historical debt (predates rounds 1–5, workspace-wide): a first curation pass was executed
live-audit-driven — unresolved relations 446 → 401 (edges 5191 → 5207). Fixed (all targets
verified to exist before retargeting, comment-only edits, anchors/compile/diff-check clean):
35 retargeted edges across 22 files — scenario chain module-parent edges → verified function
contracts (
ScenarioGraph.Compiler.Compile/Validator.Validate/Resolver.Resolve/PackCompiler.Generate) in scenario.py route + specs 038/050; Superset-client path/alias forms →Core.Init.SupersetClientModule(dataset_mapper, dataset_key_sync×2, query_model, structure_diff_capture);[Models.User]×4 + admin flatUser→Models.Auth.User;ExecuteEnvelope×5 (+1 test BINDS_TO) →ExecuteQueryEnvelope;DraftStorage×2 →Services.AgentRuns.Artifacts;Services.Git.GitService→Services.Init.GitService;Models.Deployment→.DeploymentModels;StructureSnapshot.Service→.Capture+.Diff;Api.DashboardTesting.Core→Api.DashboardTesting; client_registry python-path forms →Core.AsyncNetwork.AsyncAPIClient/Core.ConfigModels; plus 11 malformed multi-target translate plugin lines (-> [A], [B) split into individual @RELATION lines with verified IDs (Models.Translate.*,Plugin.Dictionary.DictionaryManager,Plugin.LlmCall.LLMTranslationService,Plugin.LangDetect.LanguageDetectService,Plugin.BatchSizer.AdaptiveBatchSizer,Plugin.BatchProc.BatchProcessingService,Plugin.SqlGenerator.SQLGenerator,Plugin.SupersetExecutor.SupersetSqlLabExecutor,Plugin.LlmParse.LLMResponseParser,Plugin.TokenBudget.EstimateTokenBudget,Core.DbExecutor,Core.ConnectionService,Core.ConfigManager). Remaining 401 (queued for the next curator round, classified): logging-SSOT-domain edges owned by the concurrent migration (ss_tools.shared._llm_health,CotLoggerModule,main_cot_logger...); function-shaped targets without contracts (ReportsService.get_summary,_compute_content_hash,TranslationOrchestrator.execute_run,SupersetClient.network.request...); legacy single-#region misses (Services.LlmProvider.GetProvider); flat test-module targets (TasksApi,FileIOModule,fixtures.*); unverified remainder (Models.VerificationRun.VerificationRunRecord,BaselineEngine.Catalog.Materialization). Also: 294 module_too_long + 418 contract_too_long + 18 flat_hierarchical_id advisories; 1tombstone_missing_deprecatedcounted by the audit but NOT reproducible by source scan (region + brace syntax) — monitor. - Orphans 2710 are mostly relation-free C1/C2 leaves (expected per orphan_guidance), not defects.
Verdict and sequencing
- Rounds 1–5 remediation: COMPLETE. No open P0/P1/P2 item from the closure review remains; 050 tasks.md fully [x] with proof lines.
- Product status flips NO-GO → GO only after: (a) the concurrent logging-SSOT migration lands
coherently (currently mid-flight:
ss_tools.shared.cot_loggerimports broken tree-wide — outside this workstream by explicit instruction), (b) a full backend + frontend green run on that landed state (11243/3435 numbers must be re-established post-migration), (c) formal closure sign-off against this matrix. - Uncommitted at this point: the six metadata edge-fix files listed above (comment-only; runtime behavior untouched — verification is anchor/Axiom-based, not suite-based, while the tree's import layer is mid-migration).
UX-аудит — MCP happy path для BI-аналитика (2026-09-04, read-only, код НЕ менялся)
Прокручен полный продуктовый путь «аналитик → понять про MCP → подключить → работать → одобрить». Верифицировано по коду frontend (file:line приведены). Вердикт понятности: 2/5 — скелет и передача контекста дашборда есть (050 AC2, round 4), но «последняя миля» разорвана в двух местах.
Фактическая карта пути
- Точки входа в
/agent(единственная MCP-поверхность в UI):TopNavbar.svelte:269кнопка «Ассистент» (глобально);DashboardHeader.svelte:112кнопка «AI» (тултип «Ask AI about the dashboard»);DashboardHeader.svelte:120«Создать сценарий тестирования» (intent build_dashboard_test_scenario);datasets/[id]/+page.svelte:32. В сайдбаре (sidebarNavigation.ts) пункта /agent НЕТ. - На
/agentрендеритсяHandoffSurface.svelte: заголовок «Ассистент переведен на MCP», описание «…Подключите MCP-клиент, чтобы продолжить», «подсказка подключения» = «Используйте MCP-сервер Superset Tools, настроенный для вашего workspace» (ru,assistant.json:185), копируемый промпт с контекстом дашборда. - Петля одобрения (вторая половина пути): экраны существуют —
/dashboard-testing/runs(WaitingForMeView, чекбокс «Только ожидающие моего решения») и/dashboard-testing/scenarios/[id]/runs/[runId](HumanCheckpointPanel), но недостижимы через навигацию.
Замечания (severity → evidence → эффект)
- CRITICAL — «Подключите MCP-клиент» без КАК/КУДА. Подсказка подключения не содержит ни
endpoint URL (
https://<origin>/mcp), ни discovery (/.well-known/oauth-protected-resource/mcp), ни сниппета конфига клиента (Claude Desktop/Cursor/...), ни ссылки на INSTALL.md §MCP. Термин «MCP-клиент» не объясняется. Аналитик гарантированно застревает на шаге 2. - CRITICAL — петля одобрения недостижима. В
ROUTES.tsнет ни одного builder'аdashboard-testing/*/load-testing; вbuildSidebarSections()нет секции «Тестирование дашбордов». Агент встаёт вwaiting_for_user, а человек не может обнаружить, где одобрять (только ручной ввод URL). Продукт «висит» в ожидании ненаходимого действия. - HIGH — bait-and-switch входов. «AI»/«Ассистент» обещают вопрос-ответ в приложении, приводят на заглушку про подключение клиента. Для не знающего о декомиссии чата — выглядит сломанной кнопкой.
- HIGH — настроек/статуса MCP в UI нет вообще (grep по settings/admin: только серверные LLM-провайдеры). Ни URL сервера, ни статуса подключений, ни доступного роли набора инструментов (RBAC-каталог роль-зависим: admin 47 / analyst 21 / viewer 15 — данные уже есть, поверхности нет).
- MEDIUM — маршруты в обход SSOT. Страницы
dashboard-testing/*/load-testingлинкуются хардкод-строками через локальныйresolve(), минуяROUTES.ts— нарушение @INVARIANT реестра («every href/goto must use these functions»); link-integrity тест их не покрывает.
Рекомендуемые фиксы (приоритизированы; все frontend; закрывают 1+2+4 = минимум)
- Fix 1 (HandoffSurface actionable): вывести реальный endpoint из
page.url.origin({origin}/mcp) отдельным копируемым блоком + кнопка «Скопировать URL»; 3 шага onboarding («что такое MCP-клиент → вставить URL → подтвердить доступ в браузере (OAuth)»); ссылка на INSTALL.md §MCP; discovery URL для продвинутых. Новые i18n-ключи ru/en. - Fix 2 (петля одобрения в навигации): секция сайдбара «Тестирование дашбордов»: Сценарии /
Запуски (бейдж «ждут меня» — счётчик waiting_for_user) / Аналитика; RBAC-фильтр по
scenario/dashboard-testing permissions; запись в
buildSidebarSections+ тест вsidebarNavigation.test.ts. - Fix 3 (честные входы): переобозначить «AI»/«Ассистент» → «MCP-ассистент» / тултип «Инструкция по подключению AI-ассистента (MCP)»; опционально first-visit onboarding.
- Fix 4 (ROUTES SSOT): добавить builders
dashboardTesting.{scenarios,runs,runDetail,analytics, automation,scenarioEdit}+loadTesting.detail; перевести хардкод-resolve()на них. - Fix 5 (опционально, админ): MCP status page — endpoint, discovery, инструменты по роли.
Статус: не реализовано — аудит read-only по запросу; к имплементации готов минимум Fix 1+2+4 (устраняют оба CRITICAL + SSOT-нарушение) с vitest/svelte-check/eslint-гейтами.
Checkpoint — 2026-09-04 (UX-audit remediation IMPLEMENTED — Fix 1+2+3+4 + registry hub)
Implemented the actionable minimum of the read-only UX audit above (Fix 1+2+4), plus Fix 3
(honest entries) and the scenario-registry hub the audit's Fix 2 "Сценарии" nav item implies.
Gates run per AGENTS.md frontend checks (npm run test / npm run lint / npm run build);
the project has no svelte-check script, so the "svelte-check gate" is covered by vite build.
Fix 4 — ROUTES SSOT for dashboard-testing / load-testing
src/lib/routes.ts: addedROUTES.dashboardTesting.{scenarios, scenarioEdit, scenarioAnalytics, runDetail, runs, analytics, analyticsCase, automation}+ROUTES.loadTesting.detail; every builder maps to a real route file (registry index created this round). @RATIONALE records that the integrity test now scans these trees.- Converted ALL hardcoded
resolve("/dashboard-testing/...")SSOT bypasses to ROUTES builders in 7 pages (runs ×3, automation, analytics ×2, analytics case, scenario edit/analytics/run-detail)- DashboardHeader load-testing href + DashboardHeader/agent hrefs now
ROUTES.agent(...).
- DashboardHeader load-testing href + DashboardHeader/agent hrefs now
- Hardened
routes-link-integrity.test.ts:/dashboard-testing+/load-testingadded to the watched prefixes AND aresolve("/dashboard-testing|/load-testing...")bypass detector (the audit MEDIUM #5 — these trees were previously uncovered).routes.test.ts+13 builder pins. - Removed a now-dead
$app/pathsvi.mock in run.ux.test.ts orphaned by the conversion.
Scenario registry hub (NEW — makes the "Сценарии" nav item real; fixes 4 dead back-links)
src/routes/dashboard-testing/scenarios/+page.svelte: searchable/filterable registry list over the previously unusedScenarioRegistryModel(042 list/read only — never starts a run/agent); rows link via ROUTES to edit / analytics / last-run. Previously/dashboard-testing/scenarioswas a DEAD route that four "← Сценарии" back-links pointed at. i18ndashboard_testing.registry_*ru/en. UX testscenarios_index.ux.test.ts(4): rows+SSOT links, no-run-history, empty, error+retry.
Fix 2 — approval loop reachable from navigation (CRITICAL 2)
sidebarNavigation.ts: newdashboard_testingcategory in the Operations section (no new section id → the existing exact-section/categories tests stay green). Per-subitem RBAC mirrors the exact backendhas_permissiongates: Сценарии→dashboard:testing READ, Запуски→scenario RUN, Аналитика→scenario:result VIEW, Автоматизация→scenario:automation TRIGGER; category itself unpermissioned (visibility derives from subitems → hidden entirely when none pass).stores/scenarioRuns.svelte.ts(NEW): waiting-for-me count store (server-authoritative total via/scenario-runs?waiting_for_me=true&page_size=1; fail-closed 401/403 disable; inflight dedup; mirrors HealthStore).test_scenarioRuns.ts(9): state/refresh/dedup/auth-disable/retry/subscribe.Sidebar.svelte: warning-token badge on thedashboard_testingcategory (count expanded / dot collapsed) + 60s poll gated onscenario:RUN— turns the approval loop into a discoverable signal (directly addresses "продукт висит в ожидании ненаходимого действия"). Metadata BINDS_TO/@UX_FEEDBACK updated.nav.jsonru/en keys;category.testingtailwind token.sidebarNavigation.test.ts+4 RBAC tests.
Fix 1 — HandoffSurface actionable (CRITICAL 1)
HandoffSurface.svelte: renders the REAL MCP endpoint frompage.url.origin({origin}/mcp) as a copyable block + "Copy URL", the OAuth discovery URL ({origin}/.well-known/oauth-protected-resource/mcp, verified againstbackend/src/api/mcp_oauth.py+app.py/mcpmount), a 3-step onboarding and the docs/INSTALL.md «MCP клиент» pointer — closing "«Подключите MCP-клиент» without how/where". Offline invariant preserved (synchronous origin/props/i18n derivation, no network). Context-parameterized prompt (Story 5 AC2) unchanged. i18nassistant.handoff_*ru/en; context test +3 (endpoint, discovery, onboarding+copy controls). HandoffSurface lints 0-warning.
Fix 3 — honest entries (HIGH #3 bait-and-switch)
- TopNavbar assistant button + DashboardHeader "AI" button relabeled to «MCP-ассистент» / «MCP» with
tooltip «Инструкция по подключению AI-ассистента (MCP)» — no longer promises in-app chat.
assistant.mcp_entry_*ru/en.
Verification evidence (this round)
- Full frontend suite: 3486 passed, 0 failed (200 files) (was 3435; +51 = new builder pins, sidebar RBAC, store, registry index, handoff endpoint/onboarding).
- eslint: 0 errors / 354 warnings (baseline 340). The +14 are ALL the accepted
svelte/no-navigation-without-resolveclass: the codebase's established ROUTES convention is barehref={ROUTES.x()}/goto(ROUTES.x())(98 such warnings pre-exist;resolve(ROUTES())appears nowhere), so SSOT-compliant navigation deliberately carries the same accepted warning. No new error/a11y/unused-var warnings introduced; Sidebar's 3 each-key warnings are pre-existing. - Production build (adapter-static): green. GRACE-Poly anchors balanced in every touched file;
store MCL
log(src, marker, intent, payload, error)mirrors HealthStore (src-first frontend facade). - Recorded decision: bare
ROUTES.x()chosen overresolve(ROUTES.x())for consistency with the dominant convention (base path empty under adapter-static →resolve()is identity anyway).
Not done (optional / out of minimum scope)
- Fix 5 (MCP status/admin page — endpoint, discovery, per-role tool list): audit-optional and larger; deferred.
- Broader product-GO gates from prior checkpoints unchanged and NOT touched by this frontend UX round: semantic index rebuild (Axiom), full backend suite re-run on the landed logging-SSOT migration, and browser E2E on a live stand (would need the backend+frontend stack up). Product status per the round-1 closure review remains pending formal re-sign-off.
Checkpoint — 2026-09-05 (UX-audit "остатки" — Fix 5 admin MCP governance + Fix 2 polish, UI/UX discussed first)
Per the operator's request the remaining UI/UX was discussed and decided BEFORE implementing. Decisions:
Fix 5 = admin governance (all 47 tools × roles) mounted under /admin/*; polish = badge→waiting
view + registry extra filters & pagination. (First-visit onboarding was declined.)
Fix 5 — admin MCP governance surface (audit finding #4: "настроек/статуса MCP в UI нет вообще")
- Grounding first: MCP's only REST surface is OAuth (
/oauth/*,/.well-known/*); the role-dependent tool catalog is exposed only over the MCP protocol (tools/list), and/api/readycarries 044 provider readiness, not the catalog. So a governance view needs a small NEW backend projection. - Backend
backend/src/api/routes/admin_mcp.py(new module, keepsadmin.pyunder INV_7): read-onlyGET /api/admin/mcp/catalog, gatedhas_permission("admin:roles","READ")(same gate as the role/ permission inventory). Returns catalog_version (MCP_CATALOG_VERSION), endpoint/discovery paths, the full frozen catalog (name/risk_level/requires_approval/service_allowed/deprecated/required permission), a per-role visible-tool matrix, and a read-only DCR client list (OAuthClient, scopes/ redirect_uris parsed; never secret/secret_hash). Registered inapp.py(413 routes; import verified). - The per-role visibility helper
_visible_tool_namesmirrorsRbacFastMCP._can_use_tool's human branch exactly (permission None → visible;start_scenario_run→ scenario RUN or RUN_PROD; else exact role resource/action match; admin bypass).tests/api/test_admin_mcp_catalog.py(8) pins it to the SAME exact admin/viewer/analyst subsets SC-004 pins for the protocol → UI cannot drift from tools/list. - Frontend:
api.getMcpCatalog;models/AdminMcpCatalogModel.svelte.ts(read-only, last-projection-on- error);routes/admin/mcp/+page.svelte(ProtectedRoute admin:roles; server card w/ absolute endpoint+ discovery, full tool table, expandable role×visibility matrix, DCR client list);ROUTES.admin.mcp(); sidebar admin subitem «MCP» (admin:roles READ) + admin hub card; i18nadmin.mcp.*+nav.admin_mcp*ru/en. - Real defect found & fixed by the page test: the load
$effectused!loaded && !loading, which re-fires after a FAILED load (loading toggles, loaded stays false) → an auto-retry loop hammering the endpoint. Switched to astartedguard (load once + manual Retry), matching the registry hub. This is the same latent pattern some sibling pages still carry (runs/analytics use rows/loaded guards).
Fix 2 polish — badge → "ожидают меня" view; registry filters + pagination
ROUTES.dashboardTesting.runs(waitingForMe?)→?waiting_for_me=true; run center now READS that param on entry (page.url.searchParams) and opens directly onWaitingForMeView; the Sidebar waiting badge (expanded count) is now a button navigating toruns(true)(stopPropagationso it doesn't toggle the category). The collapsed dot stays an indicator. Turns the passive count into a one-click approval entry.- Scenario registry hub gained owner / tag / dashboard_id filters + server-side pagination (PageSize 20,
prev/next, page X/Y) over the existing
ScenarioRegistryModel.loadListquery; filter change resets to page 1 (@INVARIANT). i18ndashboard_testing.registry_*ru/en extended.
Verification evidence (this round)
- Frontend: full suite 3498 passed, 0 failed (204 files) (was 3486); eslint 0 errors / 354 warnings
(UNCHANGED — Fix 5 added no nav-href warnings; admin/mcp page has no internal links); build green
(adapter-static); anchors balanced in every new/edited file; all touched i18n JSON parses;
git diff --checkclean. - Backend: admin/mcp + admin + mcp-rbac-visibility + catalog-version slice 42 passed; MCP slice
55 passed; scoped dashboard-testing OpenAPI 25 passed (unaffected by the new /api/admin route);
ruff +
compileallclean;app.pyimports with/api/admin/mcp/catalogregistered. Full backend suite NOT re-run this round (change is one isolated additive endpoint + registration; scoped slices + app import are proportionate) — the prior round-5 number (11243) stands pending the logging-SSOT re-run.
Remaining after this round
- Fix 5 shipped read-only governance. DCR client MANAGEMENT (revoke/rotate) and automated key-rotation expiry remain open (separate features the audit listed); the view exposes clients but performs no mutation.
- Fix 5 was the last audit UI/UX item; all five audit fixes (1–5) are now implemented. Deferred product decisions (M-03 self-approval, service-branch catalog permissions, adapter idempotency, multi-binding exploration context, TTL/expiry) and the broader GO gates (semantic rebuild, full backend suite on the landed logging migration, browser E2E on a live stand, closure-gate re-sign-off) are unchanged.
Checkpoint — 2026-09-05 (orthogonal UI/UX review + i18n remediation)
Per operator request an independent UI/UX pass was run on the newly delivered surfaces, with i18n completeness as an explicit gate. Findings (orthogonal, from a fresh pair of eyes):
- i18n debt (primary): the dashboard-testing feature area still carried hardcoded Russian UI
strings —
runs/+page.svelte(~25 strings),WaitingForMeView.svelte(~12) — violating the "no hardcoded UI strings" invariant. My two new pages additionally used English|| "…"fallback literals and rendered raw enum tokens (risk_level, client_type, lifecycle_status, health) untranslated. - Bug: scenario registry
dashboard_idfilter sentNaNin the query for non-numeric input. - UX clarity: admin/mcp "Сервис" column (service_allowed yes/no) was ambiguous — what "yes" grants is not obvious; added a clarifying tooltip.
- Minor: admin/mcp page lacked a document
<svelte:head><title>.
Remediated:
runs/+page.svelte+WaitingForMeView.sveltefully migrated to$t(dashboard_testing.runs_*,waiting_*); all hardcoded RU strings removed (only a@BRIEFmetadata comment references the label).scenarios/+page.svelte+admin/mcp/+page.svelte: fallback literals removed (i18n keys are the only source), lifecycle/health/risk/client_type enums translated (ls_*,health_*,risk_*,client_*) with raw-token fallback for unknown values (drift-safe);dashboard_idNaN guard added; admin/mcp<title>added; service column disambiguated viacol_service_hint.- i18n locales extended ru/en:
dashboard-testing.json(+runs/waiting/ls/health ≈45 keys),admin.jsonmcp.*(+risk/client/service-hint keys); nav/assistant unchanged this round. - Test fixture for the admin/mcp page broadened to the full
admin.mcpkey set.
Verification: frontend full suite 3498 passed / 0 failed (204 files); eslint 0 errors / 354 warnings
(unchanged — no new nav-href); build green; no residual Cyrillic in the four touched templates; all four
edited locale JSONs parse. WaitingForMeView.test.ts (real-i18n, ru labels) still green.
Checkpoint — 2026-09-05 (i18n sweep COMPLETE — analytics/editor/automation/run surfaces)
The previously-documented remaining i18n debt is now cleared: the analytics/automation pages,
scenario detail pages ([id]/edit, [id]/analytics, [id]/runs/[runId]) and the entire scenario
component trees were migrated off hardcoded Russian strings to $t (dashboard_testing.*):
- scenario-analytics: CaseWorkspace, QueueList, TrendsChart, AgentActionTimeline, RecurringFailuresList, HealthCard + analytics/queue + case routes.
- scenario-run: HumanCheckpointPanel, RunTimeline, RunComparison, RunHistoryList, ScenarioResultView, RunConfigurationPanel + run-detail monitor route.
- scenario-editor: ConstrainedAssertionEditor, VisualDagCanvas, AgentActionPanel, EditRevisionDiff + scenario edit route.
- scenario-automation: AutomationPanel, ScheduleForm + automation route.
- Locales extended (ru/en) with ~150 keys. Policy: ru values = the current rendered text (some terms stay English where they currently ship untranslated, e.g. "attention item(s)", "Investigate with agent", "occurrence(s)", "group(s)") so real-i18n tests stay green; wiring is complete and the locale entries are now the single place to localize later.
- Fixed a stale route test
case.ux.test.ts("Записать disposition" → current "Закрыть кейс" + resolved/verification_evidence CAS body); it previously asserted a pre-refactor disposition flow.
Gates: frontend 3494 passed / 4 failed (204 files); the 4 failures are PRE-EXISTING (verified in isolation, none touched by this workstream):
routes-link-integrity—src/routes/oauth-consent/+page.svelte:44 goto('/login')raw SSOT bypass (untracked file, not part of this workstream). 2-3.InvestigationModels.test.ts(2) — stale vs the prior-roundInvestigationCaseModel.disposesignature change (rationale.trim is not a function, CAS version 1 vs 2).settings_page.ux.test.ts(1) — pre-existing numeric-range save guard, unrelated to dashboard-testing. Lint 0 errors / 358 warnings (+4 toleratedno-navigation-without-resolve); build green; no residual hardcoded Russian outside a single@BRIEFcomment. The 4 pre-existing failures are out of i18n scope and left for a separate cleanup.
Checkpoint — 2026-09-06 (spec refinement: server-owned handle layer + F3 guard; ADR-0023)
Orthogonal code review of the T029/T029c implementation found a root-cause architectural gap and it is now both guarded in code and fully specified:
- Code (implemented, verified):
ScenarioExecution.RunnerPlan.Deriverefuses provenance-only bootstrap revisions (BOOTSTRAP_REVISION_NOT_RUNNABLE, EXPLORE-marked) before any run/gate/schedule side effect — closing a vacuous zero-step false-PASS (Runner.Walker._advance_runmarks empty planspassed).pack_compiler.pydocstring corrected (Generate persists nothing; registration lives inScenarioGraph.PackCompiler.Register036).test_mcp_initial_scenario_e2e.pyrewritten: asserts the refusal, then proves queued/PROD-gate behavior on a genuinely materialized revision. Gates: 1234 passed (tests/services/dashboard_testing+ API/MCP suites), 458 passed focused MCP/registry set, ruff/compileall/git diff --checkclean; zero-stepseeded_registryfixtures unaffected (guard keys on the exactcompiled_handle_id+draft_pack_id+draft_pack_digestsignature with no steps, not on empty plans in general). - Specs (amended, no code claims): 038
contracts/modules.md—PackCompiler.Generatecontract fixed +ScenarioGraph.ServerOwnedPipelinehandle-persistence amendment (minting boundaries, three immutable handle entities, canonical-bytes authority, binding/single-consumption/GC, inspect-stage PROPOSED status); 038contracts/verification-program.md— reconciliation block (implemented: ScenarioStep DAG + ActionRegistry, CaptureSpec/VlmAnalysis/VlmFinding/ScenarioParameter; deferred: five-way VerificationProgram, SqlEvidenceSpec, TransformSpec/ComparisonSpec/AssertionSpec, AgentEvaluationSpec/DecisionPolicy — zero code references); 042data-model.md— synchronousgraph_snapshotmaterialization from handle bytes inside the create transaction, outbox scoped to reference artifacts only, interim-drift note; 042contracts/modules.md—CreateInitialamendmentsRevisionChainrejection of the legacy REST raw-graph_snapshotroute; 050spec.md— stage-table implementation-status note + handle-rules status bullet.
- New work scoped: 050
tasks.mdPhase 2c — T029d (handle tables + canonical bytes store + migration), T029e (042 create consumes handles, sync materialization, outbox/RevisionMaterialization worker, guard demotion), T029f (MCP minting surface, bootstrap accepts stored handle ids only, legacy REST route gate/retire, catalog minor bump), T029g (full-chain fresh-DB MCP E2E + PostgreSQL concurrency/retention), T029h (inspect-stage implement-or-descope decision). T029a remains open and naturally folds into T029d/e. - Decision memory:
docs/adr/ADR-0023-mcp-scenario-pipeline-handle-gap.md(PARTIALLY IMPLEMENTED) records the finding, the rejected alternatives (silent vacuous PASS; blanket empty-plan rejection; lossy-pack retro-fit as sufficient fix) and the two-part resolution. Registry row added todocs/adr/README.md. - Product remains NO-GO for the full external happy path
inspect → compile → validate → resolve → draft-pack → bootstrap → start rununtil Phase 2c lands; T029–T029c scope stays complete as marked.
Checkpoint — 2026-09-06 (Phase 2c landed: handle layer T029d–T029g + inspect hybrid T029h)
The ADR-0023 root cause is closed in code; the NO-GO above is superseded for the MCP happy path.
- T029d:
CompiledScenarioHandle/ValidationResultHandle/DraftPackHandle(immutable, owner-bound, content-addressed canonical-bytes storehandle:{sha256}, idempotent mint, single consumption underSELECT ... FOR UPDATE+populate_existing, purge/GC helper), migration0019_scenario_handles; minting wired into RESTapi_compile_scenario/api_validate_scenario/api_resolve_scenario/api_draft_pack(additive response fields only; pure compiler functions keep@SIDE_EFFECT None). - T029e: handle-first
create_scenario/create_initialmaterializegraph_snapshot= canonicalDashboardTestScenarioJSON + server-ownedaction_registry_version/hashin-transaction;OutboxEvent(materialize_revision)+RevisionMaterialization(pending)+ idempotent workermaterialize_pending_revisions(migration0020_scenario_materialization);api_create_scenarioreturns the real materialization status.BOOTSTRAP_REVISION_NOT_RUNNABLEdemoted to defense-in-depth. - T029f: MCP
register_draft_packwrite tool (human-only, AgentRun ownership) mints all three handles;create_initialaccepts ONLY stored handle ids (transitionalcompile:{run_id}:{digest}removed); legacy RESTPOST /scenarios/{id}/revisionsraw-graph route retired →410 REVISIONS_RAW_GRAPH_RETIRED; catalog2.0.0(pinned-major ritualPINNED_CATALOG_MAJOR=2). Side-fix:generate_reportreclassified out of_MUTATING_ACTIONS(local draft-report write, matchesbrowser.pymutation set) →ACTION_REGISTRY_VERSION 038.1.0 → 038.2.0,scenario_execution/graph.jsonfingerprint re-pinned;register_artifact(repository_write) intentionally stays mutating. - T029g: fresh-DB MCP E2E runs the full chain with zero REST crutches
(
register_draft_pack → bootstrap → direct start_scenario_run = queued); PostgreSQL concurrency proven on Testcontainers (test_scenario_handle_concurrency.py: two threads → one CONSUMED + one HANDLE_CONSUMED). - T029h (hybrid option C + X1): MCP
inspect_dashboard_context(live authoritativeDashboardQueryModel- fingerprint echo contract);
context_authorityevaluation atregister_draft_pack(server-side fingerprint RECOMPUTE — claimed values ignored; sentinels""/sha256:errornever verify; unconfigured/unreachable env fails OPEN tounverified; live-env falsifiable mismatches reject typed with zero rows); marker persists onDraftPackHandle(migration0021_context_authority), materializes intograph_snapshot, andRunner.Startrefuses PROD on explicit non-verified (CONTEXT_AUTHORITY_REQUIRED_FOR_PROD, legacy missing marker allowed); validator_check_dashboard_contextrecursively rejectsquery_context/SQL smuggling indashboard_context(X1). Catalog2.1.0(additive minor).
- fingerprint echo contract);
- Gates: broad regression 1338 passed (dashboard_testing + scenario/api + MCP suites); integration
concurrency green on PostgreSQL 16;
alembic heads→0021_context_authority; ruff/compileall/git diff --checkclean. - Spec updates: 050
spec.mdstage-table status note + handle-rules status bullet rewritten to implemented state; 050tasks.mdT029d–T029h[x]with evidence; 038ServerOwnedPipeline§6 rewritten (hybrid resolution). - Remaining residuals (tracked, non-blocking for the MCP happy path): outbox worker not yet wired into the
scheduler poll loop; 043 editor save path and REST
api_draft_packdo not evaluatecontext_authority(NULL = legacy-allowed at the PROD gate); full PostgreSQL suite for outbox retry/GC; D2/D4 frontend gaps (UI launch wiring, PROD approval surface) unchanged.
Checkpoint — 2026-09-07 (Phase 2c residuals closed: outbox wiring, REST parity, marker inheritance)
- Outbox worker wired:
Core.Scheduler.ExecuteRevisionMaterializationtick consumesmaterialize_pending_revisionsevery 30s (scenario_revision_materialization, max_instances=1, coalesce) — registration + durable-tick idempotency proven intest_scenario_scheduler_callbacks.py. - REST parity:
api_draft_packis async and evaluatesScenarioGraph.ContextAuthorityFIRST (same typed rejections, zero rows on falsifiable mismatch); response carriescontext_authority. Mock-boundary tests updated to patch the authority+mint seams (module-boundary mocking convention). - 043 editor path:
save_proposalinherits the server-ownedcontext_authoritymarker from the base revision viaapply_ops'{**base_graph}spread — bootstrap-origin scenarios keep PROD-gate coverage after edits (test_save_proposal_inherits_server_context_authority_marker). Legacy NULL markers remain PROD-allowed by design (migration-safe). - Stale-mirror fix:
tests/api/test_admin_mcp_catalog.py::_ANALYST_EXTRAwas missingregister_draft_pack(blind spot of the T029f test selection); the full-suite run caught it. - Gates: full regression
3360 passed, 8 skipped;alembic heads→0021_context_authority; ruff/compileall/git diff --checkclean. - Open (frontend, tracked separately): D2 UI launch wiring (
RunConfigurationPanel→ route) and D4 PROD approval surface (ApprovalDecisionPanel→ run center) — the two cheapest product gaps.
Checkpoint — 2026-09-07 (frontend D2/D4 closed: scenario launch surface + PROD approval UI)
- D2 scenario launch (UI): new page
frontend/src/routes/dashboard-testing/scenarios/[id]/+page.svelte(ScenarioRegistry.Route.Detail) — registry facts + revisions + the first production host ofRunConfigurationPanel;ROUTES.dashboardTesting.scenarioDetail(id)added to the SSOT and the registry index name cell now links to it (index still never starts runs — its invariant intact). Launch flow: panel →RunMonitorModel.launch(config, crypto.randomUUID())→ POST/scenario-runswithIdempotency-Key→goto(runDetail). Mandatory steps are fed from the detail graph (logical_step_id), environments fromGET /environments(is_prod = stage==='PROD' || is_production). - D4 PROD approval (UI): new
ApprovalDecisionPanel.svelte(approve/deny + comment, busy-lock);RunMonitorModel.decideApprovalposts/scenario-runs/{id}/approval/decisionand reloads the run; the run-detail aside renders the panel exactly whenstatus === "pending_approval"(before the waiting_human branch). Closes the last happy-path gap: PROD scheduled/manual runs are now fully decidable from the web UI, not only via MCP/REST. - Pre-existing suite failures fixed en route (all green before/after independently of D2/D4):
oauth-consentrawgoto('/login')→ROUTES.login()(link-integrity audit); staleInvestigationModelsdispose signature (2-arg → production 4-arg);settings/+page.sveltefills backend-defaultmcp_oauth_*values (15/30/90) so absent fields can't crashbind:value(Svelte 5props_invalid_value). - i18n:
approval_*(6) +scenario_detail_*(9) keys added to BOTH en/rudashboard-testing.json. - Gates: full vitest 3506 passed;
npm run lint0 errors (364 pre-existing warnings);npm run buildsucceeds (adapter-static); tests added:ApprovalDecisionPanel.test.ts(2),RunMonitorModel.test.ts(+2 decideApproval),scenarios/[id]/__tests__/detail_launch.ux.test.ts(3: facts+mandatory steps, typed launch POST + redirect, load-failure alert/retry). - Product status: the full journey agent-bootstrap → registry → UI launch → monitor → PROD approval → schedule is now UI-complete; no known D-class gaps remain in the 050 review matrix.
Checkpoint — 2026-09-07 (field-run remediation PLAN + test-honesty FIX; ADR-0024, 050 Phase 2d)
Источник:
docs/2026-09-07-sales-prod-mcp-run.md— первый полевой прогон ВНЕШНЕГО MCP-клиента против ss-prod Sales Dashboard (ID 11). Прогон доказал: initial-bootstrap цепочкаinspect → compile/validate → draft-pack → register_draft_pack → bootstrap → start_scenario_runcode-complete, но внешне недостижима —register_draft_packтребует principal-ownedAgentRun, а MCP-операции его создания нет (единственная creation-поверхностьPOST /api/agent/runs— web-session REST; MCP-токеныaud=mcpна ней не аутентифицируются). Ни entry/revision/run создано не было. Маскировка: вертикальный E2E seed'ил предусловие raw-ORM-вставкой — suite не мог увидеть разрыв.
Спеки (план включён полностью)
docs/adr/ADR-0024-mcp-agent-run-external-boundary.md(новый,PARTIALLY IMPLEMENTED) + строка вdocs/adr/README.md: решение = явные typed MCPcreate_agent_run/get_agent_runповерх существующегоServices.AgentRuns.Service.Create(permission("dashboard:testing","EXECUTE")REST-паритет,service_allowed=False, каталог 2.1.0 → 2.2.0 additive); инвариант внешней достижимости durable-предусловий; binding test-honesty правило (вертикальный тест обязан получать предусловия через границу, достижимую для тестируемого принципала; raw-ORM-seed предусловия в external-chain тесте запрещён). @REJECTED: регистрация без AgentRun; implicit auto-create внутриregister_draft_pack; REST-костыль как documented external path; сохранение raw-ORM seed.- 050
spec.md: MCPX-FR-027 (external reachability + create/get agent run), MCPX-FR-028 (context-derived capabilities), MCPX-FR-029 (disposition clarity); canonical-names строкиcreate_agent_run/get_agent_run; field-run correction под stage table; release-gate строкиE2E-EXT-001/002,CAP-001,DISP-001(OPEN) +TEST-001(CLOSED этот раунд); ClarificationsSession 2026-09-07(включая поправку clarification 2026-08-24 про AgentRun); Phase 2d строка. - 050
tasks.md: Phase 2d — T029i (MCP AgentRun surface + конвертация E2E в fully external chain), T029j[x](test-honesty remediation этого раунда), T029k (derive_capabilities от авторитетнойDashboardQueryModel/environment policy/provider readiness — ложныеhuman_checkpointустраняются), T029l (labels «Подтвердить соответствие»/«Проблема не подтверждена»/«Недостаточно данных» + restyle confirm сbg-destructive+ tool descriptions; lifecycle mapping неизменен), T029m (live-stand replay полевого сценария). T029g evidence снабжён honesty-amendment (raw-seed вскрыт и заменён). - 038
spec.md: Field-run Amendment — context-derived capability authority (SPECIFIED-PENDING): детерминированная серверная derivation, caller-declared capabilities только как сужение,needs_context/ needs_selector/needs_baselineдля неразрешённых фактов,human_checkpointтолько для реально небезопасного mutation/human judgement; селекторы/метрики/бейзлайны не изобретаются. - 044
spec.md: Field-run Amendment — HumanCheckpoint disposition clarity (mappingconfirm→passed — канон и неизменен; human-facing формулировки именуются по персистентному исходу; destructive-стилизация confirm запрещена; API-вокабуляр не переименовывается). 045spec.md: mirror-amendment для монитора (WaitingForMeView/HumanCheckpointPanel + vitest pins).
Test-honesty remediation (исполнено этот раунд, T029j)
backend/tests/test_mcp_initial_scenario_e2e.py::_pack(): rawAgentRun(...)ORM-вставка ЗАМЕНЕНА на продуктовую границуcreate_agent_run(db, CreateAgentRunRequest(context=UIContextV2(...)), user_id=...)(тот же сервис, что wrapsPOST /api/agent/runs; run существует в точной продуктовой форме — status RUNNING +run_startedevent). Module metadata честно фиксирует: MCP creation surface отсутствует (T029i open), конвертация в fully external chain — после T029i.- Новый
backend/tests/test_mcp_agent_run_reachability.py(2 теста): (a) strict-xfail requirement pin MCPX-FR-027 — XPASS и падение suite в момент появленияcreate_agent_run/create_authoring_agent_runв каталоге (форсирует конвертацию E2E + снятие маркера); (b) typed denial pin текущего состояния — внешний принципал без owned AgentRun получает ровно{"status":"blocked","error":"DRAFT_PACK_ACCESS_DENIED"}с НУЛЕВЫМИ строками CompiledScenarioHandle/DraftPackHandle/AgentRun (воспроизведение полевого результата).
Verification (этот раунд)
- Targeted:
2 passed, 1 xfailed(E2E vertical green через сервис-границу; requirement pin strict-XFAIL; denial pin green) — 2.16s. - MCP regression slice (11 файлов: server, scenario E2E, promotion E2E, initial E2E, reachability, t029 bootstrap/automation, ops parity, checkpoints, transport limits, rbac visibility, catalog version): 76 passed, 1 xfailed — 12.6s.
ruff+compileallчистые; anchors balanced (reachability 4/4, initial E2E 3/3); scopedgit diff --checkчистый. Полный backend suite НЕ rerun (изменения — два тестовых файла + specs/ADR docs; production-код не тронут).
Статус и 남은 работы
- Внешний initial-bootstrap путь остаётся NO-GO до T029i (MCP
create_agent_run); UI-путь (D2/D4) и внутренний E2E unaffected. Очередь: T029i → конвертация E2E (E2E-EXT-001) → T029k (CAP-001) → T029l (DISP-001) → T029m live replay (E2E-EXT-002). - Все изменения не закоммичены (working-tree note предыдущих чекпоинтов в силе).
Checkpoint — 2026-09-07 (Phase 2d EXECUTED: T029i+T029k+T029l реализованы; первый полный suite после батчей)
T029i — MCP AgentRun surface (E2E-EXT-001 CLOSED)
- Новый
backend/src/mcp_server/tools_agent_run.py(155 LOC,McpServer.ToolsAgentRun):create_agent_run(wrapsServices.AgentRuns.Service.Create; server-pinned objectType/route/contextVersion/intent; caller владеет dashboard_id/environment_id/dashboard_name/idempotency_key → conversation-reuse; продуктовая форма run: RUNNING +run_startedevent; bounded envelope) иget_agent_run(ownership-scoped проекция + draft_count; чужой/неизвестный → typed not_found, ноль existence-oracle). - Каталог
2.1.0 → 2.2.0: +2 записи (("dashboard:testing","EXECUTE")/("dashboard:testing","READ"),service_allowed=False); seam в_build_probe_server(automation → agent-run → scenario); PINNED_CATALOG_MAJOR=2 не тронут; RBAC-зеркала (rbac_visibility/admin_mcp_catalog) правок не потребовали (derived-наборы; analyst без EXECUTE/READ). - Тесты: новый
test_mcp_agent_run_tools.py(5); catalog-pintest_mcp_server.pyрасширен;test_mcp_initial_scenario_e2e.pyконвертирован в fully external chain — AgentRun минтится вызовомcreate_agent_runчерезtools/call(ноль не-MCP seeding,_pack()удалён); strict-xfail pin вtest_mcp_agent_run_reachability.pyсработал как спроектирован — снят и закалён в hard requirement-тест; denial-pin (typed DRAFT_PACK_ACCESS_DENIED, zero side effects) сохранён.
T029k — context-derived capability authority (CAP-001 CLOSED)
- Новый
backend/src/services/dashboard_testing/scenario/capability_authority.py(236 LOC,ScenarioGraph.CapabilityAuthority):derive_capabilities— facts-only (native_filters; text_filter←STRING; time_rollover←DATE/TIME/TIME_GRAIN; table_filter+pagination←executable table-viz; xlsx_export←capabilities; dataset_field_read+has_dataset_fields←accessible dataset c колонками; browser←T040 readiness-снимок, иначе undetermined);NEVER_DERIVED(row_edit/bulk_edit/persistence_refresh/safe_test_data/safe_clock_fixture/ cross_dashboard) — unsafe-mutation автоматизация метаданными невозможна (@INVARIANT);merge_capabilities— derived-wins в обе стороны + overrides-аудит; legacy-payload → caller-declared passthrough; единый choke pointbuild_capability_authority. - Wiring: MCP
inspect_scenario(+capability_authorityсекция),inspect_dashboard_context(+derived_capabilities), RESTapi_compile_scenario(parity, +секция). - Тесты: фикстура
query_model_sales.json; unit truth-table + CAP-001 classification-fix черезmap_all(sales-shape декларации → B01–B04/T01–T03 automated, C04–C06 unsupported, B05–B09/C01–C03/C02/C07 легитимно human_checkpoint) — 7; tool-level 3; REST-parity pin вtest_scenario_routes.py. Readiness-seam запинен monkeypatch (детерминизм против глобального composition-root в full-suite).
T029l — disposition clarity (DISP-001 CLOSED)
- RU/EN: «Подтвердить соответствие»/"Confirm conformance", «Проблема не подтверждена»/"Issue not confirmed",
«Недостаточно данных»/"Insufficient data"; confirm-кнопки
bg-destructive→bg-primary(HumanCheckpointPanel + WaitingForMeView, @RATIONALE); MCPdecide_checkpointdocstring несёт immutable outcome-mapping; API-вокабуляр и lifecycle mapping не менялись; vitest-пины обновлены + новый DISP-001 style/label/dispatch pin (RunMonitorViews), WaitingForMeView.test.ts, run.ux.test.ts.
Полный suite впервые после батчей 4d5ef6be/58c5ae39 — вскрыты и закрыты 2 PRE-EXISTING регрессии HEAD
Первый full-run после этих коммитов дал 15 failed — все НЕ из кода этого раунда (app.py/test_app_lifespan.py/
mcp_oauth.py/test_mcp_client_flow_http.py в working tree не модифицированы; git show доказал самопротиворечие HEAD):
- SC-007 resource pin: код и twin-pin
test_mcp_server.py:81канонизировали post-redirect идентификатор/mcp/(Mount 307/mcp→/mcp/;4d5ef6beизменил код,58c5ae39выровнял twin), но stale-пинtest_mcp_client_flow_http.py:123endswith("/mcp")не обновлён → обновлён на/mcp/с комментарием (эмпирическая поддержка: живой внешний клиент 2026-09-07 прошёл discovery против этого значения). - Lifespan re-entry (13 тестов):
src.app.mcp_transport_app— модульный синглтон; SDKStreamableHTTPSessionManager.run()once-per-instance → второй вход в lifespan внутри pytest-процесса = RuntimeError. Продакшн стартует lifespan один раз на процесс — синглтон корректен; фикс test-side: autouse fixture_fresh_mcp_transport_app(свежийcreate_mcp_asgi_app()на тест, оригинал восстанавливается). - Capability-тест: глобальный readiness-снимок (заполняется lifespan-тестами) влиял на derivation browser —
seam-patch (выше).
Группа
test_app_lifespan + client_flow + capabilityпосле фиксов: 26 passed (было 15 failed/11 passed).
Verification (этот раунд)
- Full backend suite:
11357 passed, 243 skipped, 1 xpassed, 0 failed(7:00) — первый зелёный полный прогон после батчей. targeted-срезы: 8-файловый MCP+catalog70 passed; capability-срез67 passed. - Frontend: vitest
3507 passed(206 файлов), lint0 errors / 364 warnings(baseline),npm run buildOK. - ruff + compileall чистые; anchors balanced (все 13 touched backend-файлов + svelte-компоненты);
git diff --checkчистый. - INV_7 хвосты (зафиксированы, не ухудшать дальше; новый код уведён в новые модули 155/236 LOC):
mcp_server/tools_scenario.py508 LOC (пре-existing >400 после T029f/h; +28 этот раунд) иapi/routes/dashboard_testing/scenario.py442 — кандидаты на декомпозицию следующим gate-раундом (например, вынос InspectDashboardContext+register_draft_pack в tools_pipeline seam / разбивка route-модуля).
Статус
- 050 Phase 2d: T029i
[x], T029j[x](обновлён note о срабатывании xfail-ритуала), T029k[x], T029l[x]; release-gate rows:E2E-EXT-001/CAP-001/DISP-001/TEST-001CLOSED; единственная OPEN —E2E-EXT-002(T029m, live-stand replay полевого sales-сценария). Внешняя initial-bootstrap цепочка достижима через MCP alone. - Учёт обновлён: 050 spec.md/tasks.md, 038/044/045 amendments → IMPLEMENTED, ADR-0024 → IMPLEMENTED (+closure consequence), README-реестр, follow-up секция docs-отчёта.
- Commit (2026-09-07): всё вышеперечисленное этого и plan-раунда (ADR-0024 + спеки + T029i/k/l код/тесты +
фикс двух pre-existing HEAD-регрессий + docs-отчёт полевого прогона) закоммичено единым коммитом
feat(mcp): Phase 2d field-run remediation — ADR-0024 agent-run surface, derived capabilities, disposition clarity(hash — см.git log). Намеренно НЕ включены (не этого workstream'а, остаются в рабочем дереве): translate/migration integration-тесты и_job_routes.py(чужой незакоммиченный workstream),.kilo/agent-manager.json(live UI/recovery state), корневой снапшотspecs-036-050-20260907-111314.md. T029m (E2E-EXT-002) — единственная OPEN gate-row.
Checkpoint — 2026-09-16 (recovered re-baseline 2026-09-07 → HEAD 4746af2f)
Восстановленный чекпоинт: WORKSTATE не обновлялся 9 дней, факты ниже собраны из git log (45 коммитов) и handoff-отчётов (
docs/reports/agentic-runtime-*.md,ux10-*.md,docs/2026-09-11-sales-prod-mcp-replay.md), а не из памяти. Suite-числа этого чекпоинта — свежие прогоны на HEAD; исторические числа помечены датой источника.
2026-09-08..10 — offline agentic runtime chain + INV_7 + live canaries v1/v2
83727aa7complete offline agentic runtime chain; INV_7-декомпозиция execution-пакета:runner.py1207→121 (фасад; 6 модулей),lifecycle.py611→299,executors.py508→257; все модули < 400 LOC (a21481ea,8f057f30,3e5cc942).- D1 published-catalog source (fail-closed, 9 тестов) + D2 server-side live-binding resolution
из
settings.scenario_live_execution_bindings(12 тестов) — клиент binding не передаёт. - Live canary v1 (run
597274d3: browseropen_dashboard+ 8 durable screenshots) и v2 (REST-only, run4eebfab3: server-side binding → PROD gate → scheduler dispatch → browser → реальный LLM →AgentEvaluationpersisted → DecisionPolicy row 11). Три production-дефекта найдены и закрыты живыми прогонами (preflight deadlockba2f1f45,artifact_byte_lengths, нормализация provider-ответов3458343d). - Suite (2026-09-10, источник: evening handoff): 11527 passed / 0 failed / 244 skipped.
2026-09-11 — T029m live replay + T045 publish + 046 scheduled happy path
- E2E-EXT-002 CLOSED:
live_mcp_replay.pyпрогнал полную внешнюю цепочку на ss-prod (run110a6517…: inspect → create_agent_run → compile B01 → registercontext_authority=verified→ bootstrap → PROD gatepending_approval(идемпотентный retry = один durable gate) → approval → livecapture_screenshotpassed (8 refs) → typedBROWSER_ACTION_NOT_SUPPORTED→ честныйinconclusive); human loop живьём черезlive_mcp_human_loop.py(B05 HumanCheckpoint →waiting_human→list_checkpoints→decide_checkpoint confirm(CAS) → terminalpassed). Два fail-closed дефекта найдены и исправлены. Trace:docs/2026-09-11-sales-prod-mcp-replay.md. - 050 T045 CLOSED: gated MCP
publish_baseline_catalog+ 037 publication worker (CAS, receipts, reconcile), REST parity; live pin-from-Gitea proven (canary v4: pin stamped intoAgentEvaluation.baseline_pin, strict equality). Каталог 2.2.0 → 2.3.0. - 046 T018 CLOSED live: scheduled runs execute as schedule-owner principal; deterministic
scheduled-run idempotency key; typed scheduled-reject observability (
d0466a2c,b4148cfe). - 044 T043/T046: live baseline pin PROVEN; residuals — graph-level deterministic comparison PASS
(T043) и graph-level terminal PASS (T046, baseline-semantic policy
EVALUATION_UNAVAILABLEдля compare без bound evaluation).
2026-09-12..14 — UX-10 Wave C (закоммичено как 8522a2ee 2026-09-15)
- DG-1 реализован: run-scoped browser session (
browser_session.py501), 9 read-only действий (browser_readonly_actions.py534),apply_native_filter+ safe checkpoint (browser_native_filter.py275), admission split (browser_admission.py208). - 046 retention:
deletions.py(331) + расширениеretention.py+ миграция0024_retention_deletions(PG-verified; после инцидента с 33-char revision id добавлен guardtest_revision_ids_fit_varchar32; правило: revision id ≤ 32, цепочку проверять на PostgreSQL). - Frontend:
TerminalReasonBanner+ typed terminal-reasons;docs/mcp-client-setup.md(UX-4). - 044 T040
[x](deployment-evidence rule); T042 offline contract suites complete — остаток: live Superset query + явныйRESULT_TOO_LARGEbound. - Suite (2026-09-14, источник: ux10 handoff, HEAD
b4148cfe+ dirty tree): backend 11796 passed / 264 skipped / 1 xpassed; frontend 3585 passed / 214 файлов; alembic head0024_retention_deletions.
2026-09-15..16
8522a2eeзакоммитил дерево UX-10 (94 файла, +11014/−589): MCP automation parity tests, terminal reasons, scenario UX flow.3de0756dудалил legacy LLM dashboard validation (routes/service/schemas/ DashboardValidationPlugin; миграция0025_drop_legacy_validation; frontend validation models/routes/i18n удалены; settings/automation редиректит на scenario automation).4746af2fhealth consolidation review fixes.
Свежие прогоны на HEAD 4746af2f (этот чекпоинт)
- Backend full suite: 11269 passed, 252 skipped, 1 xpassed, 0 failed (7:42). Дельта против
11796 — удалённые legacy-validation тесты (
3de0756d), не регрессии. - Frontend: 3359 passed / 208 файлов, 0 failed (49s). Дельта против 3585 — удалённые
validation frontend-тесты.
npm run lint→ 0 errors / 333 warnings (baseline 340–364);npm run build(adapter-static) → green. ruff check .clean;compileall -q srcclean;alembic heads→ единственная голова0025_drop_legacy_validation.- Axiom rebuild (incremental, live): 10870 contracts / 5365 edges; unresolved relations
393 (было 401 на re-review 2026-09-04); исправлены 3 malformed multi-target
@RELATIONвbackend/src/core/cot_logger.py(split на individual lines с verified targets). - MCP slice (T014 re-verification): 84 passed; каталог 61 запись,
MCP_CATALOG_VERSION=2.3.0. - 037 visual slice (T044 re-verification): 61 passed.
Реконсиляция учёта (этот чекпоинт)
- 050 tasks.md: T014, T016, T042 →
[x](боксы с evidence-текстом, но не отмеченные; evidence перепроверено свежими прогонами/grep на HEAD). - 037 tasks.md: T044 →
[x](visual candidate реализован вcandidates.pyветкойkind == "visual"+ mandatory human disposition гейтом; closure note «All 47 tasks completed» теперь соответствует факту для T001–T047; T082–T084 — отдельные production rows 2026-09-08). - 050 quickstart.md: устаревшие OPEN-строки исправлены (T029m/E2E-EXT-002 CLOSED 2026-09-11; T045/T046 CLOSED; каталог 2.3.0/61).
- 050 CHK003 OPEN — не противоречие: строка покрывает и T046 (CLOSED), и frontend-boundary (T030 OPEN); оставлено OPEN корректно.
- 044 T042b остаётся OPEN осознанно: канарейки 2026-09-01 предшествуют Wave-C session
architecture (
8522a2ee) и не являются evidence для текущего кода.
Текущий фронт работ (актуализировано на HEAD)
- 044: T042b PREPROD-канарейки под Wave-C архитектуру; T044 live ACL/status/header canary;
T045 fault-injection canary; T046 graph-level terminal PASS; T043 graph-level comparison PASS;
T022 полный аудит; T042 live Superset query +
RESULT_TOO_LARGEbound. - 046: T017 lifecycle notifications (callers отсутствуют); T021 5/15/50-tab canaries; T013e бокс vs landed retention code (сверить); T013 формально открыт при покрытии.
- 050: T044 (REST-vs-MCP error-shape parity fixtures — последний OPEN production gate); CHK004 negative UI test; T029a; T030/T032 — продуктовое решение.
- 047: T017–T021 (SCAN-FR-015 atomic triage). 042: T024–T032. 043: T015/T023, T021–T022. 045: T020–T022. 037: T082–T084. 038: T041 (VLM), T060–T062.
- Skill drift:
.agents/skills/semantics-pythonиsemantics-svelteссылались на удалённыйss_tools.shared.cot_logger(ADR-0022 absorb) — исправлено в этом чекпоинте: facadesrc.core.logger(intent-first),notify()из$lib/toasts.svelte.tsвместо несуществующегоaddToast(), i18n dictionary-proxy стиль ($derived($t.migration ?? {})) вместо function-call$t("key");./scripts/sync-skills.shпрогнан,.kiloкопии byte-identical.
Product status: NO-GO — production-acceptance rows 2026-09-08 открыты во всех спеках; формальный closure re-sign-off не проводился.
Checkpoint — 2026-09-16 (T042b CLOSED: шесть canary-векторов GREEN под Wave-C архитектурой)
- Все шесть векторов T042b перепроверены живьём на owner-authorized тестовом стенде
(
https://ss-prod.bebesh.ru,SS_STAND_STAGE=PREPROD, dashboard 11) через реальную продуктовую цепочку Wave-C (run-scoped browser session, capacity admission, receipts, durable evidence). Evidence вspecs/044-dashboard-scenario-execution/evidence/browser-provider/(readonly-canary-20260916T*.json+ PNG; креды в evidence не пишутся):Вектор Результат Evidence read-only actions 3/3 passed (open_dashboard 7.2s, wait_for_state 3.9s, refresh 4.3s) 150158Zforced timeout/cleanup typed BROWSER_ACTION_TIMEOUT, 0 refs, lease released150407Zmutation row_edit+restore 82.74→mutate→restore verified, 2/2 receipts completed150518Zreconciliation sweep stale receipt resolved через live SELECT-only observation 150754Zsafe-checkpoint reconstruction attempt 2 passed,reconstruction_replay=true150908Zscheduler soak 75s 3/3 runs passed,attempts==1(CAS), 0 active leases, 3 artifacts155013Z - Harness defect найден и исправлен: в bare-script контексте DI-синглтон
SchedulerServiceзахватывал мёртвый event loop (asyncio.new_event_loop()без runner), из-за чего async-job'ы (maintenance_auto_end→AsyncJobRunner.run) блокировались на 300s safety cap, аstop()ждал их — первый soak-прогон завис. Фикс вspecs/044-dashboard-scenario-execution/prototype/browser_readonly_canary.py: soak-окно выполняется внутриasyncio.run(singleton создаётся при работающем loop),scheduler.stop()вынесен черезasyncio.to_thread(loop свободен для drain in-flight coroutines). Продуктовый код не менялся; в production loop FastAPI всегда работает, дефект — артефакт harness'а, но он же демонстрирует задокументированный drain-first компромисс (job занимает слот до cap). - PG-канареечная БД
canary_044создана в локальном postgres (recovery/soak/reconcile-режимы); read-only/timeout/mutation используют temp SQLite по умолчанию. - 044 tasks.md: T042b →
[x]с полным evidence.
Checkpoint — 2026-09-16 (T022 audit: [~] — gates зелёные, открыт INV_7 хвост Wave-C)
- T022 команды выполнены (см. текст задачи): scoped 044 suite 1590 passed; full backend
11269 / 0 failed; frontend 3359 passed + lint 0 errors / 333 warnings + build green;
свежий PostgreSQL 16
0001→0025_drop_legacy_validation+alembic checkбез drift (проверочная БДalembic_044); Axiom live rebuild — 10870 contracts / 5367 edges, unresolved 393. - НЕ
[x]: Axiom-аудитexecution/показал 8 structural warnings — Wave-C модули выросли за INV_7:providers/browser_readonly_actions.py534,providers/browser_session.py501,providers/browser.py407 (регрессия против 398 из round 4). Требуется decomposition-gate по прецеденту server.py (specs/050-mcp-interface/plans/ server-decomposition-gate.md): binding plan + behavior-neutral split с frozen contract IDs. - T044 (044): offline evidence matrix уже зелёная —
test_scenario_artifact_content_api.py11/11 (GET/HEAD same-headers, 401/403/404-cross-owner, digest/mime 409 без байт, inactive 410, range 416, oversized 413, storage 503); остаток — live ACL/status/header canary на стенде. - Skill drift закрыт (semantics-python/svelte → facade
src.core.logger,notify()из$lib/toasts.svelte.ts, dictionary-proxy i18n); sync-skills прогнан; malformed multi-target@RELATIONвcot_logger.pyразбиты (3 линии) — cot_logger.py 0 unresolved. - Ruff clean на всём backend; compileall clean; anchors сбалансированы во всех touched-файлах
(pre-existing внутри-комментария
#regionв semantics-svelte line 86 — текст, не якорь).
Checkpoint — 2026-09-17 (Wave 1.6 EXECUTED: provider decomposition gate → T042b+T022 закрыты)
- Wave 1.6 исполнен по binding-плану
specs/044-dashboard-scenario-execution/plans/ provider-decomposition-gate.md(статус PLAN → EXECUTED, полный execution log внутри). Три gated фазы, нулевой behavior diff (per-phase scoped 1590 + финальный полный suite):Модуль Было Стало Новые sibling-модули browser_readonly_actions.py534 160 browser_readonly_limits.py82,flows_nav199,flows_interact192browser_session.py501 58 browser_session_handle.py64,browser_session_checkpoint.py162,browser_session_managers.py57,browser_session_registry.py287browser.py407 398 browser_factory_helpers.py35 - Frozen contract IDs + import surface сохранены: фасады ре-экспортируют все перенесённые
публичные имена; monkeypatch-сеамы последовали за владеющими модулями (
_MAX_EXTRACT_OUTPUT_BYTES→ flows_nav,_MAX_DOWNLOAD_BYTES→ flows_interact);_register_managerвынесен вbrowser_session_managers.py(иначе registry⇄facade цикл). Все переносы verbatim. - Гейты: per-phase scoped 1590 passed ×3; финальный полный backend suite 11269 passed / 252 skipped / 1 xpassed / 0 failed (7:10) — нулевая дельта против базовой 11269 до декомпозиции; ruff + compileall clean; anchors сбалансированы во всех 9 модулях.
- Post-decomposition Axiom
audit_contracts(execution/providers): 0 module_too_long (было 3). Принятые advisory: 2 ×contract_too_long(BrowserProvider.Factory305,ScreenshotProvider.Factory200) — typed-лестницы исключений оставлены inline осознанно. - 044 tasks.md: T042b
[x](2026-09-16, шесть векторов GREEN) и T022[x](2026-09-17, zero P0/P1). Оба production-критичных бокса 044 закрыты с evidence. - Skill drift закрыт окончательно: заголовок примера semantics-svelte ссылался на удалённый
notificationStore— заменён на$lib/toasts; sync-skills прогнан.
Checkpoint — 2026-09-17 (Wave 1.2/1.3 закрыты; T046 диагностирован с исполняемым proof)
T044 CLOSED [x] — live ACL/status/header canary (11/11 GREEN)
- Новый harness
specs/044-dashboard-scenario-execution/prototype/artifact_content_canary.py(351 LOC, 5 anchors) гоняет РЕАЛЬНОЕ FastAPI-приложение (без dependency overrides) поверх реального PostgreSQL (canary_044) с реальными JWT (реальныйcreate_access_token, реальный поиск user→role→permission; безsid— session-policy пропускается) и РЕАЛЬНЫМИ live-байтами браузерного evidence (soak-прогонsoak-canary-c071fb98, PNG 104746 B, sha22b4b8c1…). - Evidence:
evidence/artifact-content/artifact-content-canary-20260917T085438Z.json— 11/11: anonymous→401AUTHENTICATION_REQUIRED(+HEAD пустой); viewer GET→200 byte-exact + согласованные Content-Length/ETag/Disposition/Cache-Control/nosniff/Accept-Ranges; viewer HEAD→200 полный header-parity + пустое тело; no-VIEW→403PERMISSION_DENIED; unknown run / unknown artifact / foreign-owned → неразличимые 404NOT_FOUNDс идентичным message; expired→410; corrupt sha→409ARTIFACT_INTEGRITY_FAILED; declared-MIME off-allowlist→409; oversized→413; missing bytes→409ARTIFACT_MISSING; Range→416RANGE_NOT_SUPPORTED+Accept-Ranges: none. - Harness-инфраструктура: добавлен knob
SS_CANARY_STORAGE_ROOT(персистентный evidence-root; ВАЖНО: выделенный root на прогон — soak-ассерт считает файлы в root). Для HTTP-канарейки нуженSTORAGE_ROOT_PATHиз одобренных корней (/app/storage, проект,../ss-tools-storage) — использован/home/busya/dev/ss-tools-storage/canary-artifact-acl. - Инциденты harness'а (исправлены): (1) повторный прогон брал собственный oversized-clone как
«good»-row → фикстуры-клоны теперь именуются
canary-*и удаляются в начале прогона; (2) soak с переиспользованным root далfailures: ["artifacts=6"]→ выделенный root.
T045 CLOSED [x] — fault-injection canary (16/16 GREEN)
- Новый harness
specs/044-dashboard-scenario-execution/prototype/fault_injection_canary.py(365 LOC, 7 anchors) поверх реального PostgreSQL (canary_044): реальный provider runtime, capacity manager, receipt CAS, cancel lifecycle. Evidence:evidence/fault-injection/fault-injection-canary-20260917T105541Z.json. - Векторы: submit до старта/истёкший deadline → типизированные
PROVIDER_LOOP_NOT_RUNNING/PROVIDER_SUBMIT_DEADLINEза bounded время; медленная корутина отменяется по caller-deadline; shutdown во время in-flight работы разворачивает вызывающего по его же deadline, loop → not-running (без зависания); crash (lease не освобождён) карантинит слот (второй claim →CAPACITY_UNAVAILABLE) доreconcile_expired_leases, который освобождает ровно эти units; unknown effect → receiptreconciliation_required/unknown→reconcile_provider_operationразрешает; late response добавляет только history (терминальный статус неизменен), reconcile по терминальному receipt отвергнут (PROVIDER_OPERATION_TERMINAL); cancel открывает drain-окно (draining), дедлайн-финализатор терминализирует (0 running steps, 0 unexpired worker leases), capacity-lease отменённого прогона согласуется по TTL, не исчезает молча; немедленный cancel терминализирует одним проходом. - Найдено при отладке:
cancel_runэкспайрит worker-leases (ScenarioStepLease), а не provider capacity-leases (CapacityLease) — последние освобождает provider вfinallyили TTL-reconcile (совпадает с семантикой «released or reconciled» из SC-008). Первый вариант ассерта это смешивал; исправлено. - Offline-референсы (зелёные):
test_provider_operations.py(CAS/late/reconcile),test_scenario_cancel_timeout.py(6),test_dispatch_capacity_lifecycle.py,test_provider_capacity.py,test_provider_contract.py,test_provider_preflight.py,test_provider_runtime.py— 41 + 47 passed.
T046 [~] — диагноз с исполняемым proof (binding, не truth table)
- Остаток T046 — walker-level binding, а не ошибка политики:
walker.pyсчитаетdecide_step_outcome(policy_inputs_from_outcome(...))ПОШАГОВО, иevaluationберётся только изstep_outcome["evaluation_input"], который эмитит единственный шагagent_evaluation(evaluation_adapter.py:379). Шагassertion compare_to_baselineвсегда видитevaluation=None;derive_runner_planвключает mandatory-режим при наличииagent_evaluationшага → каждый нормативный шаг (включая compare) резолвится вEVALUATION_UNAVAILABLE. - Proof на HEAD (pure function, без БД): (A) compare-only + mandatory →
inconclusive ["EVALUATION_UNAVAILABLE"]; (B) тот же + disabled →passed ["BASELINE_PASS"]; (C) тот же + boundevaluation_input(succeeded/pass/0.9) + mandatory →passed ["BASELINE_AND_SEMANTIC_PASS"]. - Канон (
contracts/production-chain.md§8) ставит optional evaluation МЕЖДУ comparison и pinned DecisionPolicy → решение обязано видеть оба; сейчас оно считается изолированно по шагу. - Опции закрытия (следующий пакет, truth table не меняется): (1) dependency-ordered binding — compare
зависит от покрывающего
evaluate-visual, walker инжектит persistedAgentEvaluation(agent_evaluation_ids) какevaluation_inputдля зависимого решения; (2) deferred decision — walker откладывает решение comparison-шагов до персиста покрывающей оценки и считает одно решение на агрегированных входах. Закрытие = live-рерun v4-графа с терминальнымpassed.
Открытые строки 044 после этого чекпоинта
- T043
[ ]: residual — graph-level deterministic comparison PASS (050 T045 publish закрыт). - T046
[~]: binding (см. выше) + live terminal PASS. - Ранее закрыты: T042b
[x], T044[x], T045[x], T022[x], T040[x].
Checkpoint — 2026-09-17 (DESIGN AMENDMENTS: complex scenarios, D/M/E, zero-human, tiering, investigation loop)
Спецификационный пакет (только спеки, implemented=false): пять Design Amendments из дизайн-сессий 2026-09-17 (кастомные Playwright-сценарии → типизированный реестр; R1 «метрики — только baseline»; R2 «zero-human runtime»; hot/cold evidence tiering; аудит интерфейса расследования). Код не менялся; все новые требования закрываются только executable evidence.
Внесённые amendments
| Спека | Amendment | Суть |
|---|---|---|
| 038 spec.md | Design Amendment 2026-09-17 | AGSCN-FR-016..025: R1 (числа только через baseline, валидатор-пин embedded literals), R2 (запрет эмиссии human-шагов, LLM не авторизует небезопасное, co-authoring промптов E-класса, DecisionPolicy hardening), таксономия D/M/E, set_step_inputs op, новые read-only действия (assert_dom, inspect_filter_options, navigate_tabs, wait_for_selector; входы search_text/mode/date у фильтров; click→URL-исход), дисциплина registered≠implemented, @REJECTED code-backed provider (не тронут) |
| 038 contracts/browser-actions.md | 2026-09-17 блок | additive-строки 038.5.0 design; disabled-дисциплина; navigate_dashboard = disabled: pending driver |
| 037 spec.md | Design Amendment 2026-09-17 | AGBASE-FR-015..018: двухъярусная политика (курируемый gating ≤50 записей + observatory non-gating с запретом→PASS), BaselineSelectionProposal ко-авторинг (агент предлагает facts-only координаты, оператор курирует), scenario-derived кандидаты (kind=scenario_transform, тот же lifecycle/immutability), staleness→кандидат не автосмена; visual-baseline lossless-исключение из tiering |
| 044 spec.md | Design Amendment 2026-09-17 | SCEX-FR-031..036: evidence tiering (immutable ref ≠ физический тир, chain-of-custody transcode-receipt, hot_runs=2 + holds: открытые кейсы/pending-кандидаты, форматная политика webp:lossless/lossy + XLSX→Parquet Phase 2), LLM-as-judge hardening (evidence-as-data, verdict-схема, zero tool access, LLM-INJ-001), zero-human dispatch (scheduled без касаний, LLM не авторизует PROD-мутации/гейты), registered≠implemented; гейт-строки BSC-FILT/ASSERT/METRIC/ZEROHUMAN/LLMSTAB/EVID-TIER/LLM-INJ-001 |
| 045 spec.md | Design Amendment 2026-09-17 | RUNMON-FR-015..017: waiting-me сужается до launch-гейтов, чекпоинт-панели deprecated для новых ревизий (legacy работают), tiering-совместимость EvidenceViewer, scheduled read-only наблюдение |
| 046 spec.md | Design Amendment 2026-09-17 | SCAUTO-FR-020..024: hot_runs retention + holds, zero-human schedule eligibility, уведомления wiring (NOTIFY-001), DLQ/quarantine + 72h SOAK (SCHED-SOAK-001), capacity-SLO модель (CAP-SLO-001) |
| 047 spec.md + tasks.md | Design Amendment 2026-09-17 + T022–T027 | SCAN-FR-016..022: MCP investigation-контур (list_queue/get_case read, record_case_note, propose_disposition pull-only — решение human CAS; service-principal decide denied), case index/findability (in_case видимы, human-readable заголовки), workspace-рендер linked runs/evidence/tier-маркеры, уведомления; гейты INV-MCP/FIND/ATOMIC/NOTIFY-001; tasks T022–T027 |
| 050 spec.md | Design Amendment 2026-09-17 | MCPX-FR-031..034: investigation tools в каталоге, authoring input ops (set_step_inputs/set_step_evaluation + parity), опциональный exploration_step (Phase 2), prompt dangerous-content профиль; spec-impact 047 строка обновлена |
Ключевые решения сессии (зафиксированы в amendments)
- R1: числовые бизнес-истины не встраиваются в шаги — только transform→baseline→approve→compare (AGSCN-FR-016); таксономия D/M/E (FR-017); TransformSpec — деривация, не проверка (FR-018).
- R2: ноль human в рантайме — эмиссия human-шагов запрещена (FR-019), LLM заменяет суждение, не авторизацию (FR-020: unsafe-мутации без fixture = unsupported; PROD-гейты человеческие), промпты E-класса co-authored и версионируются в content_hash (FR-021), DecisionPolicy pinned (FR-022).
- Tiering: hot=2 последних прогона оригиналы, дальше WebP-архив с chain-of-custody; baseline-сеты lossless-исключение; holds на открытые кейсы/pending-кандидаты.
- Baseline: курируемый набор (агент предлагает facts-only, оператор аппрувит) + observatory-ярус non-gating; scenario-derived кандидаты через transform-выход.
- Investigation: агентская петля через MCP (read/note/propose), решение human; findability (case index, in_case-видимость, human-readable заголовки).
Открытые точки для clarify (до имплементации)
- Семантика
clearдля multi-select фильтров (один/все чипы) и форматsearch_text-входа. - Политика E-вердиктов: single-shot vs best-of-N (порог консистентности).
- Пороговые значения
element_count/bounding_boxкак входы D-ассертов (формат pin в ревизии). - Жизненный цикл baseline_set при смене отчётного периода (подтверждение UX staleness→кандидат).
- XLSX→Parquet Phase 2: критерий включения (доля XLSX в профиле хранилища).
Связность
- Все amendments additive к существующим контрактам; ни один upstream
@REJECTEDне переопределён (code-backed provider остаётся future/unimplemented; per-step изоляция DG-1 остаётся rejected). - Гейт-строки добавлены в spec-файлы; WORKSTATE-статус продукта не менялся: NO-GO pending closure re-sign-off — новый пакет добавляет OPEN-требования, не закрывает существующие.
- Следующий шаг по пакету: speckit-цикл уточнения (clarify по 5 открытым точкам) → plan → tasks per воркстрим (WS-1 реестр/входы; WS-2 фильтры; WS-3 D-ассерты; WS-4 M-путь+037; WS-5 E-класс; WS-6 navigate_dashboard; tiering-воркстрим S-M).
Checkpoint — 2026-09-18 (Wave 1 implementation: D/M/E foundation merged)
Волновая имплементация Design Amendments 2026-09-17. Все четыре ветки независимо верифицированы (scoped-сьюты + ruff + compileall) и слиты в master без NO-GO-статус-изменений продукта: сетевые гейты живых стендов (BSC-FILT-001 live canary, EVID-TIER-001 и др.) остаются OPEN.
Слитые ветки и верификация
| Ветка | Коммиты | Скоуп | Тесты |
|---|---|---|---|
| feat/038-step-inputs | 534d488f, f123210b → merge e495f9db |
op set_step_inputs (AGSCN-FR-023), per-action валидация step_inputs.py, валидатор EMBEDDED_METRIC_LITERAL (FR-016), MCP parity; контракт-фикс: ScenarioStep.inputs остался list[Ref], аргументы — в новом optional action_inputs (канонические байты старых ревизий неизменны — пин-тест) |
602+12 green |
| feat/047-investigation-mcp | 23414f9a, 5eea72bc → merge c0224486 |
MCP-инструменты list_investigation_queue / get_investigation_case / record_case_note / propose_case_disposition (SCAN-FR-016..019); queue-проекция queued+in_case с human-readable заголовками; case: linked_run_ids + evidence_refs; каталог 2.4.0; disposition строго resolved | accepted; service-principal denied на writes |
| feat/047-investigation-ui | 6fcc8616, e7925d0e → merge 80980cee |
Queue findability (title, severity/scenario-фильтры, пагинация, in_case-бейдж), case-index маршрут /analytics/cases (SCAN-FR-020), workspace linked runs/evidence + i18n timeline (FR-021), RU/EN i18n | 3360 vitest green, lint, build |
| feat/038-registry-filters | d29959f6, 683347be → merge 5d86da54 |
Реестр 038.5.0: + disabled-строки assert_dom/inspect_filter_options/navigate_tabs/wait_for_selector; pagination/navigate_dashboard → disabled: pending driver; ACTION_DISABLED валидация до I/O (FR-025, негатив-тесты); apply_native_filter: search_text (B02) / mode=set | clear / date (DatePicker) + CLEAR_CONFLICT-гард (FR-024); clear фолдится в checkpoint/replay (_FILTER_REPLAY_KEYS); click → page_url/popup_url с закрытием попапа до page-bound |
Пост-merge master: 869 passed (registry+scenario+MCP+analytics+catalog+authoring) + ruff green.
Инцидент контаминации (2026-09-18)
Registry-агент волны 1 писал в главный репозиторий (master) вместо своего worktree:
templates/__init__.py (реестр 038.5.0) + README.md + untracked docs/dashboard-testing-guide.md.
Патч перенесён в worktree ветки и слит штатно через merge; master очищен. Урок: worktree-сессии
Agent Manager могут мутировать CWD главного репозитория при ошибке cd — контролировать
git status главного репо после каждой остановки агента. Судьба dashboard-testing-guide.md
(scope drift, вне ТЗ) — на решение оператора.
Остатки волны 2 (OPEN, не имплементировано)
Драйверы disabled-действий: assert_dom, inspect_filter_options, navigate_tabs, wait_for_selectorЗАКРЫТО 2026-09-184274e403: драйверы реализованы вbrowser_readonly_flows_observe.py(fail-closed, bounded, 8 unit-тестов), реестр включил их (disabled остались только pagination/navigate_dashboard), admission/transport/facade подключены, чекпоинты dom_asserted/filter_options_inspected/tabs_swept/selector_awaited. Живые канарейки на стенде (гейт BSC-ASSERT-001 live) — остаются OPEN до стенда.set_step_evaluationop + E-класс DecisionPolicy hardening (AGSCN-FR-021/022, WS-5).- 037 scenario-derived baseline candidates + BaselineSelectionProposal (AGBASE-FR-016/017, WS-4).
- Evidence hot/cold tiering (SCEX-FR-031..034) + 046 hot_runs retention.
- navigate_dashboard драйвер (cross-dashboard сверки).
Zero-human mapper: прекращение эмиссии human_checkpointЗАКРЫТО 2026-09-18a59f5e3c: capability_mapper больше не эмитит human_checkpoint; former-human → unsupported (unsafe mutations — никогда LLM-авторизованы, FR-020); компилятор не создаёт шагов для unsupported-кейсов; 1577 тестов green. Runtime-механика HumanCheckpoint сохранена для legacy-ревизий (deprecated). E-класс (evaluate_declared_spec) остаётся declared-only до пункта 2.- Уведомления wiring (SCAUTO-FR-022), DLQ/SOAK (FR-023/024), MCPX-FR-033 exploration_step.
Checkpoint — 2026-09-18 (LIVE SUPSERSET CANARIES: wave-2 observe + mutation restore GREEN)
Пользователь авторизовал тестовый Superset-контур
https://ss-prod.bebesh.ruкак PREPROD для read-only и mutation canary (credentials provided out-of-band; секреты не коммитятся). Dashboard 11 (Sales Dashboard) использован как структурный fixture; все evidence-retention файлы сохранены подspecs/044-dashboard-scenario-execution/evidence/browser-provider/.
Live Superset readiness
/health→ 200;/api/v1/security/login(db) → 200; dashboard catalog count=11.- Browser login через direct form fallback → authenticated redirect
/; dashboard 11 открыт на/superset/dashboard/11/?standalone=true&native_filters_key=.... - Базовая canary (
SS_CANARY_MODE=actions): open_dashboard / wait_for_state / refresh — 3/3 PASS; durable PNG refs/digests, capacity leases released, evidence JSONreadonly-canary-20260918T125144Z.json.
Wave-2 observe canary (SS_CANARY_MODE=wave2) — 4/4 PASS
| Action | Live факт | Checkpoint/evidence |
|---|---|---|
| wait_for_selector | .dashboard-content visible |
selector_awaited; PNG sha f54ea056…; 104774 B |
| navigate_tabs | visited 2 tabs: 🎯 Sales Overview, 🧭 Exploratory |
tabs_swept; PNG sha 481032a4…; 49486 B |
| assert_dom | .dashboard-content, min_count=1, actual=1 |
dom_asserted; passed=true; PNG sha f95e4b30…; 104784 B |
| inspect_filter_options | Region: 19 live options (Australia…USA) |
filter_options_inspected; PNG sha f842ea03…; 110560 B |
Evidence: readonly-canary-20260918T125451Z.json + per-action PNG. failures=[].
BSC-ASSERT-001 live-part CLOSED для четырёх wave-2 драйверов; registry disabled flags сняты
в 4274e403 обоснованно (unit 811 + live 4/4).
Mutation canary (SS_CANARY_MODE=mutation, owner-authorized PREPROD) — 2/2 PASS
- Fixture:
video_game_sales, key (Wii Sports,Wii), fieldglobal_sales. - Original 82.74 → mutate expected 82.75, provider post_rows 82.75; SQL verify value 82.74 (stand SQL path observed a cache/isolation discrepancy after mutation; provider receipt/post_rows are authoritative operation evidence — needs follow-up); restore expected/verified 82.74.
- Receipts: два durable
completed/completed(2c8a5630…,f813cc03…); оба artifact PNG; cleanup checkpointfixture_restored; final value restored 82.74;failures=[]. - Evidence:
readonly-canary-20260918T125542Z.json+ mutate/restore PNG.
Live Superset residuals
- Playwright connection tasks emit
Task was destroyed but it is pending!after provider loop stop (the result/evidence is complete, but process teardown leaks Connection.run tasks) — follow-up cleanup test/fix required before closing soak quality. - Mutation canary mismatch: mutate provider post_rows=82.75 but immediate stand SQL verify=82.74; restore returns 82.74. Investigate cache/transaction/view path; do not classify false PASS solely from post_rows until independent SQL proof is reconciled.
Checkpoint — 2026-09-18 (P0 RUNTIME SAFETY WAVE: PROVIDER-SHUTDOWN + MUT-RECON CLOSED, BSC-FILT live PASS)
Волна 1 release-последовательности матрицы (PRODUCTION-ACCEPTANCE-MATRIX-043-050.md §3):
G-PROVIDER-SHUTDOWN и G-MUTATION-RECON CLOSED, G-BROWSER-FILTERS live-векторы PASS
(date-вектор BLOCKED внешне). Full backend suite 11332 passed; ruff/compileall/validate_static
green.
SCEX-FR-037 Provider shutdown (PROVIDER-SHUTDOWN-001)
browser_session_managers.close_all_sessions(reason): weak-registry fan-out, fail-closed.ProviderEventLoop.stop(): sessions close на живом loop ДО detach (submit()ещё видитself._loop— первичный вариант с detach-рано терял fan-out, поймано unit-тестом), затемloop.stop();_run_loopfinally: bounded drain (cancel +asyncio.wait(timeout=5s)), escapee-логPROVIDER_LOOP_DRAIN_INCOMPLETE, только затемloop.close().- Unit:
test_provider_shutdown.py— session-close-on-live-loop / pending-task drain без «Task was destroyed» / idempotent stop (3 passed). - Live: actions+mutation canaries
readonly-canary-20260918T{174659,184343}Z—task_destroyed_warnings=[],thread_alive=false,leftover_sessions=0.
SCEX-FR-038 Mutation readback (MUT-RECON-001)
- Root-cause 82.75/82.74:
post_rows— честный mid-step SELECT (после UPDATE, ДО in-steprestore_fixturecleanup); внешний канареечный SQL читает ПОСЛЕ полного шага (restore уже вернул 82.74). Lifecycle-skew, не cache/transaction-дефект; оба значения корректны в своих точках. Предыдущий residual-диагноз в чекпоинте 2026-09-18 (cache/isolation) снят. browser_readback.py(новый): independent readback = тот же авторизованный page-сеанс, свежий SQL Labclient_id, SELECT-only скрипт; cleanup-aware expectation (restore→pre-image, retain→assignments); каноническая нормализация (float 6dp, key-sorted) — рендеринг float не может сфабриковать mismatch; crash → typedBrowserReadbackError(никакого выдуманного ok).- Transport: после mutation-флоу (включая in-step cleanup) — readback; divergence →
BrowserTransportReadbackMismatch→ provider:inconclusive, effectunknown, receiptreconciliation_required, retry blocked; PASS-путь: checkpointreadback_verified, receipt summary несётreadback_rows/hash/ok. - Unit:
test_browser_readback.py(expectation/normalization/evaluate ok-mismatch-crash) + provider mismatch→reconciliation; cleanup-тесты обновлены под readback-контракт (3 evaluated scripts на restore-шаге). - Live: mutation canary
readonly-canary-20260918T184343Z—readback_ok=true×2 в outcome и receipts,failures=[], shutdown clean. - INV_7 (browser.py ≤400): вынесены download-side-artifact, transport_factory (через
session_plan_box — late-bound plan), mutation_readback_summary в
browser_factory_helpers.py; итог 395 LOC.
BSC-FILT-001 live filters canary
- Новый
prototype/browser_filters_canary.py: page-based discovery по всем 11 дашбордам (REST query-model не экспонирует Superset 4.x data-mode фильтры) + acceptance-векторы в ОДНОЙ run-scoped сессии (эпемерныйnative_filters_keyне переносим между контекстами). - PASS на dashboard 11: values-apply (
applied_mode=values,chart_data_observed=true, chip Japan наблюдается inspect-ом после apply), clear (mode=clear, chart=true), clear-checkpoint replay capture (1 replayable entry). Evidencefilters-canary-20260918T195946Z.json+ PNG. - Структурные факты стенда: Region dropdown — checkbox-list БЕЗ search input (search-вектор typed-N/A здесь; search-флоу остаётся unit-доказанным); фильтры есть только на dashboards 5, 11.
- date-вектор BLOCKED внешний: ни на одном дашборде стенда нет date/time-фильтра; harness
авто-детектит
date_filtersи прогонит вектор, когда owner добавит фильтр. Строка матрицы G-BROWSER-FILTERS → PARTIAL (BLOCKED external). - Продуктовые robustness-фиксы по live-флейку (оба в
browser_native_filter.py): (1) pre-apply Escape — в reused run-сессии dropdown предыдущего шага остаётся открытым и control-click его TOGGLE-закрывает (опции исчезают); (2) bounded dropdown settle 400ms — antd slide-up анимация гонит same-frame visibility snapshot. Оба покрыты существующим сьютом (FakePage-совместимость через getattr).
Live Superset residuals (обновление)
Playwright connection tasks leak— CLOSED (SCEX-FR-037 выше).Mutation 82.75/82.74 mismatch— CLOSED (lifecycle-skew root-cause + SCEX-FR-038 readback).- NEW (external): date-фильтр на тестовом дашборде стенда — нужен owner; после добавления
перезапустить
browser_filters_canary.py(date-вектор прогонится автоматически).
Checkpoint — 2026-09-21 (WAVE 2 CORRECTNESS CHAIN: BSL-CATALOG core + BSL-DERIVED 3/4 axes CLOSED)
Волна 2 release-последовательности матрицы §3 (correctness truth chain). Full backend suite 11358 passed, ruff clean, compileall, GRACE-anchors сбалансированы.
G-BSL-CATALOG core (037 T082–T084) — catalog_revision_log.py
- Append-only JSONL revision log
*.revisions.jsonlрядом с каталогом (inter-process per-catalog lock; D11: без новой Alembic-миграции — ORM CatalogRevision отвергнут). - T082: duplicate baseline_id / approved coordinate_hash внутри одной ревизии →
CatalogDuplicateEntry(первый approval выигрывает); stale If-Match →CatalogCasConflictс указанием текущего head; два конкурентных approval → ровно один head (thread-race тест с барьером, проигравший получает typed conflict — no lost update). - T084:
append_status_transition(superseded/retired/invalidated) — аппенд, история байт-идентична (тест startswith);record_publish_outcome: publish_failed требует typed error_code, retry-receipt ложится на ту же ревизию (без дублей). - Дурэблити: append = полный префикс + новая строка через temp+os.replace — краш оставляет либо старый лог, либо полный новый, но не частичную запись.
- Статус матрицы: PARTIAL (core proven, wiring open) — интеграция в consume/materialization API-поверхность остаётся (лог — CAS-авторитет рядом с YAML-каталогами).
G-BSL-DERIVED (037 T086–T088) — 3/4 оси закрыты
- T086
scenario_transform_provenance.py: полная цепь провенанса от server-owned 044 ScenarioArtifact — run существует → artifact owner_type/owner_id/kind/is_active → sha256 == source_response_hash → transform-шаг passed → candidate_value == server-computed actual. Подделки (forged run, foreign/retired artifact, digest/value mismatch) — typed rejection до создания candidate. 8/8test_scenario_transform_provenance.py; wiring вcandidates.create_candidate. - T087
baseline_selection.py+schemas/dashboard_testing/baseline_selection.py(AGBASE-FR-016): facts-only ранжирование (cross_check → checklist → big-number/closed period → filter_sentinel → lineage_repr), hard-кап ≤50 с budget_note об исключениях, default+business filter contexts (default обязателен), tolerance rationale; review с CAS (SELECTION_CAS_CONFLICT/SELECTION_CAS_MISMATCHtyped), accept/drop/adjust; unknown rank — fail-closed. 8/8test_baseline_selection.py. - T088 observatory tier (AGBASE-FR-015): отдельная секция
observatory_entriesвcatalog-revision.schema.json(non-gating by construction); resolver-guard — observatory entry, провезенная вentry_revisions, → BASELINE_AMBIGUOUS; секция исключена из пина. 11/11test_baseline_resolver.py(2 новых observatory-edge). - Остаток G-BSL-DERIVED: transform execution→artifact capture (038 T068/T070) и период stale→blocked→candidate.
Остатки волны 2 (не в этой волне)
- G-E2E-EVALUATION (044 T051 walker binding) — следующая итерация.
- G-LLM-SECURITY (scheduled real-LLM terminal PASS) — отдельный прогон.
Checkpoint — 2026-09-21 (T051 WALKER BINDING CLOSED: E2E evaluation chain decision now observes the covering evaluation)
G-E2E-EVALUATION offline-часть закрыта. Full backend suite 11362 passed, ruff clean,
walker.py 398 LOC (INV_7 удержан через вынос capacity-block в capacity_block.py).
T051 dependency-ordered binding — execution/evaluation_binding.py
- Диагноз 2026-09-17 подтверждён и закрыт: compare-шаг видел
evaluation=None, потому чтоpolicy_inputs_from_outcomeбиндит evaluation только из шага-владельца; sibling-топология v4-графов оставляла compare без оценки → EVALUATION_UNAVAILABLE при mandatory-режиме. - Решение (вариант 1 из диагноза):
bind_covering_evaluation(db, run_id, plan, outcome, step_meta)— если план декларирует покрывающий agent_evaluation-шаг (comparison_refsна step-уровне, server-owned plan facts), walker ищет PERSISTED immutable AgentEvaluation этого run с покрывающим comparison_id и инжектит его какevaluation_inputcompare-решения. Truth table не изменена (SCEX-FR-028 инвариант сохранён). - Deferral: пока покрывающая evaluation не persisted, compare-решение откладывается —
шаг → queued,
_deferred_this_call-множество в вызове walker; release, когда все покрывающие eval-шаги terminal (persisted/failed/blocked/skipped) — без busy-loop, перенос в следующий dispatch-цикл. Отсутствие покрывающего шага в плане — прежний немедленный путь (EVALUATION_UNAVAILABLE остаётся честным ответом mandatory-режима без объявленной оценки). - Bound ids (
agent_evaluation_ids,evaluation_input) штампуются на compare-шаг outcome для трассируемости; запись остаётся принадлежащей agent_evaluation-шагу. - 4/4
test_evaluation_binding.py: (1) graph-level terminal PASS — compare stepBASELINE_AND_SEMANTIC_PASS, run passed (T046 residual closed offline); (2) deferral без ложного PASS; (3) no-covering — прежний путь; (4) disabled mode никогда не биндит.
INV_7 рефактор
capacity_block.py(новый):block_run_on_capacity+reconcile_capacity_blocked_runs+CAPACITY_RETRY_CODESвынесены из walker.py (394→398 LOC после добавления binding); импорты в runner.py/dispatch_runs.py переключены, обратная совместимость сохранена.
Остатки G-E2E-EVALUATION
- Live rerun v4-style графа с terminal passed (живой стенд + провайдер, ближайшее окно).
- Scheduled zero-human real-LLM PASS и single-shot/best-of-N policy-варианты.
- G-LLM-SECURITY
LLM-INJ-001live/runtime proof — отдельный live-прогон.
Checkpoint — 2026-09-21 (T068/T069 CLOSED: bounded transform DSL + best-of-N verdict policy)
Обе оставшиеся offline-оси волны 2 закрыты. Full backend suite 11384 passed (+22 новых), ruff clean, compileall, anchors сбалансированы.
T068 (AGSCN-FR-018) — execution/transform_dsl.py + wiring в bounded_transform
- Structured op-tree DSL (никаких строковых парсеров):
sum_column(Σ колонки declared ref, ≤10000 rows),difference(a−b),ratio(a/b, 6dp canonical, деление на ноль typed),scale(k×a). Depth ≤ 4. Decimal-арифметика с canonical 2dp/6dp рендерингом — byte-stable (тест ×25 повторов). Не-числовая ячейка → typedTRANSFORM_CELL_NON_NUMERIC(никакой тихой конверсии). Unknown op/keys → typed. Никакого SQL/Python/shell/network by construction. - Wiring:
bounded_transformпринимаетderive_valueop-tree на step; outcome несётderived_value+ normalized decimalactual(готово к scenario_transform candidate T086 через server-owned артефакт-цепь); DSL-ошибки → typed inconclusive сdsl_error. - 12/12
test_transform_dsl.py.
T069 (AGSCN-FR-022) — execution/evaluation_aggregation.py
VerdictPolicy(revision-authored элемент): single_shot | best_of_n (n∈[2..5]), confidence_threshold ∈[0..1] — валидация typed.aggregate_verdictspure function: strict majority (>N/2) по admissible голосам (succeeded-with-verdict; provider/parser/budget/cancel/timeout не голосуют); tie →EVALUATION_AGGREGATION_TIE; all-error →EVALUATION_AGGREGATION_ALL_ERROR; quorum-miss →EVALUATION_QUORUM_INPUTS_MISSING; пусто →EVALUATION_AGGREGATION_EMPTY; pass-below-threshold → inconclusiveEVALUATION_CONFIDENCE_BELOW_THRESHOLD(no quiet PASS). Mean confidence 6dp. Aggregate evaluation_id детерминированно именует когорту.- 044-binding: агрегат потребляется существующим T051
evaluation_bindingпутём. - Остаток: revision-schema поверхность (пин VerdictPolicy в AgentEvaluationSpec) — additive authoring-работа.
- 10/10
test_evaluation_aggregation.py.
T070 — уже закрыт волной 037 (T087 proposal + T088 observatory)
Матрица: G-BSL-DERIVED → PARTIAL (offline оси все закрыты; остатки — period stale→blocked →candidate lifecycle и live capture wiring transform-выхлопов в ScenarioArtifact); G-E2E-EVALUATION residual сузился до live rerun + scheduled real-LLM + revision-schema.
Остатки волны 2 (все live-зависимые)
- Live rerun v4-style графа с terminal passed (T046 closing evidence).
- Scheduled zero-human real-LLM PASS (G-LLM-SECURITY
LLM-INJ-001live proof). - Period stale→blocked→candidate + live transform capture wiring.
- Catalog revision log wiring в consume/materialization.
Checkpoint — 2026-09-21 (LIVE WAVE: E2E evaluation chain proven on the stand; terminal PASS external-blocked)
Стенд запускался run.sh --skip-install, затем uvicorn с captured stdout (/tmp/kilo/uvicorn.log).
Полный live-цикл T051-цепи выполнен на ss-prod (dashboard 11) + Gitea-published baseline envelope.
Live canary v5 — prototype/live_canary_v5_binding.py (новый)
Sibling-топология: open-dashboard → capture-evidence → {compare-to-baseline, evaluate-visual};
evaluate декларирует inputs.comparison_refs (server-owned plan fact); decision_policy
required. Полная цепь по прогону live-canary-v5-binding-20260921T164034Z.json (+eval-record):
- open-dashboard passed (браузер, PREPROD-стенд);
- capture-evidence passed BASELINE_PASS (8 durable артефактов) — semantic-фикс этой волны;
- evaluate-visual: real-LLM evaluation PERSISTED (
ce4685e7…, status succeeded, finding с evidence_artifact_ids); - compare-to-baseline: walker BINDING сработал — decision несёт
evaluation_input(persisted id) +agent_evaluation_ids; deferral-цикл прошёл (compare ждал evaluate); - policy: LOW_CONFIDENCE (честный verdict inconclusive 0.0 от text-only модели) — fail-closed.
Продуктовые фиксы, найденные live-прогоном
- Semantic overlay scope (
decision_policy.py): required-режим больше не требует собственную evaluation от evidence-продюсеров (screenshot/sql без comparison-payload) — они проходят rows 1-8 BASELINE_PASS; semantic-оверлей (rows 9-14) привязан к comparison-шагам (semantic_comparisonв StepPolicyInputs; mapper ставит по tool==assertion / comparison-payload). Соответствует production-chain («deterministic-only runs need no LLM»); инвариант truth table сохранён. - Defer deadlock guard (
walker.py::_step_feeds_evaluation): deferring шага, от которого зависит покрывающая evaluation, невозможен — живой прогон выявил цикл (capture ждёт evaluate, evaluate ждёт capture). - Test-фикстуры decision-policy обновлены под
semantic_comparison(21+1 passed).
Внешние блокеры, обнаруженные live-прогоном (2026-09-21)
- LLM gateway vision routes ALL 502: omniroute мультимодальные роуты (auto/best-vision,
auto/pro-vision, auto/vision, auto/claude-*, auto/gemini, auto/glm) возвращают upstream 502
«Cloudflare Playground browser session failed: browserType.launch: Executable doesn't
exist». Text-only
auto/best-fast(qwen3.8) работает — провайдер переключён на него сis_multimodal=false(честный text-only manifest по контракту). Terminal PASS (BASELINE_AND_SEMANTIC_PASS) требует визуальный вердикт ≥0.7 — недостижим, пока gateway vision-роут не починен; chain закрывается честно LOW_CONFIDENCE. - Операционные правки локального стенда:
llm_providers.base_urlисправленhttps://omniroute.bebesh.ru→https://omniroute.bebesh.ru/v1(OpenAI SDK строит{base}/chat/completions; без /v1 — 401/404 на каждый вызов, маскировался под EVALUATION_PROVIDER_ERROR). run.sh-стенд не передаёт stdout uvicorn — для live-дебага использован отдельный uvicorn с env из backend/.env (DRAFT_STORAGE_ROOT не нужен — settings.storage.root_path через STORAGE_ROOT_PATH).
Учёт
- T046: offline+live CLOSED (binding); terminal
passed— external-blocked (vision gateway). - T051: CLOSED (binding + deferral + deadlock guard; best-of-N через T069 aggregation).
- G-E2E-EVALUATION → PARTIAL (chain live-proven; terminal PASS blocked external).
- Evidence:
specs/044-dashboard-scenario-execution/evidence/browser-provider/ live-canary-v5-binding-20260921T164034Z.json+-eval-record.json.
Checkpoint — 2026-09-22 (WAVE 3 OPERATIONS integrated: NOTIFICATIONS, ATOMIC INVESTIGATION, POISONED-RUN, MCP PARITY)
Четыре фазы волны 3 выполнены параллельными Agent Manager worktree-сессиями (ds/deepseek-flash,
omniroute) и интегрированы в master патч-моделью (база веток была устаревшей: origin/master
= 1411c03f, −48 коммитов; rebase не применялся, патчи наложены на master с разрешением
пересечений A∩C в analytics/investigation.py и B∩D в api/routes/.../scenario_automation.py).
Коммиты:
93e97228feat(automation)— NOTIFY-001: fail-closed emit/durable домены, channel/sla, идемпотентные receipts, queue/case/baseline_stale wiring.def2bf40feat(automation)— SCHED-SOAK offline:poisoned_store.py(durable JSONL идентичных-сбоев, flock+fsync, CAS epoch, без миграций) +poisoned.pyN=3 → quarantine, один blocked DLQ-receipt, операторский CAS выхода; REST/quarantine.223296bcfeat(analytics)— INV-ATOMIC-001:locking.py(pg_advisory_xact_lock), savepoint вset_disposition,context.pyexact-context comparability; 22 unit + 7 PG-векторов.65684d6ctest(mcp)— 050 T051: единыйclassify_start_error(REST+MCP), parity-матрица 17 векторов.
Verification: полный backend suite 11472 passed, 290 skipped, 1 xpassed, 0 failed; ruff clean.
Интеграционные адаптации (дрейф устаревшей базы, внесены при слиянии):
test_mcp_rest_error_parity.py— investigation-ACL вектор переписан с устаревшей посылки «нет MCP-инструментов investigation» (на master они есть, G-INVESTIGATION-MCP CLOSED) на реальный no-bypass инвариант: чужой case схлопывается вnot_foundчерез MCP без утечки содержимого; каталог-присутствие инструмента проверяется.test_due_admission_atomicity.py— ожидания обновлены: терминальный прогон эмититfailed+queue_created, replay идемпотентен (2 receipt, не растут).
Матрица: G-NOTIFICATIONS → CLOSED, G-INVESTIGATION-ATOMIC → CLOSED, G-MCP-PARITY → CLOSED,
G-DLQ-SOAK → PARTIAL (offline закрыт; 72h soak/restart storms и 5/15/50-tab — live).
Инфра-инцидент: недоступность PostgreSQL с хоста из-за VPN-маршрута 172.19.0.0/24 dev throne-tun, перекрывавшего docker-bridge 172.19.0.0/16; маршрут исчез сам, БД восстановилась.
Checkpoint — 2026-09-22 (WAVE 4: offline P0-остатки — catalog revision-log wiring, period stale lifecycle, LLM-INJ-001 offline)
Волна выполнена в одном рабочем дереве (без Agent Manager worktrees); изменения на момент записи
НЕ закоммичены (коммиты — только по явному запросу, по одному на фазу).
План: .kilo/plans/1790082327000-wave4-offline-p0-residuals.md.
Phase A — G-BSL-CATALOG wiring (037 T082/T084) → CLOSED
backend/src/services/dashboard_testing/catalog_revision_wiring.py(новый): server-owned проекция consumed-entry на envelope лога —compute_coordinate_hash(expected value исключён),build_capture_profile_hash(из server-issued capture artifact, caller-хэш отвергнут),build_entry_revisions,settle_revision_outcome(publish outcome только на digest-matched head).materialization._commit_with_catalog_compensation(..., revision=…): append ревизии внутри того же_catalog_lockДО YAML-записи (duplicate/CAS отказ до мутаций) + компенсация log-bytes вместе с catalog-bytes при отказе DB commit — фантомная ревизия в CAS-логе невозможна. Новые хелперы:_append_revision_locked,_restore_bytes,_restore_revision_log.approvals.consume_approvalстроит ревизию из server-built entry + capture meta (единственная точка materialize).publication_worker.publish_baseline_catalog(local_catalog_path=…): settle-mirror в трёх точках (verify-reconcile, attempt success, failure) через_record_revision_outcome/_settle_published;_trace_revision_outcomeне маскирует исходную типизированную ошибку.- Тесты:
tests/services/dashboard_testing/test_catalog_revision_wiring.py(10) +execution/test_publication_worker.py(+3).
Решение (reactive micro-ADR): привязка локального лога к publish — через явный server-side
параметр local_catalog_path; hard-fail при его отсутствии отвергнут (publish-контракт владеет
Git-коммитом, локального каталога в нём нет) — отсутствие трассируется
(PUBLISH_REVISION_LOG_ABSENT), а НЕВЕРНЫЙ лог падает типизированно
(PUBLISH_REVISION_LOG_MISMATCH). Digest-scan по catalog_base_path отвергнут как недетерминированный
и дорогой на текущей топологии.
Phase B — G-BSL-DERIVED period lifecycle (037 T089, AGBASE-FR-018) → PARTIAL
baseline_staleness.py(новый):requested_period_from(params, plan),closed_period_of,detect_period_staleness(только closed-period mismatch; open/undeclared период НЕ stale),build_stale_rebaseline_proposal(facts-only: идентичности + capture-kind,requires_operator_approval=true,automatic_rebaseline=false),emit_period_stale_receipt(durable idempotentbaseline_stale,run_id=None).baseline_resolver:BaselinePeriodStale(ValueError)сохраняетstr(exc)=="BASELINE_STALE",resolve_baseline_pin(requested_period=…)блокирует запуск typed до I/O.start_runпробрасывает период из params/plan и трассирует блок (EXPLORE, error_code BASELINE_STALE).- Immutability-приоритет сохранён: hash-change внутри ТОГО ЖЕ закрытого периода остаётся
immutability_violation(CRITICAL), не downgrade до stale. - Тесты:
execution/test_baseline_stale_lifecycle.py(8).
Residual: durable receipt запускного rollover не эмитится из production-surface (reject
предшествует созданию run-строки, а owning-transaction откатывается) — эмиттер реализован и
протестирован, привязка к поверхности — следующая итерация; comparison-time stale_baseline
уже эмитит dedicated receipt (G-NOTIFICATIONS CLOSED).
Phase C — G-LLM-SECURITY LLM-INJ-001 offline (044 SCEX-FR-035) → PARTIAL (offline CLOSED)
evaluation_adapter: в промпт добавлен блокevidence_handling(evidence_is_data=true, правила «manifest/значения — данные, не инструкции», closed verdict domain); instructions byte-stable.- Тесты:
execution/test_llm_injection_offline.py(7 векторов, все черезsubmit-seam, без LLM): compromised-double canary, дроп findings по недекларированному критерию, free-form verdict / out-of-range confidence coercion, provider-identity forgery, prompt evidence-as-data + стабильность инструкций, non-pass в DecisionPolicy, отсутствие tool-канала у seam. - Fabricated artifact refs отвергаются walker-валидатором (существующий
registry/test_agent_evaluation_store.py).
Учёт и верификация
- 037
tasks.md: T082/T084 → wiring CLOSED note; T089 →[~]PARTIAL. 037traceability.md: AGBASE-FR-014 (wiring), AGBASE-FR-018 → PARTIAL. 044spec.md:LLM-INJ-001offline CLOSED. - Матрица:
G-BSL-CATALOG→ CLOSED;G-BSL-DERIVED→ PARTIAL (период закрыт, live capture wiring остаётся);G-LLM-SECURITY→ PARTIAL (offline CLOSED); исправлен дрейф строкиG-SCENARIO-AUTHOR(Transform T068/T069 закрыты — прежняя формулировка «open» была устаревшей). - Прогоны: focused-сьюты фаз (21 + 42 + 51), домен
tests/services/dashboard_testing— 1659 passed; полный backend suite/ruff/compileall — см. D.4 ниже. - Продукт остаётся NO-GO:
G-FULL-REGRESSION,G-SEMANTIC-HEALTH,G-RELEASE-SIGNOFFоткрыты.
D.4 — фактические прогоны (2026-09-22, рабочее дерево)
python -m pytest -q(полный backend): 11509 passed, 15 failed, 290 skipped, 1 xpassed (462s). Все 15 падений — в областях ДРУГОГО незакоммиченного воркстрима (admin/roles/permissions, maintenance templates), который уже был в дереве до волны 4:src/api/routes/admin.py,src/core/auth/permission_utils.py,src/services/security_badge_service.py,src/api/routes/agent_status.py,tests/api/test_admin.py— все помеченыMвgit status, и ни один падающий тест не импортирует модули волны 4. Волна 4 эти файлы не трогала.python -m pytest -q tests/services/dashboard_testing(домен): 1675 passed (101s).python -m ruff check .(весь backend): All checks passed.python -m compileall -q src: без ошибок.- Axiom
verify_after_editпо 12 путям волны:orphans delta 0,unresolved delta 0,parse_warning_count 0,naked_count 0. - Найден и исправлен в ходе верификации дефект: правка теста publication worker потеряла
закрывающий якорь
Test.ScenarioExecution.PublicationWorker.Reconcile(INV_3); пара восстановлена, повторныйverify_after_edit→ 0 parse warnings, 37 тестов волны собираются.
Ортогональное ревью волны 4 + аудит GRACE INV_1–INV_7 (2026-09-22)
Инструменты: axiom audit detect_missing_contracts (домен dashboard_testing: 176 файлов,
naked=0), axiom search read_outline, axiom search verify_after_edit, ruff --select C901,
сравнение с git show HEAD:<file> для атрибуции (моё vs унаследованное).
Аудит инвариантов
| INV | Вердикт | Доказательство |
|---|---|---|
| INV_1 (region на каждый definition) | PASS | naked=0; ложно-сработавшие container-coverage по RevisionLogMismatch/BaselinePeriodStale закрыты листовыми регионами, ID = именам символов |
| INV_2 (NEED_CONTEXT при слепоте) | n/a | слепых зависимостей не было; все контракты выведены из кода |
| INV_3 (пары region/endregion) | PASS (после фикса) | в ходе верификации был потерян #endregion ...Reconcile — восстановлен; сейчас 0 parse-warnings по всем 12 путям |
| INV_4 (теги до кода, контигуально) | PASS | теги идут сразу после якоря, код после блока метаданных |
| INV_5 (локальный workaround не отменяет ADR) | n/a | ADR не затрагивались; publish-решение задокументировано как micro-ADR |
| INV_6 (не удалять контракт с входящими рёбрами) | PASS | контракты только добавлялись; тексты существующих тегов дополнялись |
| INV_7 (модуль <400 строк; CC ≤10) | FAIL | размер: approvals 387→401, publication_worker 373→464, baseline_resolver 391→426, evaluation_adapter 390→406, start_run 411→429; CC: publish_baseline_catalog 16 (унаследовано 16), start_run 12→13 (+1) |
Находки, исправленные в этом проходе
- F1 (High, корректность): падение записи каталога после append ревизии оставляло фантомную ревизию (CAS-голова ссылалась на нематериализованную запись). Исправлено: append откатывается внутри того же лока перед пробросом ошибки; +2 регресс-теста.
- F3 (Medium, точность графа): 5 новых межмодульных зависимостей не были объявлены
@RELATION DEPENDS_ON— добавлены (materialization, approvals, baseline_resolver, start_run, publication_worker). - F2 (Low, INV_1): листовые регионы для
RevisionLogMismatchиBaselinePeriodStale. - F4 (Low, наблюдаемость): легитимное отсутствие привязанного лога логировалось как WARNING
на каждом publish → alert fatigue; переведено на
level="DEBUG". - F5 (Low): удалена мёртвая константа
_UNAVAILABLE.
Находки, оставленные сознательно (с обоснованием)
- F6 (Medium, INV_7 размер): 4 файла пересекли 400 строк,
start_runуже превышал. Не рефакторил в этом проходе (риск для верифицированного кода); рекомендация: выделить settle-хелперы publication_worker в сиблинг-модуль, вынести период-детекцию из baseline_resolver, а период-блок из start_run — отдельным PR. - F7 (Medium, CC):
publish_baseline_catalog16 — унаследовано, не мной;start_run+1. Рекомендуется извлечь ветки settle/period-block в хелперы. - F8 (Low, дизайн): digest-mismatch привязанного лога демотирует уже опубликованную операцию
в
publish_failed(Git-коммит уже есть). Самоизлечимо через reconcile (commit сохранён); рекомендуется записать в runbook. - F9 (Medium, интеграционный пробел): у launch-периода нет продюсера:
close_period(сторона closure) имеет и продюсера, и guard'ы, ноreporting_periodв params/plan никто не выставляет — детектор периода остаётся capability без end-to-end прогона. Это подтверждает PARTIAL матрицы. Следующий шаг: пробрасывать объявленный период запуска в params либо читать период из координаты запиненного плана. - F10 (Low): Phase C изменил контракт промпта — live zero-human гейт требует повторной валидации (prompt drift), offline-тестов недостаточно.
- F11 (Low): metric-entry без
capture_artifact_refмолча даёт null capture-поля (не пинится) вместо типизированного отказа на consume. Приемлемо (схема требует ref для metric), но consume-time проверка была бы строже.
Проход «правь все» — закрытие находок F6–F11 (2026-09-22)
F6/F7 — INV_7 (размер <400 строк, CC ≤10)
Разбиты 5 модулей; все затронутые теперь <400, а start_run и publish_baseline_catalog ≤10:
| Было | Стало | Как |
|---|---|---|
| approvals 401 | 399 | revision-payload собран в catalog_revision_wiring.build_revision_payload |
| publication_worker 464 | 356 | publication_errors (тип ошибки), publication_request (валидация+target), publication_mirror (settle-зеркало) |
| baseline_resolver 426 | 372 | execution/baseline_codes (D11-словарь+reject), load_published_catalog → published_catalog_source, BaselinePeriodStale → baseline_staleness |
| start_run 429 | 363 | execution/start_contract (request-hash+classify), execution/prod_guards (PROD-гварды), request-hash/classify оформлены как Tombstone c @REPLACED_BY (INV_6) |
| evaluation_adapter 406 | 368 | execution/evaluation_prompt.build_evaluation_prompt |
CC: publish_baseline_catalog 16→≤10 (валидация вынесена в publication_request.validate_publish_request);
start_run 13→≤10 (PROD-гварды вынесены). Остальные C901-хиты в репозитории (~40 функций,
включая decide_step_outcome 27 и validate_evaluation_evidence 16 в 044) — унаследованный долг,
волной 4 не затронуты; рекомендация вынесена отдельно.
F8 — зеркало publish не демотирует опубликованную операцию
publication_mirror.settle_published больше не вызывает _mark_failed: Git-коммит уже состоялся,
поэтому digest-mismatch привязанного лога теперь громкий (PUBLISH_REVISION_LOG_MISMATCH,
EXPLORE) но НЕ переписывает durable publish truth; операция остаётся published с сохранённым
commit (reconcile-путь). Тест test_publish_wrong_bound_log_is_typed обновлён соответственно.
F9 — период запуска стал контрактом, а не конвенцией
requested_period_from теперь fail-closed: присутствующая, но не non-blank-string декларация
reporting_period реджектится REPORTING_PERIOD_INVALID вместо молчаливого отключения guard'а;
контракт зафиксирован в StartRunRequest.params (REST-комментарий) и @INVARIANT start_run.
Первопартийный продюсер (UI/автоматизация) остаётся работой следующей итерации; сама capability
достижима через свободный params.reporting_period (REST и MCP используют один и тот же путь).
F10 — prompt drift учтён
Изменение промпта Phase C — событие live-гейта: offline-тестов достаточно только для LLM-INJ-001 offline; live zero-human прогон обязан быть повторён на новом промпте (записано в матрице).
F11 — типизированный отказ на gating-ревизию без capture-ref
build_entry_revisions реджектит metric/scenario_transform entry без server capture artifact
(CATALOG_REVISION_CAPTURE_REF_REQUIRED) вместо материализации не-пинируемой ревизии; visual
по-прежнему допускает null (+тест).
Верификация после прохода
python -m pytest -q(полный backend): 11534 passed, 290 skipped, 1 xpassed, 0 failed (438s). Прежние 15 падений чужого воркстрима (admin/roles/permissions) на момент прогона также отсутствуют.python -m pytest -q tests/services/dashboard_testing: 1679 passed.python -m ruff check .: All checks passed;compileall -q src: OK.axiom verify_after_edit(20 путей):parse_warnings 0,naked 0,orphans Δ0;workspace_healthпоstart_run.py:unresolved 0,orphans 0; tombstone-целиScenarioExecution.StartContract.*существуют в индексе.
Checkpoint — 2026-09-23 (WAVE UI: фронтенд-пробелы матрицы — DLQ recovery, run-monitor reload, disposition E2E-спек)
План: .kilo/plans/1790163777000-wave5-frontend-ui.md (каталог планов gitignored).
Изменения на момент записи НЕ закоммичены (коммиты — только по явному запросу).
UI-1 (P0, G-DLQ-SOAK) — операторское восстановление карантина → offline CLOSED
- Backend-REST уже существовал (
GET /quarantine,POST /quarantine/{scenario_id}/releaseс CASexpected_version; 404 NOT_QUARANTINED / 409 STALE_QUARANTINE_VERSION) — backend не менялся. frontend/src/lib/types/scenario-automation.ts:QuarantineRecord+QuarantineReleaseResult(mirrorPoisonedRunStore.Fold: scenario_id, environment_id, error_code, identical_count, version, reason, quarantined_at).frontend/src/lib/api/scenario-automation.ts:listQuarantines+releaseQuarantine.frontend/src/lib/models/AutomationQuarantineModel.svelte.ts(новый,[TYPE Model]): FSM idle/loading/loaded/confirming/releasing/error; release отсылает показанную версию; 409 → reload + noticestale_version(без молчаливого повтора); 404 → reload + noticealready_released; неизвестная ошибка → error без reload.frontend/src/lib/components/scenario-automation/QuarantinePanel.svelte(новый) + монтаж наroutes/dashboard-testing/automation/+page.svelte;$lib/ui(Button/Badge/ConfirmDialog), семантические токены, i18n (15 ключей ru+en).- Тесты:
AutomationQuarantineModel.test.ts(7) +QuarantinePanel.test.ts(4) + обновлён мокautomation_page.ux.test.ts.
UI-3 (P1, G-MONITOR) — reload/reconnect → offline CLOSED
RunMonitorModel.test.ts+2 L1-вектора: (1) свежая модель после «reload» восстанавливаетwaiting_humanс сервера, ровно один checkpoint-шаг (перезагрузка не добавляет второй); повторныйloadRunзаменяет снапшот; (2) повторныйbindEventsзакрывает предыдущий EventSource (closed), named-листенеры не стекаются (ровно один на имя).- Tiering-маркеры остаются отсутствующими (blocked на backend
G-EVIDENCE-TIER).
UI-2 (P1, G-INVESTIGATION-UI) — disposition E2E-спек → написан, исполнение заблокировано
frontend/e2e/tests/scenario-disposition.e2e.js(новый): queue→case→disposition через case-workspace (section[aria-label="Рабочее пространство кейса"], selectresolved, кнопка «Закрыть кейс») + отдельный CAS-вектор (staleexpected_version→ 409STALE_CASE).- Residual: на изолированном стенде нет пути засева очереди (backend не отдаёт API создания queue-item; сигнал создаёт падающий прогон) → спек скипается с явной причиной. Это названный остаток строки, не зелёный прогон.
UI-4/UI-5 — только объявленные контракты (BLOCKED)
Evidence.TierBadge(tier_hot/tier_cold/untiered) иScenarioBaseline.SelectionReview(accept/drop/adjust + CAS) объявлены в плане; код не писался, т.к. backend-полей/поверхности нет (иначе UI биндил бы несуществующий контракт данных). Строки матрицы обновлены пометкой «contract-declared, not implemented».
Верификация UI-волны
npm run test -- --run: 3425 passed (220 файлов);npm run lint: 0 errors (341 warningsvelte/require-each-key);npm run build: ✓.- Новые файлы: 2 модели/компонента + 1 e2e-спек + 16 новых frontend-тестов (7+4+2 сверх обновлённого page-мока).
- Матрица:
G-DLQ-SOAK(recovery UI закрыт offline),G-MONITOR(reload/reconnect закрыт),G-INVESTIGATION-UI(спек написан, остаток — seeding),G-EVIDENCE-TIER/G-MCP-BASELINE-SELECTION(blocked-пометки).
Handoff для следующего агента (2026-09-23)
Промежуточное состояние передано: specs/agent-handoffs/wave4-ui/README.md (полный инвентарь
коммитов, «мои» vs «чужие» файлы рабочего дерева, статусы строк матрицы, оставшиеся шаги по
зависимостям, команды верификации, инфра- и координационные грабли).
- Закоммичено:
5a5ac4dd(wave 4, 26 файлов). Не закоммичено: wave UI (UI-1/UI-3 + спек UI-2) и учётные правки матрицы/traceability/WORKSTATE — файлы в дереве, страховочный патчspecs/agent-handoffs/wave4-ui/wave-ui-and-specs.patch(17 файлов). - Планы (каталог
.kilo/plans/gitignored, ключевое продублировано в README):1790144880000-wave5-offline-p0-residuals.md(backend: DLQ retry-dispatcher, launch-receipt, SECURITY-OPS),1790163777000-wave5-frontend-ui.md(UI-1..UI-6). - Внимание: в дереве параллельно работает вторая frontend-сессия (модальные binding-пробы,
runes-contract, admin settings) — её файлы перечислены в README §2 и не должны попадать в коммиты;
.kilo/agent-manager.jsonне коммитить. - Остаток для GO (кратко): 10 незакрытых P0-строк (2 внешних блокера, 4 live/offline-микс, 3 релизных гейта, 1 signoff) + 8 P1; продукт остаётся NO-GO.
Wave 5 Phase A — G-DLQ-SOAK retry-dispatcher wiring (2026-09-23)
Мастер-план: .kilo/plans/1790240000000-wave5-6-execution-master.md (порядок фаз и контракты).
- Phase 0 закоммичена
c4d6faa0: wave UI (17 файлов §2 + handoff-каталог) — quarantine recovery UI, run-monitor reload/reconnect, disposition e2e spec, учёт матрицы/traceability/WORKSTATE. - Phase A (offline, закрыта в дереве):
automation/retry_policy.py(закрытый allowlist retryable infra-кодов, fail-closed;MAX_ATTEMPTS= порог карантина),execution/retry_dispatcher.py(failed→queued CAS re-queue; durable identical-failure counter = attempt-authority — у ScenarioRun нет attempt-колонки, D11 замораживает миграции; карантинные пары и потолок пропускаются), терминальный учёт вexecution/terminal_effects.py(один терминальный отказ = одна запись счётчика; automation-источники + allowlist только; passed automation-ран сбрасывает стрик; карантин чистит только операторный CAS), карантин-гейт вexecution/dispatch_runs.py(устаревшие queued-строки карантинной пары не диспатчатся до CAS-release), scheduler-тикcore/scheduler.pyкомпозирует retry→dispatch. Тесты:tests/services/dashboard_testing/automation/test_retry_dispatch.py— 10 векторов (CAS exactly-once, N=3 quarantine + одна нотификация, distinct-code reset, content/manual исключение, гейт+release, потолок, allowlist-membership). - Верификация:
pytest -q tests/services/dashboard_testing1689 passed (было 1679); полныйpytest -q— 11579 passed, 290 skipped, 1 xpassed, 0 failed;ruff check .All checks passed;compileall -q srcOK; axiomverify_after_edit(automation/, execution/, scheduler.py, automation-tests): parse_warnings 0, naked 0, orphans Δ0, unresolved Δ0;make docs-nav— контракты и рёбра обновлены (новые:ScenarioAutomation.RetryPolicy,ScenarioExecution.RetryDispatcher,ScenarioExecution.Runner.QuarantineGate,ScenarioExecution.Runner.PoisonedAccounting,ScenarioExecution.Runner.PoisonedStreakReset). - Матрица:
G-DLQ-SOAKPARTIAL — offline + recovery UI + retry wiring CLOSED; остаток только live (72h soak, 3 restart storms, 5/15/50-tab canaries). 046 tasks.md T023/T024 и traceability SCAUTO-FR-023/024 обновлены. - Следующие фазы (по мастер-плану): B (launch-receipt, после A — теперь разблокирована), C (SECURITY-OPS, ∥ B), S (UI-2 queue-seed, ∥ B/C), затем P1-волна и релизные гейты.
- Коммит Phase A — по явному запросу (файлы в дереве, чужие файлы второй сессии не тронуты).
Wave 5 Phase B + C — G-BSL-DERIVED launch-receipt + G-SECURITY-OPS offline (2026-09-23)
- Phase B (launch-receipt, закрыта в дереве):
BaselineEngine.BaselineStaleness.LaunchReceipt(emit_launch_period_stale_receipt) — durablebaseline_stalereceipt на отдельной закоммиченной сессии после отката транзакции запуска; fail-safe (провал эмита — EXPLORE-трейс, транспортный ответ не маскируется). Вызов ровно один на транспорт: RESTscenario_runs.py(BaselinePeriodStale до generic ValueError → 422 BASELINE_STALE, ответ байт-идентичен прежнему контракту) и MCPtools_scenario.py(blocked + classify_start_error, 050 T044 parity). Idempotency-key стора схлопывает двойной вызов REST+MCP в одну строку. Тесты:tests/.../test_start_run_period_receipt.py— 6 векторов (REST/MCP receipt, exactly-once, digest-mismatch без receipt, отсутствие run-строки, провал эмита не маскирует блок); parity-матрица 17 векторов зелёная. - Phase C (SECURITY-OPS offline, закрыта в дереве):
- C.1:
Core.EncryptionKey.KeyAge— fail-closed вердикт возраста ключа по mtime .env (неизвестный возраст = warning, не ok; превышение интервала = age_exceeded); тест ротации (_reencrypt_value+Fernet): старый шифротекст нечитаем новым ключом, после re-encrypt читаем новым и нечитаем старым; legacy-значения репортятся, не конвертируются. - C.2: redaction на границе записи evidence —
ScenarioExecution.EvaluationAdapter.Raw+_redact_tree(per-string-leaf через pluginRedactionService; сериализованную строку редактить нельзя — регексы не JSON-aware и ломают структуру, @REJECTED); findings и структура JSON сохраняются. Второй реализации redaction нет (reuse единственного модуля). - C.3: retention-политика evidence объявлена в
GlobalSettings(scenario_evidence_retention_daysdefault 90 +scenario_evidence_archive_target), falsifiable config-check в тесте; enforcement-sweep не реализован (scope-забор → G-STORAGE-BUDGET, P1.e).
- C.1:
- Верификация (волна 5 целиком): полный
pytest -q— 11589 passed, 290 skipped, 1 xpassed, 0 failed; dashboard_testing-сьют 1699 passed;ruff check .All checks passed;compileallOK; axiomverify_after_edit: parse_warnings 0, orphans Δ0, unresolved Δ0; naked=0 на моих файлах (11 унаследованных naked в чужих тест-файлах test_comparison/test_normalization/ test_visual_baseline_staleness — вне скоупа волны, INV_1-долг на P1/релизную гигиену). - Матрица:
G-BSL-DERIVEDPARTIAL (launch-receipt CLOSED; остаток live capture wiring);G-SECURITY-OPSPARTIAL (offline CLOSED: rotation/age/redaction/retention-policy; остаток live drill).G-DLQ-SOAK— см. checkpoint Phase A выше. - Следующие фазы (мастер-план): S (UI-2 queue-seed для e2e), P1-волна (P1.a G-EVIDENCE-TIER backend → UI-4; P1.b MCP-surface → UI-5; P1.c/d/e/f), релизные гейты R.1–R.3.
- Коммит волны 5 (A+B+C одним или тремя коммитами) — по явному запросу; файлы в дереве, чужие файлы второй сессии не тронуты.
Post-wave-5 P0 contract closure — transform evidence + retention enforcement (2026-09-23)
- G-BSL-DERIVED: успешный bounded transform теперь сохраняет канонические JSON-байты результата
через существующий DraftStorage и передаёт ref/digest/MIME/length в outcome. Walker регистрирует
их как ScenarioArtifact с владельцем
scenario_run, logical_step_id и attempt; provenance-проверка кандидата связывает ref, digest и вычисленное значение. Подмена ref/digest/value отвергается. Отдельный bridge-тест проверяет bytes → artifact → candidate. Offline-векторы зелёные; live canary остаётся открытым. - G-SECURITY-OPS: регистрация evidence назначает raw VLM класс
raw_vlm(7 дней), изображениямscreenshots(30 дней), прочим step-артефактамartifacts(30 дней). Плановый retention job находит истёкшие ScenarioArtifact, идемпотентно создаёт deletion receipts и коммитит их до удаления байтов; существующий sweep соблюдает baseline/active-run/open-case holds, проверяет отсутствие байтов перед tombstone и на момент sweep защищает новый артефакт с тем же content_ref. Проверены границы 8/31/91 дней, прерывание удаления после mark и повторный sweep.scenario_evidence_retention_daysуправляет только явно назначенным классомscenario_evidence(текущие producer-ы его не назначают);scenario_evidence_archive_targetпока не архивирует байты. Live rotation/PII/retention drill остаётся открытым; storage tiering/budget — P1. - Верификация: независимый focused pytest — 73 passed; Ruff для затронутых backend-файлов
чистый;
git diff --checkчистый. Полный release regression на замороженном коммите не выполнялся. - Live-стенд:
./run.sh --skip-installне дошёл до backend: PostgreSQL отвечает внутри Docker (SELECT 1), но host TCP на опубликованном порту 5432 обрывается до PostgreSQL;/api/readyнедоступен. Вbackend/.envтакже обнаружены legacy multiline PEM-секции, из-за которых Docker Compose не может разобрать env-файл. Секреты и данные стенда не изменялись. - Матрица сохраняет
G-BSL-DERIVEDиG-SECURITY-OPSкак PARTIAL до retained live evidence.
Business Analyst UX acceptance target — 2026-09-23
- Новая обязательная release-планка
G-BA-UX-READINESSзафиксирована в unified matrix (P0): каждый из семи критических сценариев аналитика должен получить отдельную оценку ≥4/5; среднее не маскирует провал отдельного сценария. Приёмка — 5 бизнес-аналитиков, ≥80% шагов без модератора, без false PASS/approval/case closure, с keyboard/a11y и retained browser evidence на одном release candidate. - Контракты и backlog связаны из 039 (authoring + scoring), 045 (launch/result/monitor), 046 (schedule operations), 047 (queue→case/disposition). Текущий аудит оставляет gate OPEN: baseline review UI отсутствует; editor неполон; evidence/comparison не подключены; queue→case ID/pagination расходятся с API; resolution не требует ссылку на evidence; cron/ID — в форме расписания. Реальный browser/user study не проводился.
- Подробная шкала и протокол находятся в
specs/039-dashboard-scenario-ui/spec.md→ Business Analyst Readiness Amendment; тесты/build являются supporting evidence, не UX sign-off.
047 investigation seeded browser proof — 2026-09-24
- В изолированном Compose-проекте
ss-tools-e2eдобавлен одноразовый seed после миграций: фиксированные scenario/run/artifact IDs, failed terminal run, active JSON evidence bytes в общем storage и сигнал черезemit_terminal_run_signal. Повторный seed сбрасывает только собственную фикстуру. - Playwright
scenario-disposition.e2e.js: queue→правильный case через UI→linked run/evidence viewer→resolvedс проверяемым ref→закрытие queue→старыйdecision_versionполучает 409STALE_CASE. Chromium: 1 passed, 0 skipped; trace/screenshot сохранены вfrontend/test-results/scenario-disposition.e2e.j-cc566-dence-and-rejects-stale-CAS-chromium/. - Прогон выявил реальный HTTP 405: модель case disposition использовала GET-обёртку
fetchApi. Исправлено наpostApi; focused Vitest 10/10, frontend lint/build, backend Ruff/compileall, синтаксис JS, Compose config иgit diff --checkпрошли. Playwright runner обновлён до 1.60.0;ENCRYPTION_KEYберётся из игнорируемого.env.e2e, аbackend/.envс multiline PEM больше не передаётся Compose. - T027/T032 закрыты на уровне изолированного browser proof. Axiom
verify_after_editпо четырём затронутым source/test путям: success, parse warnings 0, naked 0, orphan/unresolved Δ0. Контейнеры удалены с volumes; коммита нет.G-INVESTIGATION-UIимеет scoped implementation proof, ноG-BA-UX-READINESS, полный regression на release commit и остальные release gates остаются открыты.
Correction — live browser filters and multimodal provider evidence (2026-09-25)
- В
browser_filters_canary.pyдобавлена поддержкаSS_STAND_NATIVE_FILTERS_KEY, scope discovery при этом ограничен dashboard-владельцем ключа. Browser session сохраняетnative_filters_keyв URL. Удалены вручную выставлявшиеся глобальныеSec-Fetch-*заголовки: они применялись к JS/CSS subresources и мешали dashboard React-приложению загрузиться в Playwright, хотя обычный браузер открывал страницу. browser_native_filter.pyтеперь связывает Ant Design label с.select-containerфильтра. Для labelCountryэто устранило miss: live canary применилCanada, дождался 5 charts, затем очистил фильтр и дождался 6 charts.- Evidence
specs/044-dashboard-scenario-execution/evidence/browser-provider/filters-canary-20260925T191720Z.json(+ PNG): inspect/apply/inspect/clear/inspect и replay checkpoint прошли; shutdownleftover_sessions=0,task_destroyed_warnings=[].G-BROWSER-FILTERSостаётся PARTIAL только из-за отсутствующего date-picker fixture; search structurally N/A для выбранного dropdown. - Проверено, что Omniroute
deepseek-flashпринимает screenshot: OpenAI-compatibleimage_urlrequest с живым PNG завершился HTTP 200 иIMAGE_CANARY_OK; фактически gateway сообщил route modelgpt-5.6-sol. Это опровергает прежний blanket вывод о text-only fallback/неработающих всех vision routes. Это отдельный provider-capability probe, не полный VLM scenario PASS. - Для
G-E2E-EVALUATIONостаётся полный scenario rerun с реальным evaluation spec, persisted AgentEvaluation → walker binding → terminal policy result, а также scheduled zero-human и injection proof. Общий продуктовый статус остаётся NO-GO; релизные регрессия, semantic health, BA readiness и signoff не закрыты.
Evidence для browser vectors: filters-canary-20260925T191720Z.json. Matrix и 044 quickstart/traceability отражают correction; предыдущие исторические checkpoints выше сохраняются как датированные наблюдения, но не как актуальный вывод о мультимодальной способности.