Files
busya bcc69f4bbe feat(dashboard-testing): add sampled traversal and ClickHouse test lab
Add versioned metric graph authority, owned browser evidence, paginated and all-tab traversal, deterministic sampling policies, and analyst-facing run inspection.

Provision the DEV/PREPROD/PROD Superset, Gitea and million-row ClickHouse lab; retain reproducible lifecycle evidence and explicit incomplete-traversal limits.
2026-10-02 10:54:42 +03:00

56 KiB
Raw Permalink Blame History

050-mcp-interface — Tasks

Правило: [ ] не начата; [~] в работе; [x] только с доказательством (команда + вывод).

Phase 0 — OAuth skeleton + /mcp probe

  • T001 Зависимость MCP SDK/FastMCP в backend/requirements.txt; каркас backend/src/mcp_server/ с монтированием /mcp в FastAPI app.
  • T002 Authorization Server на Authlib: /oauth/authorize (поверх существующей web-сессии: local password + ADFS OIDC), JWKS, AS metadata (RFC 8414).
  • T003 /oauth/token: authorization_code + PKCE S256 (обязателен для публичных клиентов), refresh_token с ротацией, client_credentials для service-principal.
  • T004 DCR (RFC 7591): регистрация публичных клиентов с first-party scope bound; список/отзыв в Admin.
  • T005 Resource Server: RFC 9728 metadata для /mcp, 401 + WWW-Authenticate, валидация access-JWT (aud=mcp, jti/blacklist) на каждый запрос; resolve_principal() адаптер.
  • T005a Transport and domain input limits: server-owned request byte/JSON-depth/collection/string limits, ScenarioGraph and SQL/DSL size limits, and per-session rate limits; rejected payloads create no ToolInvocationRecord, gate, draft, run, or mutation.
  • T006 Refresh rotation + reuse-detection: повтор ротированного refresh токена отзывaет семейство (расширение TokenBlacklist); fault-injection тесты.
  • T007 2–3 пробных инструмента (read-only: list_environments, get_health_summary, search_dashboards) с явной регистрацией и курируемыми схемами.
  • T008 CI: скриптованный клиент проходит discovery → DCR → PKCE → token → tools/list без ручных шагов (SC-007); подключение MCP Inspector; документация в README/INSTALL. Доказательство (remediation round 1, 2026-09-03): pytest tests/test_mcp_client_flow_http.py -q → 2 passed — оба потока целиком поверх реального HTTP: (a) machine-поток: RFC 9728 discovery → AS metadata (grant client_credentials, client_secret_post) → confidential DCR с одноразовым client_secret → client_credentials → signed aud=mcp токен с principal_type=service → /mcp initialize → tools/list (gated/human-only тулы невидимы) → tools/call; (b) пользовательский поток: public DCR → PKCE S256 GET /oauth/authorize с Bearer веб-сессии (документированный SPA-mediated consent-путь) → 302 code → обмен → identity-only токен → /mcp tools/list с live-RBAC (RUN_PROD-тул скрыт для RUN-юзера); wrong secret → invalid_client. Реализовано: грант client_credentials (services/mcp_oauth.py), oauth_clients.secret_hash (миграция 0017, idempotent-guard), service-principal short-circuit в McpTokenVerifier, AS metadata расширена. Документация: INSTALL.md §«MCP клиент». MCP Inspector подключается по тому же discovery-контракту (Streamable HTTP /mcp). Ограничение браузера-без-Bearer (cookie-consent HTML): РЕШЕНО 2026-09-04 — не строим в 050. Документированный и протестированный путь — SPA-mediated authorize: аутентифицированный фронтенд вызывает GET /oauth/authorize с Bearer веб-сессии (SC-007 «zero manual token management» выполнен: DCR → PKCE → code → токены автоматизированы полностью). Отдельная интерактивная consent-страница (login+consent на cookie-сессии) — новая frontend-поверхность со своим UX/i18n/security-ревью, непропорциональная first-party perimeter scope (MCPX-FR-020: только локальные клиенты). Fallback для browser-only клиента описан в INSTALL.md §«MCP клиент». Если browser-only third-party клиенты станут требованием — это отдельное продуктовое решение через architecture amendment (MCPX-FR-020 stand).
  • T008a DCR abuse controls: rate limits, redirect-URI validation, reviewed first-party scope bounds, admin audit/revocation, and tests proving registration cannot grant permissions or escalate scopes.
  • T008b Local-perimeter deployment tests and docs: reject or disable non-local MCP/LLM/VLM endpoint configuration by default; assert PII is permitted only for configured enterprise-local providers while credentials/cookies/tokens/secrets remain rejected everywhere. (Closure review 2026-09-03: тестов на отклонение non-local эндпоинтов не найдено; секрет-гигиена provenance подтверждена — PARTIAL.) Доказательство (remediation round 5, 2026-09-04): deny-by-default guard Core.EndpointLocality (src/core/utils/endpoint_locality.py) в choke points LLMProviderService.create_provider/update_provider — типизированный отказ endpoint_not_local:<reason> ДО персистенции (create не добавляет row, update не мутирует), route-слой маппит в HTTP 400 (не 500). Локальность: loopback/local-литералы, RFC1918/ULA private IP, enterprise DNS-суффиксы (.local/.internal/.lan/.corp/.intranet, расширяемо), DNS-имена с полностью приватным резолвом; нерезолвимое = fail closed; substring-spoofing (https://api.openai.com/localhost) отклоняется по hostname; пустой base_url (публичное SDK-облако по умолчанию) отклоняется. Escape hatches env-документированы (LLM_ALLOW_NONLOCAL_ENDPOINTS, LLM_NONLOCAL_ENDPOINT_ALLOWED_HOSTS, LLM_LOCAL_HOST_SUFFIXES), по умолчанию закрыты. MCP-транспорт: server-owned JSON-depth limit (typed 400 pre-dispatch) + per-session rate limit с Retry-After (E6) + существующие body-limit/DNS-rebinding defaults. Docs: INSTALL.md §«Локальный периметр». Тесты: tests/test_endpoint_locality.py (22) + tests/test_mcp_transport_limits.py (5) + provider/route slices → pytest tests/test_endpoint_locality.py tests/services/test_llm_provider.py tests/api/test_llm.py tests/api/test_encryption_health.py -q → 107 passed. PII-часть (секрет-гигиена provenance) подтверждена closure review round 1 и E13-отклонениями authoring chain.

Phase 1 — Parity catalog (37 tools)

  • T010 Каталог-реестр инструментов по доменам: env/health/tasks, git/deploy/migration/backup/maintenance, superset ops, baseline, scenario. Каждый инструмент декларирует required_permission(resource, action) из канонического словаря; единая серверная политика поглощает _tool_filter.
  • T011 Инструменты env/health/tasks (list_environments, get_health_summary, get_task_status, maintenance CRUD, llm status).
  • T012 Инструменты git/deploy/migration/backup. Доказательство: pytest tests/test_mcp_ops_parity.py tests/test_mcp_server.py tests/test_mcp_approvals.py tests/test_mcp_maintenance.py tests/test_mcp_oauth.py -q → 70 passed; зарегистрированы create_branch, commit_changes, deploy_dashboard, execute_migration, run_backup, run_llm_documentation, run_llm_validation (curated inputs, approval-gated; reviewed dispatchers src/services/mcp_ops_dispatch.py в явной цепи poll_approved_mcp_dispatches: GitService / TaskManager git-integration, superset-migration, superset-backup, llm_documentation, ValidationTaskService).
  • T013 Superset-инструменты (databases/explore/sql/format/permissions/dashboard+dataset CRUD) — SQL-класс помечен risk-классом и отдельным правом. Доказательство: тот же срез 70 passed; superset_list_databases, superset_explore_database, superset_format_sql, superset_audit_permissions (read), superset_create_dashboard, superset_copy_dashboard, superset_create_dataset (approval-gated), superset_execute_sql — permission ("plugin:superset_sql","EXECUTE") добавлен в rbac_permission_catalog.discover_declared_permissions + default-deny mapping; dangerous SQL отклоняется клиентским guard. Доп. свидетельство (орт. аудит 2026-09-02): guard расширен (safety.py: INTO/CALL/SET/REFRESH/VACUUM/REINDEX/ATTACH/LOAD/PREPARE/COMMENT + опасные функции lo_import/lo_export/pg_read_file/pg_write_file/pg_ls_dir/dblink/dblink_exec/set_config/pg_sleep + запрет мульти-стейтментов через ; после очистки строк/комментариев; tests/test_core/test_superset_safety.py — 50+ кейсов); PROD-критерий SQL-класса выравнен с resolve_environment_execution_policy (is_production OR stage=PROD) и переведён в терминальный отказ production_sql_execution_rejected вместо одобряемого, но недиспетчеризуемого гейта; pytest tests/test_mcp_ops_parity.py tests/test_core/test_superset_safety.py -q → 85 passed; полный срез аудита 214 passed; полный бекенд 11202 passed.
  • T014 Baseline-домен: capture_baseline_candidate, request/decide/consume_baseline_approval, create_verification_run (паритет fixtures с tools_037). Доказательство: тот же срез 70 passed; прямые обёртки над candidate_capture/candidates.request|decide_approval/verification_service.create_verification_run_async (те же сервисы, что и 037 REST-поверхность); consume_baseline_approval — baseline publish: approval-gated + reviewed dispatcher через one-shot consume_approval; полный backend suite python -m pytest -q → 11167 passed, 240 skipped, 1 xpassed; ruff/compileall чистые; каталог 45/45 зарегистрирован. Re-verified 2026-09-16: срез pytest -q tests/test_mcp_ops_parity.py tests/test_mcp_server.py tests/test_mcp_approvals.py tests/test_mcp_maintenance.py tests/test_mcp_oauth.py tests/api/test_mcp_parity_baseline_037.py → 84 passed; каталог вырос до 61 записи (version 2.3.0), baseline-домен на месте; consume_baseline_approval дополнен gated publish_baseline_catalog (T045, CLOSED 2026-09-11).
  • T015 Scenario-домен: scenario_compile/scenario_validate/scenario_resolve/generate_draft_pack/request_save/activate_revision/start_scenario_run (паритет с tools_038 и 042/044; save только из server-stored draft, activation отдельная CAS-операция).
  • T016 Контрактные тесты паритета: выводы MCP-инструментов сопоставлены с legacy-обёртками на общих фикстурах (SC-002). Доказательство: pytest tests/api/test_mcp_parity_baseline_037.py -q → 3 passed (re-run 2026-09-16 на HEAD 4746af2f → 3 passed in 1.79s): request/decide baseline-approval — MCP-gate валидируется моделью ApprovalGateResponse и совывает с REST по operation/risk_level/required_permission/status/reason_required/target_paths + actor parity на decide; create_verification_run — идентичные overall_status/category statuses/evidence_refs/created_by против REST /verification-runs; superset_format_sql — идентичная строка против legacy REST /api/agent/superset/sqllab/format (единственный дабл — [EXT:Superset] клиент). Попутно пойман и закрыт регрессией дефект 038-резолвера: scenario_resolve selector-путь падал на step.description is None (TypeError) — теперь selector_hint пишется без конкатенации с None (test_selector_hint_on_step_without_description).
  • T017 Bounded-response дисциплина: лимит инлайн-ответа, артефакты как ref+digest.
  • T018 Hidden/gated матрица: admin/analyst/viewer × каталог — отсутствие права скрывает инструмент из tools/list; role-change виден на следующем вызове без re-consent (SC-004, SC-009).

Phase 2 — Gates & provenance over MCP

  • T020 AgentAction-provenance на каждый вызов (ToolInvocationRecord): principal, tool, digest аргументов, исход, связь с AgentRun.
  • T021 approval_required конверт для негelegированных действий; durable ActionApprovalGate без side effects до решения (SC-003).
  • T022 Gate-инструменты как первичный путь полного цикла в клиенте: list_pending_approvals(filter) (+гейты от автоматизации) и decide_approval(gate_id, confirm|deny, reason) с CAS, обязательным reason для high-risk confirm, отказом service-principals и typed-ошибками на expired/replayed; web gate cards — равноправный рендер тех же строк (SC-003).
  • T023 E2E-walkthrough: внешний клиент создаёт сценарий фикстурного дашборда end-to-end → revision в registry (SC-001). Доказательство: pytest tests/test_mcp_scenario_e2e.py -q → 2 passed: один человеческий принципал с реальным scenario:EDIT/scenario:RUN (без RBAC-стабов) проходит tools/list-видимость → create_authoring_session → propose_graph_revision (server-derived op) → get_graph_diff → promote_to_scenario (awaiting_user_review) → request_save → ScenarioRevision(candidate) в registry → 038 inspect_scenario/scenario_resolve/validate_scenario поверх MCP → activate_revision → entry.current_revision_id продвинут отдельным CAS; второй тест — принципал без scenario:EDIT не видит и не может вызвать save/activate. Транспортный уровень (initialize → mcp-session-id → tools/list → tools/call) покрыт скриптовым клиентом T008.
  • T024 Gated-вызовы по контексту: PROD-окружение и baseline publish возвращают approval_required при видимом инструменте (hidden-vs-gated, MCPX-FR-018).
  • T024a Context revalidation: environment-policy and provider-binding fingerprints are rechecked at dispatcher admission and before provider I/O; target reclassification or binding drift invalidates a prior gate without I/O (SCEX-FR-026).
  • T025 Checkpoint-инструменты: list_checkpoints + decide_checkpoint(run_id, disposition, expected_version) через тот же CAS/аудит, что и монитор 045; user-principal only (service → permission_denied); тест что автоматизация не имеет пути к чекпоинтам (MCPX-FR-019). Доказательство (после closure-review داунгрейда, 2026-09-03): pytest tests/test_mcp_checkpoints.py -q → 3 passed: каталог ("scenario","RUN") + service_allowed=False на оба инструмента; сервис-принципал не видит и не может вызвать; list-проекция pending/decided чекпоинтов рана; decide повторяет REST /human/decision (pending-резолв на сервере, CAS decision_version, continue_after_human_decision, step→passed/HUMAN_*, run→queued/executing); stale expected_version → conflict без потребления; отсутствие pending/рана → checkpoint_not_found. (Closure review 2026-09-03 зафиксировал прежний фиктивный [x]; закрыт реальной реализацией.)
  • T026 AgentAuthoringWorkspace operations: create_authoring_session, propose_test_plan, start_exploration, get_exploration_result, propose_graph_revision, get_graph_diff, promote_to_scenario; persistent server-owned state, provenance, idempotency and workspace CAS.
  • T027 Authoring sandbox contract tests: isolated runtime, origin/API/action allowlists, no shell/credential/filesystem escape, limits, cancellation, receipts, artifact ownership and zero production side effects; retain exploratory traces/screenshots/diagnostics as bounded refs.
  • T028 Authoring promotion E2E: sandbox output -> typed proposal -> deterministic 038 compile/validate -> user diff review -> 042 handle-based save -> immutable revision; reject raw code/URLs/cookies/secrets/paths/caller digests and keep code-backed production execution unimplemented. Доказательство: pytest tests/test_mcp_authoring_promotion_e2e.py -q → 8 passed: полный прогон через tools/call — start_exploration теперь ставит queued при зарегистрированном раннере (wiring provider_available=get_registered_runner() is not None; регистрация default_runner+deployment context в bootstrap_live_execution_composition) → реальный execute_scheduled_exploration_dispatch из потокового контекста планировщика → exploration_passed с observations/proposed_graph/evidence_ref (draft:exploration-*) при [EXT:Browser]-дабле; get_exploration_result остаётся ограниченной проекцией без утечки наблюдений; значение observed_dashboard_title в сохранённой ревизии выводится только из наблюдения; propose→promote→diff→request_save→activate завершается current-ревизией без продвижения до отдельного CAS-активации. Негативные ветки: typed ops с SQL/..//\\/drop отклоняются до предложения и CAS; exploration-spec с code-токенами и незарегистрированными действиями не персистируется; GraphRevisionInput/PromoteScenarioInput/RequestSaveInput/ActivateRevisionInput/ExplorationInput структурно отвергают digest/content_hash. Дополнительно: диспетчерский soak (test_mcp_ops_parity.py) — 3 цикла поллера, каждое одобренное действие диспетчеризуется ровно один раз (14 passed). Полный срез: 83 passed (8 файлов MCP-вертикали); полный backend suite 11182 passed.

Phase 2b — Initial scenario and automation parity

  • T029 Implement ScenarioRegistry.CreateInitial and bootstrap_authoring_scenario; prove a fresh external MCP client creates a first registry scenario without a seeded base revision. Evidence: tests/test_mcp_initial_scenario_e2e.py creates real AgentRun/DraftArtifact rows and drives bootstrap_authoring_scenario without a seeded registry; replay and conflicting idempotency are covered by tests/test_mcp_t029_bootstrap_automation.py.
  • [~] T029a Partial: strict profile preview derives from fresh server inspection and returns case coverage, deterministic metric coordinates and typed questions; runtime catalog 2.6.0 exposes propose_test_pack_profile and resolve_test_pack_profile as authenticated-human read previews (service principals denied). Digest-CAS accepts exact-step selector hints and server-issued metric coordinate choices, re-inspects/recompiles, and transitions chosen metric to needs_baseline without expected values or a pin. Accepted amendments now persist owner-scoped profile/CAS/idempotency state, bind a server-owned eligibility receipt to the canonical handle chain, reject disallowed environments and mismatched agent-run handles, and keep the Alembic chain linear/idempotent. Evidence: test_test_pack_profile.py, test_compiler.py, test_handles.py, test_mcp_t029_bootstrap_automation.py, profile MCP E2E in test_mcp_scenario_e2e.py; profile-focused tests 9 passed, selected MCP/profile run 9 passed and 2 failed. Still open: durable authoritative-context and published-baseline resolution, fully proven save-eligible handle minting, fresh external-client profile→resolve→register→bootstrap E2E, and the remaining baseline/context and external-client vectors. Do not mark complete until those vectors pass. Working-tree continuation (2026-09-30): catalog 2.7.0 adds optional profile_handle_id to register_draft_pack; the server rebuilds a save-eligible graph from an owner-scoped profile, rechecks context and binds session/CAS/digest in the receipt. Bootstrap checks the bound receipt. Sequential profile CAS and unrelated metric blocking were corrected; unbound required launch parameters now warn without blocking authoring, per 038/044. The two named regressions pass (2/2); targeted MCP/registry selection passes 217 with one stalled HTTP initialize test deselected. A synthetic isolated in-process MCP test reaches selector-only eligible registration and current bootstrap without a caller graph, and rejects changed inspection fingerprint before writes. Superset inspection is mocked; full external wire-client, baseline/context and release proofs remain open. Ruff, compileall, diff-check and prior AXIOM verify passed. Current checkpoint (2026-09-30): a separate locator-only needs_baseline answer now stores a server-derived exact-entry/release/filter selection under owner CAS and canonical digest. The URL remains transient; the question stays needs_baseline and the profile stays preview_only. Graph/receipt binding requires a typed 038 v2 coordinate on the metric producer plus same-case comparison-edge validation; the compiler's baseline.default cannot yet identify the selected coordinate. The user-requested sales dashboard release is also open: Superset slug resolves to dashboard 11, but this local release ledger has no deployment/release and no configured Git PREPROD stage. ss-prod is the configured Git PROD target; the earlier canary PREPROD label did not configure that release stage. See the binding plan and release checkpoint. Operational update (2026-10-01): The separate supported M01 v2 path now completes the isolated Docker DEV → Gitea → PREPROD → PROD release and published-baseline chain, fresh external MCP owner/CAS profile → register/bootstrap, and actual manual PASS → changed-data FAIL → restored PASS → independent cron PASS. Its table_text_v1 extension also completed owned browser table/metric evidence, real Omniroute text evaluation and the same four verdicts; exact runs and independent receipt checks are in Stage 6 status. T029a remains partial for the full catalog, external dashboard 11, release acceptance and unsupported contexts; the 2026-09-30 lines above are historical checkpoints, not current M01 runtime status.
  • T029b Expose 046 schedule, trigger-rule, policy and automation-metrics management through curated MCP tools with REST-equivalent RBAC, idempotency and PROD gate behavior. Evidence: curated tools and catalog entries in src/mcp_server/tools_automation.py/rbac_server.py, idempotency migration alembic/versions/0018_automation_idempotency.py, and tests/test_mcp_t029_bootstrap_automation.py plus RBAC catalog tests.
  • T029c Add fresh-DB MCP E2E for bootstrap → visible registry entry → revision activation → manual run, plus scheduler eligibility/PROD-gate integration evidence. Evidence: tests/test_mcp_initial_scenario_e2e.py verifies fresh bootstrap, current revision, refusal to run an un-promoted bootstrap revision (BOOTSTRAP_REVISION_NOT_RUNNABLE), a pinned manual run against a genuinely materialized revision, human-step schedule rejection with zero side effects, and an idempotent PROD approval gate; focused MCP regression set: 145 passed (2026-09-06).

Phase 2c — Server-owned handle layer (root-cause closure of ADR-0023; contracts: 038 ScenarioGraph.ServerOwnedPipeline amendment 2026-09-06, 042 ScenarioRegistry.DataModel amendment, 050 stage table status note)

  • T029d Persist the 038 handle layer: CompiledScenarioHandle/ValidationResultHandle/DraftPackHandle tables (immutable, append-only, owner_principal+agent_run binding, digest/content_hash columns per 038 amendment) + Alembic migration + server content-store persistence of the canonical DashboardTestScenario bytes (ScenarioGraph.Models.CanonicalBytes) behind canonical_bytes_ref. Minting happens ONLY at the persisted REST boundaries (api_compile_scenario/api_validate_scenario/api_resolve_scenario/api_draft_pack); pure compiler functions keep @SIDE_EFFECT None. Single-consumption (consumed_by_revision_id), cross-binding rejection, GC via 036 draft-retention. Evidence: src/models/scenario_handles.py, src/services/dashboard_testing/scenario/handles.py (idempotent mint / verify / consume / purge), migration 0019_scenario_handles, REST minting in api/routes/dashboard_testing/scenario.py, authority tests tests/services/dashboard_testing/scenario/test_handles.py (3 passed).
  • T029e 042 create consumes handles: create_scenario/create_initial re-verify handle ownership/binding/save_eligible, materialize graph_snapshot = canonical DashboardTestScenario JSON from handle bytes inside the create transaction, and write OutboxEvent(type=materialize_revision) + RevisionMaterialization(pending); implement the idempotent reference-artifact worker (scenario.yaml/reference runner.plan.json) or record an explicit descope decision. BOOTSTRAP_REVISION_NOT_RUNNABLE demotes from primary guard to defense-in-depth; api_create_scenario stops hardcoding materialization_status="materialized". Evidence: handle-first branch in registry/create.py (materializes full canonical graph + server-owned action_registry_version/hash), src/models/scenario_materialization.py + migration 0020_scenario_materialization, worker registry/materialize.py::materialize_pending_revisions (idempotent, pending→materialized/failed), api_create_scenario returns the actual RevisionMaterialization.status.
  • T029f MCP minting surface: new register_draft_pack write tool returns bounded server-issued handle ids; bootstrap_authoring_scenario accepts ONLY stored handle ids (transitional compile:{run_id}:{digest} string check removed from create_initial); catalog bumped 1.0.0 → 2.0.0 (breaking bootstrap handle semantics) with pinned-major ritual PINNED_CATALOG_MAJOR=2. Legacy REST POST /scenarios/{id}/revisions raw-graph_snapshot path retired (returns 410 GONE REVISIONS_RAW_GRAPH_RETIRED; the server-side create_revision service remains for the 043 editor save path). Also reclassified generate_report out of _MUTATING_ACTIONS (local draft report, no external side effect), bumping ACTION_REGISTRY_VERSION 038.1.0 → 038.2.0 (fingerprint + scenario_execution/graph.json fixture updated).
  • T029g Full-chain fresh-DB MCP E2E without REST crutches and without hand-seeded revisions: register_draft_pack→bootstrap_authoring_scenario→direct start_scenario_run on a bootstrapped current revision (now materialized) returns queued. Evidence: tests/test_mcp_initial_scenario_e2e.py rewrote _pack() to drive the MCP register_draft_pack tool. PostgreSQL concurrency: verify_handle_chain uses SELECT ... FOR UPDATE + populate_existing() to serialize handle consumption, proven by tests/integration/test_scenario_handle_concurrency.py (one winner + one HANDLE_CONSUMED across two threads on a real PostgreSQL container; --run-integration green). (Amended 2026-09-07, Doc.Adr.ADR0024/T029j: the claim "without REST crutches" did not cover the AgentRun prerequisite — _pack() inserted the AgentRun row with a raw ORM constructor unreachable for any external MCP client; the field run of 2026-09-07 proved register_draft_pack therefore denied every real client. _pack() now creates the run through the production create_agent_run service boundary with honest test metadata; full external reachability closes at T029i and is pinned by tests/test_mcp_agent_run_reachability.py (strict xfail).)
  • T029h Inspect-stage decision: resolved as hybrid option C (2026-09-06) — no persisted InspectionContextHandle; instead (a) new MCP read tool inspect_dashboard_context exposes the existing live resolver (BaselineEngine.QueryModel.Inspect via GET /query-model service) returning the full authoritative DashboardQueryModel + fingerprint for the agent to echo into compile; (b) the register boundary (register_draft_pack) evaluates context_authority by RECOMPUTING the fingerprint from the client-carried dashboard_context.query (claimed fingerprints are ignored — C2) against a live inspection, with sentinel/degraded fingerprints ("", sha256:error) never verifying (C1), unreachable/unconfigured environments failing OPEN to unverified (C3), and falsifiable model-shape claims on live environments rejecting typed (CONTEXT_FINGERPRINT_MISMATCH, CONTEXT_QUERY_MODEL_REQUIRED, CONTEXT_ENVIRONMENT_MISMATCH) — C5 inspect-first; (c) the marker persists on DraftPackHandle.context_authority (migration 0021_context_authority), materializes into graph_snapshot at 042 create, and ScenarioExecution.Runner.Start refuses PROD dispatch on an explicit non-verified marker (CONTEXT_AUTHORITY_REQUIRED_FOR_PROD, zero side effects; missing marker = legacy, allowed); (d) X1 smuggling closure: validator _check_dashboard_context recursively rejects query_context keys and SQL text inside dashboard_context. Catalog minor bump 2.0.0 → 2.1.0 (additive, pinned major untouched). Evidence: tests/services/dashboard_testing/scenario/test_context_authority.py (7), validator X1 tests (+2), test_prod_start_enforces_context_authority_marker, E2E marker-propagation + PROD-block assertions; broad regression 1338 passed; alembic heads → 0021_context_authority. Residual: editor promote/save path (request_save revisions) does not yet carry the marker — PROD gate treats it as legacy-allowed; tightening requires authority evaluation in the 043 save boundary (follow-up, not scoped by T029h).

(Orthogonal code review 2026-09-06, ADR-0023; specs amended the same day and the durable handle layer implemented — 038 contracts/modules.md, 042 data-model.md/contracts/modules.md, 050 spec.md): the bootstrap→run happy path is realizable through MCP alone (register_draft_pack → bootstrap_authoring_scenario → direct start_scenario_run on a materialized current revision). Phase 2c residuals now closed: api_create_scenario returns real materialization_status, legacy REST POST /scenarios/{id}/revisions retired (410 GONE), PostgreSQL single-consumption concurrency proven (SELECT ... FOR UPDATE + populate_existing(); tests/integration/test_scenario_handle_concurrency.py green on real PostgreSQL). T029h resolved as the hybrid (see above); follow-up residuals closed 2026-09-07: outbox worker wired into the scheduler poll loop (scenario_revision_materialization, 30s, singleton/coalesced, durable-tick test), REST api_draft_pack evaluates context_authority at parity with the MCP register boundary (async, typed 422 on falsifiable mismatch), and the 043 editor save path provably inherits the server-owned marker (test_save_proposal_inherits_server_context_authority_marker). Known-accepted limitation: legacy/NULL markers stay PROD-allowed by design; full-suite regression 3360 passed.)

Phase 2d — Field-run remediation (2026-09-07 sales-prod external MCP run; ADR-0024; spec: MCPX-FR-027/028/029, release-gate rows E2E-EXT-001/002, CAP-001, DISP-001, TEST-001)

Источник: docs/2026-09-07-sales-prod-mcp-run.md — первый полевой прогон внешнего MCP-клиента против ss-prod Sales Dashboard (ID 11). Все звенья цепи реализованы, но external-клиент упёрся в DRAFT_PACK_ACCESS_DENIED: register_draft_pack требует principal-owned AgentRun, а MCP-операции создания AgentRun не существует (единственная creation-поверхность POST /api/agent/runs — REST/web-session, MCP-токены aud=mcp на ней не аутентифицируются). Вертикальный E2E маскировал разрыв raw-ORM-seed'ом.

  • T029i MCP AgentRun surface (MCPX-FR-027): typed create_agent_run + get_agent_run tools wrapping the existing Services.AgentRuns.Service.Create/snapshot boundary — curated bounded input (UIContext v2: objectType=dashboard, numeric objectId, envId, route, contextVersion=2, intent=build_dashboard_test_scenario; extra=forbid; optional idempotency_key), permission ("dashboard:testing","EXECUTE") at REST parity, service_allowed=False, idempotent active-run reuse, bounded snapshot projection (get_agent_run returns the ownership-scoped AgentRunSnapshot fields only). Catalog additive bump 2.1.0 → 2.2.0; update SC-004 exact-set fixtures (tests/test_mcp_rbac_visibility.py) and tests/api/test_admin_mcp_catalog.py mirrors. Closure evidence: convert tests/test_mcp_initial_scenario_e2e.py to the FULLY external chain (create_agent_run → register_draft_pack → bootstrap_authoring_scenario → start_scenario_run), remove the strict=True xfail marker in tests/test_mcp_agent_run_reachability.py and make it green (E2E-EXT-001). Доказательство (2026-09-07): новый seam src/mcp_server/tools_agent_run.py (McpServer.ToolsAgentRun, 155 LOC) — create_agent_run (server-pinned objectType/route/contextVersion/intent; caller владеет только dashboard/environment/name/idempotency_key → conversation-reuse; продуктовая форма run: status RUNNING + run_started event) и get_agent_run (bounded-проекция + draft_count; чужой/неизвестный id → typed not_found без existence oracle); catalog +2 (EXECUTE/READ, human-only), MCP_CATALOG_VERSION=2.2.0 (PINNED_CATALOG_MAJOR=2 не тронут); registration seam в _build_probe_server (automation → agent-run → scenario). Тесты: tests/test_mcp_agent_run_tools.py (5: продуктовая форма+envelope keys, idempotent replay + env-scoping, structural rejection, ownership read + not_found, RBAC denial by name); test_mcp_initial_scenario_e2e.py конвертирован в fully external chain (create_agent_run через tools/call, ноль не-MCP seeding); xfail-маркер в test_mcp_agent_run_reachability.py снят — requirement pin зелёный на живом каталоге; catalog-pin test_mcp_server.py +2 имени +4 assertion. RBAC-зеркала правок не потребовали (analyst без EXECUTE/READ —derived-наборы). Gate: 8-файловый срез 70 passed; полный backend suite 11357 passed.
  • T029j Test-honesty remediation (ADR-0024 §3, binding rule): vertical/E2E tests obtain every prerequisite through a boundary the principal under test can reach; raw-ORM seeding of a chain prerequisite in a test claiming external reachability is forbidden. Executed: test_mcp_initial_scenario_e2e.py::_pack() replaced the raw AgentRun(...) insert with the production create_agent_run service boundary with honest test metadata; new tests/test_mcp_agent_run_reachability.py pins (a) the external-reachability REQUIREMENT as strict=True xfail (flips the suite red the moment create_agent_run lands without spec follow-through) and (b) the current typed zero-side-effect denial (DRAFT_PACK_ACCESS_DENIED, no handle rows). Evidence: see TEST-001 row in spec.md release gates (targeted pytest green 2026-09-07). (Update 2026-09-07, T029i closure: строгий xfail-пин сработал как спроектирован — конвертирован в обычный requirement-тест (unmarked, green) вместе с конвертацией вертикали в fully external chain; denial-pin сохранён.)
  • T029k Context-authority-derived capability map (MCPX-FR-028; 038 amendment 2026-09-07): server-owned derivation of capabilities/has_dataset_fields from the authoritative DashboardQueryModel (the same live inspection context_authority binds), environment policy and provider readiness — exposed through inspect_dashboard_context (derived-capability section) and consumed by the persisted compile boundaries (MCP scenario_compile + REST api_compile_scenario); caller-declared capabilities remain accepted only as an explicit subset narrowing, never as an authority that downgrades verifiable facts into human_checkpoint; unresolved facts stay needs_context/needs_selector/needs_baseline; genuinely unsafe PROD mutation contexts keep human_checkpoint; HumanStep revisions remain automation-ineligible (SCEX-FR-004a stands). Evidence: CAP-001 row + capability-derivation unit/E2E tests. Доказательство (2026-09-07): новый src/services/dashboard_testing/scenario/capability_authority.py (236 LOC, ScenarioGraph.CapabilityAuthority): derive_capabilities — только верифицируемые факты (native_filters; text_filter по STRING-фильтру; time_rollover по DATE/TIME/TIME_GRAIN; table_filter+pagination по executable table-viz; xlsx_export из capabilities-флага; dataset_field_read+has_dataset_fields по accessible dataset с колонками; browser — только при readiness-факте из T040-снимка composition-root), NEVER_DERIVED = row_edit/bulk_edit/persistence_refresh/safe_test_data/safe_clock_fixture/cross_dashboard (unsafe-mutation автоматизация метаданными невозможна — module @INVARIANT); merge_capabilities — derived-wins в ОБЕ стороны (declared-false не понижает проверяемую правду; declared-true не фабрикует) + overrides-аудит; legacy/unparseable payload → дословный caller-declared passthrough (register-time context_authority остаётся жёстким гейтом). Wiring: MCP inspect_scenario (+additive capability_authority секция), inspect_dashboard_context (+derived_capabilities), REST api_compile_scenario (parity через общий build_capability_authority choke point + additive секция). Тесты: фикстура query_model_sales.json (форма sales-стенда); unit truth-table tests/services/dashboard_testing/scenario/test_capability_authority.py (7, включая CAP-001 classification-fix через map_all: полевого shape декларации → B01–B04/T01–T03 automated, C04–C06 unsupported, B05–B09/C01–C03/C02/C07 легитимно human_checkpoint); tool-level tests/test_mcp_capability_authority.py (3: derived-wins overrides ["text_filter","xlsx_export"], legacy caller_declared, derived_capabilities в inspect-ответе); REST-parity pin test_scenario_routes.py::test_capability_authority_derived_wins. Браузер-readiness seam запинен monkeypatch (детерминизм против глобального composition-root состояния в full-suite). Gate: 5-файловый срез 67 passed; полный suite 11357 passed.
  • T029l Disposition vocabulary clarity (MCPX-FR-029; 044/045 amendment 2026-09-07): the 044 lifecycle mapping (confirm→passed, false_positive/inconclusive→inconclusive, aliases pass/fail) is immutable and stays; human-facing wording aligns to the persisted outcome — RU labels «Подтвердить соответствие» / «Проблема не подтверждена» / «Недостаточно данных» (+ EN parity), confirm-button restyled from bg-destructive to a positive token, decide_checkpoint/decide_approval MCP tool descriptions and 045 monitor docs carry the outcome table, vitest pins updated (WaitingForMeView.test.ts, HumanCheckpointPanel tests). Evidence: DISP-001 row. Доказательство (2026-09-07): waiting_disposition_confirm/false_positive в ru/en dashboard-testing.json именуют персистентный исход (confirm→passed: «Подтвердить соответствие»/"Confirm conformance"; false_positive: «Проблема не подтверждена»/"Issue not confirmed"; inconclusive без изменений); confirm-кнопки HumanCheckpointPanel.svelte и WaitingForMeView.svelte переведены bg-destructive → bg-primary (прецедент approve-кнопки ApprovalDecisionPanel) + @RATIONALE в контрактах компонентов; API-вокабуляр не переименован (continuity аудита/CAS); MCP decide_checkpoint docstring (видим в tools/list) несёт immutable outcome-mapping и правило «confirm = проверка пройдена, никогда не подтверждение дефекта»; vitest-пины обновлены (RunMonitorViews.test.ts — новый DISP-001 pin лейбла+стиля+dispatch "confirm", WaitingForMeView.test.ts, run.ux.test.ts). Lifecycle-маппинг не менялся — backend-тесты decision-пути зелёные без правок. Gate: frontend 3507 passed (206 файлов), lint 0 errors / 364 warnings (baseline), npm run build OK.
  • T029m Live-stand replay of the field run: after T029i/T029k/T029l, re-run the 2026-09-07 sales scenario externally against the live stand — inspect_dashboard_context → derived-capability compile/validate → create_agent_run → register_draft_pack → bootstrap_authoring_scenario → start_scenario_run (PROD gate observed, no unexpected high-risk approvals) → list_checkpoints/decide_checkpoint human loop with aligned labels; retain the run report beside docs/2026-09-07-sales-prod-mcp-run.md. Evidence: E2E-EXT-002 row. Status (2026-09-11): CLOSED. Committed replay client specs/044-dashboard-scenario-execution/prototype/live_mcp_replay.py (transport-only adaptation of the proven scripted OAuth/PKCE + tool-chain flows, no new client machinery) replayed the FULL external chain on the live stand with zero non-MCP seeding: inspect (derived capabilities) → create_agent_run a6168c7d… → compile B01 (capability_authority derived) → validate → draft-pack save_eligible → register_draft_pack (context_authority verified) → bootstrap (scenario e2687f03…, current revision) → PROD start pending_approval (identical retry = one durable gate) → approval → live execution to an honest typed terminal (capture_screenshot passed with 8 durable refs; apply_native_filter typed BROWSER_ACTION_NOT_SUPPORTED — no synthesized PASS). Full trace + two fail-closed defects the replay exposed and fixed (binding resolution for identity-less compiled steps; per-step target-identity stamping at bootstrap materialization) + D5 binding-admin live exercise: docs/2026-09-11-sales-prod-mcp-replay.md. The human-checkpoint loop was exercised live by the companion live_mcp_human_loop.py (compiled B05 HumanCheckpoint → PROD gate → waiting_human → MCP list_checkpoints decision_version 1 → decide_checkpoint confirm decision_version 2 CAS → terminal passed), so every element of this task's formulation is covered. The baseline-pinned variant remains blocked on the Gitea PAT (050 T045 / T046-pin).

Phase 3 — Frontend decommission (flag-driven)

  • T030 Replace historical HandoffSurface/agent entry with ordinary manual scenario editor; remove copy-prompt, agent launch/workspace and proposal-generation UI. Previous handoff implementation is history, not current acceptance.
  • T031 Скрыть /agent, AssistantChatPanel, кнопку «Ассистент» в TopNavbar за флагом; обновить link-integrity тесты.
  • T032 Retire /api/assistant/* и прокси /api/agent/gradio; retention-настройки assistant скрыть. (Closure review 2026-09-03: на HEAD роутеры /api/assistant и agent_* смонтированы безусловно (app.py), retention-UI в SystemSettings жив, frontend api/assistant.ts в поставке — фиктивный [x]. Флаг удалён в T041, поэтому retirement теперь безусловный.)
  • T033 vitest/build/link-integrity зелёные при включённом флаге демонтажа.

Phase 4 — Removal

  • T040 Удалить сервис agent/ из run.sh и compose-профилей (порт 7860); обновить AGENTS.md/INSTALL.md. Доказательство: rg "agent|7860" run.sh build.sh docker-compose.yml docker-compose.enterprise-clean.yml → пусто (кроме исторического секьюрити-комментария); bash -n/yaml.safe_load чистые; стенд ./run.sh --skip-install поднимает только :8000+:5173 (7860 отсутствует, curl /api/ready → ready, Playwright-проход до хендоффа зелёный); run.sh/build.sh/AGENTS.md/INSTALL.md обновлены; из build.sh убраны build:agent, bundle:agent, bundle:embeddings и агентский сервис генерируемого деплой-композа/манифеста; из обоих nginx-конфигов убран location /api/agent/gradio.
  • T041 Удалить код чата: agent/src (app, langgraph_setup, tools*.py, _confirmation, middleware...), frontend agent/assistant компоненты, i18n, типы. Доказательство: agent/ удалён целиком; удалены docker/Dockerfile.agent, docker/agent.entrypoint.sh, backend/tests/test_gradio_proxy_config.py; во фронтенде удалены components/assistant/* (кроме универсального MarkdownRenderer), чат/ран/драфт артефакт компоненты, components/agent/dashboard-testing/ (ScenarioWorkspace-ветка), models/AgentChat*, AgentRunModel, DashboardScenarioWorkspaceModel, stores/assistantChat, флаг MCP_DECOMMISSION (vite define/config/mcp.ts/global.d.ts), gradio-прокси из vite.config.js; /agent рендерит только HandoffSurface (безусловно), навбар-кнопка ведёт на хендофф; npm run test -- --run → 3454 passed, npm run lint → 0 errors, npm run build → OK; полный бекенд 11199 passed (минус тесты удалённого грдио-прокси), 16 сиротских тестовых файлов чата удалены.
  • T042 Финальные правки спек 036–047: перенести drift-amendments из статуса «planned» в «done» со ссылками на доказательства. Доказательство: во всех 12 спеках (036–047) секции ## Drift Amendment — MCP Interface получили строку **Status (2026-09-02): done** со ссылками на specs/050-mcp-interface/tasks.md (T012–T028), T030–T033 и чекпоинты specs/WORKSTATE-043-047.md. Re-verified 2026-09-16 (grep): 11 spec-файлов несут секцию Drift Amendment; 040/041 — дословный маркер Status (2026-09-02): done; 036/037/038/042–047 — пост-closure-review формулировка «reported done (2026-09-02); not production-readiness evidence» с теми же ссылками (более строгая, принята как эквивалент).
  • T043 Полный прогон backend/frontend suites + стенд без 7860 (SC-005). Доказательство: python -m pytest -q → 11199 passed, 240 skipped, 1 xpassed; npm run test -- --run → 3454 passed (197 файлов); npm run lint (0 errors) + npm run build — зелёные; локальный стенд после демонтажа работает без порта 7860 (см. T040) и отдаёт рабочий хендофф на /agent. Browser E2E (изолированный стенд): docker compose -p ss-tools-e2e --env-file .env.e2e -f docker-compose.e2e.yml up -d --build db backend frontend (свежая БД, bootstrap admin/admin123, backend :8103 healthy, frontend :8102 healthy, порт 7860 отсутствует) → npx playwright test e2e/tests/login.e2e.js e2e/tests/agent.e2e.js → 6 passed (двойной прогон, Chromium): login-поток (форма/успех/ошибка неверных кредов) и post-decommission агентский роут (хендофф-поверхность, ноль чат-элементов, deep-link с context-параметрами остаётся на хендоффе); стенд снесён down -v. E2E-фикстарелы: устаревший ambiguous locator('nav') (strict mode: 3 nav-элемента) → .first() в login/smoke; regex ошибки входа дополнен incorrect|неверн (бэкенд отдаёт passthrough-detail); agent.e2e.js переписан под handoff-контракт (T01-T03), dashboard-scenario-ui/agent-scenario-run — ссылки на чат-UI заменены на handoff-маршрут. Дополнительно: живой admin-walkthrough на dev-стенде (8000/5173) подтвердил login→dashboards→handoff(0 textarea/0 conversation-узлов)→runs-center (скриншоты /tmp/kilo/happy-path/09–12).

Production readiness — 2026-09-08 (MCPX-FR-030)

Historical [x] rows above retain only their dated local/transport evidence; they do not prove current production readiness. Reopened rows were contradicted by the audited gaps. Removed frontend/agent paths are historical, not implementation prerequisites. New acceptance is implemented=false / OPEN.

Contract: Contract-complete public parity.

  • T044 [P0/P1/P2] REST/MCP lifecycle/read/auth errors and disabled automation validation are identical; service principal cannot decide human gate. Implement at the existing 050 domain boundary; verify with independent hardcoded fixtures and retain command/evidence references in traceability.md. Status (2026-09-11): offline parity tests green; live REST-only canary v2 exercised start→gate→approve→dispatch (4eebfab3); the live MCP-external replay (T029m CLOSED, docs/2026-09-11-sales-prod-mcp-replay.md) exercised the same lifecycle through /mcp with the identical typed start/gate/terminal outcomes — remaining: dedicated REST-vs-MCP error-shape parity fixtures for this matrix.
  • T045 [P0/P1/P2] consume/publish failures return typed errors/pending state with no legacy fallback; every prerequisite is externally MCP-reachable. Implement at the existing 050 domain boundary; verify with independent hardcoded fixtures and retain command/evidence references in traceability.md. Status (2026-09-11): CLOSED. Consume path typed fail-closed (bbbd4ccf, live-proven); publication is now externally MCP-reachable through the curated gated tool publish_baseline_catalog (human-only, scenario RUN_PROD, requires_approval) executing the 037 publication-worker contract via ScenarioExecution.PublicationWorker + durable publication_operations (migration 0023): idempotency by key, envelope validation (resolver-canonical), branch-head CAS, commit/published receipts, failure retention + reconcile of the same commit; REST parity POST/GET /api/catalog-publications. LIVE evidence (docs/reports/agentic-runtime-live-mcp-publish-t045-2026-09-11.md): MCP gate approval_required → decide_approval → dispatch → published (commit abce63db…, receipt head 3f782574→abce63db) + moved-HEAD canary → typed PUBLISH_HEAD_MOVED with a durable publish_failed row. Offline: tests/services/dashboard_testing/execution/test_publication_worker.py, tests/api/test_catalog_publications_api.py, tests/test_mcp_ops_parity.py.
  • T046 [P0/P1/P2] Fresh external-client chain preserves authoritative context and complete baseline pin without raw ORM/REST repair; no frontend agent controls/routes/requests. Implement at the existing 050 domain boundary; verify with independent hardcoded fixtures and retain command/evidence references in traceability.md. Status (2026-09-11): CLOSED. Context authority: fresh REST chain (canary v2 4eebfab3) and fresh MCP chain (T029m) preserved the server-resolved binding and verified context (context_authority=verified). Complete baseline pin: (a) the MCP-external chain with --baseline (live_mcp_replay.py --baseline, no raw ORM seeding) compiled the baseline capability, started with the published-catalog selector, and carried a full runner_plan.baseline_pin (set/version/digest from the published envelope); (b) the REST canary v4 (ace916a0…, then re-run adeabe63…) proved the walker stamps that same pin into the persisted AgentEvaluation.baseline_pin (plan pin == record pin, strict equality asserted) — docs/reports/agentic-runtime-live-canary-v4-baseline-pin-2026-09-11.md. No raw ORM/REST pin repair; frontend agent controls remain absent (2026-09-10 evidence).

Design Amendment implementation tasks — 2026-09-17/18

  • T047 [P0] Investigation MCP tools + RBAC/human-only writes + catalog 2.4.0. CLOSED: 23414f9a + 5eea72bc, 73 focused green.
  • T048 [P0] set_step_inputs and set_step_evaluation through propose_graph_revision/EditOperation parity. CLOSED: 534d488f/f123210b + 87a90f6f.
  • T049 [P1] propose_baseline_selection MCP surface for curated coordinates/review handle; observatory non-gating.
  • T050 [P1] exploration_step read-only sandbox action with CAS/receipts; reuse 044 browser session; no second Playwright stack.
  • T051 [P0] Complete REST/MCP lifecycle/error parity (prior T044 residual), including disabled-action/investigation tools. CLOSED 2026-09-22: shared classify_start_error in start_run.py consumed by REST and MCP (removes the raw-text vs RUN_START_CONFLICT divergence); tests/test_mcp_rest_error_parity.py 12+ vectors (env/baseline codes/idempotency/conflict/disabled action/disabled automation/investigation RBAC/unauthenticated + service-principal human-gate denial + MCP no-bypass on the case ACL); 17 green.
  • T052 [P0, MCPX-FR-035] Expose the server-owned reference-URL resolver and selection/capture path to external MCP clients with REST parity; require URL in new capture contracts, preserve parsed identity through catalog/pin, reject foreign dashboards and caller-supplied dataMask/expected values. Prove both supplied ss-dev URL forms and URL→scheduled-run replay without re-fetching key. Contract recorded 2026-09-24; implementation and acceptance OPEN.

Frontend boundary for this package: manual CRUD/editor, human review/approval, monitoring and read-only evidence/evaluation only; all agent interaction is external MCP. No agent chat/prompt/assistant editing/proposal generation/workspace/start/handoff controls. Runtime removal is OPEN, not performed by this spec refresh. Optional approved performance baseline is outside scope.