Files
ss-tools/specs/050-mcp-interface/spec.md

61 KiB
Raw Blame History

#region McpInterface.Spec [C:3] [TYPE ADR] [SEMANTICS mcp,interface,tools,agent,dashboard-testing,decommission] @BRIEF Simple MCP interface to ss-tools that replaces the Gradio chat agent as the primary agentic surface, including full dashboard-test creation. @RELATION DEPENDS_ON -> [Doc.Adr.ADR0001] @RELATION DEPENDS_ON -> [Doc.Adr.ADR0005] @RELATION DEPENDS_ON -> [Doc.Adr.ADR0006] @RELATION DEPENDS_ON -> [AgentTestStabilization.Spec] @RATIONALE ss-tools already concentrates truth and policy server-side: every agent tool is a thin forwarder over backend services (RBAC guard -> httpx -> truncated result), and 042-047 demand that any delegated actor works through the same server-owned contracts as an analyst. A conversational runtime coupled to a bespoke Gradio UI adds a parallel transport, a parallel HITL mechanism (LangGraph interrupt + confirmation module) and triple bookkeeping per capability (route + LangChain wrapper + allowlist). MCP externalizes reasoning to standard clients while the backend keeps deterministic validators, ActionApprovalGate, AgentAction provenance and InvestigationSignal semantics unchanged. @REJECTED Keeping the Gradio chat agent long-term — rejected 2026-08-24: the chat surface duplicates policy transport, blocks external clients, and every new spec (042-047) multiplies hand-written wrappers. @REJECTED Personal access tokens as the primary authentication mechanism — deferred 2026-08-24 in favor of full OAuth 2.1 (authorization code + PKCE) so that standard clients connect through discovery without manual token management; static tokens MAY return later as a convenience feature. @REJECTED Embedding roles or permissions into access tokens — rejected because rights must be revocable in real time; tokens carry identity only and permissions resolve live from the database. @REJECTED Adopting the vendored research/mcp-superset server as the tool surface — rejected because it talks to Superset directly, bypassing the ss-tools policy layer; direct SQL and raw Superset mutations violate 037 chart-truth boundaries and 038/044 executor contracts. It remains a reference implementation only. @REJECTED Replacing 37 tools with one generic api_call tool — rejected because it destroys curated schemas/discoverability, degrades small-model tool selection, and is not MCP. @REJECTED Auto-generating the tool catalog from the OpenAPI dump — rejected because curated descriptions, bounded response discipline and per-tool risk semantics are part of the safety contract; generation is allowed only as a scaffold for explicitly reviewed registrations.

Navigation (DSA Indexer keywords)

@SEMANTICS: spec, requirements, mcp, interface, tools, catalog, approval, provenance, decommission, gradio

Feature Branch: 050-mcp-interface Created: 2026-08-24 | Status: Partially implemented; production acceptance OPEN (refresh 2026-09-08) Input: "Убрать веб-интерфейс агента (Gradio, кнопка «Ассистент» и другие точки входа) и заменить его простым MCP-интерфейсом к ss-tools, через который внешние клиенты создают все тесты. Полноценный внешний доступ; первый срез — паритет текущих 37 инструментов; продуктовый frontend не содержит агентских взаимодействий; агент работает только во внешнем MCP-клиенте."

Decisions (session 2026-08-24)

  • Direction: remove the Gradio agent (agent/ service, port 7860, /agent route, assistant drawer/API); its place is taken by ONE simple MCP server mounted in the FastAPI backend.
  • Hosting: /mcp inside the existing FastAPI app (FastMCP / official SDK, Streamable HTTP). No separate process, port or auth stack; tools call services in-process.
  • External access: full, via OAuth 2.1 from Phase 0 — the ss-tools backend acts as the Authorization Server (authorization code + PKCE on top of the existing session login incl. ADFS OIDC), the MCP endpoint acts as a Resource Server. Dynamic client registration for public clients; client_credentials/SERVICE_JWT fallback for machine clients.
  • Deployment boundary: "external MCP client" means a process external to the ss-tools backend, not an Internet-hosted client. Until a later ADR explicitly changes this, MCP clients and every LLM/VLM provider are enterprise-local deployments; PII exposure inside that perimeter is an accepted residual risk.
  • First slice: parity with the current 37 LangChain @tool wrappers before any new capability lands exclusively on MCP.
  • Dashboard entry: ordinary manual scenario creation/editor only. External MCP clients are configured outside this product workflow; no agent launch, prompt, copy-prompt or handoff control.

Authorization Model

The backend already owns the canonical RBAC shape (Role(is_admin) → Permission(resource, action), pure predicate user_has_permission() in core/auth/permission_utils.py) and ships Authlib + cryptography. The drift reuses both; no parallel policy vocabulary is created.

Roles and endpoints

Part Role Endpoints / behavior
ss-tools backend OAuth 2.1 Authorization Server /oauth/authorize (requires existing web session: local password or ADFS OIDC), /oauth/token (authorization_code + PKCE S256, refresh_token, client_credentials), AS metadata (RFC 8414), JWKS
/mcp OAuth 2.1 Resource Server Protected-resource metadata (RFC 9728), 401 + WWW-Authenticate discovery hint, token validation per request
External client Public/confidential OAuth client Discovers AS via RFC 9728 → registers via DCR (RFC 7591) → authorization code + PKCE in the user's browser session → calls tools

Token and identity rules

  • Access tokens are short-lived signed JWTs (aud=mcp) carrying identity only (sub, scope, jti, session_id). Roles/permissions are NEVER embedded; they resolve live from the DB per request, so role changes take effect immediately.
  • Refresh tokens rotate on every use; reuse detection revokes the token family (extends the existing TokenBlacklist mechanism).
  • Machine clients authenticate with client_credentials or the existing SERVICE_JWT; both map to a service principal with its own scope set — never to a user's roles.
  • Authentication is a swappable adapter resolve_principal(credentials) -> McpSessionPrincipal; the tool catalog never sees transport credentials.

Tool catalog binding (hidden vs gated)

  1. Every tool registration declares required_permission=("resource","action") from the canonical vocabulary already used by specs (scenario:view/create/edit/archive, dataset:lineage:refresh, dashboard:loadtest:*, …). This is the single source of truth; the agent-side _TOOL_PERMISSIONS shadow is absorbed and deleted.
  2. tools/list filters the catalog through user_has_permission(principal.user, resource, action); is_admin=True sees everything. A tool the principal has no right for is absent from the listing (HIDDEN).
  3. Every invocation re-checks the same predicate plus service-level policies (defense in depth against stale client catalogs).
  4. Contextual risk does NOT hide a tool: a visible tool whose current invocation is risky (PROD environment, baseline publish) returns approval_required{gate_id,...} instead of failing silently (GATED).

Gates live entirely inside MCP

The MCP session MUST be self-sufficient for the whole approval loop: a principal operating only through MCP clients completes creation, gating, running and investigation without opening the ss-tools web UI. Rule set:

  • Gated invocations return approval_required{gate_id, targets, risk, reason_required, expires_at} synchronously — the client resolves it in-session.
  • list_pending_approvals(filter?) returns all gates visible to the principal (including gates raised by automation or other surfaces), each with target diff/provenance refs.
  • decide_approval(gate_id, confirm|deny, reason) consumes the gate atomically (CAS), records the reason, and unblocks the original workflow deterministically; the caller may then retry or consume per the workflow contract. Expired/already-decided gates return typed errors, never silent success.
  • Decision tools require an authenticated USER principal; service principals can list but never decide. Reasons are mandatory for confirm on high-risk classes (baseline publish, PROD dispatch, policy change) per AGSTAB-FR-005.
  • There is exactly ONE durable gate store (ActionApprovalGate). Web gate cards (042/043/045 product pages) and MCP decision tools are two renderers of the same rows — neither creates parallel state, and a decision through either path is immediately visible to the other.

Human checkpoints through MCP

HumanCheckpoint disposition (waiting_human runs, 044) is also completable inside MCP, amending SCEX-FR-013 for the post-chat world:

  • list_checkpoints(filter?) exposes waiting checkpoints visible to the principal with their evidence context; decide_checkpoint(run_id, disposition=confirm|false_positive|inconclusive, expected_version) consumes the checkpoint via the same CAS contract as the monitor.
  • Guards preserve the real invariant (no automated false PASS): decision tools accept USER principals only, CAS version mismatch is a typed error, every decision is audited with principal+reason, and no scheduled/deploy/API origin can ever reach these tools. A service principal calling them gets permission_denied.
  • The web monitor (045) remains an equal renderer of the same checkpoint rows; dispositions from either surface are immediately visible in both.

User Scenarios

Story 1 — Connect an External Client (P1)

Why P1: The MCP endpoint is the product surface; connecting must be boring and secure.

Independent Test: Point MCP Inspector or Claude Desktop at {backend}/mcp with a personal token and verify tools/list returns exactly the role-permitted catalog.

Acceptance:

  1. Given a valid user JWT When the client connects and lists tools Then the catalog contains only tools permitted for that user's role, with curated descriptions and JSON schemas.
  2. Given an invalid/expired token When the client connects Then it receives a typed authentication error and zero tool metadata beyond the public surface.
  3. Given a service JWT When a machine client lists tools Then machine-scoped tools resolve per service policy, indistinguishable in contract shape from user sessions.

Story 2 — Create a Dashboard Test Scenario End-to-End (P1)

Why P1: "Создавать все тесты через MCP" is the core value replacing the chat flow.

Independent Test: From an external client, drive inspect → compile → validate → resolve → draft-pack → save for a fixture dashboard and verify the saved immutable revision appears in the 042 registry.

Acceptance:

  1. Given a dashboard id and environment When the client calls the inspection and scenario-authoring tools Then the chain produces a validated graph with needs_context/needs_selector markers instead of invented data (038 semantics).
  2. Given delegated policy permits the actor When save is requested Then a server-stored WorkingDraft produces a new immutable revision with full provenance; the client never uploads a full graph for save (SCEDIT-FR-009 inheritance).
  3. Given policy does not delegate the save When save is requested Then the tool returns a typed approval_required envelope and nothing persists until the gate resolves.

Story 3 — Approval Round-Trip Fully Inside the Client (P1)

Why P1: HITL must survive the chat removal WITHOUT forcing a context switch to the web UI; the MCP session is self-sufficient.

Independent Test: Trigger a gated baseline-approval action via MCP, approve it via decide_approval in the same client session, and verify the subsequent consume succeeds; repeat with deny and with expiry.

Acceptance:

  1. Given a gated action When invoked via MCP Then the result is approval_required{gate_id, targets, risk, reason_required, expires_at} and no side effect occurs.
  2. Given the analyst calls list_pending_approvals in the same or a later session When gates exist (including gates raised by automation) Then they are listed with target diff/provenance refs.
  3. Given decide_approval(gate_id, confirm|deny, reason) When submitted by a user principal Then the gate is consumed atomically, the reason is recorded, and the original workflow continues deterministically (retry/consume).
  4. Given a denial, expiry, or a replayed decision When any path retries Then the gate refuses with the recorded outcome; no duplicate side effect.
  5. Given the same gate row When rendered as a web gate card OR decided via MCP from another session Then both surfaces show identical state — one durable store, two renderers.

Story 4 — Baseline and Verification Lifecycle via MCP (P2)

Why P2: 037 capture/approval/verification flows were the first agent tools and must reach parity first.

Independent Test: Execute capture_baseline_candidate → request_baseline_approval → decide → consume and create_verification_run via MCP against fixtures, matching legacy wrapper outputs.

Acceptance:

  1. Given fixture release/dashboard/metric inputs When capture runs Then the server computes hashes and creates the candidate exactly as the HTTP route does (no client hash logic).
  2. Given the full approval lifecycle When driven via MCP Then outcomes are byte-comparable with the legacy tool results on the same fixtures.

Story 5 — Decommission the Chat Surface (P2)

Why P2: The drift ends with the Gradio stack gone, not duplicated.

Independent Test: Flip the removal flag, rebuild frontend and backend, and verify no /agent references, no agent process, and ordinary manual editor entry points remain and no agent interaction controls remain.

Acceptance:

  1. Given parity is proven When the flag removes the chat Then run.sh/compose no longer start the agent service; port 7860 disappears from profiles.
  2. Given the dashboard page When the user clicks «Создать сценарий тестирования» Then the manual editor opens; no agent prompt, handoff, workspace or invocation occurs.
  3. Given the frontend build When link-integrity tests run Then zero references to /agent, assistant API or Gradio proxy remain.

Edge & Failure Cases

# Scenario Expected Behavior Recovery
E1 JWT expires mid-session Typed auth error on next call; no privileged retry Client reconnects with fresh token
E2 Stale cached tools/list after upgrade Unknown tool call returns typed TOOL_NOT_FOUND + re-list hint Client refreshes catalog
E3 Oversized tool response Bounded payload; large artifacts returned as ref + digest, never raw dumps Client fetches artifact by ref
E4 Concurrent scenario saves 409 revision conflict with current hash (042 semantics) Reload / compare
E5 SQL-class tool requested in restricted context Server-side denial regardless of client claims Elevated permission required
E6 Rate/abuse from external client Standard throttling honored (Retry-After) Back off
E7 Client attempts raw graph upload to save Rejected; save only from server-stored draft Recompile server-side
E8 Rotated refresh token replayed Whole token family revoked; client re-authorizes Re-run consent flow
E9 ADFS unavailable at authorize time Local-password session login still completes the flow Retry SSO later
E10 Role changed after tools/list cached by client Invocation re-check denies with typed permission error + re-list hint Client refreshes catalog
E11 MCP request exceeds transport or domain limit Typed payload-limit error; no tool dispatch or mutable state Client reduces request / paginates
E12 Local LLM/VLM receives capture containing PII Accepted residual risk inside enterprise perimeter; no masking gate Optional display/export masking
E13 Local provider payload/log includes a credential, cookie, token, secret, or raw storage path Typed secret-exposure rejection; event is audit-safe Reconfigure producer; retry with sanitized metadata

Requirements

Functional

  • MCPX-FR-001: The system MUST expose exactly one MCP server at {backend}/mcp (Streamable HTTP) inside the FastAPI application; stdio transport MAY exist for local development only.
  • MCPX-FR-002: Tools MUST be registered explicitly with curated name, description and input schema; OpenAPI-derived bulk generation MUST NOT ship.
  • MCPX-FR-003: The /mcp endpoint MUST authenticate every request through the Authorization Model above; tools/list MUST filter by caller RBAC and every invocation MUST re-enforce permissions server-side (single server-side policy source; the agent-side _tool_filter/_guard_tool_permission logic is absorbed here).
  • MCPX-FR-004: Every tool invocation MUST persist AgentAction provenance (principal, tool, arguments digest, outcome, linked AgentRun where applicable) and MUST pass deterministic ACL/environment/capacity checks before execution, inheriting AGSTAB-FR-011.
  • MCPX-FR-005: Non-delegated risky actions MUST return the typed approval_required envelope bound to a durable ActionApprovalGate. The approval loop MUST be completable entirely within MCP (list_pending_approvals + decide_approval); web gate cards (042/043/045 product pages) remain an equal renderer of the same durable gates. Denial MUST record cancellation with zero side effects; decision tools MUST reject service principals and expired/decided gates with typed errors.
  • MCPX-FR-006: The scenario-authoring chain (inspect query model, compile, validate, resolve, draft-pack, initial bootstrap, request-save) MUST be fully available via MCP; all mutable state MUST live server-side (WorkingDraft, revisions, candidates).
  • MCPX-FR-007: Tool responses MUST be bounded; artifacts larger than the inline limit MUST be returned as typed refs with digests resolvable through existing artifact APIs.
  • MCPX-FR-007a: MCP transport and every tool schema MUST enforce server-owned request byte, JSON-depth, collection-count, string-length, graph-size, and per-session rate limits before dispatch. Rejection is typed and creates no mutable state; bounded responses alone are insufficient.
  • MCPX-FR-008: The initial catalog MUST provide 1:1 parity with the current 37 tools across domains: environments/health/tasks, git/deploy/migration/backup/maintenance, Superset operations (read + admin CRUD), baseline capture/approval/verification, scenario compile/validate/resolve/draft-pack/save. show_capabilities is retired in favor of standard tools/list.
  • MCPX-FR-009: Decommission MUST be flag-driven and ordered: parity proven → chat hidden → agent service dropped from run.sh/compose → code deletion. Frontend entry points MUST switch to ordinary manual CRUD/editor routes. Agent chat, prompt textareas, assistant editing, typical-operation selectors that invoke an agent, proposal-generation and agent launch/handoff controls MUST be removed, not renamed.
  • MCPX-FR-010: The catalog MUST carry a version; breaking changes (rename/schema change/removal) MUST bump the major version and remain listed with a deprecation marker for one minor cycle.
  • MCPX-FR-011: MCP tools MUST reuse existing services and contracts; a second Playwright/LLM/SQL/execution stack is forbidden (SCEX-FR-009 inheritance). No MCP path may bypass 038 validation, 037 baseline rules, 044 executors or a required gate.
  • MCPX-FR-012: InvestigationSignals and queue items MUST NOT trigger any MCP activity automatically; MCP is pull-only from the client side (SCAN-FR-001 inheritance).
  • MCPX-FR-013: The backend MUST implement the Authorization Server endpoints of the Authorization Model table; PKCE S256 MUST be mandatory for public clients, and /oauth/authorize MUST reuse the existing web session (local password or ADFS OIDC) without a second credential prompt inside an active session.
  • MCPX-FR-014: The MCP endpoint MUST publish RFC 9728 protected-resource metadata and answer unauthenticated calls with 401 + WWW-Authenticate so a compliant client completes discovery → registration → authorization → tool listing without manual token pasting.
  • MCPX-FR-015: Dynamic client registration (RFC 7591) MUST be supported for public clients with first-party-bounded scopes; registered clients MUST be visible and revocable in Admin.
  • MCPX-FR-016: Refresh tokens MUST rotate on use, and replay of a rotated refresh token MUST revoke the whole token family (extending TokenBlacklist).
  • MCPX-FR-017: Access tokens MUST carry identity only (sub, scope, jti, session_id, aud=mcp); permission resolution MUST hit live DB state per request so role changes apply to the next call without re-consent.
  • MCPX-FR-018: Catalog visibility MUST follow the hidden-vs-gated rule: missing permission hides the tool from tools/list; contextual risk (PROD environment, baseline publish, non-delegated mutation) keeps the tool listed but returns approval_required.
  • MCPX-FR-019: HumanCheckpoint disposition MUST be available via list_checkpoints + decide_checkpoint under the same CAS/audit contract as the 045 monitor, restricted to authenticated user principals; no automated origin may create, consume or bypass a checkpoint, and manual-run-only revisions remain ineligible for automation (SCEX-FR-004a stands).
  • MCPX-FR-020: Until an explicit architecture amendment states otherwise, MCP clients and all LLM/VLM providers are locally deployed inside the enterprise trust perimeter. PII exposure to those local providers is an accepted residual risk and is not blocked by masking. Credentials, cookies, access/refresh/service tokens, raw secrets, and raw storage paths MUST remain absent from tool payloads, telemetry, provenance, and artifact references.
  • MCPX-FR-021: Dynamic client registration MUST be rate-limited, fully audited, redirect-URI validated, constrained to reviewed first-party scopes, and immediately revocable. DCR must not grant database permissions or allow scope escalation beyond the principal's live RBAC.
  • MCPX-FR-022: The catalog MUST provide the persistent AgentAuthoringWorkspace operations create_authoring_session, propose_test_plan, start_exploration, get_exploration_result, propose_graph_revision, get_graph_diff, and promote_to_scenario (or an explicitly versioned equivalent) with server-owned state, provenance, idempotency and CAS.
  • MCPX-FR-023: Authoring exploration MUST use an isolated, allowlisted, time/size/network-bounded and cancellable sandbox with durable receipts and artifact ownership; shell, credential, filesystem/network escape and production side effects MUST be rejected. Sandbox readiness is a release gate, not an implementation claim.
  • MCPX-FR-024: MCP MUST pass authoring output only as typed 038 action candidates/graph proposals through deterministic compile/validate, user-reviewed diff, and 042 handle-based save. Raw code, browser URLs, cookies, secrets, filesystem paths and caller digests MUST never establish authority; code-backed production execution is a separate unimplemented contract.
  • MCPX-FR-025: MCP MUST provide bootstrap_authoring_scenario for a first scenario. It accepts only a bounded InitialScenarioIntent and server-issued valid compiled/draft-pack handles; atomically creates the registry entry, initial current revision, provenance/outbox and a workspace bound to that revision. It MUST NOT require or accept a client-created base revision, graph, content hash or revision ID.
  • MCPX-FR-026: MCP MUST expose the SCAUTO-FR-017 automation operations with identical REST policy/RBAC behavior. Schedule mutations require an explicit active revision, environment and idempotency key; they reject HumanStep revisions before any durable side effect and preserve the 046 PROD gate before dispatch.
  • MCPX-FR-027: External reachability (ADR-0024, field run 2026-09-07): every durable prerequisite of an MCP write tool MUST be obtainable by the same principal class through the MCP catalog itself. Concretely, the catalog MUST provide create_agent_run and get_agent_run wrapping the existing Services.AgentRuns.Service create/snapshot boundary (permission ("dashboard:testing","EXECUTE") at REST parity, service_allowed=False, curated UIContext-v2-bounded input, idempotent active-run reuse), so that register_draft_pack → bootstrap_authoring_scenario → start_scenario_run is completable by an external client without any non-MCP transport. A pinned catalog-introspection test MUST fail when a write tool gains a durable prerequisite that is not MCP-creatable.
  • MCPX-FR-028: Context-derived capabilities (field run 2026-09-07): the MCP-carried compile path MUST consume a server-owned capability authority derived from the live inspect_dashboard_context result (authoritative DashboardQueryModel + environment policy + provider readiness), not only caller-declared booleans. Facts the server can verify truthfully (dataset-field availability, native filters, export capability) MUST NOT degrade into human_checkpoint; unresolved facts remain needs_context/needs_selector/needs_baseline; human_checkpoint is reserved for genuinely unsafe mutation contexts, human judgement, or unavailable automation. Contract: 038 amendment 2026-09-07.
  • MCPX-FR-029: Disposition clarity (field run 2026-09-07): every human-facing surface of the HumanCheckpoint disposition vocabulary (MCP tool descriptions, monitor UI labels, docs) MUST name the action by its persisted step outcome — confirm → passed, false_positive/inconclusive → inconclusive (044 lifecycle mapping is immutable) — and MUST NOT present a passing choice with defect-confirming wording or destructive styling. Contract: 044/045 amendment 2026-09-07.

Canonical Scenario MCP Names

The following snake_case names are the canonical MCP names for the complete authoring, persistence, activation, and run-target chain. The operationIds in the final column are transport-specific REST aliases only; they are not MCP names and MUST NOT be exposed by tools/list.

Canonical MCP name Responsibility Legacy REST operationId alias(es)
create_agent_run Create the durable principal-owned AgentRun prerequisite for draft-pack registration (ADR-0024) Api.AgentRuns.Create
get_agent_run Read the ownership-scoped AgentRun snapshot Api.AgentRuns.GetSnapshot
inspect_dashboard_context Inspect dashboard/query context inspectDashboardQueryModel
bootstrap_authoring_scenario Atomically create a first scenario, initial current revision and bound workspace from server-owned handles scenarioRegistry.create
create_authoring_session Create persistent authoring workspace none; MCP-only workspace operation
propose_test_plan Record typed test-plan intent none; MCP-only workspace operation
start_exploration Start bounded authoring exploration none; MCP-only workspace operation
get_exploration_result Read exploration result/artifact refs none; MCP-only workspace operation
propose_graph_revision Submit typed graph proposal scenarioEditor.agentPropose, scenarioEditor.apply
get_graph_diff Read server-computed proposal/revision diff scenarioRegistry.revisionDiff
promote_to_scenario Create/update validated server handles or a save request; never activate scenarioEditor.proposalSave, scenarioRegistry.create
scenario_compile Compile canonical ScenarioGraph compileDashboardScenario, compileScenario
scenario_validate Validate compiled graph validateDashboardScenario, validateScenario
scenario_resolve Resolve typed authoring inputs into a new compiled handle resolveDashboardScenario, resolveScenario
generate_draft_pack Generate server-owned draft pack compileScenarioDraftPack, buildScenarioDraftPack
request_save Create immutable candidate revision, or eligible initial current revision editor.save, scenarioRegistry.create
activate_revision CAS-activate an eligible materialized revision scenarioRegistry.activateCurrentRevision
start_scenario_run Start a run pinned to an explicit promoted revision scenarioRun.start
list_scenario_schedules / upsert_scenario_schedule / delete_scenario_schedule Manage schedule projections through 046 policy automation.listSchedules, automation.upsertSchedule, automation.deleteSchedule
list_scenario_trigger_rules / upsert_scenario_trigger_rule / delete_scenario_trigger_rule Manage event trigger rules through 046 policy automation.listTriggerRules, automation.upsertTriggerRule, automation.deleteTriggerRule
get_scenario_automation_policy / upsert_scenario_automation_policy / get_scenario_automation_metrics Read or manage automation policy and observability automation.getPolicy, automation.upsertPolicy, automation.metrics

promote_to_scenario does not activate and does not advance current_revision. It creates or updates server-owned validated handles or a server-stored save request. request_save creates a candidate; only the separate activate_revision operation can advance current_revision, after eligibility, materialization, policy, required approval, and CAS checks. Initial creation is the sole explicit exception when the initial eligible revision is created as current. start_scenario_run must receive an explicit promoted revision_id and verified content_hash, never a candidate or an implicit latest/current target.

Key Entities

  • McpToolCatalog: Versioned, explicitly registered set of tools grouped by domain; source of truth for names, schemas, risk class and required_permission(resource, action).
  • McpSessionPrincipal: Authenticated caller identity (user or service), role set resolved live from DB per invocation, and delegation scope.
  • OAuthClientRecord: Dynamically or statically registered client (public/confidential), first-party scope bound, owner visibility, revocation state.
  • ApprovalRequiredEnvelope: Typed result binding a refused-or-deferred action to its durable gate (gate_id, targets, risk, reason_required).
  • ToolInvocationRecord: AgentAction-provenance row per MCP tool call (principal, tool, argument digest, outcome, run linkage).
  • HandoffSurface: Retired frontend concept (2026-09-08). Connection/security administration may show endpoint configuration, but no prompt, agent launch or proposal workflow belongs in product UI.
  • AgentAuthoringWorkspace: Persistent server-owned MCP co-authoring session with CAS state, proposals, exploration results, artifacts, review and promotion references.
  • AuthoringArtifact: Bounded server-owned source/patch/trace/screenshot/diagnostic/receipt reference used for review, never an execution program or production authority.

@{ McpInterface.ScenarioPipeline [C:5] [TYPE ADR]

@BRIEF Normative ownership and provenance contract for the MCP-to-scenario-run path. @RELATION DEPENDS_ON -> [ScenarioGraph.Compiler.Compile] @RELATION DEPENDS_ON -> [ScenarioRegistry.RevisionChain] @RELATION DEPENDS_ON -> [ScenarioExecution.RunnerPlan.Derive]

AgentAuthoringWorkspace MCP operations

The MCP catalog MUST expose bootstrap_authoring_scenario for initial creation and the following target contract for persistent co-authoring of an existing revision: create_authoring_session, propose_test_plan, start_exploration, get_exploration_result, propose_graph_revision, get_graph_diff, and promote_to_scenario. Names MAY be versioned during implementation, but the operation semantics and ownership boundary are normative.

All mutable workspace, proposal, exploration, artifact, review and promotion state is server-owned. Mutating operations require authenticated principal/delegation, idempotency key and expected workspace CAS version; replays return the original result and changed requests or stale versions return typed 409 without partial mutation. Read operations enforce object ACL and return bounded refs/digests for large traces/screenshots.

start_exploration uses only the isolated sandbox contract: allowlisted origins/APIs/actions, no shell/credential/filesystem escape, bounded time/size/network, cancellation, operation receipts, artifact ownership and no production side effects. The MCP server does not generate arbitrary Playwright code; it exposes reviewed templates/actions and consumes sandbox traces/candidates. Raw code, URLs, cookies, secrets, paths and caller digests are rejected as authority.

The initial chain is inspect -> compile/validate -> draft-pack -> bootstrap_authoring_scenario -> immutable current revision. The existing-scenario chain is session -> plan -> exploration -> typed proposal -> 038 compile/validate -> user-reviewed diff -> 042 save/promotion -> immutable revision. An unbound workspace is not a creation path and cannot propose or save a graph. promote_to_scenario cannot skip user review or required approval. 044 accepts only the resulting promoted revision. Code-backed production execution is a separate future, unimplemented contract.

The following table is the single cross-spec stage contract. A handle is an opaque server-issued reference; a digest is computed and verified by the owner named in the table. A client may pass a handle or request digest for correlation, but it cannot establish authority with a graph, caller-computed digest, path, or file.

Implementation status (audit 2026-09-06, Doc.Adr.ADR0023; Phase 2c T029d–T029h in tasks.md). The deterministic cores of compile/validate/resolve/draft-pack are IMPLEMENTED as pure 038 functions. The durable handle layer is IMPLEMENTED: CompiledScenarioHandle/ValidationResultHandle/DraftPackHandle entities (migrations 0019_scenario_handles, 0020_scenario_materialization, 0021_context_authority) are minted at the persisted boundaries — REST api_compile_scenario/api_validate_scenario/api_resolve_scenario/ api_draft_pack and the MCP register_draft_pack write tool — with content-addressed canonical-bytes storage, owner binding, single consumption (SELECT ... FOR UPDATE, PostgreSQL-proven) and binding/digest/staleness rejections (ScenarioGraph.Handles). bootstrap_authoring_scenario and 042 CreateScenario/CreateInitial accept ONLY stored handle ids; the transitional caller-composed compile:{run_id}:{digest} string path is removed for bootstrap. 042 create materializes graph_snapshot = the canonical DashboardTestScenario JSON (plus server-owned action-registry identity and context_authority) inside the create transaction, so a freshly bootstrapped current revision is directly runnable/schedulable; BOOTSTRAP_REVISION_NOT_RUNNABLE survives as defense-in-depth for provenance-only legacy rows. The OutboxEvent/RevisionMaterialization reference-artifact worker exists (materialize_pending_revisions); wiring it into the scheduler poll loop is the remaining operational step. Residuals: the 043 editor promote/save path and the REST-only api_draft_pack flow do not yet evaluate context_authority (persisted as NULL = legacy-allowed at the PROD gate). The inspect/context row is resolved as a hybrid (T029h, 2026-09-06): no persisted InspectionContextHandle is minted; the MCP read tool inspect_dashboard_context exposes the live authoritative DashboardQueryModel resolver, and the register_draft_pack boundary binds the client-carried context by recomputing its fingerprint against a live inspection (context_authority = verified/unverified; sentinel fingerprints never verify; unreachable environments fail open; PROD dispatch refuses an explicit non-verified marker via CONTEXT_AUTHORITY_REQUIRED_FOR_PROD; the validator recursively rejects query_context/SQL smuggling inside dashboard_context). The table remains the normative target; the InspectionContextHandle row records the declined persisted-handle alternative, superseded by the hybrid binding above.

Field-run correction (2026-09-07, Doc.Adr.ADR0024; Phase 2d in tasks.md). The live external run against ss-prod (docs/2026-09-07-sales-prod-mcp-run.md) proved the initial bootstrap row is NOT externally reachable end-to-end: register_draft_pack requires a principal-owned AgentRun, and no MCP operation creates one since the chat decommission removed the only creation surface. The row's test_mcp_initial_scenario_e2e evidence seeded that prerequisite with a raw-ORM insert, which masked the gap (test-honesty rule, ADR-0024 §3). Closed the same day (Phase 2d): MCPX-FR-027 landed as MCP create_agent_run/ get_agent_run (catalog 2.2.0, T029i) — the vertical is now the fully external chain with zero non-MCP seeding and the reachability pin is hard-green (E2E-EXT-001 CLOSED); MCPX-FR-028 landed as ScenarioGraph.CapabilityAuthority (T029k, CAP-001 CLOSED); MCPX-FR-029 landed as outcome-based labels/descriptions (T029l, DISP-001 CLOSED). Remaining: none for Phase 2d — the live-stand replay E2E-EXT-002 (T029m) CLOSED 2026-09-11 (see the gate row below).

Stage Input Authoritative owner Output handle Digest Persistence Failure / no-side-effect rule Test id
inspect/context dashboard/environment request 038 context/capability resolver, with server ACL InspectionContextHandle server context fingerprint server request record only unresolved facts become needs_context, needs_selector, or needs_baseline; no graph/revision/run Test.Scenario.Capability, Test.Scenario.Resolver
compile bounded intent + server context 038 ScenarioGraph.Compiler CompiledScenarioHandle server content_hash of canonical graph/program immutable server compile result deterministic validation first; errors persist no mutable artifact Test.Scenario.Compiler, Test.Scenario.Serializer
validate CompiledScenarioHandle or server graph 038 ScenarioGraph.Validator ValidationResultHandle server validation-result digest bound to compiled hash immutable server validation result invalid/blocking findings produce no draft, revision, gate, or run Test.Scenario.Validator, Test.Scenario.Validator.Edge
resolve server compiled handle + typed resolution + expected base 038 resolver; 042 owns revision CAS when persisted new CompiledScenarioHandle new server content_hash, parent hash link immutable server compile result; no in-place graph edit stale base is 409 STALE_REVISION; no partial graph or revision mutation Test.Scenario.Resolver, Test.Scenario.Resolver.Edge
draft-pack valid compiled/validation handles + registered templates 038 ScenarioGraph.PackCompiler DraftPackHandle server draft_pack_digest over rendered manifest/bytes server draft registry, save_eligible or preview_only unresolved/error/unsafe input is preview_only; no registry revision or launch side effect Test.Scenario.Pack, Test.Scenario.Pack.Security
initial bootstrap InitialScenarioIntent + valid compiled/draft-pack handles + idempotency key 042 ScenarioRegistry.CreateInitial ScenarioRegistryEntry + initial current ScenarioRevision + bound workspace server-verified content hash and draft digest PostgreSQL registry/revision/workspace + outbox invalid, stale, unauthorized or replay-conflicting input creates no partial row; no synthetic base revision test_mcp_initial_scenario_e2e
request-save compiled/draft handles + expected base + idempotency key 042 registry transaction ScenarioRevision (candidate, or initial current) server-verified graph/content hash and draft digest PostgreSQL revision + outbox; materialization async CAS/idempotency failure is atomic; no row/outbox on rejection; save never activates test_scenario_registry, test_revisions
run preflight selected ScenarioRevision + typed bindings + server target snapshot 044 RunPreflight RunPreflightHandle server preflight digest over revision/target/bindings/policy durable preflight record attached to launch request missing/invalid context, selector, baseline, policy, or revision mismatch blocks before dispatch/I/O test_scenario_runner, test_scenario_automation_api
runner-plan eligible preflight + immutable revision 044 RunnerPlan.Derive RunnerPlanHandle server descriptor/plan digest persisted on ScenarioRun; git runner.plan.json is reference only altered/unknown descriptor rejects before lease, execution, or provider I/O test_scenario_runner_plan, test_scenario_dispatch
launch request RunnerPlanHandle + idempotency key 044 server runner ScenarioRun server launch/request digest and pinned plan digest PostgreSQL run, queue/gate, audit same key+hash replays same run; changed request is 409 IDEMPOTENCY_KEY_REUSED; no duplicate dispatch test_scenario_queued_dispatch, test_scenario_scheduler_callbacks
dispatch/evidence pinned plan descriptor + capacity lease 044 executor/provider and evidence owner EvidenceRef / StepOutcome provider-verified receipt SHA-256 Artifact(owner_type=scenario_run) + immutable receipt + step outcome missing ownership/digest/ref is non-pass; late/unknown effect reconciles or terminalizes non-pass test_scenario_terminal_signals, test_provider_contract
signal terminal run + immutable evidence provenance 044 producer; 047 queue/episode projection idempotent InvestigationSignal / queue input server signal identity digest durable signal; 047 projection only failed/blocked/inconclusive emits once; passed emits none; never opens chat/action automatically test_scenario_terminal_signals

Handle and Digest Rules

  • CompiledScenarioHandle is minted only after 038 canonical serialization. Its content_hash is SHA-256 of the executable canonical Verification Program graph; it excludes timestamps, display fields, ParameterBindings, and runtime identity. 038 never mints scenario_id or revision_id.
  • ValidationResultHandle is valid only for the exact compiled-handle hash and validator/schema versions used to produce it. A validation digest does not become graph authority and cannot be supplied by a caller as proof of validity.
  • DraftPackHandle is minted by the server pack registry from registered, versioned templates. Its digest covers the server-rendered manifest and bytes; save_eligible requires valid validation and no unresolved required marker.
  • 042 alone mints scenario_id, revision_id, parent links, and server-owned revision records. Save verifies the compiled hash and draft digest against stored handles inside one transaction; client graph uploads are rejected.
  • RunPreflightHandle and RunnerPlanHandle are derived by 044 from persisted revision data, server target/policy context, and typed bindings. Their digests are server recomputed; a caller-supplied runner or digest is correlation data.
  • runner.plan.json and scenario.yaml are materialized reference artifacts. They can be regenerated and are never authority for validation, revision, preflight, dispatch, retry, recovery, or evidence.
  • Status (2026-09-06, updated after Phase 2c): CompiledScenarioHandle, ValidationResultHandle and DraftPackHandle ARE minted and stored (038 ScenarioGraph.Handles, migrations 0019–0021); the rules in this section are their implemented acceptance criteria. InspectionContextHandle is NOT persisted — T029h resolved the inspect stage as a live fingerprint binding (context_authority) instead. The authority that 042 materializes into ScenarioRevision.graph_snapshot is the persisted canonical DashboardTestScenario serialization behind CompiledScenarioHandle.canonical_bytes_ref — explicitly NOT the lossy pack templates above (038 ScenarioGraph.ServerOwnedPipeline, amendment 2026-09-06).

Resolution and Continuation Semantics

038 owns inspection and semantic resolution. needs_context means a required dashboard/query/change-request fact is absent; needs_selector means a browser target/action selector is unknown; needs_baseline means an approved baseline reference is absent or stale. None may be guessed or silently defaulted. 038 may compile a WorkingDraft/preview, but only 042 can save a revision and only 044 can reject launch bindings at RunPreflight.

Save is draft/working -> validated save request -> candidate, with the initial revision as the sole exception (candidate -> current at creation). Activation is a distinct CAS operation: candidate -> current, guarded by eligibility, materialization, policy and required approval. Approval continuation is pending_approval -> queued; deny/expiry is pending_approval -> blocked. Every continuation consumes its gate/checkpoint exactly once. Expected revision, gate, checkpoint and idempotency versions are CAS inputs; stale values return a typed 409 and create no side effect. Same idempotency key plus the same canonical request hash replays the existing result; the same key plus a different hash is 409 IDEMPOTENCY_KEY_REUSED.

The execution chain is strictly: ScenarioRevision -> RunPreflight -> RunnerPlan -> ScenarioRun -> queued|pending_approval -> dispatch -> EvidenceRef/StepOutcome -> 047 signal. ScenarioRevision owns immutable program input, RunPreflight owns launch eligibility, RunnerPlan owns descriptor order/policy, ScenarioRun owns the execution snapshot, 044 owns step/evidence receipts, and 047 owns queue/episode projection. No client graph, caller digest, or materialized plan can replace these owners.

@{ McpInterface.ScenarioPipeline.ReleaseGates [C:5] [TYPE Block]

@BRIEF Conjunctive release gates for the coordinated pipeline contract.

Release is GO only when every applicable gate passes: 038 fixture/schema and deterministic compiler/validator/serializer/pack checks; 042 registry create, revision CAS, idempotency, outbox and materialization checks; 044 preflight, plan, queued approval, exact dispatch, evidence ownership and terminal signal checks; 050 parity, bounded transport, hidden-vs-gated and end-to-end MCP checks; and structural anchor/Markdown audit. Partial green is NO-GO and does not mark implementation tasks complete.

Evidence commands:

cd backend && source .venv/bin/activate && python -m pytest -q tests/services/dashboard_testing/scenario/test_*.py tests/api/test_scenario_runs_api.py tests/api/test_scenario_automation_api.py tests/api/test_scenario_analytics_api.py
cd backend && source .venv/bin/activate && python -m ruff check src/services/dashboard_testing/execution src/api/routes/dashboard_testing/scenario_runs.py
python specs/044-dashboard-scenario-execution/prototype/validate_static.py
cd frontend && npm run test -- --run && npm run lint && npm run build

The required MCP parity and end-to-end evidence is test_mcp_* plus the 050 SC-001/SC-002 walkthrough; absent or partial evidence keeps the gate NO-GO.

Authoring E2E Traceability

These are required evidence rows for the authoring promotion path. They remain open until executable evidence is retained; unchecked rows keep this release gate NO-GO and do not mark T023 or T028 complete.

Evidence row Required trace Evidence required Status
E2E-AUTH-001 create_authoring_session -> propose_test_plan -> start_exploration -> get_exploration_result External MCP client trace proves persistent server-owned workspace, bounded exploration, durable receipt, and typed artifact refs [x] CLOSED 2026-09-03 — tests/test_mcp_authoring_promotion_e2e.py::test_sandbox_output_promotes_to_immutable_revision (plan → queued exploration → sandbox dispatch → exploration_passed, evidence draft:exploration-*, bounded projection)
E2E-AUTH-002 propose_graph_revision -> scenario_compile -> scenario_validate -> get_graph_diff External MCP client trace proves typed proposal, deterministic 038 validation, server-computed diff, and explicit user review boundary [x] CLOSED 2026-09-03 — tests/test_mcp_scenario_e2e.py (propose → promote validation → 038 inspect_scenario/scenario_resolve/validate_scenario over MCP → get_graph_diff; canonical tool names inspect_scenario/validate_scenario per FR-008 rename clause)
E2E-AUTH-003 promote_to_scenario -> request_save -> activate_revision -> start_scenario_run External MCP client trace proves handle-based save creates candidate, activation separately passes eligibility/materialization/policy/approval/CAS, and the run pins promoted revision_id + content_hash [x] CLOSED 2026-09-03 — tests/test_mcp_scenario_e2e.py extended: post-activation start_scenario_run creates a queued run pinned to scenario_revision_id + scenario_content_hash of the promoted revision

Field-run remediation gates (2026-09-07, Doc.Adr.ADR0024)

Required evidence rows produced by the live external run against ss-prod (docs/2026-09-07-sales-prod-mcp-run.md). Unchecked rows keep this release gate NO-GO and do not mark the referenced tasks complete.

Evidence row Required trace Evidence required Status
E2E-EXT-001 create_agent_run -> register_draft_pack -> bootstrap_authoring_scenario -> start_scenario_run External MCP-only chain on a fresh DB with ZERO non-MCP prerequisite seeding; the strict-xfail pin tests/test_mcp_agent_run_reachability.py flipped to green and unmarked (MCPX-FR-027, T029i) [x] CLOSED 2026-09-07 — tests/test_mcp_initial_scenario_e2e.py converted (AgentRun minted via MCP create_agent_run tools/call; no service/ORM seeding remains); reachability pin unmarked+green; tests/test_mcp_agent_run_tools.py (5) covers create/replay/read/RBAC; slice 70 passed, full suite 11357 passed
E2E-EXT-002 live stand replay of the 2026-09-07 sales run inspect_dashboard_context -> derived-capability compile -> create_agent_run -> register -> bootstrap -> queued/gated run on ss-prod + list_checkpoints/decide_checkpoint human loop, with retained evidence (T029m) [x] CLOSED 2026-09-11 — committed replay client specs/044-dashboard-scenario-execution/prototype/live_mcp_replay.py replayed the full external chain on the live stand (run 110a6517…: inspect derived capabilities → create_agent_run → compile B01 → register context_authority=verified → bootstrap → PROD gate pending_approval (identical retry = one durable gate) → approval → live capture_screenshot passed (8 refs) + typed BROWSER_ACTION_NOT_SUPPORTED → honest inconclusive); the human loop was exercised live by live_mcp_human_loop.py (compiled B05 HumanCheckpoint → waiting_human → MCP list_checkpoints (v1) → decide_checkpoint confirm (v2 CAS) → terminal passed). Two fail-closed defects found and fixed (binding resolution for identity-less compiled steps; per-step target-identity stamping at bootstrap). Trace: docs/2026-09-11-sales-prod-mcp-replay.md
CAP-001 inspect_dashboard_context -> compile capabilities Server-derived capability map proven: verifiably-available dataset fields/native filters classify automated, not human_checkpoint; unsafe-mutation cases stay human_checkpoint (MCPX-FR-028, 038 amendment, T029k) [x] CLOSED 2026-09-07 — ScenarioGraph.CapabilityAuthority (derivation/derived-wins merge/boundary choke point) wired into MCP inspect_scenario+inspect_dashboard_context and REST api_compile_scenario; CAP-001 classification-fix test through map_all on the sales-shape fixture (B/T automated, C04–C06 unsupported, mutation cases legitimately human) + tool/REST parity pins; slice 67 passed
DISP-001 disposition label/vocabulary audit RU/EN labels and MCP tool descriptions name the persisted outcome (confirm→passed); no defect-confirming wording or destructive styling on a passing choice; mapping pinned by tests (MCPX-FR-029, 044/045 amendment, T029l) [x] CLOSED 2026-09-07 — ru/en labels renamed to outcome-based wording; confirm buttons bg-destructive→bg-primary (HumanCheckpointPanel + WaitingForMeView); decide_checkpoint docstring carries the immutable outcome table; vitest DISP-001 style/label/dispatch pin; frontend 3507 passed, lint 0 errors, build OK
TEST-001 vertical-test honesty remediation test_mcp_initial_scenario_e2e.py::_pack() creates the AgentRun through the production create_agent_run service boundary (no raw-ORM prerequisite seeding); reachability requirement pinned as strict=True xfail; typed zero-side-effect denial pinned (ADR-0024 §3, T029j) [x] CLOSED 2026-09-07 — pytest tests/test_mcp_initial_scenario_e2e.py tests/test_mcp_agent_run_reachability.py green (1 passed + 1 xfailed(strict) + denial pin); metadata records the open T029i gap. (Superseded the same day by T029i closure: the strict-xfail flip ritual executed as designed — pin unmarked and hardened, vertical converted to the fully external chain; denial pin retained.)

@} McpInterface.ScenarioPipeline.ReleaseGates

@} McpInterface.ScenarioPipeline

Success Criteria

  • SC-001: An external client completes the full scenario creation chain for a fixture dashboard and the revision appears in the registry with correct provenance.
  • SC-002: 100% of parity tests pass: MCP tool outcomes match legacy @tool wrapper outputs on shared fixtures.
  • SC-003: 100% of gated invocations produce zero side effects before approval; denial/expiry paths refuse cleanly.
  • SC-004: tools/list is RBAC-exact for admin/analyst/viewer roles in fixture tests.
  • SC-005: After decommission, builds and link-integrity suites pass with zero /agent, assistant-API or Gradio-proxy references; the stack starts without port 7860.
  • SC-006: No catalog path reaches arbitrary SQL or raw Superset mutation outside the governed tools; scenario contexts cannot invoke SQL-class tools.
  • SC-007: A compliant external client completes discovery (RFC 9728) → DCR → authorization code + PKCE → tools/list in one automated flow, with zero manual token management.
  • SC-008: Refresh-token replay revokes the token family in 100% of fault-injection cases; pre-revocation access tokens die at expiry, not silently extended.
  • SC-009: A role change is reflected in the next tools/list and the next invocation without new consent; a revoked permission hides the tool and denies cached-catalog calls.

Clarifications

Session 2026-08-24

  • Q: Separate MCP process or mounted in backend? → A: Mounted in the FastAPI app; simplest deployment, in-process service reuse, shared auth middleware.
  • Q: Does removing chat remove HITL? → A: No, and it does not force the web UI either (session 2026-08-24): gates are server-durable and fully resolvable inside MCP (list_pending_approvals + decide_approval); web gate cards on product pages remain an equal renderer of the same rows for browser-first users.
  • Q: Is the vendored mcp-superset adopted? → A: No — reference only; it bypasses the policy layer.
  • Q: What happens to AgentRun/DraftArtifact/InvestigationSignal contracts? → A: They persist unchanged; MCP invocations create the same provenance rows. Only the conversational transport and its UI retire.
  • Q: OAuth now or later? → A: Now (session 2026-08-24 decision). The backend becomes a full OAuth 2.1 Authorization Server in Phase 0 — Authlib==1.6.6 is already a dependency; personal access tokens are explicitly deferred, not rejected forever.
  • Q: Where do rights live? → A: In the existing DB RBAC (Role/Permission, user_has_permission). Tokens never embed roles; the catalog binds tools to canonical resource:action pairs and filters both listing and invocation through the same predicate.

Session 2026-09-07 (field-run remediation, Doc.Adr.ADR0024)

  • Q: The 2026-08-24 clarification said AgentRun contracts "persist unchanged" — who creates an AgentRun after the chat retirement? → A: The clarification preserved persistence but silently retired the only creation entry point. Corrected: MCP gets an explicit typed creation/read surface create_agent_run/get_agent_run over the same service (MCPX-FR-027, T029i). Rejected alternatives (no-AgentRun registration boundary, implicit auto-create, REST crutch, raw-ORM test seeds) are recorded in ADR-0024.
  • Q: May a vertical E2E seed a chain prerequisite directly into the DB? → A: No. A vertical test MUST obtain every prerequisite through a boundary the principal under test can reach; where that boundary does not exist yet, the test uses the production service boundary with honest metadata AND a strict=True xfail reachability pin (test-honesty rule, ADR-0024 §3; executed as T029j).
  • Q: Why did the B01 preview classify verifiably automatable checks as human_checkpoint? → A: Caller-declared capability booleans are not an authority. Server-derived capabilities from the live inspection are required (MCPX-FR-028, 038 amendment 2026-09-07, T029k); human_checkpoint stays reserved for genuinely unsafe mutation contexts, human judgement, or unavailable automation, and HumanStep revisions remain automation-ineligible.
  • Q: Is the UI label «Подтвердить проблему» for disposition confirm correct? → A: No — confirm persists step outcome passed (044 lifecycle mapping is immutable and correct); the label inverts the meaning. Labels/tool descriptions must name the persisted outcome (MCPX-FR-029, 044/045 amendment 2026-09-07, T029l).

Phases

Phase Scope Exit evidence
0 OAuth AS+RS skeleton: authorize/token/DCR/JWKS/metadata endpoints, /mcp validation middleware, 2–3 probe tools, Inspector + scripted-client connectivity SC-007 green in CI
1 Parity catalog for the 37 tools + contract tests mirroring wrapper tests; permission-bound hidden/gated matrix SC-002, SC-004, SC-009
2 Gates/provenance over MCP; end-to-end scenario creation walkthrough; refresh rotation/reuse tests SC-001, SC-003, SC-008
2d Field-run remediation (ADR-0024): MCP AgentRun surface, external reachability pin, context-derived capabilities, disposition clarity, live stand replay E2E-EXT-001/002, CAP-001, DISP-001 (TEST-001 closed)
3 Frontend decommission; manual editor entry; assistant interaction removal SC-005 partial
4 Delete agent/ service; run.sh/compose updates; spec amortization closed SC-005 full

Spec Impact & Amortization Map

Spec Amendment
036 Transport rejection recorded; durable-runtime contracts carry over to MCP provenance
037 Baseline tools join the MCP catalog unchanged (thin forwarders)
038 Authoring entry becomes MCP-callable; validator remains sole gateway
039 AGUI-FR-001..013 rewritten for manual authoring/review; no agent entry
040 Delegated load experiments readable as MCP-driven, policy unchanged
041 Blast-radius explanation consumed by external clients; index contracts unchanged
042 Delegated actors explicitly include MCP principals; registry contracts unchanged
043 External MCP proposal authoring only; frontend manual editing/review; SCEDIT-FR-009 stands
044 Runner unaffected; authoring/investigation boundary inherited by MCP clients
045 Manual investigation case open/review only; external agents pull MCP, no UI launch
046 Schedule management exercisable through MCP tools under same policy
047 Case threads continue in external clients; server keeps evidence/timeline

Production contract refresh — 2026-09-08

Frontend boundary (user decision 2026-09-08): All agent interaction is external MCP only. Product frontend MUST NOT contain agent chat, prompt/request textarea, assistant editing, typical-operation-to-agent selector, proposal-generation, agent workspace/start or handoff controls/routes. Ordinary manual CRUD/editor, human approval/review, monitoring and read-only evidence/evaluation are permitted. AgentEvaluationCard is read-only, with no prompt/retry-agent/provider controls. Existing agent proposal UI is runtime drift; removal/negative DOM-route-network acceptance remains OPEN in this spec-only change.

MCPX-FR-030 — Contract-complete public parity: Curated versioned tool schemas MUST expose baseline capture/review/request/decide/consume/explicit publish/lifecycle and run result/evidence operations with the same services, permissions, errors, CAS and idempotency as REST. No fallback success for consume/publication. Start/run/automation retain full server-resolved BaselineSelectionPin, never caller digests. External clients remain enterprise-local; live visual authoring/execution readiness remains open.

Normative contract: Contract-complete public parity. New requirements are specified, implemented=false / acceptance OPEN until executable evidence closes the linked tasks/checklist/traceability rows. Historical local tests and the manual inconclusive ss-prod run do not prove browser/capture/baseline/LLM production readiness. The refresh scope is the audited P0/P1/P2 agentic E2E and baseline gaps; an approved ExecutionPerformanceBaseline is not introduced.

#endregion McpInterface.Spec