61 KiB
#region McpInterface.Spec [C:3] [TYPE ADR] [SEMANTICS mcp,interface,tools,agent,dashboard-testing,decommission]
@BRIEF Simple MCP interface to ss-tools that replaces the Gradio chat agent as the primary agentic surface, including full dashboard-test creation.
@RELATION DEPENDS_ON -> [Doc.Adr.ADR0001]
@RELATION DEPENDS_ON -> [Doc.Adr.ADR0005]
@RELATION DEPENDS_ON -> [Doc.Adr.ADR0006]
@RELATION DEPENDS_ON -> [AgentTestStabilization.Spec]
@RATIONALE ss-tools already concentrates truth and policy server-side: every agent tool is a thin forwarder over backend services (RBAC guard -> httpx -> truncated result), and 042-047 demand that any delegated actor works through the same server-owned contracts as an analyst. A conversational runtime coupled to a bespoke Gradio UI adds a parallel transport, a parallel HITL mechanism (LangGraph interrupt + confirmation module) and triple bookkeeping per capability (route + LangChain wrapper + allowlist). MCP externalizes reasoning to standard clients while the backend keeps deterministic validators, ActionApprovalGate, AgentAction provenance and InvestigationSignal semantics unchanged.
@REJECTED Keeping the Gradio chat agent long-term — rejected 2026-08-24: the chat surface duplicates policy transport, blocks external clients, and every new spec (042-047) multiplies hand-written wrappers.
@REJECTED Personal access tokens as the primary authentication mechanism — deferred 2026-08-24 in favor of full OAuth 2.1 (authorization code + PKCE) so that standard clients connect through discovery without manual token management; static tokens MAY return later as a convenience feature.
@REJECTED Embedding roles or permissions into access tokens — rejected because rights must be revocable in real time; tokens carry identity only and permissions resolve live from the database.
@REJECTED Adopting the vendored research/mcp-superset server as the tool surface — rejected because it talks to Superset directly, bypassing the ss-tools policy layer; direct SQL and raw Superset mutations violate 037 chart-truth boundaries and 038/044 executor contracts. It remains a reference implementation only.
@REJECTED Replacing 37 tools with one generic api_call tool — rejected because it destroys curated schemas/discoverability, degrades small-model tool selection, and is not MCP.
@REJECTED Auto-generating the tool catalog from the OpenAPI dump — rejected because curated descriptions, bounded response discipline and per-tool risk semantics are part of the safety contract; generation is allowed only as a scaffold for explicitly reviewed registrations.
Navigation (DSA Indexer keywords)
@SEMANTICS: spec, requirements, mcp, interface, tools, catalog, approval, provenance, decommission, gradio
Feature Branch: 050-mcp-interface
Created: 2026-08-24 | Status: Partially implemented; production acceptance OPEN (refresh 2026-09-08)
Input: "Убрать веб-интерфейс агента (Gradio, кнопка «Ассистент» и другие точки входа) и заменить его простым MCP-интерфейсом к ss-tools, через который внешние клиенты создают все тесты. Полноценный внешний доступ; первый срез — паритет текущих 37 инструментов; продуктовый frontend не содержит агентских взаимодействий; агент работает только во внешнем MCP-клиенте."
Decisions (session 2026-08-24)
- Direction: remove the Gradio agent (
agent/service, port 7860,/agentroute, assistant drawer/API); its place is taken by ONE simple MCP server mounted in the FastAPI backend. - Hosting:
/mcpinside the existing FastAPI app (FastMCP / official SDK, Streamable HTTP). No separate process, port or auth stack; tools call services in-process. - External access: full, via OAuth 2.1 from Phase 0 — the ss-tools backend acts as the Authorization Server (authorization code + PKCE on top of the existing session login incl. ADFS OIDC), the MCP endpoint acts as a Resource Server. Dynamic client registration for public clients;
client_credentials/SERVICE_JWTfallback for machine clients. - Deployment boundary: "external MCP client" means a process external to the ss-tools backend, not an Internet-hosted client. Until a later ADR explicitly changes this, MCP clients and every LLM/VLM provider are enterprise-local deployments; PII exposure inside that perimeter is an accepted residual risk.
- First slice: parity with the current 37 LangChain
@toolwrappers before any new capability lands exclusively on MCP. - Dashboard entry: ordinary manual scenario creation/editor only. External MCP clients are configured outside this product workflow; no agent launch, prompt, copy-prompt or handoff control.
Authorization Model
The backend already owns the canonical RBAC shape (Role(is_admin) → Permission(resource, action), pure predicate user_has_permission() in core/auth/permission_utils.py) and ships Authlib + cryptography. The drift reuses both; no parallel policy vocabulary is created.
Roles and endpoints
| Part | Role | Endpoints / behavior |
|---|---|---|
| ss-tools backend | OAuth 2.1 Authorization Server | /oauth/authorize (requires existing web session: local password or ADFS OIDC), /oauth/token (authorization_code + PKCE S256, refresh_token, client_credentials), AS metadata (RFC 8414), JWKS |
/mcp |
OAuth 2.1 Resource Server | Protected-resource metadata (RFC 9728), 401 + WWW-Authenticate discovery hint, token validation per request |
| External client | Public/confidential OAuth client | Discovers AS via RFC 9728 → registers via DCR (RFC 7591) → authorization code + PKCE in the user's browser session → calls tools |
Token and identity rules
- Access tokens are short-lived signed JWTs (
aud=mcp) carrying identity only (sub,scope,jti,session_id). Roles/permissions are NEVER embedded; they resolve live from the DB per request, so role changes take effect immediately. - Refresh tokens rotate on every use; reuse detection revokes the token family (extends the existing
TokenBlacklistmechanism). - Machine clients authenticate with
client_credentialsor the existingSERVICE_JWT; both map to a service principal with its own scope set — never to a user's roles. - Authentication is a swappable adapter
resolve_principal(credentials) -> McpSessionPrincipal; the tool catalog never sees transport credentials.
Tool catalog binding (hidden vs gated)
- Every tool registration declares
required_permission=("resource","action")from the canonical vocabulary already used by specs (scenario:view/create/edit/archive,dataset:lineage:refresh,dashboard:loadtest:*, …). This is the single source of truth; the agent-side_TOOL_PERMISSIONSshadow is absorbed and deleted. tools/listfilters the catalog throughuser_has_permission(principal.user, resource, action);is_admin=Truesees everything. A tool the principal has no right for is absent from the listing (HIDDEN).- Every invocation re-checks the same predicate plus service-level policies (defense in depth against stale client catalogs).
- Contextual risk does NOT hide a tool: a visible tool whose current invocation is risky (PROD environment, baseline publish) returns
approval_required{gate_id,...}instead of failing silently (GATED).
Gates live entirely inside MCP
The MCP session MUST be self-sufficient for the whole approval loop: a principal operating only through MCP clients completes creation, gating, running and investigation without opening the ss-tools web UI. Rule set:
- Gated invocations return
approval_required{gate_id, targets, risk, reason_required, expires_at}synchronously — the client resolves it in-session. list_pending_approvals(filter?)returns all gates visible to the principal (including gates raised by automation or other surfaces), each with target diff/provenance refs.decide_approval(gate_id, confirm|deny, reason)consumes the gate atomically (CAS), records the reason, and unblocks the original workflow deterministically; the caller may then retry or consume per the workflow contract. Expired/already-decided gates return typed errors, never silent success.- Decision tools require an authenticated USER principal; service principals can list but never decide. Reasons are mandatory for confirm on high-risk classes (baseline publish, PROD dispatch, policy change) per AGSTAB-FR-005.
- There is exactly ONE durable gate store (ActionApprovalGate). Web gate cards (042/043/045 product pages) and MCP decision tools are two renderers of the same rows — neither creates parallel state, and a decision through either path is immediately visible to the other.
Human checkpoints through MCP
HumanCheckpoint disposition (waiting_human runs, 044) is also completable inside MCP, amending SCEX-FR-013 for the post-chat world:
list_checkpoints(filter?)exposes waiting checkpoints visible to the principal with their evidence context;decide_checkpoint(run_id, disposition=confirm|false_positive|inconclusive, expected_version)consumes the checkpoint via the same CAS contract as the monitor.- Guards preserve the real invariant (no automated false PASS): decision tools accept USER principals only, CAS version mismatch is a typed error, every decision is audited with principal+reason, and no scheduled/deploy/API origin can ever reach these tools. A service principal calling them gets
permission_denied. - The web monitor (045) remains an equal renderer of the same checkpoint rows; dispositions from either surface are immediately visible in both.
User Scenarios
Story 1 — Connect an External Client (P1)
Why P1: The MCP endpoint is the product surface; connecting must be boring and secure.
Independent Test: Point MCP Inspector or Claude Desktop at {backend}/mcp with a personal token and verify tools/list returns exactly the role-permitted catalog.
Acceptance:
- Given a valid user JWT When the client connects and lists tools Then the catalog contains only tools permitted for that user's role, with curated descriptions and JSON schemas.
- Given an invalid/expired token When the client connects Then it receives a typed authentication error and zero tool metadata beyond the public surface.
- Given a service JWT When a machine client lists tools Then machine-scoped tools resolve per service policy, indistinguishable in contract shape from user sessions.
Story 2 — Create a Dashboard Test Scenario End-to-End (P1)
Why P1: "Создавать все тесты через MCP" is the core value replacing the chat flow.
Independent Test: From an external client, drive inspect → compile → validate → resolve → draft-pack → save for a fixture dashboard and verify the saved immutable revision appears in the 042 registry.
Acceptance:
- Given a dashboard id and environment When the client calls the inspection and scenario-authoring tools Then the chain produces a validated graph with
needs_context/needs_selectormarkers instead of invented data (038 semantics). - Given delegated policy permits the actor When save is requested Then a server-stored WorkingDraft produces a new immutable revision with full provenance; the client never uploads a full graph for save (SCEDIT-FR-009 inheritance).
- Given policy does not delegate the save When save is requested Then the tool returns a typed
approval_requiredenvelope and nothing persists until the gate resolves.
Story 3 — Approval Round-Trip Fully Inside the Client (P1)
Why P1: HITL must survive the chat removal WITHOUT forcing a context switch to the web UI; the MCP session is self-sufficient.
Independent Test: Trigger a gated baseline-approval action via MCP, approve it via decide_approval in the same client session, and verify the subsequent consume succeeds; repeat with deny and with expiry.
Acceptance:
- Given a gated action When invoked via MCP Then the result is
approval_required{gate_id, targets, risk, reason_required, expires_at}and no side effect occurs. - Given the analyst calls
list_pending_approvalsin the same or a later session When gates exist (including gates raised by automation) Then they are listed with target diff/provenance refs. - Given
decide_approval(gate_id, confirm|deny, reason)When submitted by a user principal Then the gate is consumed atomically, the reason is recorded, and the original workflow continues deterministically (retry/consume). - Given a denial, expiry, or a replayed decision When any path retries Then the gate refuses with the recorded outcome; no duplicate side effect.
- Given the same gate row When rendered as a web gate card OR decided via MCP from another session Then both surfaces show identical state — one durable store, two renderers.
Story 4 — Baseline and Verification Lifecycle via MCP (P2)
Why P2: 037 capture/approval/verification flows were the first agent tools and must reach parity first.
Independent Test: Execute capture_baseline_candidate → request_baseline_approval → decide → consume and create_verification_run via MCP against fixtures, matching legacy wrapper outputs.
Acceptance:
- Given fixture release/dashboard/metric inputs When capture runs Then the server computes hashes and creates the candidate exactly as the HTTP route does (no client hash logic).
- Given the full approval lifecycle When driven via MCP Then outcomes are byte-comparable with the legacy tool results on the same fixtures.
Story 5 — Decommission the Chat Surface (P2)
Why P2: The drift ends with the Gradio stack gone, not duplicated.
Independent Test: Flip the removal flag, rebuild frontend and backend, and verify no /agent references, no agent process, and ordinary manual editor entry points remain and no agent interaction controls remain.
Acceptance:
- Given parity is proven When the flag removes the chat Then
run.sh/compose no longer start the agent service; port 7860 disappears from profiles. - Given the dashboard page When the user clicks «Создать сценарий тестирования» Then the manual editor opens; no agent prompt, handoff, workspace or invocation occurs.
- Given the frontend build When link-integrity tests run Then zero references to
/agent, assistant API or Gradio proxy remain.
Edge & Failure Cases
| # | Scenario | Expected Behavior | Recovery |
|---|---|---|---|
| E1 | JWT expires mid-session | Typed auth error on next call; no privileged retry | Client reconnects with fresh token |
| E2 | Stale cached tools/list after upgrade | Unknown tool call returns typed TOOL_NOT_FOUND + re-list hint |
Client refreshes catalog |
| E3 | Oversized tool response | Bounded payload; large artifacts returned as ref + digest, never raw dumps | Client fetches artifact by ref |
| E4 | Concurrent scenario saves | 409 revision conflict with current hash (042 semantics) | Reload / compare |
| E5 | SQL-class tool requested in restricted context | Server-side denial regardless of client claims | Elevated permission required |
| E6 | Rate/abuse from external client | Standard throttling honored (Retry-After) |
Back off |
| E7 | Client attempts raw graph upload to save | Rejected; save only from server-stored draft | Recompile server-side |
| E8 | Rotated refresh token replayed | Whole token family revoked; client re-authorizes | Re-run consent flow |
| E9 | ADFS unavailable at authorize time | Local-password session login still completes the flow | Retry SSO later |
| E10 | Role changed after tools/list cached by client | Invocation re-check denies with typed permission error + re-list hint | Client refreshes catalog |
| E11 | MCP request exceeds transport or domain limit | Typed payload-limit error; no tool dispatch or mutable state | Client reduces request / paginates |
| E12 | Local LLM/VLM receives capture containing PII | Accepted residual risk inside enterprise perimeter; no masking gate | Optional display/export masking |
| E13 | Local provider payload/log includes a credential, cookie, token, secret, or raw storage path | Typed secret-exposure rejection; event is audit-safe | Reconfigure producer; retry with sanitized metadata |
Requirements
Functional
- MCPX-FR-001: The system MUST expose exactly one MCP server at
{backend}/mcp(Streamable HTTP) inside the FastAPI application; stdio transport MAY exist for local development only. - MCPX-FR-002: Tools MUST be registered explicitly with curated name, description and input schema; OpenAPI-derived bulk generation MUST NOT ship.
- MCPX-FR-003: The
/mcpendpoint MUST authenticate every request through the Authorization Model above;tools/listMUST filter by caller RBAC and every invocation MUST re-enforce permissions server-side (single server-side policy source; the agent-side_tool_filter/_guard_tool_permissionlogic is absorbed here). - MCPX-FR-004: Every tool invocation MUST persist AgentAction provenance (principal, tool, arguments digest, outcome, linked AgentRun where applicable) and MUST pass deterministic ACL/environment/capacity checks before execution, inheriting AGSTAB-FR-011.
- MCPX-FR-005: Non-delegated risky actions MUST return the typed
approval_requiredenvelope bound to a durable ActionApprovalGate. The approval loop MUST be completable entirely within MCP (list_pending_approvals+decide_approval); web gate cards (042/043/045 product pages) remain an equal renderer of the same durable gates. Denial MUST record cancellation with zero side effects; decision tools MUST reject service principals and expired/decided gates with typed errors. - MCPX-FR-006: The scenario-authoring chain (inspect query model, compile, validate, resolve, draft-pack, initial bootstrap, request-save) MUST be fully available via MCP; all mutable state MUST live server-side (WorkingDraft, revisions, candidates).
- MCPX-FR-007: Tool responses MUST be bounded; artifacts larger than the inline limit MUST be returned as typed refs with digests resolvable through existing artifact APIs.
- MCPX-FR-007a: MCP transport and every tool schema MUST enforce server-owned request byte, JSON-depth, collection-count, string-length, graph-size, and per-session rate limits before dispatch. Rejection is typed and creates no mutable state; bounded responses alone are insufficient.
- MCPX-FR-008: The initial catalog MUST provide 1:1 parity with the current 37 tools across domains: environments/health/tasks, git/deploy/migration/backup/maintenance, Superset operations (read + admin CRUD), baseline capture/approval/verification, scenario compile/validate/resolve/draft-pack/save.
show_capabilitiesis retired in favor of standardtools/list. - MCPX-FR-009: Decommission MUST be flag-driven and ordered: parity proven → chat hidden → agent service dropped from
run.sh/compose → code deletion. Frontend entry points MUST switch to ordinary manual CRUD/editor routes. Agent chat, prompt textareas, assistant editing, typical-operation selectors that invoke an agent, proposal-generation and agent launch/handoff controls MUST be removed, not renamed. - MCPX-FR-010: The catalog MUST carry a version; breaking changes (rename/schema change/removal) MUST bump the major version and remain listed with a deprecation marker for one minor cycle.
- MCPX-FR-011: MCP tools MUST reuse existing services and contracts; a second Playwright/LLM/SQL/execution stack is forbidden (SCEX-FR-009 inheritance). No MCP path may bypass 038 validation, 037 baseline rules, 044 executors or a required gate.
- MCPX-FR-012: InvestigationSignals and queue items MUST NOT trigger any MCP activity automatically; MCP is pull-only from the client side (SCAN-FR-001 inheritance).
- MCPX-FR-013: The backend MUST implement the Authorization Server endpoints of the Authorization Model table; PKCE S256 MUST be mandatory for public clients, and
/oauth/authorizeMUST reuse the existing web session (local password or ADFS OIDC) without a second credential prompt inside an active session. - MCPX-FR-014: The MCP endpoint MUST publish RFC 9728 protected-resource metadata and answer unauthenticated calls with
401+WWW-Authenticateso a compliant client completes discovery → registration → authorization → tool listing without manual token pasting. - MCPX-FR-015: Dynamic client registration (RFC 7591) MUST be supported for public clients with first-party-bounded scopes; registered clients MUST be visible and revocable in Admin.
- MCPX-FR-016: Refresh tokens MUST rotate on use, and replay of a rotated refresh token MUST revoke the whole token family (extending
TokenBlacklist). - MCPX-FR-017: Access tokens MUST carry identity only (
sub,scope,jti,session_id,aud=mcp); permission resolution MUST hit live DB state per request so role changes apply to the next call without re-consent. - MCPX-FR-018: Catalog visibility MUST follow the hidden-vs-gated rule: missing permission hides the tool from
tools/list; contextual risk (PROD environment, baseline publish, non-delegated mutation) keeps the tool listed but returnsapproval_required. - MCPX-FR-019: HumanCheckpoint disposition MUST be available via
list_checkpoints+decide_checkpointunder the same CAS/audit contract as the 045 monitor, restricted to authenticated user principals; no automated origin may create, consume or bypass a checkpoint, and manual-run-only revisions remain ineligible for automation (SCEX-FR-004a stands). - MCPX-FR-020: Until an explicit architecture amendment states otherwise, MCP clients and all LLM/VLM providers are locally deployed inside the enterprise trust perimeter. PII exposure to those local providers is an accepted residual risk and is not blocked by masking. Credentials, cookies, access/refresh/service tokens, raw secrets, and raw storage paths MUST remain absent from tool payloads, telemetry, provenance, and artifact references.
- MCPX-FR-021: Dynamic client registration MUST be rate-limited, fully audited, redirect-URI validated, constrained to reviewed first-party scopes, and immediately revocable. DCR must not grant database permissions or allow scope escalation beyond the principal's live RBAC.
- MCPX-FR-022: The catalog MUST provide the persistent
AgentAuthoringWorkspaceoperationscreate_authoring_session,propose_test_plan,start_exploration,get_exploration_result,propose_graph_revision,get_graph_diff, andpromote_to_scenario(or an explicitly versioned equivalent) with server-owned state, provenance, idempotency and CAS. - MCPX-FR-023: Authoring exploration MUST use an isolated, allowlisted, time/size/network-bounded and cancellable sandbox with durable receipts and artifact ownership; shell, credential, filesystem/network escape and production side effects MUST be rejected. Sandbox readiness is a release gate, not an implementation claim.
- MCPX-FR-024: MCP MUST pass authoring output only as typed 038 action candidates/graph proposals through deterministic compile/validate, user-reviewed diff, and 042 handle-based save. Raw code, browser URLs, cookies, secrets, filesystem paths and caller digests MUST never establish authority; code-backed production execution is a separate unimplemented contract.
- MCPX-FR-025: MCP MUST provide
bootstrap_authoring_scenariofor a first scenario. It accepts only a bounded InitialScenarioIntent and server-issued valid compiled/draft-pack handles; atomically creates the registry entry, initial current revision, provenance/outbox and a workspace bound to that revision. It MUST NOT require or accept a client-created base revision, graph, content hash or revision ID. - MCPX-FR-026: MCP MUST expose the SCAUTO-FR-017 automation operations with identical REST policy/RBAC behavior. Schedule mutations require an explicit active revision, environment and idempotency key; they reject HumanStep revisions before any durable side effect and preserve the 046 PROD gate before dispatch.
- MCPX-FR-027: External reachability (ADR-0024, field run 2026-09-07): every durable prerequisite of an MCP write tool MUST be obtainable by the same principal class through the MCP catalog itself. Concretely, the catalog MUST provide
create_agent_runandget_agent_runwrapping the existingServices.AgentRuns.Servicecreate/snapshot boundary (permission("dashboard:testing","EXECUTE")at REST parity,service_allowed=False, curated UIContext-v2-bounded input, idempotent active-run reuse), so thatregister_draft_pack→bootstrap_authoring_scenario→start_scenario_runis completable by an external client without any non-MCP transport. A pinned catalog-introspection test MUST fail when a write tool gains a durable prerequisite that is not MCP-creatable. - MCPX-FR-028: Context-derived capabilities (field run 2026-09-07): the MCP-carried compile path MUST consume a server-owned capability authority derived from the live
inspect_dashboard_contextresult (authoritativeDashboardQueryModel+ environment policy + provider readiness), not only caller-declared booleans. Facts the server can verify truthfully (dataset-field availability, native filters, export capability) MUST NOT degrade intohuman_checkpoint; unresolved facts remainneeds_context/needs_selector/needs_baseline;human_checkpointis reserved for genuinely unsafe mutation contexts, human judgement, or unavailable automation. Contract: 038 amendment 2026-09-07. - MCPX-FR-029: Disposition clarity (field run 2026-09-07): every human-facing surface of the HumanCheckpoint disposition vocabulary (MCP tool descriptions, monitor UI labels, docs) MUST name the action by its persisted step outcome —
confirm→passed,false_positive/inconclusive→inconclusive(044 lifecycle mapping is immutable) — and MUST NOT present a passing choice with defect-confirming wording or destructive styling. Contract: 044/045 amendment 2026-09-07.
Canonical Scenario MCP Names
The following snake_case names are the canonical MCP names for the complete authoring, persistence, activation, and run-target chain. The operationIds in the final column are transport-specific REST aliases only; they are not MCP names and MUST NOT be exposed by tools/list.
| Canonical MCP name | Responsibility | Legacy REST operationId alias(es) |
|---|---|---|
create_agent_run |
Create the durable principal-owned AgentRun prerequisite for draft-pack registration (ADR-0024) | Api.AgentRuns.Create |
get_agent_run |
Read the ownership-scoped AgentRun snapshot | Api.AgentRuns.GetSnapshot |
inspect_dashboard_context |
Inspect dashboard/query context | inspectDashboardQueryModel |
bootstrap_authoring_scenario |
Atomically create a first scenario, initial current revision and bound workspace from server-owned handles | scenarioRegistry.create |
create_authoring_session |
Create persistent authoring workspace | none; MCP-only workspace operation |
propose_test_plan |
Record typed test-plan intent | none; MCP-only workspace operation |
start_exploration |
Start bounded authoring exploration | none; MCP-only workspace operation |
get_exploration_result |
Read exploration result/artifact refs | none; MCP-only workspace operation |
propose_graph_revision |
Submit typed graph proposal | scenarioEditor.agentPropose, scenarioEditor.apply |
get_graph_diff |
Read server-computed proposal/revision diff | scenarioRegistry.revisionDiff |
promote_to_scenario |
Create/update validated server handles or a save request; never activate | scenarioEditor.proposalSave, scenarioRegistry.create |
scenario_compile |
Compile canonical ScenarioGraph | compileDashboardScenario, compileScenario |
scenario_validate |
Validate compiled graph | validateDashboardScenario, validateScenario |
scenario_resolve |
Resolve typed authoring inputs into a new compiled handle | resolveDashboardScenario, resolveScenario |
generate_draft_pack |
Generate server-owned draft pack | compileScenarioDraftPack, buildScenarioDraftPack |
request_save |
Create immutable candidate revision, or eligible initial current revision | editor.save, scenarioRegistry.create |
activate_revision |
CAS-activate an eligible materialized revision | scenarioRegistry.activateCurrentRevision |
start_scenario_run |
Start a run pinned to an explicit promoted revision | scenarioRun.start |
list_scenario_schedules / upsert_scenario_schedule / delete_scenario_schedule |
Manage schedule projections through 046 policy | automation.listSchedules, automation.upsertSchedule, automation.deleteSchedule |
list_scenario_trigger_rules / upsert_scenario_trigger_rule / delete_scenario_trigger_rule |
Manage event trigger rules through 046 policy | automation.listTriggerRules, automation.upsertTriggerRule, automation.deleteTriggerRule |
get_scenario_automation_policy / upsert_scenario_automation_policy / get_scenario_automation_metrics |
Read or manage automation policy and observability | automation.getPolicy, automation.upsertPolicy, automation.metrics |
promote_to_scenario does not activate and does not advance current_revision. It creates or updates server-owned validated handles or a server-stored save request. request_save creates a candidate; only the separate activate_revision operation can advance current_revision, after eligibility, materialization, policy, required approval, and CAS checks. Initial creation is the sole explicit exception when the initial eligible revision is created as current. start_scenario_run must receive an explicit promoted revision_id and verified content_hash, never a candidate or an implicit latest/current target.
Key Entities
- McpToolCatalog: Versioned, explicitly registered set of tools grouped by domain; source of truth for names, schemas, risk class and
required_permission(resource, action). - McpSessionPrincipal: Authenticated caller identity (user or service), role set resolved live from DB per invocation, and delegation scope.
- OAuthClientRecord: Dynamically or statically registered client (public/confidential), first-party scope bound, owner visibility, revocation state.
- ApprovalRequiredEnvelope: Typed result binding a refused-or-deferred action to its durable gate (gate_id, targets, risk, reason_required).
- ToolInvocationRecord: AgentAction-provenance row per MCP tool call (principal, tool, argument digest, outcome, run linkage).
- HandoffSurface: Retired frontend concept (2026-09-08). Connection/security administration may show endpoint configuration, but no prompt, agent launch or proposal workflow belongs in product UI.
- AgentAuthoringWorkspace: Persistent server-owned MCP co-authoring session with CAS state, proposals, exploration results, artifacts, review and promotion references.
- AuthoringArtifact: Bounded server-owned source/patch/trace/screenshot/diagnostic/receipt reference used for review, never an execution program or production authority.
@{ McpInterface.ScenarioPipeline [C:5] [TYPE ADR]
@BRIEF Normative ownership and provenance contract for the MCP-to-scenario-run path. @RELATION DEPENDS_ON -> [ScenarioGraph.Compiler.Compile] @RELATION DEPENDS_ON -> [ScenarioRegistry.RevisionChain] @RELATION DEPENDS_ON -> [ScenarioExecution.RunnerPlan.Derive]
AgentAuthoringWorkspace MCP operations
The MCP catalog MUST expose bootstrap_authoring_scenario for initial creation and the following target contract for persistent co-authoring of an existing revision: create_authoring_session, propose_test_plan, start_exploration, get_exploration_result, propose_graph_revision, get_graph_diff, and promote_to_scenario. Names MAY be versioned during implementation, but the operation semantics and ownership boundary are normative.
All mutable workspace, proposal, exploration, artifact, review and promotion state is server-owned. Mutating operations require authenticated principal/delegation, idempotency key and expected workspace CAS version; replays return the original result and changed requests or stale versions return typed 409 without partial mutation. Read operations enforce object ACL and return bounded refs/digests for large traces/screenshots.
start_exploration uses only the isolated sandbox contract: allowlisted origins/APIs/actions, no shell/credential/filesystem escape, bounded time/size/network, cancellation, operation receipts, artifact ownership and no production side effects. The MCP server does not generate arbitrary Playwright code; it exposes reviewed templates/actions and consumes sandbox traces/candidates. Raw code, URLs, cookies, secrets, paths and caller digests are rejected as authority.
The initial chain is inspect -> compile/validate -> draft-pack -> bootstrap_authoring_scenario -> immutable current revision. The existing-scenario chain is session -> plan -> exploration -> typed proposal -> 038 compile/validate -> user-reviewed diff -> 042 save/promotion -> immutable revision. An unbound workspace is not a creation path and cannot propose or save a graph. promote_to_scenario cannot skip user review or required approval. 044 accepts only the resulting promoted revision. Code-backed production execution is a separate future, unimplemented contract.
The following table is the single cross-spec stage contract. A handle is an opaque server-issued reference; a digest is computed and verified by the owner named in the table. A client may pass a handle or request digest for correlation, but it cannot establish authority with a graph, caller-computed digest, path, or file.
Implementation status (audit 2026-09-06, Doc.Adr.ADR0023; Phase 2c T029d–T029h
in tasks.md). The deterministic cores of compile/validate/resolve/draft-pack
are IMPLEMENTED as pure 038 functions. The durable handle layer is IMPLEMENTED:
CompiledScenarioHandle/ValidationResultHandle/DraftPackHandle entities
(migrations 0019_scenario_handles, 0020_scenario_materialization,
0021_context_authority) are minted at the persisted boundaries — REST
api_compile_scenario/api_validate_scenario/api_resolve_scenario/
api_draft_pack and the MCP register_draft_pack write tool — with
content-addressed canonical-bytes storage, owner binding, single consumption
(SELECT ... FOR UPDATE, PostgreSQL-proven) and binding/digest/staleness
rejections (ScenarioGraph.Handles). bootstrap_authoring_scenario and 042
CreateScenario/CreateInitial accept ONLY stored handle ids; the transitional
caller-composed compile:{run_id}:{digest} string path is removed for bootstrap.
042 create materializes graph_snapshot = the canonical DashboardTestScenario
JSON (plus server-owned action-registry identity and context_authority) inside
the create transaction, so a freshly bootstrapped current revision is directly
runnable/schedulable; BOOTSTRAP_REVISION_NOT_RUNNABLE survives as
defense-in-depth for provenance-only legacy rows. The
OutboxEvent/RevisionMaterialization reference-artifact worker exists
(materialize_pending_revisions); wiring it into the scheduler poll loop is the
remaining operational step. Residuals: the 043 editor promote/save path and the
REST-only api_draft_pack flow do not yet evaluate context_authority
(persisted as NULL = legacy-allowed at the PROD gate). The inspect/context row
is resolved as a hybrid (T029h, 2026-09-06): no persisted
InspectionContextHandle is minted; the MCP read tool inspect_dashboard_context
exposes the live authoritative DashboardQueryModel resolver, and the
register_draft_pack boundary binds the client-carried context by recomputing
its fingerprint against a live inspection (context_authority =
verified/unverified; sentinel fingerprints never verify; unreachable
environments fail open; PROD dispatch refuses an explicit non-verified marker
via CONTEXT_AUTHORITY_REQUIRED_FOR_PROD; the validator recursively rejects
query_context/SQL smuggling inside dashboard_context). The table remains the
normative target; the InspectionContextHandle row records the declined
persisted-handle alternative, superseded by the hybrid binding above.
Field-run correction (2026-09-07, Doc.Adr.ADR0024; Phase 2d in tasks.md). The
live external run against ss-prod (docs/2026-09-07-sales-prod-mcp-run.md)
proved the initial bootstrap row is NOT externally reachable end-to-end:
register_draft_pack requires a principal-owned AgentRun, and no MCP operation
creates one since the chat decommission removed the only creation surface. The
row's test_mcp_initial_scenario_e2e evidence seeded that prerequisite with a
raw-ORM insert, which masked the gap (test-honesty rule, ADR-0024 §3).
Closed the same day (Phase 2d): MCPX-FR-027 landed as MCP create_agent_run/
get_agent_run (catalog 2.2.0, T029i) — the vertical is now the fully external
chain with zero non-MCP seeding and the reachability pin is hard-green
(E2E-EXT-001 CLOSED); MCPX-FR-028 landed as ScenarioGraph.CapabilityAuthority
(T029k, CAP-001 CLOSED); MCPX-FR-029 landed as outcome-based labels/descriptions
(T029l, DISP-001 CLOSED). Remaining: none for Phase 2d — the live-stand replay E2E-EXT-002
(T029m) CLOSED 2026-09-11 (see the gate row below).
| Stage | Input | Authoritative owner | Output handle | Digest | Persistence | Failure / no-side-effect rule | Test id |
|---|---|---|---|---|---|---|---|
| inspect/context | dashboard/environment request | 038 context/capability resolver, with server ACL | InspectionContextHandle |
server context fingerprint | server request record only | unresolved facts become needs_context, needs_selector, or needs_baseline; no graph/revision/run |
Test.Scenario.Capability, Test.Scenario.Resolver |
| compile | bounded intent + server context | 038 ScenarioGraph.Compiler |
CompiledScenarioHandle |
server content_hash of canonical graph/program |
immutable server compile result | deterministic validation first; errors persist no mutable artifact | Test.Scenario.Compiler, Test.Scenario.Serializer |
| validate | CompiledScenarioHandle or server graph |
038 ScenarioGraph.Validator |
ValidationResultHandle |
server validation-result digest bound to compiled hash | immutable server validation result | invalid/blocking findings produce no draft, revision, gate, or run | Test.Scenario.Validator, Test.Scenario.Validator.Edge |
| resolve | server compiled handle + typed resolution + expected base | 038 resolver; 042 owns revision CAS when persisted | new CompiledScenarioHandle |
new server content_hash, parent hash link |
immutable server compile result; no in-place graph edit | stale base is 409 STALE_REVISION; no partial graph or revision mutation |
Test.Scenario.Resolver, Test.Scenario.Resolver.Edge |
| draft-pack | valid compiled/validation handles + registered templates | 038 ScenarioGraph.PackCompiler |
DraftPackHandle |
server draft_pack_digest over rendered manifest/bytes |
server draft registry, save_eligible or preview_only |
unresolved/error/unsafe input is preview_only; no registry revision or launch side effect |
Test.Scenario.Pack, Test.Scenario.Pack.Security |
| initial bootstrap | InitialScenarioIntent + valid compiled/draft-pack handles + idempotency key | 042 ScenarioRegistry.CreateInitial |
ScenarioRegistryEntry + initial current ScenarioRevision + bound workspace | server-verified content hash and draft digest | PostgreSQL registry/revision/workspace + outbox | invalid, stale, unauthorized or replay-conflicting input creates no partial row; no synthetic base revision | test_mcp_initial_scenario_e2e |
| request-save | compiled/draft handles + expected base + idempotency key | 042 registry transaction | ScenarioRevision (candidate, or initial current) |
server-verified graph/content hash and draft digest | PostgreSQL revision + outbox; materialization async | CAS/idempotency failure is atomic; no row/outbox on rejection; save never activates | test_scenario_registry, test_revisions |
| run preflight | selected ScenarioRevision + typed bindings + server target snapshot |
044 RunPreflight |
RunPreflightHandle |
server preflight digest over revision/target/bindings/policy | durable preflight record attached to launch request | missing/invalid context, selector, baseline, policy, or revision mismatch blocks before dispatch/I/O | test_scenario_runner, test_scenario_automation_api |
| runner-plan | eligible preflight + immutable revision | 044 RunnerPlan.Derive |
RunnerPlanHandle |
server descriptor/plan digest | persisted on ScenarioRun; git runner.plan.json is reference only |
altered/unknown descriptor rejects before lease, execution, or provider I/O | test_scenario_runner_plan, test_scenario_dispatch |
| launch request | RunnerPlanHandle + idempotency key |
044 server runner | ScenarioRun |
server launch/request digest and pinned plan digest | PostgreSQL run, queue/gate, audit | same key+hash replays same run; changed request is 409 IDEMPOTENCY_KEY_REUSED; no duplicate dispatch |
test_scenario_queued_dispatch, test_scenario_scheduler_callbacks |
| dispatch/evidence | pinned plan descriptor + capacity lease | 044 executor/provider and evidence owner | EvidenceRef / StepOutcome |
provider-verified receipt SHA-256 | Artifact(owner_type=scenario_run) + immutable receipt + step outcome |
missing ownership/digest/ref is non-pass; late/unknown effect reconciles or terminalizes non-pass | test_scenario_terminal_signals, test_provider_contract |
| signal | terminal run + immutable evidence provenance | 044 producer; 047 queue/episode projection | idempotent InvestigationSignal / queue input |
server signal identity digest | durable signal; 047 projection only | failed/blocked/inconclusive emits once; passed emits none; never opens chat/action automatically | test_scenario_terminal_signals |
Handle and Digest Rules
CompiledScenarioHandleis minted only after 038 canonical serialization. Itscontent_hashis SHA-256 of the executable canonical Verification Program graph; it excludes timestamps, display fields,ParameterBindings, and runtime identity. 038 never mintsscenario_idorrevision_id.ValidationResultHandleis valid only for the exact compiled-handle hash and validator/schema versions used to produce it. A validation digest does not become graph authority and cannot be supplied by a caller as proof of validity.DraftPackHandleis minted by the server pack registry from registered, versioned templates. Its digest covers the server-rendered manifest and bytes;save_eligiblerequires valid validation and no unresolved required marker.- 042 alone mints
scenario_id,revision_id, parent links, and server-owned revision records. Save verifies the compiled hash and draft digest against stored handles inside one transaction; client graph uploads are rejected. RunPreflightHandleandRunnerPlanHandleare derived by 044 from persisted revision data, server target/policy context, and typed bindings. Their digests are server recomputed; a caller-supplied runner or digest is correlation data.runner.plan.jsonandscenario.yamlare materialized reference artifacts. They can be regenerated and are never authority for validation, revision, preflight, dispatch, retry, recovery, or evidence.- Status (2026-09-06, updated after Phase 2c):
CompiledScenarioHandle,ValidationResultHandleandDraftPackHandleARE minted and stored (038ScenarioGraph.Handles, migrations0019–0021); the rules in this section are their implemented acceptance criteria.InspectionContextHandleis NOT persisted — T029h resolved the inspect stage as a live fingerprint binding (context_authority) instead. The authority that 042 materializes intoScenarioRevision.graph_snapshotis the persisted canonicalDashboardTestScenarioserialization behindCompiledScenarioHandle.canonical_bytes_ref— explicitly NOT the lossy pack templates above (038ScenarioGraph.ServerOwnedPipeline, amendment 2026-09-06).
Resolution and Continuation Semantics
038 owns inspection and semantic resolution. needs_context means a required
dashboard/query/change-request fact is absent; needs_selector means a browser
target/action selector is unknown; needs_baseline means an approved baseline
reference is absent or stale. None may be guessed or silently defaulted. 038
may compile a WorkingDraft/preview, but only 042 can save a revision and only
044 can reject launch bindings at RunPreflight.
Save is draft/working -> validated save request -> candidate, with the initial
revision as the sole exception (candidate -> current at creation). Activation
is a distinct CAS operation: candidate -> current, guarded by eligibility,
materialization, policy and required approval. Approval continuation is
pending_approval -> queued; deny/expiry is pending_approval -> blocked.
Every continuation consumes its gate/checkpoint exactly once. Expected revision,
gate, checkpoint and idempotency versions are CAS inputs; stale values return a
typed 409 and create no side effect. Same idempotency key plus the same canonical
request hash replays the existing result; the same key plus a different hash is
409 IDEMPOTENCY_KEY_REUSED.
The execution chain is strictly:
ScenarioRevision -> RunPreflight -> RunnerPlan -> ScenarioRun -> queued|pending_approval -> dispatch -> EvidenceRef/StepOutcome -> 047 signal.
ScenarioRevision owns immutable program input, RunPreflight owns launch
eligibility, RunnerPlan owns descriptor order/policy, ScenarioRun owns the
execution snapshot, 044 owns step/evidence receipts, and 047 owns queue/episode
projection. No client graph, caller digest, or materialized plan can replace
these owners.
@{ McpInterface.ScenarioPipeline.ReleaseGates [C:5] [TYPE Block]
@BRIEF Conjunctive release gates for the coordinated pipeline contract.
Release is GO only when every applicable gate passes: 038 fixture/schema and
deterministic compiler/validator/serializer/pack checks; 042 registry create,
revision CAS, idempotency, outbox and materialization checks; 044 preflight,
plan, queued approval, exact dispatch, evidence ownership and terminal signal
checks; 050 parity, bounded transport, hidden-vs-gated and end-to-end MCP checks;
and structural anchor/Markdown audit. Partial green is NO-GO and does not mark
implementation tasks complete.
Evidence commands:
cd backend && source .venv/bin/activate && python -m pytest -q tests/services/dashboard_testing/scenario/test_*.py tests/api/test_scenario_runs_api.py tests/api/test_scenario_automation_api.py tests/api/test_scenario_analytics_api.py
cd backend && source .venv/bin/activate && python -m ruff check src/services/dashboard_testing/execution src/api/routes/dashboard_testing/scenario_runs.py
python specs/044-dashboard-scenario-execution/prototype/validate_static.py
cd frontend && npm run test -- --run && npm run lint && npm run build
The required MCP parity and end-to-end evidence is test_mcp_* plus the 050
SC-001/SC-002 walkthrough; absent or partial evidence keeps the gate NO-GO.
Authoring E2E Traceability
These are required evidence rows for the authoring promotion path. They remain
open until executable evidence is retained; unchecked rows keep this release
gate NO-GO and do not mark T023 or T028 complete.
| Evidence row | Required trace | Evidence required | Status |
|---|---|---|---|
E2E-AUTH-001 |
create_authoring_session -> propose_test_plan -> start_exploration -> get_exploration_result |
External MCP client trace proves persistent server-owned workspace, bounded exploration, durable receipt, and typed artifact refs | [x] CLOSED 2026-09-03 — tests/test_mcp_authoring_promotion_e2e.py::test_sandbox_output_promotes_to_immutable_revision (plan → queued exploration → sandbox dispatch → exploration_passed, evidence draft:exploration-*, bounded projection) |
E2E-AUTH-002 |
propose_graph_revision -> scenario_compile -> scenario_validate -> get_graph_diff |
External MCP client trace proves typed proposal, deterministic 038 validation, server-computed diff, and explicit user review boundary | [x] CLOSED 2026-09-03 — tests/test_mcp_scenario_e2e.py (propose → promote validation → 038 inspect_scenario/scenario_resolve/validate_scenario over MCP → get_graph_diff; canonical tool names inspect_scenario/validate_scenario per FR-008 rename clause) |
E2E-AUTH-003 |
promote_to_scenario -> request_save -> activate_revision -> start_scenario_run |
External MCP client trace proves handle-based save creates candidate, activation separately passes eligibility/materialization/policy/approval/CAS, and the run pins promoted revision_id + content_hash |
[x] CLOSED 2026-09-03 — tests/test_mcp_scenario_e2e.py extended: post-activation start_scenario_run creates a queued run pinned to scenario_revision_id + scenario_content_hash of the promoted revision |
Field-run remediation gates (2026-09-07, Doc.Adr.ADR0024)
Required evidence rows produced by the live external run against ss-prod
(docs/2026-09-07-sales-prod-mcp-run.md). Unchecked rows keep this release gate
NO-GO and do not mark the referenced tasks complete.
| Evidence row | Required trace | Evidence required | Status |
|---|---|---|---|
E2E-EXT-001 |
create_agent_run -> register_draft_pack -> bootstrap_authoring_scenario -> start_scenario_run |
External MCP-only chain on a fresh DB with ZERO non-MCP prerequisite seeding; the strict-xfail pin tests/test_mcp_agent_run_reachability.py flipped to green and unmarked (MCPX-FR-027, T029i) |
[x] CLOSED 2026-09-07 — tests/test_mcp_initial_scenario_e2e.py converted (AgentRun minted via MCP create_agent_run tools/call; no service/ORM seeding remains); reachability pin unmarked+green; tests/test_mcp_agent_run_tools.py (5) covers create/replay/read/RBAC; slice 70 passed, full suite 11357 passed |
E2E-EXT-002 |
live stand replay of the 2026-09-07 sales run | inspect_dashboard_context -> derived-capability compile -> create_agent_run -> register -> bootstrap -> queued/gated run on ss-prod + list_checkpoints/decide_checkpoint human loop, with retained evidence (T029m) |
[x] CLOSED 2026-09-11 — committed replay client specs/044-dashboard-scenario-execution/prototype/live_mcp_replay.py replayed the full external chain on the live stand (run 110a6517…: inspect derived capabilities → create_agent_run → compile B01 → register context_authority=verified → bootstrap → PROD gate pending_approval (identical retry = one durable gate) → approval → live capture_screenshot passed (8 refs) + typed BROWSER_ACTION_NOT_SUPPORTED → honest inconclusive); the human loop was exercised live by live_mcp_human_loop.py (compiled B05 HumanCheckpoint → waiting_human → MCP list_checkpoints (v1) → decide_checkpoint confirm (v2 CAS) → terminal passed). Two fail-closed defects found and fixed (binding resolution for identity-less compiled steps; per-step target-identity stamping at bootstrap). Trace: docs/2026-09-11-sales-prod-mcp-replay.md |
CAP-001 |
inspect_dashboard_context -> compile capabilities |
Server-derived capability map proven: verifiably-available dataset fields/native filters classify automated, not human_checkpoint; unsafe-mutation cases stay human_checkpoint (MCPX-FR-028, 038 amendment, T029k) |
[x] CLOSED 2026-09-07 — ScenarioGraph.CapabilityAuthority (derivation/derived-wins merge/boundary choke point) wired into MCP inspect_scenario+inspect_dashboard_context and REST api_compile_scenario; CAP-001 classification-fix test through map_all on the sales-shape fixture (B/T automated, C04–C06 unsupported, mutation cases legitimately human) + tool/REST parity pins; slice 67 passed |
DISP-001 |
disposition label/vocabulary audit | RU/EN labels and MCP tool descriptions name the persisted outcome (confirm→passed); no defect-confirming wording or destructive styling on a passing choice; mapping pinned by tests (MCPX-FR-029, 044/045 amendment, T029l) |
[x] CLOSED 2026-09-07 — ru/en labels renamed to outcome-based wording; confirm buttons bg-destructive→bg-primary (HumanCheckpointPanel + WaitingForMeView); decide_checkpoint docstring carries the immutable outcome table; vitest DISP-001 style/label/dispatch pin; frontend 3507 passed, lint 0 errors, build OK |
TEST-001 |
vertical-test honesty remediation | test_mcp_initial_scenario_e2e.py::_pack() creates the AgentRun through the production create_agent_run service boundary (no raw-ORM prerequisite seeding); reachability requirement pinned as strict=True xfail; typed zero-side-effect denial pinned (ADR-0024 §3, T029j) |
[x] CLOSED 2026-09-07 — pytest tests/test_mcp_initial_scenario_e2e.py tests/test_mcp_agent_run_reachability.py green (1 passed + 1 xfailed(strict) + denial pin); metadata records the open T029i gap. (Superseded the same day by T029i closure: the strict-xfail flip ritual executed as designed — pin unmarked and hardened, vertical converted to the fully external chain; denial pin retained.) |
@} McpInterface.ScenarioPipeline.ReleaseGates
@} McpInterface.ScenarioPipeline
Success Criteria
- SC-001: An external client completes the full scenario creation chain for a fixture dashboard and the revision appears in the registry with correct provenance.
- SC-002: 100% of parity tests pass: MCP tool outcomes match legacy
@toolwrapper outputs on shared fixtures. - SC-003: 100% of gated invocations produce zero side effects before approval; denial/expiry paths refuse cleanly.
- SC-004:
tools/listis RBAC-exact for admin/analyst/viewer roles in fixture tests. - SC-005: After decommission, builds and link-integrity suites pass with zero
/agent, assistant-API or Gradio-proxy references; the stack starts without port 7860. - SC-006: No catalog path reaches arbitrary SQL or raw Superset mutation outside the governed tools; scenario contexts cannot invoke SQL-class tools.
- SC-007: A compliant external client completes discovery (RFC 9728) → DCR → authorization code + PKCE →
tools/listin one automated flow, with zero manual token management. - SC-008: Refresh-token replay revokes the token family in 100% of fault-injection cases; pre-revocation access tokens die at expiry, not silently extended.
- SC-009: A role change is reflected in the next
tools/listand the next invocation without new consent; a revoked permission hides the tool and denies cached-catalog calls.
Clarifications
Session 2026-08-24
- Q: Separate MCP process or mounted in backend? → A: Mounted in the FastAPI app; simplest deployment, in-process service reuse, shared auth middleware.
- Q: Does removing chat remove HITL? → A: No, and it does not force the web UI either (session 2026-08-24): gates are server-durable and fully resolvable inside MCP (
list_pending_approvals+decide_approval); web gate cards on product pages remain an equal renderer of the same rows for browser-first users. - Q: Is the vendored mcp-superset adopted? → A: No — reference only; it bypasses the policy layer.
- Q: What happens to AgentRun/DraftArtifact/InvestigationSignal contracts? → A: They persist unchanged; MCP invocations create the same provenance rows. Only the conversational transport and its UI retire.
- Q: OAuth now or later? → A: Now (session 2026-08-24 decision). The backend becomes a full OAuth 2.1 Authorization Server in Phase 0 —
Authlib==1.6.6is already a dependency; personal access tokens are explicitly deferred, not rejected forever. - Q: Where do rights live? → A: In the existing DB RBAC (
Role/Permission,user_has_permission). Tokens never embed roles; the catalog binds tools to canonicalresource:actionpairs and filters both listing and invocation through the same predicate.
Session 2026-09-07 (field-run remediation, Doc.Adr.ADR0024)
- Q: The 2026-08-24 clarification said AgentRun contracts "persist unchanged" — who creates an AgentRun after the chat retirement? → A: The clarification preserved persistence but silently retired the only creation entry point. Corrected: MCP gets an explicit typed creation/read surface
create_agent_run/get_agent_runover the same service (MCPX-FR-027, T029i). Rejected alternatives (no-AgentRun registration boundary, implicit auto-create, REST crutch, raw-ORM test seeds) are recorded in ADR-0024. - Q: May a vertical E2E seed a chain prerequisite directly into the DB? → A: No. A vertical test MUST obtain every prerequisite through a boundary the principal under test can reach; where that boundary does not exist yet, the test uses the production service boundary with honest metadata AND a
strict=Truexfail reachability pin (test-honesty rule, ADR-0024 §3; executed as T029j). - Q: Why did the B01 preview classify verifiably automatable checks as
human_checkpoint? → A: Caller-declared capability booleans are not an authority. Server-derived capabilities from the live inspection are required (MCPX-FR-028, 038 amendment 2026-09-07, T029k);human_checkpointstays reserved for genuinely unsafe mutation contexts, human judgement, or unavailable automation, and HumanStep revisions remain automation-ineligible. - Q: Is the UI label «Подтвердить проблему» for disposition
confirmcorrect? → A: No —confirmpersists step outcomepassed(044 lifecycle mapping is immutable and correct); the label inverts the meaning. Labels/tool descriptions must name the persisted outcome (MCPX-FR-029, 044/045 amendment 2026-09-07, T029l).
Phases
| Phase | Scope | Exit evidence |
|---|---|---|
| 0 | OAuth AS+RS skeleton: authorize/token/DCR/JWKS/metadata endpoints, /mcp validation middleware, 2–3 probe tools, Inspector + scripted-client connectivity |
SC-007 green in CI |
| 1 | Parity catalog for the 37 tools + contract tests mirroring wrapper tests; permission-bound hidden/gated matrix | SC-002, SC-004, SC-009 |
| 2 | Gates/provenance over MCP; end-to-end scenario creation walkthrough; refresh rotation/reuse tests | SC-001, SC-003, SC-008 |
| 2d | Field-run remediation (ADR-0024): MCP AgentRun surface, external reachability pin, context-derived capabilities, disposition clarity, live stand replay | E2E-EXT-001/002, CAP-001, DISP-001 (TEST-001 closed) |
| 3 | Frontend decommission; manual editor entry; assistant interaction removal | SC-005 partial |
| 4 | Delete agent/ service; run.sh/compose updates; spec amortization closed |
SC-005 full |
Spec Impact & Amortization Map
| Spec | Amendment |
|---|---|
| 036 | Transport rejection recorded; durable-runtime contracts carry over to MCP provenance |
| 037 | Baseline tools join the MCP catalog unchanged (thin forwarders) |
| 038 | Authoring entry becomes MCP-callable; validator remains sole gateway |
| 039 | AGUI-FR-001..013 rewritten for manual authoring/review; no agent entry |
| 040 | Delegated load experiments readable as MCP-driven, policy unchanged |
| 041 | Blast-radius explanation consumed by external clients; index contracts unchanged |
| 042 | Delegated actors explicitly include MCP principals; registry contracts unchanged |
| 043 | External MCP proposal authoring only; frontend manual editing/review; SCEDIT-FR-009 stands |
| 044 | Runner unaffected; authoring/investigation boundary inherited by MCP clients |
| 045 | Manual investigation case open/review only; external agents pull MCP, no UI launch |
| 046 | Schedule management exercisable through MCP tools under same policy |
| 047 | Case threads continue in external clients; server keeps evidence/timeline |
Production contract refresh — 2026-09-08
Frontend boundary (user decision 2026-09-08): All agent interaction is external MCP only. Product frontend MUST NOT contain agent chat, prompt/request textarea, assistant editing, typical-operation-to-agent selector, proposal-generation, agent workspace/start or handoff controls/routes. Ordinary manual CRUD/editor, human approval/review, monitoring and read-only evidence/evaluation are permitted. AgentEvaluationCard is read-only, with no prompt/retry-agent/provider controls. Existing agent proposal UI is runtime drift; removal/negative DOM-route-network acceptance remains OPEN in this spec-only change.
MCPX-FR-030 — Contract-complete public parity: Curated versioned tool schemas MUST expose baseline capture/review/request/decide/consume/explicit publish/lifecycle and run result/evidence operations with the same services, permissions, errors, CAS and idempotency as REST. No fallback success for consume/publication. Start/run/automation retain full server-resolved BaselineSelectionPin, never caller digests. External clients remain enterprise-local; live visual authoring/execution readiness remains open.
Normative contract: Contract-complete public parity. New requirements are specified, implemented=false / acceptance OPEN until executable evidence closes the linked tasks/checklist/traceability rows. Historical local tests and the manual inconclusive ss-prod run do not prove browser/capture/baseline/LLM production readiness. The refresh scope is the audited P0/P1/P2 agentic E2E and baseline gaps; an approved ExecutionPerformanceBaseline is not introduced.
#endregion McpInterface.Spec