Files
ss-tools/docs/2026-09-11-sales-prod-mcp-replay.md

7.0 KiB
Raw Blame History

Sales Dashboard MCP replay (T029m / E2E-EXT-002) — 2026-09-11

Objective

Close gate row E2E-EXT-002: replay the exact external MCP scenario chain on the live stand — inspect_dashboard_context → derived-capability compile → create_agent_run → register_draft_pack → bootstrap_authoring_scenario → gated start_scenario_run → approval → live execution to terminal — with zero non-MCP seeding, using the committed replay client specs/044-dashboard-scenario-execution/prototype/live_mcp_replay.py.

The client is a transport-only adaptation of the already-proven scripted flows (tests/test_mcp_client_flow_http.py OAuth/PKCE + tests/test_mcp_initial_scenario_e2e.py tool chain) against http://127.0.0.1:8000 (httpx + SSE); no new client machinery was built.

Execution facts

Stage Result
REST login + OAuth discovery → DCR (public client) → PKCE approve → token ok
/mcp initialize + session + tools/list 61 tools
inspect_dashboard_context (ss-prod, dashboard 11) ok; fingerprint sha256:16838a47…c7; derived: native_filters/text_filter/time_rollover/table_filter/pagination/dataset_field_read/browser = true, xlsx_export = false
create_agent_run ok; a6168c7d-aa93-439e-8c2e-161f585b211f
inspect_scenario (B01, capabilities = derived ∪ {screenshot}, no baseline) 2 steps: apply_native_filter (browser), capture_screenshot (screenshot); capability_authority status derived, has_dataset_fields true
validate_scenario valid=true, no blockers
generate_draft_pack save_eligible
register_draft_pack save_eligible; context_authority = verified (echoed live query model matched); handles dae2edbe… / bbfa3400…, digest 84b00d23…18e17
bootstrap_authoring_scenario scenario e2687f03-f6d3-45f7-b114-d27b023ae085, revision fd126da0-eb22-4da6-be73-387589762689 (current)
start_scenario_run (ss-prod) pending_approval; identical retry → same single durable gate (idempotent)
PROD gate decision (REST, admin) approved
Live execution (scheduler + providers) terminal inconclusive — see below

Final run: 110a6517-e433-4860-9d77-231521acba19.

Step outcomes (honest, fail-closed)

Step Status Detail
phase-5-B01-capture_screenshot (screenshot) passed 8 durable artifact refs — real live evidence captured through the MCP-authored graph
phase-2-B01-apply_native_filter (browser) inconclusive typed BROWSER_ACTION_NOT_SUPPORTED: the isolated browser transport admits only open_dashboard/wait_for_state/refresh read-only actions (plus contract-gated mutations)

The run terminal is therefore inconclusive — a typed non-pass surfaced as designed; nothing was synthesized into PASS. This is the correct fail-closed behavior, not a defect.

Production defects found by the replay (each fixed this session)

  1. Binding resolution missed compiled graphs. _resolve_configured_live_binding matched dashboards only from per-step dashboard_id, which compiled (MCP/bootstrap) steps do not carry — the run started with no binding and both steps failed typed *_BINDING_MISSING. Fix: the resolver now also matches the scenario registry entry's server-owned dashboard_id (start_run.py, scenario_dashboard_id), pinned by test_resolution_uses_registry_dashboard_for_identity_less_steps.
  2. Bootstrap materialization dropped per-step target identity. The isolated browser/screenshot admission requires step-level environment_id/dashboard_id to equal the adopted binding (fail-closed *_TARGET_MISMATCH otherwise). Fix: ScenarioRegistry.Create.Register (registry/create.py, handle path) stamps the verified dashboard_context identity onto every materialized step — authoritative because context_authority=verified is evaluated at the register boundary; steps that already declare identity are never overwritten.
  3. Stale binding principal. The deployment binding ss-prod-d11-live-001 carried an execution_principal_fingerprint from an unknown registration-time SS_STAND_USER, so the authoritative sha256(actor) guard rejected every start. Fixed operationally by re-registering the binding through the new admin surface (PUT /api/scenario-live-bindings/ss-prod-d11-live-001, principal = sha256("admin")) — which also served as the first live exercise of the D5 binding admin API (register → server-side resolution → run adoption).

Baseline pin

The compiled B01 graph declares no baseline refs (baseline capability deliberately off), so the gated start legally ran with a null pin. The live baseline-pinned variant of this replay remains blocked by the Gitea PAT (published catalog unavailable); 050 T045 and the pin part of T046 stay OPEN for that reason.

Gate row

E2E-EXT-002 (T029m) → CLOSED by this trace. Remaining MCP-surface work is limited to the PAT-blocked baseline-pinned variant (T045/T046-pin) and Expected.threshold (038 schema authority, separate architect decision).

Human-checkpoint loop (T029m formulation completeness)

live_mcp_replay.py's B01 graph contains no HumanStep, so the task's list_checkpoints/decide_checkpoint element was exercised by a companion replay specs/044-dashboard-scenario-execution/prototype/live_mcp_human_loop.py (same scripted client):

Stage Result
inspect_scenario B05 with live derived capabilities one step: phase-5-B05-human_checkpoint-1 (tool=human) — the script refuses any graph with a non-human step, so no mutation case is started
validate → draft-pack → register_draft_pack valid, save_eligible, context_authority=verified
bootstrap → start_scenario_run (ss-prod, manual) pending_approval → approved → run 6dcb9b8b-4f96-4491-9480-a3c74a67ee42
run state waiting_human
MCP list_checkpoints(run_id) pending checkpoint 35039c17-6469-4cdf-90a1-753a357c0e08, decision_version=1
MCP decide_checkpoint(confirm, expected_version=1) ok; checkpoint decided, decision_version=2 (CAS consumed)
terminal passed — the immutable 044 mapping confirm → passed names the human's conformance observation, not a synthesized PASS

This closes the aligned-label human loop live; combined with the B01 chain above, every element of the T029m task text is now backed by a live trace.

Review residuals (recorded, not defects)

  • The compile input declared screenshot=true as a caller capability; screenshot is not a server-derived dashboard fact, so merge_capabilities passes it through verbatim. The derived-wins axis that FR-028 protects (verifiable facts never downgraded to human_checkpoint) is intact and capability_authority.status=derived; the caller-key nuance belongs to the 038/T029k authority-boundary backlog, not to this replay.
  • live_mcp_replay.py now asserts the honest typed non-pass terminal (a passed terminal on the B01 chain raises UNEXPECTED_SYNTHETIC_PASS) and asserts the idempotent retry reuses the same run_id.