7.0 KiB
Sales Dashboard MCP replay (T029m / E2E-EXT-002) — 2026-09-11
Objective
Close gate row E2E-EXT-002: replay the exact external MCP scenario chain on the live stand —
inspect_dashboard_context → derived-capability compile → create_agent_run →
register_draft_pack → bootstrap_authoring_scenario → gated start_scenario_run →
approval → live execution to terminal — with zero non-MCP seeding, using the committed
replay client specs/044-dashboard-scenario-execution/prototype/live_mcp_replay.py.
The client is a transport-only adaptation of the already-proven scripted flows
(tests/test_mcp_client_flow_http.py OAuth/PKCE + tests/test_mcp_initial_scenario_e2e.py
tool chain) against http://127.0.0.1:8000 (httpx + SSE); no new client machinery was built.
Execution facts
| Stage | Result |
|---|---|
| REST login + OAuth discovery → DCR (public client) → PKCE approve → token | ok |
/mcp initialize + session + tools/list |
61 tools |
inspect_dashboard_context (ss-prod, dashboard 11) |
ok; fingerprint sha256:16838a47…c7; derived: native_filters/text_filter/time_rollover/table_filter/pagination/dataset_field_read/browser = true, xlsx_export = false |
create_agent_run |
ok; a6168c7d-aa93-439e-8c2e-161f585b211f |
inspect_scenario (B01, capabilities = derived ∪ {screenshot}, no baseline) |
2 steps: apply_native_filter (browser), capture_screenshot (screenshot); capability_authority status derived, has_dataset_fields true |
validate_scenario |
valid=true, no blockers |
generate_draft_pack |
save_eligible |
register_draft_pack |
save_eligible; context_authority = verified (echoed live query model matched); handles dae2edbe… / bbfa3400…, digest 84b00d23…18e17 |
bootstrap_authoring_scenario |
scenario e2687f03-f6d3-45f7-b114-d27b023ae085, revision fd126da0-eb22-4da6-be73-387589762689 (current) |
start_scenario_run (ss-prod) |
pending_approval; identical retry → same single durable gate (idempotent) |
| PROD gate decision (REST, admin) | approved |
| Live execution (scheduler + providers) | terminal inconclusive — see below |
Final run: 110a6517-e433-4860-9d77-231521acba19.
Step outcomes (honest, fail-closed)
| Step | Status | Detail |
|---|---|---|
phase-5-B01-capture_screenshot (screenshot) |
passed | 8 durable artifact refs — real live evidence captured through the MCP-authored graph |
phase-2-B01-apply_native_filter (browser) |
inconclusive | typed BROWSER_ACTION_NOT_SUPPORTED: the isolated browser transport admits only open_dashboard/wait_for_state/refresh read-only actions (plus contract-gated mutations) |
The run terminal is therefore inconclusive — a typed non-pass surfaced as designed; nothing was
synthesized into PASS. This is the correct fail-closed behavior, not a defect.
Production defects found by the replay (each fixed this session)
- Binding resolution missed compiled graphs.
_resolve_configured_live_bindingmatched dashboards only from per-stepdashboard_id, which compiled (MCP/bootstrap) steps do not carry — the run started with no binding and both steps failed typed*_BINDING_MISSING. Fix: the resolver now also matches the scenario registry entry's server-owneddashboard_id(start_run.py,scenario_dashboard_id), pinned bytest_resolution_uses_registry_dashboard_for_identity_less_steps. - Bootstrap materialization dropped per-step target identity. The isolated browser/screenshot
admission requires step-level
environment_id/dashboard_idto equal the adopted binding (fail-closed*_TARGET_MISMATCHotherwise). Fix:ScenarioRegistry.Create.Register(registry/create.py, handle path) stamps the verifieddashboard_contextidentity onto every materialized step — authoritative becausecontext_authority=verifiedis evaluated at the register boundary; steps that already declare identity are never overwritten. - Stale binding principal. The deployment binding
ss-prod-d11-live-001carried anexecution_principal_fingerprintfrom an unknown registration-timeSS_STAND_USER, so the authoritativesha256(actor)guard rejected every start. Fixed operationally by re-registering the binding through the new admin surface (PUT /api/scenario-live-bindings/ss-prod-d11-live-001, principal = sha256("admin")) — which also served as the first live exercise of the D5 binding admin API (register → server-side resolution → run adoption).
Baseline pin
The compiled B01 graph declares no baseline refs (baseline capability deliberately off), so the
gated start legally ran with a null pin. The live baseline-pinned variant of this replay
remains blocked by the Gitea PAT (published catalog unavailable); 050 T045 and the pin part of
T046 stay OPEN for that reason.
Gate row
E2E-EXT-002 (T029m) → CLOSED by this trace. Remaining MCP-surface work is limited to the
PAT-blocked baseline-pinned variant (T045/T046-pin) and Expected.threshold (038 schema authority,
separate architect decision).
Human-checkpoint loop (T029m formulation completeness)
live_mcp_replay.py's B01 graph contains no HumanStep, so the task's
list_checkpoints/decide_checkpoint element was exercised by a companion replay
specs/044-dashboard-scenario-execution/prototype/live_mcp_human_loop.py (same scripted client):
| Stage | Result |
|---|---|
inspect_scenario B05 with live derived capabilities |
one step: phase-5-B05-human_checkpoint-1 (tool=human) — the script refuses any graph with a non-human step, so no mutation case is started |
validate → draft-pack → register_draft_pack |
valid, save_eligible, context_authority=verified |
bootstrap → start_scenario_run (ss-prod, manual) |
pending_approval → approved → run 6dcb9b8b-4f96-4491-9480-a3c74a67ee42 |
| run state | waiting_human |
MCP list_checkpoints(run_id) |
pending checkpoint 35039c17-6469-4cdf-90a1-753a357c0e08, decision_version=1 |
MCP decide_checkpoint(confirm, expected_version=1) |
ok; checkpoint decided, decision_version=2 (CAS consumed) |
| terminal | passed — the immutable 044 mapping confirm → passed names the human's conformance observation, not a synthesized PASS |
This closes the aligned-label human loop live; combined with the B01 chain above, every element of the T029m task text is now backed by a live trace.
Review residuals (recorded, not defects)
- The compile input declared
screenshot=trueas a caller capability;screenshotis not a server-derived dashboard fact, somerge_capabilitiespasses it through verbatim. The derived-wins axis that FR-028 protects (verifiable facts never downgraded tohuman_checkpoint) is intact andcapability_authority.status=derived; the caller-key nuance belongs to the 038/T029k authority-boundary backlog, not to this replay. live_mcp_replay.pynow asserts the honest typed non-pass terminal (apassedterminal on the B01 chain raisesUNEXPECTED_SYNTHETIC_PASS) and asserts the idempotent retry reuses the samerun_id.