Source: live external MCP run against ss-prod Sales Dashboard (docs/2026-09-07-sales-prod-mcp-run.md) proved the initial-bootstrap chain externally unreachable: register_draft_pack requires a principal-owned AgentRun but no MCP operation created one after the chat decommission; the vertical E2E masked the gap with a raw-ORM prerequisite seed.
T029i: MCP create_agent_run/get_agent_run (mcp_server/tools_agent_run.py) over Services.AgentRuns.Service.Create — REST-parity EXECUTE/READ permissions, human-only, server-pinned UIContext, idempotency-key replay; catalog 2.1.0->2.2.0; MCPX-FR-027 external-reachability invariant pinned; initial-scenario E2E converted to the fully external chain (zero non-MCP seeding); strict-xfail pin flipped as designed, unmarked and hardened (E2E-EXT-001 CLOSED).
T029k: ScenarioGraph.CapabilityAuthority — truthful capability facts derived from the authoritative DashboardQueryModel (mutation-context capabilities never derived), derived-wins merge over caller declarations, single choke point wired into MCP inspect_scenario / inspect_dashboard_context and REST api_compile_scenario; CAP-001 classification-fix test on the sales-shape fixture (B02-B04/T01-T03 automated, C04-C06 unsupported, unsafe-mutation cases legitimately human).
T029l: disposition vocabulary clarity — RU/EN labels name the persisted outcome (confirm->passed), confirm restyled bg-destructive->bg-primary, decide_checkpoint description carries the immutable outcome table; lifecycle mapping and API vocabulary unchanged (DISP-001 CLOSED).
Decision memory: ADR-0024 (IMPLEMENTED, 4 rejected alternatives incl. no-AgentRun boundary and implicit auto-create) + README registry; 050 MCPX-FR-027/028/029 + release-gate rows + Clarifications session 2026-09-07; 038/044/045 field-run amendments -> IMPLEMENTED; WORKSTATE checkpoints (plan round + execution round).
Pre-existing HEAD regressions surfaced by the first full-suite rerun since 4d5ef6be/58c5ae39 and fixed: (1) stale SC-007 resource pin — canonical identifier is the post-redirect /mcp/ (code + twin pin aligned since the batches; test_mcp_client_flow_http pin updated with rationale); (2) app-lifespan tests re-entered the run-once StreamableHTTPSessionManager module singleton — autouse fresh-transport-app fixture (production lifespan runs once per process; singleton stays correct there).
Gates: full backend suite 11357 passed / 243 skipped / 1 xpassed / 0 failed (first green full run since the batches); MCP+catalog slice 70 passed; capability slice 67 passed; frontend vitest 3507 passed (206 files), lint 0 errors (364 baseline warnings), build OK; ruff/compileall clean; anchors balanced; scoped git diff --check clean. INV_7 watch: tools_scenario.py 508 LOC and routes scenario.py 442 LOC flagged for the next decomposition pass (new code lives in new modules 155/236 LOC).
OPEN: T029m / E2E-EXT-002 — live-stand replay of the sales scenario through the full external chain. Not included (foreign uncommitted workstream): translate/migration integration tests, _job_routes.py, .kilo/agent-manager.json, specs-036-050-20260907-111314.md.
uvicorn/uvicorn.error adopt the shared CotJsonFormatter handlers (own plain-text handlers replaced, propagate=False); uvicorn.access silenced by default — middleware JSON framing already narrates every non-polling request, and the plain 'INFO: host - GET ...' access lines were the duplicate-console-format source under run.sh. LoggingConfig.uvicorn_access_log=True restores access output as JSON, never plain text.
Unmarked library records map deterministically in CotJsonFormatter: ERROR+ -> EXPLORE with exc_info traceback captured into the error field (<=1000 chars), below -> REASON. Explicit marker always wins.
httpx/httpcore demoted to WARNING via LoggingConfig.http_client_log_level ('HTTP Request: ...' per outbound Superset call is polling-amplified noise; our CoT layer narrates those requests). Git NO_REPO branch: 10 interpolated REASON lines per batch poll -> one invariant-intent DEBUG line with payload{dashboard_id,result} (ADR-0021 intent-invariance).
Tests: uvicorn takeover x3, httpx demotion x2, formatter marker mapping x3 — logger/formatter suites green (42), ruff clean. NOT included (entangled with a parallel session's WIP in the same files, stays in working tree): app.py POST-batch framing suppression + its middleware test, and the Run Center runaway-fetch-loop fix + regression test.
Closure-gate re-review of rounds 1-5 with the Axiom index live (full rebuild 10534 contracts): remediation matrix all CLOSED, SC-001..SC-009 walkthrough carried current executable pins, verdict recorded in specs/WORKSTATE-043-047.md.
Curator pass (comment-only, zero runtime change; every target verified against the live index before retargeting): mcp_server zone 6 edges fixed (Services.AgentAuthoringWorkspace.Service, McpServer.Package→McpServer, phantom SupersetClient.DashboardWrite → three verified Core.DashboardsWrite/Datasets contracts, AgentSuperset.SqlFormat) — scoped audit now 0 unresolved; plus 35 historical retargets (scenario chain → function-level contracts, Superset-client alias/path forms → Core.Init.SupersetClientModule, Models.User→Models.Auth.User ×5, ExecuteEnvelope→ExecuteQueryEnvelope ×6, GitService/Deployment/StructureSnapshot/DashboardTesting.Core, client_registry python-path forms) and 11 malformed multi-target translate-plugin lines split into individual @RELATION lines with verified IDs. Workspace unresolved relations 446→401; remainder classified and queued (logging-SSOT zone owned by the concurrent shared→backend migration, function-shaped targets, legacy single-# regions).
UX audit (read-only, findings + prioritized fix plan persisted in WORKSTATE): MCP happy path for a BI analyst scores 2/5 — CRITICAL: handoff surface gives no endpoint URL/discovery/client onboarding; approval loop (/dashboard-testing/runs WaitingForMeView, HumanCheckpointPanel) unreachable via navigation (absent from ROUTES.ts and sidebar). HIGH: AI/Ассистент buttons promise chat but land on a decommission stub; no MCP settings/status surface anywhere. MEDIUM: dashboard-testing/load-testing routes bypass the ROUTES SSOT.
- .kilo/kilo.jsonc: instructions=["docs/api/nav/root.map"] — the
semantic module digest (doc-gen --nav) is auto-injected into the
starting context of every agent session (~14 KB: areas, per-module
purpose/kw/deps, collapsed tests).
- AGENTS.md: documented the L0->L3 navigation protocol (root.map ->
<Name>.map -> nodes/<Contract>.md -> source via FILE:line), nav_id.map
fallback, and the re-read/regenerate freshness rule (the injected map
is a session-start snapshot).
- .kilo/setup-script: generate docs/api/nav for fresh Agent Manager
worktrees (gitignored artifact; uses $REPO_PATH/../axiom-mcp binary,
cold index build ~60s, graceful skip when the binary is missing).
Requires ../axiom-mcp doc-gen with the semantic root (see axiom-mcp
"feat(docs): semantic module digest root.map" commit); regenerate via
`make docs-nav`.
FR-019 (was a false [x]): MCP list_checkpoints + decide_checkpoint over the
044 CAS lifecycle (server-resolved pending checkpoint, decision_version CAS,
continue_after_human_decision, typed conflict/not_found); catalog scenario:RUN
and service_allowed=False — automation has no path to checkpoints.
tests/test_mcp_checkpoints.py: 3 passed.
FR-013/SC-007: client_credentials grant for machine clients — confidential DCR
(one-time client_secret, sha256-only via migration 0017_oauth_client_secret
with idempotent guard), signed identity-only service-principal tokens
(principal_type=service, aud=mcp, no refresh), McpTokenVerifier service
short-circuit, AS metadata grants/auth-methods, INSTALL.md §MCP-client docs.
tests/test_mcp_client_flow_http.py: full scripted-client flows over real HTTP
(machine: discovery→DCR→client_credentials→/mcp initialize→tools/list→call;
user: DCR→PKCE S256 authorize via Bearer web session→exchange→/mcp live-RBAC
listing). 2 passed; oauth suite 10 passed.
PRODUCTION DEFECTS fixed en route: call_tool argument-inspection gates parsed
only the FLAT shape while FastMCP delivers {"request": {...}} —
start_scenario_run was uncallable over MCP and the PROD-SQL terminal denial was
bypassable by the wrapped shape. Both gates now unwrap via _gate_arguments;
regression pinned by flat x wrapped PROD matrix (6 combos).
E2E-AUTH-001..003 release-gate rows closed with evidence: T028 chain gained
propose_test_plan (single-trace 001); T023 extended with post-activation
start_scenario_run on the canonical runner-shaped fixture graph, asserting the
queued run pins promoted revision_id + content_hash (003); 002 already proven.
Record honesty: tasks.md T025/T008/T008b/T032 downgraded at review time;
T025+T008 re-closed with executable evidence; T032/T008b/T005a remain open in
the round-2 queue (recorded in WORKSTATE checkpoint with the full P1/P2 list:
MCL intent repair, dispatch/sandbox EXPLORE traces, INV_6 tombstones,
SC-005 remnants, FR-010 versioning, depth/rate limits).
Evidence: combined MCP slice 95 passed; alembic single head 0017; ruff and
compileall clean; anchors balanced in all touched files.
Stop preview/env fallbacks when a configured datasource is gone, skip the
duplicate final insert after streaming, and surface retry/scheduler/LLM
edge cases as FAILED instead of COMPLETED.
The send-path repair branch set _request_result but not
_request_error_code, so the AGENT_REQUEST_FAILED lifecycle event fell back
to the misleading 'agent lifecycle failure'. Set THREAD_REPAIRED so logs
and middleware show the real outcome.
The repair paths (send-path _repair_pending_tool_calls and resume fallback)
called graph.update_state — the SYNC method — which internally invokes
AsyncPostgresSaver.get_tuple() and raises InvalidStateError ('Synchronous
calls to AsyncPostgresSaver are only allowed from a different thread').
The repair therefore always failed (surfacing as 'Event loop is closed' +
a dangling aget_tuple coroutine) and broken threads stayed broken.
Switch both repair sites to agent.aupdate_state (async checkpointer
interface) and align the mocked agent in tests. Verified live: aupdate_state
repaired the broken checkpoint of conversation 69651ca1 (pending calls → 0).
Two recurring failures from live logs (conversations 69651ca1 / a8c0dff8):
1. A checkpoint whose AI messages carry tool_calls without ToolMessages
(run crashed after the LLM emitted a call) makes every send raise
INVALID_CHAT_HISTORY with no recovery. The send path now repairs the
thread via _repair_pending_tool_calls: pending calls are answered with
synthetic error ToolMessages (THREAD_REPAIRED) so the user can retry.
2. Scenario tools scheduled from plain chat (no build_dashboard_test_scenario
UI intent) had no durable AgentRun, so the resume fallback refused with
SCENARIO_RUN_REQUIRED. _ensure_scenario_run now lazily creates the run
from the tool args (dashboard_context from scenario_json for
validate/resolve), mirroring the UIContextV2 scenario contract.
Observed in a live checkpoint: the pending scenario_validate tool call
carried a scenario_json that the LLM stream cut mid-document (ended
inside the unclosed outer object, len 3837). _parse_json_value failed
with a generic 'Could not extract JSON value' that neither the operator
nor the LLM could act on. Add _looks_truncated (unbalanced structure /
unterminated string at end) and report 'truncated/incomplete JSON' with
input length, so the model regenerates the full document.
The backend route binds the search query to `search`; a bare `q` param is
not bound, so search_dashboards and prefetch_dashboards silently returned
the unfiltered catalog (e.g. query 'Sales' listed all 11 dashboards).
Switch both to `search` and update the URL-contract test.
- search_dashboards: call /api/dashboards with page_context=other and
page_size=100 so the profile 'My Dashboards Only' filter can no longer
hide the whole catalog; parse available_total/effective_profile_filter
and report hidden-by-filter instead of a false 'no dashboards' answer
- prefetch_dashboards: same full-catalog context; fix dead code where
data=resp.json() sat after return '' inside the error branch, making
every 200 response raise NameError and the prefetch always return ''
- llm-status: ?force=1 bypasses the 30s health cache so the 'Retry now'
button performs a fresh probe instead of re-reading the stale status;
frontend keeps a single retry interval (previously stacked intervals
decayed the countdown faster than 1/s and fired duplicate probes)
- tests: agent tool/prefetch, backend route bypass + force param,
frontend retry/force coverage
- Replace the global env <select> in TopNavbar with an expandable
EnvironmentStatsWidget showing per-env total/mine/published/drafts
and health status (latency, unreachable), preserving env switching.
- Add GET /api/environments/stats: per-env counts (profile-actor matched)
+ lightweight health probe, gathered concurrently with an 8s probe
timeout and a process-local TTL cache (30s, coalescing) so the full
Superset dashboard catalog is not re-fetched on every dropdown open.
- Add available_total to GET /api/dashboards so grids can show how many
dashboards exist when the profile-default filter hides everything.
- Share ProfileFilterBanner across the dashboards hub and validation
task form: 'showing X of Y' reference info + explicit Show all /
Restore filter actions.
- Russian plural forms for dashboard counts (pluralRu helper) and
compact 'Опубл.' label; i18n keys en/ru.
- Ignore :memory:test_* SQLite test artifacts and drop them from the index.
- Tests: env stats endpoint (incl. caching), widget, model fallback,
plural helper, api client, integration.
- handleSearchInput now receives the selected envId, so the debounced
search actually fires API requests instead of hitting the !envId guard
and clearing results immediately.
- The dashboard section of the global search sends page_context=other,
apply_profile_default=false, override_show_all=true (same pattern as
DashboardHubModel.loadDashboardSearchOptions), so the "show only my
dashboards" profile filter no longer zeroes out dashboards that lack
owner metadata.
- Updated unit tests: debounce now asserts API calls + profile-off flags.