Source: live external MCP run against ss-prod Sales Dashboard (docs/2026-09-07-sales-prod-mcp-run.md) proved the initial-bootstrap chain externally unreachable: register_draft_pack requires a principal-owned AgentRun but no MCP operation created one after the chat decommission; the vertical E2E masked the gap with a raw-ORM prerequisite seed.
T029i: MCP create_agent_run/get_agent_run (mcp_server/tools_agent_run.py) over Services.AgentRuns.Service.Create — REST-parity EXECUTE/READ permissions, human-only, server-pinned UIContext, idempotency-key replay; catalog 2.1.0->2.2.0; MCPX-FR-027 external-reachability invariant pinned; initial-scenario E2E converted to the fully external chain (zero non-MCP seeding); strict-xfail pin flipped as designed, unmarked and hardened (E2E-EXT-001 CLOSED).
T029k: ScenarioGraph.CapabilityAuthority — truthful capability facts derived from the authoritative DashboardQueryModel (mutation-context capabilities never derived), derived-wins merge over caller declarations, single choke point wired into MCP inspect_scenario / inspect_dashboard_context and REST api_compile_scenario; CAP-001 classification-fix test on the sales-shape fixture (B02-B04/T01-T03 automated, C04-C06 unsupported, unsafe-mutation cases legitimately human).
T029l: disposition vocabulary clarity — RU/EN labels name the persisted outcome (confirm->passed), confirm restyled bg-destructive->bg-primary, decide_checkpoint description carries the immutable outcome table; lifecycle mapping and API vocabulary unchanged (DISP-001 CLOSED).
Decision memory: ADR-0024 (IMPLEMENTED, 4 rejected alternatives incl. no-AgentRun boundary and implicit auto-create) + README registry; 050 MCPX-FR-027/028/029 + release-gate rows + Clarifications session 2026-09-07; 038/044/045 field-run amendments -> IMPLEMENTED; WORKSTATE checkpoints (plan round + execution round).
Pre-existing HEAD regressions surfaced by the first full-suite rerun since 4d5ef6be/58c5ae39 and fixed: (1) stale SC-007 resource pin — canonical identifier is the post-redirect /mcp/ (code + twin pin aligned since the batches; test_mcp_client_flow_http pin updated with rationale); (2) app-lifespan tests re-entered the run-once StreamableHTTPSessionManager module singleton — autouse fresh-transport-app fixture (production lifespan runs once per process; singleton stays correct there).
Gates: full backend suite 11357 passed / 243 skipped / 1 xpassed / 0 failed (first green full run since the batches); MCP+catalog slice 70 passed; capability slice 67 passed; frontend vitest 3507 passed (206 files), lint 0 errors (364 baseline warnings), build OK; ruff/compileall clean; anchors balanced; scoped git diff --check clean. INV_7 watch: tools_scenario.py 508 LOC and routes scenario.py 442 LOC flagged for the next decomposition pass (new code lives in new modules 155/236 LOC).
OPEN: T029m / E2E-EXT-002 — live-stand replay of the sales scenario through the full external chain. Not included (foreign uncommitted workstream): translate/migration integration tests, _job_routes.py, .kilo/agent-manager.json, specs-036-050-20260907-111314.md.
Discovery of virtual datasets now works, but a runtime blocker remained: any
virtual dataset whose SQL exceeds sqlparse's MAX_GROUPING_TOKENS (10000 tokens)
raised SQLParseError 'Maximum number of tokens exceeded (10000)' from
extract_tables_from_sql_span, which is called unguarded in the scan loop — one
oversized virtual dataset aborted the whole maintenance preview/start.
- extract_tables_from_sql_span now wraps sqlparse.parse + token walk in
try/except and falls back to regex-only extraction (keeping all schema.table
matches) instead of raising, so huge SQL no longer fails the scan.
- Tier-1 virtual filter uses value "" (not None) so the sql is_not_null filter
passes Superset's rison schema instead of always falling back to a full scan.
Tests: huge-SQL fallback (extractor) and huge-virtual-dataset scan resilience
(scanner). ADR-0020 updated with Decision 3.
Virtual (SQL) datasets were never matched, so maintenance discovery returned
0 affected dashboards. Two defects fixed:
- find_affected_dashboards filtered by is_sqllab_view, which is NOT a
filterable column in Superset's dataset list API (absent from search_columns),
so the query was rejected. Now discover virtual datasets via the filterable
sql column: primary server-side 'sql is_not_null' filter with a client-side
non-empty-sql scan as fallback (best-effort vs pagination cap), dedupe by id.
- AsyncAPIClient.request never called raise_for_status(), so rejected filters
(HTTP 400) were returned as bodies without a 'result' key and surfaced as
'Found 0 datasets', dead-coding the filtered->full-scan fallback. request()
now raises on non-2xx via the existing error mapper.
Tests cover both virtual-scan tiers, all fallback paths, the raise behavior,
and an end-to-end match with the real sql_table_extractor on production SQL.
Documented in ADR-0020.
Enhance the reliability and observability of the task execution engine
by introducing retry mechanisms, idempotency, and structured progress
tracking.
- Implement centralized retry logic with exponential backoff support
in `JobLifecycle`.
- Add `retry_task` API endpoint and `TaskManager` method for manual
task restarts.
- Introduce task idempotency using `_idempotency_key` to prevent
duplicate executions.
- Add `retry_count`, `max_retries`, `last_error`, and `progress` fields
to the `Task` model and ensure persistence via `TaskPersistenceService`.
- Upgrade `SchedulerService` to use differential synchronization with
the persistent `SQLAlchemyJobStore` for better job durability.
- Implement structured heartbeat logging to support real-time progress
updates.
- Update project documentation and ADRs to reflect the new plugin
runtime and task resilience patterns.
- Add comprehensive unit and integration tests for the new task
lifecycle features.