Commit Graph

9 Commits

Author SHA1 Message Date
9c1a1e093c feat(mcp): Phase 2d field-run remediation — ADR-0024 agent-run surface, derived capabilities, disposition clarity
Source: live external MCP run against ss-prod Sales Dashboard (docs/2026-09-07-sales-prod-mcp-run.md) proved the initial-bootstrap chain externally unreachable: register_draft_pack requires a principal-owned AgentRun but no MCP operation created one after the chat decommission; the vertical E2E masked the gap with a raw-ORM prerequisite seed.

T029i: MCP create_agent_run/get_agent_run (mcp_server/tools_agent_run.py) over Services.AgentRuns.Service.Create — REST-parity EXECUTE/READ permissions, human-only, server-pinned UIContext, idempotency-key replay; catalog 2.1.0->2.2.0; MCPX-FR-027 external-reachability invariant pinned; initial-scenario E2E converted to the fully external chain (zero non-MCP seeding); strict-xfail pin flipped as designed, unmarked and hardened (E2E-EXT-001 CLOSED).

T029k: ScenarioGraph.CapabilityAuthority — truthful capability facts derived from the authoritative DashboardQueryModel (mutation-context capabilities never derived), derived-wins merge over caller declarations, single choke point wired into MCP inspect_scenario / inspect_dashboard_context and REST api_compile_scenario; CAP-001 classification-fix test on the sales-shape fixture (B02-B04/T01-T03 automated, C04-C06 unsupported, unsafe-mutation cases legitimately human).

T029l: disposition vocabulary clarity — RU/EN labels name the persisted outcome (confirm->passed), confirm restyled bg-destructive->bg-primary, decide_checkpoint description carries the immutable outcome table; lifecycle mapping and API vocabulary unchanged (DISP-001 CLOSED).

Decision memory: ADR-0024 (IMPLEMENTED, 4 rejected alternatives incl. no-AgentRun boundary and implicit auto-create) + README registry; 050 MCPX-FR-027/028/029 + release-gate rows + Clarifications session 2026-09-07; 038/044/045 field-run amendments -> IMPLEMENTED; WORKSTATE checkpoints (plan round + execution round).

Pre-existing HEAD regressions surfaced by the first full-suite rerun since 4d5ef6be/58c5ae39 and fixed: (1) stale SC-007 resource pin — canonical identifier is the post-redirect /mcp/ (code + twin pin aligned since the batches; test_mcp_client_flow_http pin updated with rationale); (2) app-lifespan tests re-entered the run-once StreamableHTTPSessionManager module singleton — autouse fresh-transport-app fixture (production lifespan runs once per process; singleton stays correct there).

Gates: full backend suite 11357 passed / 243 skipped / 1 xpassed / 0 failed (first green full run since the batches); MCP+catalog slice 70 passed; capability slice 67 passed; frontend vitest 3507 passed (206 files), lint 0 errors (364 baseline warnings), build OK; ruff/compileall clean; anchors balanced; scoped git diff --check clean. INV_7 watch: tools_scenario.py 508 LOC and routes scenario.py 442 LOC flagged for the next decomposition pass (new code lives in new modules 155/236 LOC).

OPEN: T029m / E2E-EXT-002 — live-stand replay of the sales scenario through the full external chain. Not included (foreign uncommitted workstream): translate/migration integration tests, _job_routes.py, .kilo/agent-manager.json, specs-036-050-20260907-111314.md.
2026-09-07 16:52:38 +03:00
58c5ae39cb feat(scenario): server-owned handle pipeline T029d–h + MCP bootstrap/automation + UI launch/approval
Backend (Phase 2c, closes ADR-0023):
- T029b/c: 9 automation MCP tools (REST-parity RBAC, scheduler registration, idempotency), migration 0018; fresh-DB MCP E2E.
- Guard BOOTSTRAP_REVISION_NOT_RUNNABLE in derive_runner_plan (provenance-only revisions never queue a vacuous zero-step PASS); demoted to defense-in-depth after materialization landed.
- T029d: CompiledScenarioHandle/ValidationResultHandle/DraftPackHandle (immutable, owner-bound, content-addressed canonical-bytes store), migrations 0019/0021; minting at REST compile/validate/resolve/draft-pack boundaries; single consumption under SELECT...FOR UPDATE + populate_existing, proven on PostgreSQL (Testcontainers).
- T029e: handle-first create_scenario/create_initial materialize canonical graph_snapshot (+ server-owned action_registry identity) in-transaction; OutboxEvent + RevisionMaterialization (0020) with idempotent worker wired into the scheduler poll loop (30s tick).
- T029f: MCP register_draft_pack write tool; bootstrap accepts only stored handle ids (transitional compile:{run}:{digest} removed); legacy REST POST /scenarios/{id}/revisions retired -> 410; catalog 2.0.0 (pinned-major ritual); ActionRegistry 038.2.0 — generate_report reclassified non-mutating (local draft write), register_artifact/row_edit/bulk_edit stay mutating.
- T029h (hybrid C+X1): inspect_dashboard_context MCP tool (live DashboardQueryModel resolver); context_authority evaluation — server recomputes client-context fingerprint (claimed values ignored), sentinel fingerprints never verify, unreachable env fails open to unverified, live-env mismatches reject typed with zero rows; marker persists on DraftPackHandle, materializes into graph_snapshot; PROD start refuses explicit non-verified (CONTEXT_AUTHORITY_REQUIRED_FOR_PROD); validator recursively rejects query_context/SQL smuggling in dashboard_context; catalog 2.1.0. 043 editor path provably inherits the marker.

Frontend:
- D2: scenario detail route scenarios/[id] — first production host of RunConfigurationPanel (typed 044 launch + Idempotency-Key + redirect to run monitor); ROUTES.scenarioDetail SSOT; registry index links to detail (launch stays off the index).
- D4: ApprovalDecisionPanel + RunMonitorModel.decideApproval — PROD gates decidable from the web UI (run-detail aside on pending_approval).
- Pre-existing suite repairs: oauth-consent raw goto -> ROUTES.login(); InvestigationModels stale dispose signature; settings mcp_oauth_* bind:value undefined crash (backend defaults 15/30/90).
- i18n: approval_* + scenario_detail_* keys in en+ru.

Specs/docs (amendments 2026-09-06/07):
- 038: PackCompiler.Generate contract corrected (pure manifest); ServerOwnedPipeline handle-persistence amendment; verification-program reconciliation (implemented vs PROPOSED IR entities).
- 042: synchronous graph materialization + outbox scope; CreateInitial amendments; RevisionChain @REJECTED for raw client graph_snapshot.
- 050: stage-table implementation-status note, handle-rules status, T029–T029h evidence; WORKSTATE checkpoints; ADR-0023 -> IMPLEMENTED.

Gates: backend full regression 3360 passed (+ PostgreSQL integration green), alembic single head 0021; frontend vitest 3506 passed, lint 0 errors, build OK.
2026-09-07 11:26:34 +03:00
65121cac6b feat(logging): self-diagnosing EXPLORE + shared/ absorption + belief analytics (ADR-0021/0022)
T0: absorb shared/ into backend — cot_logger→src/core, CotJsonFormatter→src/core/cot_formatter.py, _llm_http/_llm_health/ssl→src/core/utils; imports rewritten (26 prod + tests, patch targets); run.sh/backend.Dockerfile/requirements/.axiom source_dirs/semantic_health/AGENTS/INSTALL cleaned; ADR-0022 supersedes ADR-0015; fixed latent CI defects (ss_tools ImportError, record.message in logger tests, same-name test-module collision).

ADR-0021 wire enrichment (additive): contract_id/claim/error_code/loc fields; _contract_id ContextVar + resolve_contract_id (explicit > belief_scope > declared-src mirror, derived src never mirrors); EXPLORE auto-loc via single frame walk; facade error auto-fill; 2KB payload cap with payload_truncated/payload_bytes markers; migrated 85 error="CODE" sites to error_code= (12 files); pilot editor/load.py; superset preview payload-bomb inlined bodies removed.

Analytics SSOT src/core/log_stats.py (bond transition matrix, orphan-EXPLORE ratio, REFLECT pairing, intent families, coverage, insufficient-sample flag); pretty_cot.py --stats/--digest/--trajectory/--story over one engine; log_gap_service three-tier ground-truth triangulation (FAILED w/o EXPLORE etc.) + GET /api/reports/log-stats|task-log-gaps (polling-suppressed); scripts/cot_audit.py CLI; enriched fields persisted into task_logs.payload for tier queries.

Frontend: ReportsAnalyticsModel + AnalyticsStatsPanel (Logs tab) + TaskGapPanel and per-row T1/T2/T3 gap badges (Tasks tab); cot-logger.ts ADR-0021 opts; i18n en/ru. Scheduler console spam fixed: apscheduler logger demoted to WARNING via LoggingConfig.scheduler_log_level. .axiom belief patterns -> $OBJ.* (alias undercount). molecular-cot-logging skill updated (fields, decision rules, tie-break, CLI) and synced.

Reviewed orthogonally: F1 cot_span contract pollution, F2 cap boundary accounting, F3 digest over-dedup, F4 trace-state bound, F5 tier metadata — fixed with regression tests. Validation: backend 11287 passed + ruff + compileall; frontend 3446 passed + lint + build; CLI smoke on live app.log.
2026-09-04 20:56:41 +03:00
02a97bfc9c fix(maintenance): survive sqlparse token cap on huge virtual dataset SQL
Discovery of virtual datasets now works, but a runtime blocker remained: any
virtual dataset whose SQL exceeds sqlparse's MAX_GROUPING_TOKENS (10000 tokens)
raised SQLParseError 'Maximum number of tokens exceeded (10000)' from
extract_tables_from_sql_span, which is called unguarded in the scan loop — one
oversized virtual dataset aborted the whole maintenance preview/start.

- extract_tables_from_sql_span now wraps sqlparse.parse + token walk in
  try/except and falls back to regex-only extraction (keeping all schema.table
  matches) instead of raising, so huge SQL no longer fails the scan.
- Tier-1 virtual filter uses value "" (not None) so the sql is_not_null filter
  passes Superset's rison schema instead of always falling back to a full scan.

Tests: huge-SQL fallback (extractor) and huge-virtual-dataset scan resilience
(scanner). ADR-0020 updated with Decision 3.
2026-08-03 23:48:01 +07:00
52e909a2da fix(maintenance): discover virtual (SQL) datasets in dashboard scanner
Virtual (SQL) datasets were never matched, so maintenance discovery returned
0 affected dashboards. Two defects fixed:

- find_affected_dashboards filtered by is_sqllab_view, which is NOT a
  filterable column in Superset's dataset list API (absent from search_columns),
  so the query was rejected. Now discover virtual datasets via the filterable
  sql column: primary server-side 'sql is_not_null' filter with a client-side
  non-empty-sql scan as fallback (best-effort vs pagination cap), dedupe by id.

- AsyncAPIClient.request never called raise_for_status(), so rejected filters
  (HTTP 400) were returned as bodies without a 'result' key and surfaced as
  'Found 0 datasets', dead-coding the filtered->full-scan fallback. request()
  now raises on non-2xx via the existing error mapper.

Tests cover both virtual-scan tiers, all fallback paths, the raise behavior,
and an end-to-end match with the real sql_table_extractor on production SQL.
Documented in ADR-0020.
2026-08-03 22:21:00 +07:00
fdb6541372 docs(adr): ADR-0019 — механизм импорта дашбордов Superset
Зафиксированы архитектурные решения миграции:
- UUID-трансформация БД через _transform_database_yaml() (вместо strip_databases)
- Cross-filter patching через IdMappingService + mapping_service в dry-run
- Password injection flow: await_input → wait_for_input → retry import
- Разделение dry-run (read-only) и execute (запись)
- Парсинг имён YAML-файлов БД с точками в имени

Задокументированы исправления production-багов 2026-07-16:
- add_log_callback в await_input (менеджер управляет сам)
- mapping_service=None в dry-run (лишал cross-filter patching)
- strip_databases=True → каскадный сбой 1010
2026-07-17 19:08:28 +03:00
c3ad0afc17 refactor: remove rejected dataset review feature 2026-07-14 15:56:31 +03:00
a39a76c87f feat(agent-centric-logging): consolidate CoT infra in shared, close REASON→REFLECT chains
- shared/cot_logger.py is SSOT; backend/cot_logger.py deleted
- elapsed_ms timing in all REFLECT markers
- Frontend: REASON→REFLECT/EXPLORE in all fetch/post/delete/requestApi
- Dynamic src: route.GET.api.plugins instead of hardcoded api.request_handler
- trace_id generated immediately (no 'no-trace'), X-Trace-ID in both directions
- Global error handlers (window error + unhandledrejection + error.svelte)
- Fixed duplicate logging (shared/logger.py double StreamHandler)
- propagate=False in configure_logger (was in ConfigManager = duplicated startup logs)
- belief_scope: 'Coherence OK' → '{anchor}: completed' + elapsed_ms
- Fixed 28 pre-existing test failures (scheduler sig, DB columns, DRAFT validation, etc)
2026-07-12 19:30:57 +03:00
24d3b7d1f9 refactor(task-manager): implement task resilience and execution lifecycle improvements
Enhance the reliability and observability of the task execution engine
by introducing retry mechanisms, idempotency, and structured progress
tracking.

- Implement centralized retry logic with exponential backoff support
  in `JobLifecycle`.
- Add `retry_task` API endpoint and `TaskManager` method for manual
  task restarts.
- Introduce task idempotency using `_idempotency_key` to prevent
  duplicate executions.
- Add `retry_count`, `max_retries`, `last_error`, and `progress` fields
  to the `Task` model and ensure persistence via `TaskPersistenceService`.
- Upgrade `SchedulerService` to use differential synchronization with
  the persistent `SQLAlchemyJobStore` for better job durability.
- Implement structured heartbeat logging to support real-time progress
  updates.
- Update project documentation and ADRs to reflect the new plugin
  runtime and task resilience patterns.
- Add comprehensive unit and integration tests for the new task
  lifecycle features.
2026-07-12 15:31:56 +03:00