Commit Graph

1138 Commits

Author SHA1 Message Date
093f7f600f fix(maintenance): harden banner lifecycle and guarded migrations
- Guard maintenance Alembic operations for create_all-only tables on clean DBs
- Add guarded verification_runs.fanout_plan_id backfill migration
- Improve maintenance banner rendering, chart management, orchestration, and API routes
- Expand assistant maintenance tool and edge-case coverage

Tests: cd backend && source .venv/bin/activate && python -m pytest -q tests/test_maintenance_api.py tests/test_maintenance_service.py tests/api/test_assistant_tool_maintenance.py tests/api/test_maintenance_routes_edge.py (77 passed)
2026-08-09 08:22:16 +03:00
ffa4456a62 docs(specs): add machine contract reconciliation gate
- Validate and repair OpenAPI YAML contracts for 038, 043, and 046
- Canonicalize 038 JSON schema and fixtures around scenario_key,
  content_hash, and logical_step_id; validate all fixtures with jsonschema
- Regenerate 038 validation evidence for compiler scope only
- Add reconcile_contracts.py for repeatable OpenAPI/JSON/fixture checks
- Replace raw CreateScenario payload with server-owned handles and document
  transactional outbox/materialization saga for Registry-to-git persistence
- Record reconciliation outcome in REVIEW-042-047-CLOSURE.md
2026-08-09 08:19:02 +03:00
f7e539440e feat(tooling): rewrite merge_spec.py — batch spec merging + new package support
Rewrite merge_spec.py to merge one or many feature spec packages into a
single review file.

Batch modes:
- single number: python merge_spec.py 038
- inclusive range: python merge_spec.py 036-041
- explicit list: python merge_spec.py 036 038 044
- by dir name: python merge_spec.py 042-dashboard-scenario-registry
- all: python merge_spec.py all
- custom output: python merge_spec.py 036-041 -o out.md

Handles the new spec package structure that plain *.md merging missed:
- includes contracts/openapi.yaml (YAML), contracts/ux/* (decisions.md),
  prototype/index.html + prototype/manifest.md
- skips .json/.py/.zip/.pyc and __pycache__ (fixtures/code/binaries)
- canonical per-feature order: spec -> ux_reference -> checklists ->
  UX contracts -> plan -> research -> data-model -> modules -> openapi ->
  quickstart -> traceability -> tasks -> prototype
- missing numbers warn+skip; dedup; per-feature grouping in one output

Verified: 043 (14 files), 036-041 (6 features/104 files), 036-047 (12/186),
all (50/572) with no .json/.zip/.pyc leakage.
2026-08-09 11:06:36 +07:00
4858992e15 docs(specs): cross-spec canonicalization pass — reconcile 038 core with 042-047
Reconcile the stale 038 compiler-layer model with the 042-047 lifecycle and
its normative documents (not just data-model).

038 -> clean IR/compiler layer:
- identity: scenario_id slug -> scenario_key (semantic); scenario_id (UUID)
  and revision_id (UUID) assigned by 042 at Save; compiler emits content_hash
- ScenarioStep: add logical_step_id (immutable UUID) + step_key/position/
  step_content_hash; runtime VlmFinding/HumanDisposition moved to 044
- VlmAnalysisSpec/ScreenshotCaptureSpec stay (WHAT); runtime capture/VLM/
  disposition endpoints marked deprecated -> 410 MOVED_TO_044
- CompileRequest: agent_run_id no longer required (optional provenance,
  source_type: agent_run|editor|migration|api)
- runtime evidence = Artifact(owner_type=scenario_run), never authoring DraftPack
- validation.md PASS nullified (self-contradictory COMPLETE vs OPEN);
  refocused as compiler-layer PASS only; T057-T059 moved to 044; T046 rewritten

042/043/044/045 normative (spec/research/checklists/ux/prototype/tasks):
- replace revision_hash/parent_revision_hash/scenario_revision_hash with
  revision_id/content_hash/parent_revision_id everywhere
- 044: HumanCheckpoint (confirm/false_positive/inconclusive) distinct from
  ActionApprovalGate; runner pins revision_id+content_hash
- 047: triage split (investigation_status/classification/resolution), false_positive
  vocabulary; /scenarios/{id}/health|trends|recurring-failures

Update REVIEW-042-047-CLOSURE.md with canonicalization pass status.
2026-08-09 10:55:53 +07:00
9889e09d87 docs(specs): renumber 042-043 to 048-049, add scenario lifecycle specs 042-047
- Renumber: 042-rls-management-workspace -> 048, 043-idm -> 049
  (internal refs updated; RLS research '043 Explainability' corrected)
- Add 042 Scenario Registry & Lifecycle: persistence, list/detail,
  immutable revisions (revision_id/content_hash), CreateScenario
  (Save->Register), lifecycle state machine, staleness via 037/041, health
- Add 043 Scenario Editor UX: hybrid edit model C, WorkingDraft save
  (no arbitrary-draft bypass), SetParameter/AddStep/RemoveStep ops,
  constrained assertions, visual DAG, agent edit, Revalidate migration
- Add 044 Scenario Execution Engine: ScenarioRun/StepRun, deterministic
  runner, RunnerPlan derived from revision (not stored source of truth),
  ActionApprovalGate vs HumanCheckpoint, generic artifact owner, worker
  lease/idempotency, logical_step_id, retry closure, result aggregation
- Add 045 Run Monitor & Results UX: config, live SSE monitor, human
  actions, result+provenance, history/compare, Global Run Operations Center
- Add 046 Automation & Operations: schedules/triggers/API trigger, CRUD,
  scheduler semantics, notification events, layered retention tiers, UI
- Add 047 Triage & Analytics: strict flakiness, immutable fingerprint,
  triage split, /scenarios/{id}/health|trends|recurring-failures
- Add specs/REVIEW-042-047-CLOSURE.md mapping all review gaps to fixes
- Update PRODUCT_ROADMAP for 042-049

Each spec: spec/data-model/research/plan/tasks/ux/traceability/quickstart/
checklists + contracts/modules + openapi + interactive prototype.
2026-08-07 18:30:32 +07:00
7487887e61 docs(examples): add maintenance API spec and Russian usage instructions to example scripts 2026-08-07 17:10:25 +07:00
869997554e fix(frontend): use /content endpoint for draft download 2026-08-07 16:53:15 +07:00
b367c3e4e6 fix(038): map LLM-invented selected_case_ids to registered catalog ids
Scenario compile returned 422 VALIDATION_ERROR because the LLM sent
human-readable case names (smoke, data_integrity, filter_propagation) as
selected_case_ids, but the compiler requires registered catalog ids
(B01-B09, C01-C07, T01-T03) and raised KeyError on unknown ones.

- tools_038._compile_objective: resolve case ids through _resolve_case_ids,
  which (1) passes through registered ids case-insensitively, (2) maps
  human-readable synonyms to closest catalog cases, (3) drops unresolvable
  tokens — a free-form name can never reach the compiler as a KeyError.
- CompileScenarioInput.objective_json description now enumerates the valid
  catalog id ranges and gives an example so the LLM stops inventing names.
- Tests: human-readable mapping, unknown-id dropping, dedupe/first-registered
  order (test_tools_038_parse.py, 21 passed).

Verification: ruff clean, 21 tests pass.
2026-08-07 16:39:21 +07:00
066dfe3a35 fix(alembic): merge two parallel heads from o1p2q3r4s5t6
Runtime migrations failed with 'Multiple head revisions are present' because
037 T081 (p2q3r4s5t6u7 -> verification_runs.dashboard_id) and a concurrent
session-activity change (a1b2c3d4e5f7) both branched from o1p2q3r4s5t6.
The failed 'upgrade head' left dashboard_id unapplied, causing
'column verification_runs.dashboard_id does not exist' on
GET /verification/history.

Add a no-op merge revision (015281bd7759) collapsing both into a single head
so 'upgrade head' applies the verification_runs.dashboard_id column.

Verified: ScriptDirectory.get_heads() == ['015281bd7759'].
2026-08-07 16:33:30 +07:00
0fc03d0e8f docs(041): confirm T047 fully closed (fleet-report UI + tests) 2026-08-07 16:30:14 +07:00
38e5f31eca feat(041): fleet-report panel + close H2 (T047 UI + tests)
Code-review H2: markDeprecated/recordMigration/loadFleetReport had no tests and
loadFleetReport was dead code (not bound to UI). Close it.

- api/lineage.ts: add getFanoutReport(planId) -> GET /lineage/fanout/{plan}/report
- types/lineage.ts: add optional run_status to FleetReportDTO (backend returns it)
- Datasets.LineageModel.loadFleetReport: use lineageApi.getFanoutReport (typed)
- LineagePanel: new 'Fan-out report' section (plan id input + load button +
  per-dashboard status/unresolved-impact rendering), bound to model.loadFleetReport
- i18n: en/ru lineage.json add fleet_report_title, load_fleet_report, plan

Verification: Datasets.LineageModel.test.ts 6 passed (3 new: markDeprecated,
recordMigration, loadFleetReport incl. error path), lineage_panel 4 passed,
vite build OK, eslint clean.
2026-08-07 16:29:47 +07:00
c9664dfabc fix(040/038): address code-review criticals (C1-C4, H1, M1, M3)
QA review of the 036-041 closure range returned FAIL with 3 criticals, all
confirmed. Fixes:

C1 - breaker dead: on_result=persist_batch is now wired into RunnerPool
  (breaker.record() fed per result); added test_breaker_abort_persists_partials
  proving CIRCUIT_BREAKER_ABORT reachability + partial persistence.
C2 - index-based result mapping corrupted data under concurrency: results now
  map by execution_id to their source item; test uses two distinct payloads
  and asserts chart->digest pairing (previously masked by identical fixtures).
C3 - double-acquire of the shared client semaphore (deadlock invariant):
  RunnerPool no longer manually acquires the client semaphore; capacity is
  enforced by worker count, the client bounds total concurrency.
C4 - duplicated ScenarioGraph.Vlm.Analyze region: outer region renamed
  ScenarioGraph.Vlm [TYPE Module].
H1 - _default_submit stub removed: analyze_screenshot requires submit=; no
  silent empty-findings fallback.
M1 - test_capture_dispatch.py region closed.
M3 - capture.py raw_sha256 bypass removed: digest always derived from real
  capture_bytes (no caller-supplied hash).

Verification: load_testing (77) + scenario (103) = 180 passed; ruff clean;
all region pairs balanced.
2026-08-07 16:18:29 +07:00
c162f6ee3a chore(036-041): final validation reconciliation + 039 dashboard verification binding
- 038/039/040/041 validation.md: update PASS status to reflect completed
  runtime closure (T057-T059, T054-T057, T075-T079, T045-T048); regenerate
  039/040/041 digest tables; 039 T058 documented as the sole open task
- frontend/src/routes/dashboards/[id]/+page.svelte: include the 039 T057
  VerificationHistoryList binding (was created in the 039 commit but the
  page-level wiring was left unstaged)

All closure tasks across 036-041 are now complete except 039 T058 (blocked:
no repository_id in dashboard metadata; no PREPROD deployment page).
2026-08-07 15:31:39 +07:00
37bfe12a93 docs(041): mark T045-T048 closed, add Runtime Closure Status
Record 041 frontend + opt-in closure in tasks.md/spec.md/quickstart:
T045-T048 done (LineagePanel binding, deprecation surface, fleet-report,
opt-in rationale). 041 now reflects resolved state.
2026-08-07 15:28:15 +07:00
0a445fb170 feat(041): lineage frontend blast-radius + deprecation + opt-in rationale (T045-T048)
Close the 041 frontend/opt-in gaps found in the audit: no /lineage UI existed
and lineage_index stayed opt-in without documented rationale.

T045 - bind the existing Datasets.LineagePanel (T031) onto /datasets/[id]
  (blast-radius dependents + stale_index notice); deleted my transient
  duplicate to respect component reuse.
T046 - DatasetsLineageModel.markDeprecated()/recordMigration() + LineagePanel
  deprecation lifecycle section (grace window, successor, migration uuid);
  consumes existing api/lineage.ts + lineage.json i18n keys.
T047 - DatasetsLineageModel.loadFleetReport() for fan-out fleet report.
T048 - config_models.py: lineage_index_enabled default stays FALSE with an
  explicit rationale (post-sync Superset detail-call cost; flip after live
  indexer stability proof); consumers treat disabled index as empty read-model.

Verification: lineage + api vitest = 236 passed; vite build OK; eslint clean
for changed code (pre-existing ruff/require-each-key warnings untouched).
2026-08-07 15:27:24 +07:00
31d45a9d64 feat(039): REST scenario binding + verification pipeline views (T054-T057)
Close the 039 REST-binding and pipeline-view gaps found in the audit: the
scenario API client was never imported and pipeline views were not bound to
any page.

T054/T056 - dashboard-testing.ts gains compileScenario/validateScenario/
  resolveScenario (requestApi POST); WorkspaceModel.compileFromRest /
  validateFromRest / resolveFromRest give an agent-free REST preview path;
  DashboardScenarioWorkspaceModel.rest.test.ts (3 tests).
T055 - capture/vlm/disposition REST surface already landed on backend (038
  T057/T058); EvidencePanel autonomous binding deferred to follow.
T057 - DashboardDetailModel.loadVerificationRuns() + VerificationHistoryList
  bound on /dashboards/[id], consuming 037 T081 GET /verification/history;
  DashboardDetailModel.test.ts = 67 passed.

T058 (verify action) intentionally left open: VerificationRunRequest needs a
repository_id which dashboard metadata does not expose, and there is no
PREPROD deployment page in the frontend. Documented as a blocker in tasks.md.

Verification: 79 vitest passed (REST + detail model + api), vite build OK;
eslint clean for changed code (pre-existing URLSearchParams lint on old line
left untouched).
2026-08-07 15:15:01 +07:00
410afdf40e feat(037): verification pipeline automation + GET read-API (T080-T081)
Close the 037 pipeline-automation and read-API gaps found in the audit:
deploy/release hooks did not create VerificationRun, and GET endpoints for
history/detail were absent even though 039 UI and client call them.

T080 - _release_routes.py: create_release now fires best-effort
  _trigger_release_verification -> VerificationRun with trigger=release_create
  (metric+structure); verification scheduling failures never roll back the
  release transaction.
T081 - verification.py: add GET /verification/history (dashboard_id +
  environment_id filters, newest-first) and GET /verification/{run_id}
  (404 RUN_NOT_FOUND); reuse _record_to_response.
  - verification_run.py + alembic migration p2q3r4s5t6u7: nullable indexed
    dashboard_id populated from structure/visual/metric category_params.
  - verification_service.py: _derive_dashboard_id helper.

Verification: release routes (32) + verification API (8) + persistence (21)
= 53 passed; ruff clean for changed code (pre-existing RUF012/UP017 on old
lines left untouched).
2026-08-07 14:21:25 +07:00
d1e15904e1 feat(038): wire real VLM submit + real capture bytes (T057-T059)
Close the 038 MVP runtime gaps found in the audit: VLM analysis previously
returned empty findings with no provider call, and capture registered a
synthetic sha256 derived from run/step ids instead of real image bytes.

T057 - vlm.py: replace _default_submit stub with real submit_screenshot that
  resolves a multimodal provider via LLMProviderService (decrypted key,
  multimodal-required gate) and calls Plugin.Service.LLMClient.get_json_completion
  with the masked screenshot; analyze_screenshot is now async.
T058 - capture.py: dispatch_capture now REQUIRES real capture_bytes/masked_bytes
  and computes sha256 from the actual image bytes (synthetic hashes forbidden);
  scenario API accepts base64 capture/masked bytes.
T059 - test_scenario_vlm_e2e.py: capture -> VLM -> disposition end-to-end with
  real bytes (masked bytes reach the provider; digest matches sha256 of bytes).

Verification: tests/services/dashboard_testing/scenario/ = 103 passed, ruff clean
(existing B008 on pre-existing draft-pack route lines untouched).
2026-08-07 14:08:04 +07:00
a1b20bf2cf feat(040): wire RunnerPool into run_load_run — real load execution (T075-T079)
Close the 040 MVP runtime gap: run_load_run previously only slept through
RAMP->STEADY->DRAIN->COMPLETED with zero Superset requests. Now it actually
executes load-test chart queries through the 037 client.

- executor.py: 037-native adapter (execute_superset_chart via
  execute_dashboard_query_envelope; build_execution_items with stable ids)
- persistence.py: write_load_executions batched LoadExecution persistence
- plugin load_testing.py: run_load_run -> _resolve_execution_scope ->
  _run_bounded_pool (RunnerPool + shared client_registry.get_semaphore +
  CircuitBreaker -> circuit_breaker_abort)
- test_executor_runtime.py: 5 tests proving executor delegation, item
  determinism, real-execution persistence (2 items -> 2 LoadExecution rows,
  run -> COMPLETED), and lifecycle-only degradation on missing env

Verification: tests/services/load_testing/ = 76 passed, ruff clean.
2026-08-07 13:37:46 +07:00
60345cc126 docs(specs): 041 reconciliation pass + regenerate 038-041 validation evidence
Address external review findings on the 036-041 package:

- 041 LIN-FR-016/Q2: sync with research R2 — SQL-expression parsing uses
  the in-repo sqlparse-based extractor (sqlglot rejected as new dep);
  exact-confidence bounded to authoritative column/metric refs; note that
  sqlparse is intentionally non-validating (no SQL AST guarantees)
- 041 LIN-FR-002/data-model: snapshot pinning documented as an optimistic
  consistency token, not a historical edge-set store (mismatch -> stale
  notice, no edge-set restore)
- 041 research: second R9 renamed R10 (collision with R9 labels/metrics)
- 041 tasks T017: '11 matrix rows' -> '12 data-rows' (actual matrix count)
- 041 plan: storage counts 5 -> 6 new tables + 1 additive FK (matches
  data-model); decision-memory R1-R8 -> R1-R10
- 041 spec Status: Draft -> Ready for Implementation; CHK021 reworded
  (cycles impossible by construction per LIN-FR-015)
- 038/039/040/041 validation.md: regenerate digest tables; withdraw
  038 'IMPLEMENTATION COMPLETE' claim and clarify each PASS certifies
  spec/contract validity only while runtime closure tasks stay open
2026-08-07 13:22:44 +07:00
7e57b7fb39 gitignore 2026-08-07 13:03:06 +07:00
ac95beb1a0 docs(specs): record 036-041 MVP runtime gaps as open closure phases
Fact-check the dashboard-testing spec packages against actual code and
amend the full speckit document set (spec, plan, research, tasks,
traceability, quickstart) so documented status matches reality:

- 037: discrete metric tools work, but deploy hooks do not create
  VerificationRun and GET read-API endpoints are missing (T080-T081)
- 038: compiler/validator work; VLM _default_submit and capture dispatch
  remain runtime stubs that never call LLMClient/ScreenshotService
  (T057-T059)
- 039: UI components exist, but api/dashboard-testing.ts is unbound and
  pipeline views are not wired to pages; depends on 037 read-API
  (T054-T058)
- 040: run_load_run never invokes RunnerPool, so load runs execute zero
  Superset requests; must wire 037 executor (T075-T079)
- 041: backend index works, but no /lineage frontend and lineage_index
  stays opt-in (T045-T048)
- 036: confirmed operational, relations to reused plugin modules fixed

19 open closure tasks total; region pairs balanced.
2026-08-07 12:56:40 +07:00
c1c35e5855 fix(scenario): route draft storage through StorageService with legacy fallback, persist run marker for post-restart recovery 2026-08-07 10:24:32 +07:00
a5e69de5ea fix(scenario): complete save flow with auto-created repo, robust approval gate, idempotent auto-start 2026-08-07 01:56:36 +07:00
6705437acc refactor(agent): compact chat header, remove duplicated status info 2026-08-07 01:55:50 +07:00
fa35514285 chore(agent): remove redundant PRODUCTION banner from chat 2026-08-06 22:37:46 +07:00
9cb5717a78 fix(agent): auto-start scenario chat, robust HITL resume, strict service auth
- frontend: fix auto-start on dashboards->/agent navigation (undefined params
  ReferenceError), route initial connect through ConnectionManager with auto-retry,
  reset runModel on objectId change and failed recovery
- agent: fix closure-over-loop-variable bug in _inject_env_id_into_tools (env now
  resolved from request-local ContextVar; idempotent wrapping), make
  execute_dashboard_result.result_key optional, resilient checkpoint resume with
  ToolMessage repair + direct-tool fallback, remove dead fast-path, consolidate
  tool_call parsing in _tool_resolver, context-safe ContextVar resets
- backend: llm-config gated by strict service-only auth (no user-JWT fallback),
  tighten idempotent run reuse (dashboard/env/intent match + 6h staleness),
  terminal event transitions run.status to COMPLETED/FAILED/CANCELLED,
  null-safe metric parsing in dashboard query model
- run.sh/docker-compose: require SERVICE_JWT (random per-run secret) instead of
  public default
2026-08-06 18:26:50 +07:00
b820b8b47c fix(maintenance): validate environment synchronously on start
start_maintenance accepted any environment_id, creating a stuck PENDING
event that never transitioned for unknown environments. Add synchronous
404 guard (mirrors preview_dashboards), inject config_manager via Depends,
and cover with a regression test proving no event row is created.

Also fix mock_task_manager to await broadcast_maintenance_event (AsyncMock),
aligning the fixture with the production route's awaited call.
2026-08-06 17:37:36 +07:00
6d19ecf81a fix: harden feature security and e2e integrations 2026-08-06 13:47:28 +07:00
6bd050f458 fix(run.sh): apply Alembic migrations before backend start
- Adds alembic upgrade head to start_backend (parity with docker entrypoint).
- Robust 3-way detection: alembic_version present -> upgrade head;
  schema present via create_all (no stamp) -> stamp head;
  empty DB -> upgrade head.
- run.sh-launched DB now gets lineage + load-testing tables automatically.
2026-08-05 22:40:48 +07:00
df837dbb73 feat: 039 complete — 18-step matrix, responsive, evidence a11y
- T018/T022: 18-step fixture renders without collapse; all 19 automation-status rows.
- T038: 1366px responsive assertion.
- T052/T053: evidence a11y (disposition focus order) + 3 findings distinct severities.
- 039-dashboard-scenario-ui now 0 open / 52 done.
- All three specs (041/040/039) fully closed.
2026-08-05 22:35:27 +07:00
170345af0a feat: 040 Superset/Testcontainers integration tests — spec complete (0 open)
- T066: test_dashboard_load_testing_superset.py — real Superset chart-data
  preserves source_response_hash + cache metadata (LOAD-FR-018/019).
- T067: test_load_testing_client_capacity.py — shared semaphore wiring,
  reserve slots, multiple runs, fairness.
- Ran against a real Apache Superset Testcontainers container
  (proxy-bypassed NO_PROXY for localhost).
- 040-dashboard-load-testing now 0 open / 74 done.
2026-08-05 22:30:31 +07:00
9e71b2a38a feat: 039 fixtures + recovery + a11y; 040 a11y
- 039 T001: materialize 038 scenario fixtures into __fixtures__/dashboard-testing.
- 039 T011: WorkspaceModel.setDomainError recovery (permission/missing-env, no AgentRun).
- 039 T039 / 040 T068: a11y assertions (labeled inputs, ARIA progress strip).
- 039 open -> 6, 040 open -> 2.
2026-08-05 21:56:43 +07:00
f57d92739b feat: e2e tests for 039 scenario UI + 040 load testing
- 039 T040: dashboard-scenario-ui.e2e.js (entry v2 intent, missing-env, reload recovery).
- 040 T065: dashboard-load-testing.e2e.js (entry, matrix preview, stop/reconnect).
- Playwright chrome channel now available; all three prototypes browser-validated.
2026-08-05 21:53:30 +07:00
edaaa18d39 chore: browser-validate all three prototypes (039/040/041)
Playwright chrome channel now available; validated workspace, artifacts,
evidence/VLM, pipeline views (039), editor/monitor states (040), and
lineage dependents (041) via state switchers. Marked manifests DONE.
Browser validation screenshots recorded.
2026-08-05 21:52:06 +07:00
78042174e9 chore: 039 discovery-candidate test 2026-08-05 21:47:01 +07:00
90afb78a79 feat: 039 entry/lifecycle/preview tests + verification
- T006: DashboardHeader scenario entry tests (contextVersion=2, env missing, no stale id).
- T027: parameter resolution never restarts inspect.
- T034: preview/draft never marks persisted.
- T048: ArtifactPreview evidence/ branch + disposition summary.
- T041/T042/T043: SQL-language scan (clean), regressions, lint+build verified.
- 039 open down to 11 (browser-dependent + fixtures + recovery).
2026-08-05 21:46:39 +07:00
ab92459923 feat: 039 HITL -> 036 gate, discovery candidates, evidence tree branch
- T030: discovery-candidate flow (no direct approval/catalog mutation).
- T035/T036/T037: scenario save/baseline delegate to 036 pending gate via
  WorkspaceModel.requestDurableAction; deny/blank-reason covered.
- T049: ArtifactPreviewPanel evidence/ subtree + disposition summary.
- 23 tests green, build passes.
2026-08-05 21:42:24 +07:00
128871065d feat: 039 evidence + scenario summary/coverage views
- T007: DashboardDetailModel scenarioHref (contextVersion=2 + intent).
- T019: ScenarioSummary + ScenarioCoverage views (with StepTable/ProgressStrip).
- T045/T044: WorkspaceModel evidence[] + updateFindingDisposition + tests.
- T050: EvidencePanel wired into ScenarioWorkspace.
- T051: VlmProvenanceFooter component.
- 45 frontend tests green, build passes.
2026-08-05 21:34:24 +07:00
0d9f3e09f3 chore: close 041/040 tails — contract test, lifecycle test, fixtures, integration test
- 041 T038: 040 blast-radius consumer contract test (R6 pinned snapshot shape).
- 040 T003: materialize load-testing fixtures into frontend __fixtures__.
- 040 T029: run lifecycle test (ramp/steady/terminal immutability) + T038 aggregate-progress.
- 040 T051/T059: LoadResults + LoadComparison L2 UX tests.
- 040 T064: LoadModels integration test (profile -> gate -> run -> reconnect -> results).
- Check off verified tasks; remaining 040 open: e2e (T065), integration (T066/T067), a11y (T068).
2026-08-05 21:28:54 +07:00
9de74baa08 feat: dashboard testing suite — scenario UI, load testing, dataset lineage
- 039-dashboard-scenario-ui: agent workspace (WorkspaceModel, scenario
  views, parameters/baselines, artifact preview, HITL save/approval,
  evidence/VLM review, pipeline verification views) + contextVersion=2
  scenario intent + prototype.
- 040-dashboard-load-testing: capacity/matrix/profile, runner pool,
  timing, circuit breaker, PROD gate, cache/bounded/consistency,
  comparison/schedule, API + frontend + prototype.
- 041-dataset-lineage-blast-radius: usage index, severity/schema-diff,
  propagation, deprecation, fanout, API + frontend + prototype (R9
  labels/metrics/recreate).
- 036/037 amendments: dataset_updated trigger + fanout_plan_id, lineage
  deprecation gate, 040 read-model traceability.
- Includes pre-existing 042-rls-management-workspace and
  043-idm-account-integration work.
2026-08-05 18:15:30 +07:00
f9a15a0a7b feat(maintenance): opt-in auto-end at end_time via scheduler scan
Add auto_end flag to maintenance_events: when set with an end_time, a
60s APScheduler scan dispatches the end task automatically. The scan
survives restarts and is de-duped by task_id; end_time alone stays
informational. Includes alembic migration, route/schema wiring, Svelte
checkbox with validation, examples, and backend + frontend tests.
2026-08-04 16:38:20 +07:00
4d543a3a0a agents 2026-08-04 15:38:23 +07:00
b1fcf6017a docs(specs): unify-frontend-style audit + product roadmap
- Update 001-unify-frontend-style doc package (spec, tasks, data-model,
  quickstart, plan) to reflect fact-checked ~70% implementation status.
- Mark 20/37 tasks as implemented on disk, document 4 deferred exceptions
  (StateBlock, tasks route, UX walkthrough, conformance checklist).
- Add PRODUCT_ROADMAP.md: cross-spec implementation audit + timeline.
2026-08-04 15:29:18 +07:00
f0e923a40c docs(specs): 042 plan package, R13 dynamic rules, tasks (56)
- research.md: R1-R13 (R13 dynamic filter-based rules in rls_roles_filter
  accepted; materialized apply/sync rejected and retained as decision memory)
- plan.md: filled template, constitution check PASS, ADR continuity
- contracts/modules.md: SaveDefinition/Preview C5, Api.Rls.SaveRule,
  permission declarations; ATTN-1..4 compliant
- data-model.md: entities, 18 DTO pairs, 4 screen models
- fixtures: 29 canonical JSON (preview/save/deactivate/push/snapshot/binding)
- traceability.md: 35 rows, coverage gate CLOSED (tasks linked)
- tasks.md: 56 tasks across 7 phases, C3+ contracts inlined
- spec.md/checklists: FR-023..025 (idempotency, dynamic rules, schema
  stability), edge cases for reference drift
2026-08-04 14:48:23 +07:00
022f2f6e2b docs(specs): add 042 rls-management-workspace package
- spec.md: 4 user stories (script versioning, dataset audit, bi_users
  audit via IDM, custom rule builder), 23 FR, RBAC roles
  rls_operator/rls_script_dev, clarify session 2026-08-04
- ux_reference.md: dual persona, 4 screens, failure matrix (23 classes)
- checklists/requirements.md: 46 checks tied to FR/AC/SC
- prototype: 4 screens, 19 contract states, recovery paths, design-token
  audit 57/57 hex from tailwind.config.js
- research/rls: RLS repository analysis + IDM mock server (reference)
2026-08-04 12:57:57 +07:00
03684fd445 fix(maintenance): start idempotency returns real 409, not documented-but-200
The start endpoint declared 409 in OpenAPI responses but returned the
already_active idempotency hit as a plain 200. Now returns HTTP 409
Conflict with the declared MaintenanceAlreadyActiveResponse body
{maintenance_id, status: 'already_active'}.

Consumers updated to treat 409 already_active as idempotent success:
- bash example: 409 case in api_call
- python example: 409 branch in start_maintenance
- frontend form: info toast instead of error

New test: TestStartIdempotency verifies 409 + body + no new task
dispatched (naive datetimes to match SQLite tz-stripping).
2026-08-04 11:57:41 +07:00
d54f903660 feat(maintenance): settings (timezone, height), snapshot restore, UX polish
Settings:
- display_timezone defaults to Europe/Moscow everywhere (model, service
  fallbacks, settings form); migration flips untouched 'UTC' default row.
- New banner_height setting (1-200 grid units, NULL = auto): threaded from
  settings through start/rebuild flows into MARKDOWN insert/content updates;
  settings panel gains Auto/Manual radio + number input.

Banner removal:
- update_dashboard_layout returns the pre-mutation position_json; stored on
  maintenance_dashboard_banners.original_position_json at creation.
- Removal restores the snapshot verbatim when the current layout matches the
  deterministic replay of the insert (normalize + y-shift + banner keys);
  diverged layouts (user edits during maintenance) and legacy banners fall
  back to surgical removal. Fixes layout drift from y-shift and unreverted
  ROOT->TABS -> ROOT->GRID normalization.

UX:
- StartMaintenanceForm: recent-tables chips from event history, Enter-to-
  submit, schema.table format hint, inline task progress panel (status +
  progress bar + task-log link), templates section removed, Button atoms.
- Maintenance page: environment selector + env context init.
- Events table: Active/Completed tabs with count badges.
- Backend: plugins/services report monotonic progress via
  context.logger.progress (per-dashboard) for start/end/end-all.

Tests: 203 backend (matcher, restore/fallback, height, progress) + 3645
frontend pass; migration E2E-verified (upgrade/downgrade).
2026-08-04 11:08:11 +07:00
80dad15458 fix(maintenance): render banner as native MARKDOWN element, not a chart
Two defects fixed:
- ensure_banner_chart created an orphan 'Maintenance Banner' markdown chart
  (polluted Charts menu and dashboard exports). The banner is now a native
  MARKDOWN element in position_json; chart_id is a synthetic layout key.
- insert_banner_markdown_at_top blindly targeted GRID_ID; on ROOT->TABS
  dashboards (FI-0085) GRID_ID is orphaned and the banner never rendered.
  The layout is now normalized to ROOT->GRID_ID->[ROW-banner, ...] with
  recursive parents update, matching the proven-working manual example.

Review-driven hardening:
- liveness check verifies reachability from ROOT (children graph), so
  dashboards corrupted by the old bug self-heal on the next start.
- insert removes stale ROW-banner-*/MARKDOWN-banner-* keys (single banner).
- decision memory (@RATIONALE/@REJECTED) added to both modules.
- new ops script scripts/cleanup_maintenance_banner_charts.py (dry-run by
  default) deletes already-created bogus banner charts in prod.
2026-08-04 08:46:22 +07:00
4d282b43e2 perf(maintenance): skip sqlparse on oversized virtual-dataset SQL
sqlparse raises SQLParseError above MAX_GROUPING_TOKENS=10000 tokens
(~25KB of typical SQL). The try/except fallback already handled it, but paid
~1s per oversized SQL for a parse doomed to fail. Add _SQLPARSE_SKIP_THRESHOLD
(30k chars) to bypass sqlparse for oversized text (~15x faster, 1.2s->0.08s for
a 212KB SQL) while keeping literal filtering for SQL under the threshold.

Tests: oversized-SQL skip-threshold behavior.
2026-08-03 23:52:03 +07:00