Commit Graph

1169 Commits

Author SHA1 Message Date
e7925d0e26 feat(analytics): investigation queue findability, case index, linked runs and evidence refs in workspace 2026-09-18 10:48:19 +03:00
f123210bb2 fix(scenario): keep ScenarioStep.inputs as canonical ref list; pin action args in action_inputs 2026-09-18 09:51:04 +03:00
5eea72bc46 fix(mcp): align investigation catalog contract 2026-09-18 09:17:06 +03:00
23414f9af6 feat(mcp): add investigation case tools 2026-09-17 22:11:28 +03:00
6fcc861677 feat(analytics): start investigation queue refresh 2026-09-17 21:26:27 +03:00
534d488f73 feat(scenario): add typed step inputs 2026-09-17 21:19:15 +03:00
fcf4e6eb87 docs(specs): design amendments 2026-09-17 — complex browser scenarios, D/M/E taxonomy, zero-human runtime, evidence tiering, investigation loop 2026-09-17 19:41:49 +03:00
ec917ac677 feat(scenario): 044 provider INV_7 decomposition + T044/T045 live canaries; evidence-backed spec status
The 044 provider lifecycle is decomposed to satisfy INV_7 and closed with live canaries;
the remaining acceptance row (T046 graph-level terminal PASS) is diagnosed with executable proof.

- split browser_readonly_actions (534->160), browser_session (501->58) and browser.py (407->398)
  into sibling modules with frozen contract IDs and frozen import surface (facades re-export every
  moved public name); monkeypatch seams follow the owning modules
- gates per phase: scoped 044 suite 1590 passed and full backend 11269 passed / 0 failed, no delta
- T042b: six PREPROD browser canaries GREEN under the Wave-C run-scoped session architecture;
  harness gains SS_CANARY_STORAGE_ROOT and fixes the dead-loop scheduler drain hang
- T044: new artifact-content live HTTP canary 11/11 (real app, PostgreSQL, JWTs, live-captured bytes)
- T045: new fault-injection canary 16/16 (bounded loop refusals, deadline cancellation, shutdown
  drain, capacity quarantine + TTL reconcile, receipt CAS, late response, cancel drain)
- T046: walker-level evaluation->comparison binding diagnosed (compare-only + mandatory ->
  EVALUATION_UNAVAILABLE; bound evaluation -> BASELINE_AND_SEMANTIC_PASS); T043/T046 stay open
- refresh 044 spec/traceability/checklist/quickstart/SESSION_STATE/contracts/plan/data-model and the
  037/050 rows to evidence-backed states; split three malformed multi-target @RELATION lines
- fix skill drift in semantics-python/semantics-svelte (logger facade, notify(), i18n proxy)
2026-09-17 18:05:45 +03:00
4746af2f3e fix(health): address consolidation review findings 2026-09-16 16:11:02 +03:00
3de0756d4f refactor(validation): remove legacy LLM dashboard validation mechanism
Scenario-based dashboard testing (042/044) and scenario automation (046)
replace the legacy policy/run based validation surface.

- drop backend validation routes, service, schemas and DashboardValidationPlugin
- add migration 0025_drop_legacy_validation removing legacy tables
- redirect legacy settings/automation page to scenario automation surface
- remove frontend validation models, routes, components and i18n bundles
- repair dangling semantic relations and dead plugin-id fixtures
2026-09-16 11:13:50 +03:00
63beea9694 chore(git): ignore docs/openapi external-API reference snapshots 2026-09-15 10:32:35 +03:00
8522a2ee3c feat(scenario): complete dashboard testing UX flow 2026-09-15 10:12:17 +03:00
22ffdb5ba5 chore(tools): add read-only ADFS context collector for closed-contour debugging 2026-09-15 10:01:52 +03:00
b4148cfe5f fix(automation): scheduled runs execute as schedule-owner principal; MCP upsert signature parity — live scheduled happy path proven (046 T018 CLOSED) 2026-09-11 20:22:47 +03:00
d0466a2c1e fix(automation): deterministic scheduled-run idempotency key + typed scheduled-reject observability (046 happy path) 2026-09-11 20:01:53 +03:00
578a6110e7 fix(scenario): T045 review hardening — RUN_PROD REST gate, fail-marked typed errors, marker-bound adoption, unique idempotency, path guard 2026-09-11 19:05:20 +03:00
72018383c4 test(mcp): catalog list includes publish_baseline_catalog (050 T045) 2026-09-11 18:14:30 +03:00
6dde7e15c3 feat(scenario): 050 T045 — gated MCP publish_baseline_catalog + 037 publication worker (CAS, receipts, reconcile), REST parity 2026-09-11 18:13:09 +03:00
e925d4a85c chore(kilo): retire qa-tester and security-auditor runtime agents 2026-09-11 17:27:50 +03:00
f37a46e366 docs(reports): add ss-prod agentic E2E gap analysis and task notes
Add 2026-09-08 production gap/spec-coverage/refresh-plan and dashboard
scenario E2E plan/report, axiom-mcp feedback, and 2026-09-11 data-team
maintenance dev-test run notes.
2026-09-11 17:27:50 +03:00
605c5d0552 docs(specs): refresh 017/036/037/038/039 production contracts
Add production contract refresh sections and normative contracts: 017/050
capture-reuse trust, 036 evidence promotion, 037 catalog lifecycle/revision,
038 browser actions/decision policy, 044-047 production chain/evidence/atomic
triage, and 050 MCP interface artifacts. Requirements marked OPEN pending
executable evidence.
2026-09-11 17:27:48 +03:00
6da072b9d4 test(database): cover DB_SCHEMA isolation and prepare_database reset
- env_settings schema resolution/search_path args incl. invalid-name rejection
- dedicated schema pins engine search_path
- e2e: no-public-privilege role, missing-schema bootstrap/denied, object-kind
  reset (sequences/views/domains)
- enterprise template asserts DB_SCHEMA
2026-09-11 17:27:45 +03:00
7cb02f4fe0 test(scenario): DEF-01 promotion graph validation regressions
Structure-first validation: prose double-hyphen passes, SQL in title rejected,
needs_baseline is not a promotion gate, legacy snapshot stays free-text scan.
2026-09-11 17:27:45 +03:00
7291169343 feat(scenario): durable 037 baseline catalog publication
- PublicationOperation model + 0023 migration with idempotent inspector guard
- publication_worker: reviewed-gate binding, branch-head CAS, commit/receipt
  reconcile on retry (never duplicate/force-push)
- REST /api/catalog-publications (RUN_PROD) sharing the worker with the gated
  publish_baseline_catalog MCP tool (050 T045 / MCPX-FR-030 parity)
- MCP op input schema, catalog registration and approved-dispatch adapter
- tests: worker, REST surface, MCP gate/dispatch parity
2026-09-11 17:27:43 +03:00
71135c0822 fix(maintenance): single-banner multi-event render and per-env end-all
- multi-event builds exactly one banner (dedup messages <br>, min-start/max-end
  envelope, open-ended keeps 'уточняется') instead of nested invalid HTML
- event.environment_id is authoritative for states/banner rows (settings target
  only legacy fallback)
- end_all resolves a SupersetClient per event environment via memoized factory
- structured reason/explore logging for fan-out target resolution
- suppress toast for unconfigured /maintenance/settings 404
2026-09-11 17:27:39 +03:00
5758ae4a83 fix(database): object-level reset and optional DB_SCHEMA isolation
Safe one-shot reset, fail-fast on unknown revisions, object-level drops for non-owner corporate PG, rollback before advisory unlock, configurable DB_SCHEMA via search_path, e2e migration matrix tests.
2026-09-11 16:17:11 +03:00
e334863771 build(backend): declare python-dotenv explicitly (review nit — direct import) 2026-09-11 15:10:24 +03:00
de5e098a1b fix(scenario): review hardening — run.sh key reuse, resolver-canonical publisher envelope, strict pin asserts, honest T045/T046 statuses 2026-09-11 15:09:16 +03:00
7d3fb8a770 feat(scenario): live baseline pin via published Gitea catalog — publisher POST/PUT contract, run.sh key-wipe fix, T045/T046 closed 2026-09-11 14:07:20 +03:00
1411c03f9b fix(scenario): orthogonal review hardening — binding fingerprint guard, strict image digest, typed publisher errors, T029m human-loop closure 2026-09-11 10:10:01 +03:00
1d48625b10 test(maintenance): fix 5s first-test timeout flake — static import + vi.hoisted mocks
Module load of maintenance.svelte.js happened inside the first test via
dynamic import and exceeded the 5000ms testTimeout under full-suite
parallel load. Move to a single hoisted static import (mock factories
now resolve vi.fn()s created via vi.hoisted) and drop a dead no-op
placeholder test. 39/39 in isolation, 3563/3563 full suite.
2026-09-11 09:49:52 +03:00
d04927bda7 feat(scenario): T029m live MCP replay — binding identity fixes, Gitea catalog publisher, 050 statuses 2026-09-11 08:40:37 +03:00
792bb1251d feat(scenario): live eval trust — MIME sniffing, multimodal evidence, baseline pin stamping, binding admin surface 2026-09-11 08:40:19 +03:00
bac002bcbc feat(settings): derive project-sections toggles from sidebar nav registry
- sidebarNavigation.ts: static SIDEBAR_SECTIONS registry (i18n labelKey) + getTogglableFeatureNodes() with per-flag dedup
- projectSections.json: shared feature-id manifest; vitest guard pins nav ids to it, pytest guard pins it to FeaturesConfig fields
- FeaturesSettings: toggles derived and grouped by sidebar section/category, extra block for non-nav flags
- features.svelte.ts: reactive flags store; sidebar rebuilds live after settings save, untracked health polling
2026-09-11 07:35:26 +03:00
d2444f404b docs(specs): session-2 handoff + evidence-based 044/050 task statuses and session-state update 2026-09-10 20:49:50 +03:00
edc9a98a86 fix(scenario): code-review hardening — drop unknown criterion findings, fail closed on unsupported verdict, remove dead branch 2026-09-10 20:29:27 +03:00
bf8abff753 docs(specs): live canary v2 trace + restore operation_id in browser evidence details 2026-09-10 20:14:49 +03:00
2532702478 feat(migration): localized risk messages, advanced options collapse, names over UUIDs
Iteration 2 of the migration UX plan:

Risk message localization (params + frontend templates):
- MigrationRiskItem.params (additive) populated at all 13 emission sites
  (orchestrator, dashboard_outcomes, risk_assessor); English message kept
  for assistant/logs backward compatibility
- migration.riskMessages.ts renders risk_msg_* en/ru templates with
  drift details (schema: «marts» → «public») and graceful fallback

Names over UUIDs:
- _read_target_databases caches uuid->name map; drift/missing-datasource
  messages and params carry dataset titles (table_name) and DB names
- compare_dataset_contracts: live None = no evidence, never drift
  (fixes false positives where target LIST omits database_uuid)
- UUID rendered as ds:{8-char} chip with full UUID in title tooltip

Step 2 UX:
- sed rule, composite keys (+mutation server), rescan IDs moved into
  collapsed Advanced options; selected-count chip near Dry Run CTA;
  target-mutation hint + SOURCE-mutation warning

P0 hygiene:
- task view Cancel -> Close view + background-continues hint; own i18n
  keys for Dry Run/step labels; dead Quick Actions block removed;
  PasswordPrompt alert() -> inline validationError

Decomposition (no behavior change):
- run() 330 -> ~55 lines via _process_dashboard/_resolve_effective_
  mapping/_transform_with_fallback/_assemble_report
- MigrationReviewPanel 303 -> 151 lines via MigrationOutcomeList +
  MigrationRiskAssessment

Tests: params assertions + None-skip case (backend), 5 riskMessages
cases (frontend). 116 pytest + 270 vitest passed, lint/build clean.
2026-09-10 20:13:45 +03:00
3458343d89 fix(scenario): live agent-evaluation chain — manifest byte-lengths, response normalization, prompt contract 2026-09-10 20:02:09 +03:00
bbbd4ccfa6 feat(scenario): published-catalog source fallback + server-side live binding resolution at start 2026-09-10 18:59:15 +03:00
043a14558f docs(specs): record live provider canary evidence in quickstart statuses 2026-09-10 18:32:25 +03:00
fe11a636f1 docs(scenario): live provider canary trace (T7/T8 passed) + two production defects closed 2026-09-10 18:27:46 +03:00
ba2f1f45b2 fix(scenario): browser preflight probe deadlocked lifespan loop thread (300s startup hang) 2026-09-10 18:12:59 +03:00
3e5cc942dd refactor(scenario): wrap naked helpers in GRACE regions (INV_1) 2026-09-10 16:35:13 +03:00
befa03e34f feat(migration): dashboard-first dry-run review, compact logs, second timestamps
Backend:
- dashboard_outcomes module: per-dashboard new/overwrite/identical outcome,
  composite-key drift prediction with desired-vs-live values, read-only
- dry-run: single cached live dataset-contract fetch (was O(dashboards)),
  literal_find/literal_replace reach every transform_zip call
- catalog read failures degrade to warnings (target_database/dataset_catalog_
  unavailable, risk_analysis_unavailable) instead of fabricated blockers
- MigrationDryRunResult + dashboards[]; MigrationRiskItem + dataset_uuid

Frontend:
- MigrationReviewPanel: sorted outcomes (blocked -> overwrite -> new ->
  identical), show-all toggle, risk assessment separated from technical
  details, full en/ru localization
- execute failure preserves dry-run and returns to step 3; submit guard
  hides Back while POST is in flight; confirm dialog before migration start
- task logs: dense one-line rows with expand chevron, compact toolbar
  (filters/debug/errors/autoscroll/count in one row)
- migration dashboard grid timestamps truncated to seconds via
  formatDateTimeSeconds

Tests: +6 backend regression (literal kwargs, outcomes, drift severity,
missing-db blocker, catalog degradation, single live fetch), +4 frontend
(sorting, legacy fallback, execute recovery). 116 pytest + 261 vitest passed.
2026-09-10 16:34:02 +03:00
1b79597536 test(mcp): adapt initial-scenario e2e to fail-closed baseline resolver (Front 5-R) 2026-09-10 16:31:01 +03:00
a21481ea6d refactor(scenario): INV_7 split runner into walker/start/dispatch/crash-recovery facade 2026-09-10 16:13:44 +03:00
8f057f3043 refactor(scenario): INV_7 split executors + lifecycle, repair orphan relations, fix serializer fixture 2026-09-10 15:40:17 +03:00
b024e60884 fix(scenario): commit missing Slice G router + eligibility service 2026-09-10 15:38:59 +03:00
83727aa7f9 feat(scenario): complete offline agentic runtime chain 2026-09-10 14:14:40 +03:00