Commit Graph

1169 Commits

Author SHA1 Message Date
81ba94684f feat: sed-мутации датасетов, правки миграции и фиксы
- Dataset viewer: sed-замена (find→replace) по всем/выбранным/текущим датасетам
  с обязательным визуальным предпросмотром и целями (sql/metrics/yaml); сохранение
  и выбор именованных правил (sedRules store + SedRuleEditor).
- Migration: literal-замена в dataset YAML при переносе; рескан ID чартов/датасетов
  перед миграцией + метрика средней длительности синка (GET /migration/sync-stats);
  логическая группировка опций (маппинги БД, сервер изменений, rescan) под галочками.
- Fix: дашборды хаба скрывались дефолтным профиль-фильтром show_only_slug_dashboards=true
  (теперь false); мастер миграции не грузил дашборды предзаполненного источника;
  потеря терминального task_status из-за утечки pending-корутин в /ws/logs (панель
  результата не появлялась); горизонтальный скролл в MappingTable; чистая per-env
  ошибка в mapping coverage.
- Backend: record_sync_duration + alembic-миграция sync-duration; literal_replace модуль.
2026-08-19 10:29:28 +03:00
a5764eb008 fix(orchestration): make KV-cache rules explicit; workers long-lived everywhere
- skill §11 rewritten: one invariant (byte-identical prefix) + explicit
  preserve/invalidate lists + discipline (load self-orchestration once;
  persona/toolFilter/model fixed for a worker's whole life)
- worker skills: "disposable context" -> "long-lived context"; role lines
  now say "leaf, long-lived, refined in place via send_message"
2026-08-18 16:24:19 +03:00
e3de06832a feat(orchestration): long-lived workers across all roles + KV-cache economics
- contracts: Self.Worker.Verify/Curate now carry the long-lived invariant
  and "refined in place" brief (matching Implement)
- skill §9: cycle says "read freely" + "spawn long-lived worker, refine"
- skill §11 (new): token & KV-cache economics — long-lived send_message
  keeps the prefix byte-identical (warm cache), fork invalidates it,
  spawn starts cold; merge envelopes; keep surface lean
2026-08-18 16:15:23 +03:00
b232c347fd feat(orchestration): make workers long-lived, not disposable
Workers are continuable long-lived children: spawn once, refine in place
via send_message as the feature evolves; a worker's own session persists
and compacts independently. Re-spawn only when the context is poisoned.

- skill §3: delegation tree rewritten — "refine, don't re-spawn"
- skill §5: "refine, don't re-spawn" coordination rule
- skill §7: stalled worker → send_message (keep context), fresh spawn
  only for poisoned context
- contracts: Self.Worker.Implement is "long-lived, refined in place";
  add long-lived invariant
2026-08-18 16:11:07 +03:00
8a2a12964c fix(orchestration): architect reads freely; delegate execution, not reading
Drop the "never read raw files" rule — it was a false economy. The
architect's context is large and auto-compacting, so reading is cheap and
necessary for decomposition and verification. The real boundary is
EXECUTION (edits/builds/tests) stays with workers, not READING.

- skill §1/§6/§10: "read freely" replaces "never read raw files";
  prefer read_outline/search_contracts to locate, read/grep/glob to understand
- @RATIONALE reframed: reading is cheap; only DECISIONS must not live
  solely in context (they go to files)
- contract: invariant becomes "reads freely for decisions, delegates execution"
2026-08-18 15:52:06 +03:00
ca1761f490 fix(orchestration): harden thin-context protocol against observed drift
Session-log review found the orchestrator diverging from its own design.
Encode every divergence as a rule + invariant:

- no polling: get_goal/list_agents are state tools, not completion checks
  (settlement/report are the signals) — was 101 get_goal + 98 list_agents
- no raw reads: structure via read_outline/search_contracts; file CONTENT
  is delegated (was 77 read vs 1 read_outline)
- <RESULT> enforcement: a worker result with no envelope is blocked
  (was 4/19 envelopes)
- closed role taxonomy: only Implement/Verify/Curate (was ad-hoc
  "code reviewer"/"auditor"/"adversarial")
- fork semantics: fork inherits MY context, not a stalled worker's;
  no "takeover" via fork
- interrupt only to redirect; let workers finish (was 7 interrupts)
- decision memory written by me to files (was zero file writes)
- skill hygiene: load self-orchestration once; never load worker skills
- cross-workspace: point Axiom at the target or delegate all reading
2026-08-18 14:51:50 +03:00
f1ee96fda8 fix(orchestration): stop workers from inheriting the orchestrator role
Child subagents join the parent's preset composition, so by default they
inherited the orchestrator persona AND the delegation tools — and drifted
into orchestrating instead of working (observed: a "worker" called
list_agents/get_goal/send_message and spawned its own grandchildren).

Fixes:
- preset: subagent/subagent_fork now carry a role-agnostic worker `persona`,
  `toolFilter.deny` for all delegation/coordination tools, and `maxDepth: 1`
  (children cannot spawn grandchildren) — hard enforcement at the boundary
- skill: every worker prompt opens with a mandatory role-reset block
- contracts: Self.Worker.Implement/Verify gain a LEAF invariant (no subagents,
  no delegation tools)
2026-08-18 14:41:37 +03:00
dff97e58da chore: gitignore generated semantic-index and bundle artifacts 2026-08-18 12:45:06 +03:00
2a0f334717 chore: remove hardcoded Fernet key and client cert
- ENCRYPTION_KEY: smoke test generates a fresh key inline; templates use a placeholder
- drop RUSAL_ROOT.cer (client-specific public cert) + gitignore it
2026-08-18 12:32:06 +03:00
a995e6f269 chore: remove tracked junk and add gitignore rules
Drop service/debug files accidentally tracked in the previous checkpoint:
- session.jsonl (agent transcript)
- artifacts/ (integration-test debug logs)
- research/paper.pdf (binary blob)
- .npmrc (machine-local config)

Add .gitignore rules for session.jsonl, artifacts/, .npmrc, *.pdf
so they cannot be re-committed.
2026-08-18 12:27:03 +03:00
977f3d6d75 chore: checkpoint working tree onto master
Carried over from 042-dashboard-scenario-registry:
- dashboard/migration backend changes + tests (dataset_key_sync)
- specs updates; drop generated doxygen artifacts
- research notes, integration artifacts, session log
2026-08-18 12:20:42 +03:00
604b3dc706 feat(orchestration): add implementer and verifier workers
Complete the role graph with the two remaining leaf workers.

Skills:
- self-implementation: implement inside @PRE/@POST/@INVARIANT guardrails,
  verifiable edit loop, decision-memory preservation, <RESULT> envelope
- self-verification: orthogonal falsifiable verification, hardcoded
  fixtures, @TEST_INVARIANT traceability, anti-tautology

Presets (staged under docs/design/*-preset; installed to ~/.dsh/.agent-presets):
- implementer: native wire, leaf, bash for the verifier
- verifier: native wire, leaf, bash for pytest/vitest

Contracts (self-orchestration-contracts.md, now 13 contracts):
- Self.Implement.{EditLoop,DecisionMemory}
- Self.Verify.{Traceability,AntiTautology}
- worker edges: Implement/Verify DISPATCHES -> their sub-contracts
2026-08-18 09:17:48 +03:00
e6c77cc3da feat(orchestration): self-orchestration flow — orchestrator + curator
Add the thin-context orchestration protocol as loadable skills,
agent presets, and verifiable GRACE-Poly contracts.

Skills:
- self-orchestration: architect protocol (memory hierarchy, delegation
  decision tree, <RESULT> envelope, park-don't-poll, anti-loop)
- semantic-curation: curator protocol (audit → one-file repair → verify →
  rebuild → health report; anti-corruption invariants)

Presets (staged under docs/design/*-preset; installed to ~/.dsh/.agent-presets):
- orchestrator: native wire, tuned compaction (0.75/0.20), full toolset
- curator: native wire, leaf (no delegation), bash reserved for git rollback

Contracts (docs/design/self-orchestration-contracts.md, indexed + audited):
- Self.Orchestrator, Self.Worker.{Implement,Verify,Curate}
- Self.Curation.{Loop,HealthReport,AntiCorruption}
- Self.Contract.{ResultEnvelope,DecisionTree}
2026-08-18 09:07:14 +03:00
1dd78ca548 fix integration test failures 2026-08-17 16:09:03 +03:00
ea05a42c81 fix frontend environment fallbacks and remove deprecated code 2026-08-17 15:08:34 +03:00
2e8628b2c8 tasks 2026-08-13 18:49:37 +03:00
c32b7ef509 feat(translate): extend run metrics with observed flow stats; test hardening
- backend: aggregate cache_hits/observed_runs/source_records_read/eligible/
  translated/same_language_skipped/insert_rows_* preserving NULL for
  historical runs; drop legacy translate plugin module
- frontend: history page metric cards + RunOutcomeCompact per-run summary,
  totals with observed-scope notice; tabular numbers
- tests: fix banner date-format expectations, api-key env-scope fixture,
  rate-limiter cache pollution pinning, chart/candidates guards
- specs: sync dashboard-testing openapi contract
2026-08-13 08:23:43 +03:00
6336de9c24 fix(translate): correct preview responses and target DB selection 2026-08-11 12:34:00 +03:00
e7d33ce4c8 chore(db): drop orphaned dataset-review and connection_configs tables
Remove dead schema left behind by removed features:
- dataset-review family (dataset_review_sessions, dataset_profiles, and
  related children) from c3ad0afc — its non-cascading FKs broke environment
  deletion with ForeignKeyViolation
- connection_configs from 74e64622

Both are unreachable from the app (no models register them).
2026-08-11 10:55:52 +03:00
3f12d52fbc docs: reconcile verification program architecture 2026-08-11 09:02:53 +03:00
1145a1922c docs: close scenario agentic workflow contracts 2026-08-10 20:06:27 +03:00
e52c5777ba feat(maintenance): fan out API starts to prod 2026-08-10 15:56:03 +03:00
9450559da5 feat(maintenance): improve event history and templates 2026-08-10 15:05:20 +03:00
7924ec5b10 fix(logs): reduce production log spam — agent llm-config polling, scheduler plumbing
- middleware: suppress structured REASON/REFLECT framing for high-frequency
  pollers (/api/agent/llm-config, /api/tasks/{id}, health/summary,
  session/activity, settings/consolidated); fixes tasks/{id} never matching
- agent: _fetch_llm_config treats 401/403 as terminal (no retry, log once),
  bounded backoff 5s/15s/60s on connect/timeout/5xx; langgraph_setup logs
  auth failure once per process
- scheduler: auto-end plumbing lines (executed/scan triggered) -> DEBUG
- thumbnail: Superset 4xx rejections logged at DEBUG instead of EXPLORE
2026-08-10 12:28:07 +03:00
4bc244c228 fix(examples): maintenance API scripts — error handling, JSON safety, docs
- bash: propagate api_call failures (exit 1 on 400/401/403/404/network),
  write diagnostics to stderr, escape message for JSON safety, help without
  API key
- python: argparse options after subcommand (parents), single error message
  per failure, network errors without traceback, idempotent already_completed
- move scripts to examples/maintenance/ with README instructions
- backend: correct stale envelope-shape comment in maintenance schemas
2026-08-10 12:04:50 +03:00
db255ea4e6 feat(maintenance): ui/ux audit improvements for BI analyst persona
- read-only access to /maintenance for analysts (sidebar + hidden management)
- hub badge: message tooltip, link to events, accessible aria-label, localized end
- keep hub badge fresh via shared maintenance WS (init on dashboards page)
- events table: message column, auto-end indicator, localized statuses
- confirm dialog before starting maintenance with affected-dashboards summary
- surface load errors inline; localize store toasts
- auto-end discoverability hints in form and table
- form: multiple tables, end>start validation, timezone note
- status colors: active -> warning; completed tab dashboards expandable
- settings: timezone select, fieldset, localized units/aria labels
- backend: expose auto_end in event items, message in banner states
2026-08-10 12:04:43 +03:00
512e9223e3 feat(maintenance): configurable date format for banner timestamps 2026-08-10 11:22:48 +03:00
71a76cc784 fix(git): address dashboards by stable slug, hide git actions for slugless dashboards 2026-08-10 11:22:41 +03:00
d733a1a745 chore: remove legacy semantic skills 2026-08-10 09:19:07 +03:00
3421484347 chore: synchronize remaining workspace updates 2026-08-10 07:29:08 +03:00
1c577e8561 docs(semantics): normalize code contracts 2026-08-10 07:28:05 +03:00
093f7f600f fix(maintenance): harden banner lifecycle and guarded migrations
- Guard maintenance Alembic operations for create_all-only tables on clean DBs
- Add guarded verification_runs.fanout_plan_id backfill migration
- Improve maintenance banner rendering, chart management, orchestration, and API routes
- Expand assistant maintenance tool and edge-case coverage

Tests: cd backend && source .venv/bin/activate && python -m pytest -q tests/test_maintenance_api.py tests/test_maintenance_service.py tests/api/test_assistant_tool_maintenance.py tests/api/test_maintenance_routes_edge.py (77 passed)
2026-08-09 08:22:16 +03:00
ffa4456a62 docs(specs): add machine contract reconciliation gate
- Validate and repair OpenAPI YAML contracts for 038, 043, and 046
- Canonicalize 038 JSON schema and fixtures around scenario_key,
  content_hash, and logical_step_id; validate all fixtures with jsonschema
- Regenerate 038 validation evidence for compiler scope only
- Add reconcile_contracts.py for repeatable OpenAPI/JSON/fixture checks
- Replace raw CreateScenario payload with server-owned handles and document
  transactional outbox/materialization saga for Registry-to-git persistence
- Record reconciliation outcome in REVIEW-042-047-CLOSURE.md
2026-08-09 08:19:02 +03:00
f7e539440e feat(tooling): rewrite merge_spec.py — batch spec merging + new package support
Rewrite merge_spec.py to merge one or many feature spec packages into a
single review file.

Batch modes:
- single number: python merge_spec.py 038
- inclusive range: python merge_spec.py 036-041
- explicit list: python merge_spec.py 036 038 044
- by dir name: python merge_spec.py 042-dashboard-scenario-registry
- all: python merge_spec.py all
- custom output: python merge_spec.py 036-041 -o out.md

Handles the new spec package structure that plain *.md merging missed:
- includes contracts/openapi.yaml (YAML), contracts/ux/* (decisions.md),
  prototype/index.html + prototype/manifest.md
- skips .json/.py/.zip/.pyc and __pycache__ (fixtures/code/binaries)
- canonical per-feature order: spec -> ux_reference -> checklists ->
  UX contracts -> plan -> research -> data-model -> modules -> openapi ->
  quickstart -> traceability -> tasks -> prototype
- missing numbers warn+skip; dedup; per-feature grouping in one output

Verified: 043 (14 files), 036-041 (6 features/104 files), 036-047 (12/186),
all (50/572) with no .json/.zip/.pyc leakage.
2026-08-09 11:06:36 +07:00
4858992e15 docs(specs): cross-spec canonicalization pass — reconcile 038 core with 042-047
Reconcile the stale 038 compiler-layer model with the 042-047 lifecycle and
its normative documents (not just data-model).

038 -> clean IR/compiler layer:
- identity: scenario_id slug -> scenario_key (semantic); scenario_id (UUID)
  and revision_id (UUID) assigned by 042 at Save; compiler emits content_hash
- ScenarioStep: add logical_step_id (immutable UUID) + step_key/position/
  step_content_hash; runtime VlmFinding/HumanDisposition moved to 044
- VlmAnalysisSpec/ScreenshotCaptureSpec stay (WHAT); runtime capture/VLM/
  disposition endpoints marked deprecated -> 410 MOVED_TO_044
- CompileRequest: agent_run_id no longer required (optional provenance,
  source_type: agent_run|editor|migration|api)
- runtime evidence = Artifact(owner_type=scenario_run), never authoring DraftPack
- validation.md PASS nullified (self-contradictory COMPLETE vs OPEN);
  refocused as compiler-layer PASS only; T057-T059 moved to 044; T046 rewritten

042/043/044/045 normative (spec/research/checklists/ux/prototype/tasks):
- replace revision_hash/parent_revision_hash/scenario_revision_hash with
  revision_id/content_hash/parent_revision_id everywhere
- 044: HumanCheckpoint (confirm/false_positive/inconclusive) distinct from
  ActionApprovalGate; runner pins revision_id+content_hash
- 047: triage split (investigation_status/classification/resolution), false_positive
  vocabulary; /scenarios/{id}/health|trends|recurring-failures

Update REVIEW-042-047-CLOSURE.md with canonicalization pass status.
2026-08-09 10:55:53 +07:00
9889e09d87 docs(specs): renumber 042-043 to 048-049, add scenario lifecycle specs 042-047
- Renumber: 042-rls-management-workspace -> 048, 043-idm -> 049
  (internal refs updated; RLS research '043 Explainability' corrected)
- Add 042 Scenario Registry & Lifecycle: persistence, list/detail,
  immutable revisions (revision_id/content_hash), CreateScenario
  (Save->Register), lifecycle state machine, staleness via 037/041, health
- Add 043 Scenario Editor UX: hybrid edit model C, WorkingDraft save
  (no arbitrary-draft bypass), SetParameter/AddStep/RemoveStep ops,
  constrained assertions, visual DAG, agent edit, Revalidate migration
- Add 044 Scenario Execution Engine: ScenarioRun/StepRun, deterministic
  runner, RunnerPlan derived from revision (not stored source of truth),
  ActionApprovalGate vs HumanCheckpoint, generic artifact owner, worker
  lease/idempotency, logical_step_id, retry closure, result aggregation
- Add 045 Run Monitor & Results UX: config, live SSE monitor, human
  actions, result+provenance, history/compare, Global Run Operations Center
- Add 046 Automation & Operations: schedules/triggers/API trigger, CRUD,
  scheduler semantics, notification events, layered retention tiers, UI
- Add 047 Triage & Analytics: strict flakiness, immutable fingerprint,
  triage split, /scenarios/{id}/health|trends|recurring-failures
- Add specs/REVIEW-042-047-CLOSURE.md mapping all review gaps to fixes
- Update PRODUCT_ROADMAP for 042-049

Each spec: spec/data-model/research/plan/tasks/ux/traceability/quickstart/
checklists + contracts/modules + openapi + interactive prototype.
2026-08-07 18:30:32 +07:00
7487887e61 docs(examples): add maintenance API spec and Russian usage instructions to example scripts 2026-08-07 17:10:25 +07:00
869997554e fix(frontend): use /content endpoint for draft download 2026-08-07 16:53:15 +07:00
b367c3e4e6 fix(038): map LLM-invented selected_case_ids to registered catalog ids
Scenario compile returned 422 VALIDATION_ERROR because the LLM sent
human-readable case names (smoke, data_integrity, filter_propagation) as
selected_case_ids, but the compiler requires registered catalog ids
(B01-B09, C01-C07, T01-T03) and raised KeyError on unknown ones.

- tools_038._compile_objective: resolve case ids through _resolve_case_ids,
  which (1) passes through registered ids case-insensitively, (2) maps
  human-readable synonyms to closest catalog cases, (3) drops unresolvable
  tokens — a free-form name can never reach the compiler as a KeyError.
- CompileScenarioInput.objective_json description now enumerates the valid
  catalog id ranges and gives an example so the LLM stops inventing names.
- Tests: human-readable mapping, unknown-id dropping, dedupe/first-registered
  order (test_tools_038_parse.py, 21 passed).

Verification: ruff clean, 21 tests pass.
2026-08-07 16:39:21 +07:00
066dfe3a35 fix(alembic): merge two parallel heads from o1p2q3r4s5t6
Runtime migrations failed with 'Multiple head revisions are present' because
037 T081 (p2q3r4s5t6u7 -> verification_runs.dashboard_id) and a concurrent
session-activity change (a1b2c3d4e5f7) both branched from o1p2q3r4s5t6.
The failed 'upgrade head' left dashboard_id unapplied, causing
'column verification_runs.dashboard_id does not exist' on
GET /verification/history.

Add a no-op merge revision (015281bd7759) collapsing both into a single head
so 'upgrade head' applies the verification_runs.dashboard_id column.

Verified: ScriptDirectory.get_heads() == ['015281bd7759'].
2026-08-07 16:33:30 +07:00
0fc03d0e8f docs(041): confirm T047 fully closed (fleet-report UI + tests) 2026-08-07 16:30:14 +07:00
38e5f31eca feat(041): fleet-report panel + close H2 (T047 UI + tests)
Code-review H2: markDeprecated/recordMigration/loadFleetReport had no tests and
loadFleetReport was dead code (not bound to UI). Close it.

- api/lineage.ts: add getFanoutReport(planId) -> GET /lineage/fanout/{plan}/report
- types/lineage.ts: add optional run_status to FleetReportDTO (backend returns it)
- Datasets.LineageModel.loadFleetReport: use lineageApi.getFanoutReport (typed)
- LineagePanel: new 'Fan-out report' section (plan id input + load button +
  per-dashboard status/unresolved-impact rendering), bound to model.loadFleetReport
- i18n: en/ru lineage.json add fleet_report_title, load_fleet_report, plan

Verification: Datasets.LineageModel.test.ts 6 passed (3 new: markDeprecated,
recordMigration, loadFleetReport incl. error path), lineage_panel 4 passed,
vite build OK, eslint clean.
2026-08-07 16:29:47 +07:00
c9664dfabc fix(040/038): address code-review criticals (C1-C4, H1, M1, M3)
QA review of the 036-041 closure range returned FAIL with 3 criticals, all
confirmed. Fixes:

C1 - breaker dead: on_result=persist_batch is now wired into RunnerPool
  (breaker.record() fed per result); added test_breaker_abort_persists_partials
  proving CIRCUIT_BREAKER_ABORT reachability + partial persistence.
C2 - index-based result mapping corrupted data under concurrency: results now
  map by execution_id to their source item; test uses two distinct payloads
  and asserts chart->digest pairing (previously masked by identical fixtures).
C3 - double-acquire of the shared client semaphore (deadlock invariant):
  RunnerPool no longer manually acquires the client semaphore; capacity is
  enforced by worker count, the client bounds total concurrency.
C4 - duplicated ScenarioGraph.Vlm.Analyze region: outer region renamed
  ScenarioGraph.Vlm [TYPE Module].
H1 - _default_submit stub removed: analyze_screenshot requires submit=; no
  silent empty-findings fallback.
M1 - test_capture_dispatch.py region closed.
M3 - capture.py raw_sha256 bypass removed: digest always derived from real
  capture_bytes (no caller-supplied hash).

Verification: load_testing (77) + scenario (103) = 180 passed; ruff clean;
all region pairs balanced.
2026-08-07 16:18:29 +07:00
c162f6ee3a chore(036-041): final validation reconciliation + 039 dashboard verification binding
- 038/039/040/041 validation.md: update PASS status to reflect completed
  runtime closure (T057-T059, T054-T057, T075-T079, T045-T048); regenerate
  039/040/041 digest tables; 039 T058 documented as the sole open task
- frontend/src/routes/dashboards/[id]/+page.svelte: include the 039 T057
  VerificationHistoryList binding (was created in the 039 commit but the
  page-level wiring was left unstaged)

All closure tasks across 036-041 are now complete except 039 T058 (blocked:
no repository_id in dashboard metadata; no PREPROD deployment page).
2026-08-07 15:31:39 +07:00
37bfe12a93 docs(041): mark T045-T048 closed, add Runtime Closure Status
Record 041 frontend + opt-in closure in tasks.md/spec.md/quickstart:
T045-T048 done (LineagePanel binding, deprecation surface, fleet-report,
opt-in rationale). 041 now reflects resolved state.
2026-08-07 15:28:15 +07:00
0a445fb170 feat(041): lineage frontend blast-radius + deprecation + opt-in rationale (T045-T048)
Close the 041 frontend/opt-in gaps found in the audit: no /lineage UI existed
and lineage_index stayed opt-in without documented rationale.

T045 - bind the existing Datasets.LineagePanel (T031) onto /datasets/[id]
  (blast-radius dependents + stale_index notice); deleted my transient
  duplicate to respect component reuse.
T046 - DatasetsLineageModel.markDeprecated()/recordMigration() + LineagePanel
  deprecation lifecycle section (grace window, successor, migration uuid);
  consumes existing api/lineage.ts + lineage.json i18n keys.
T047 - DatasetsLineageModel.loadFleetReport() for fan-out fleet report.
T048 - config_models.py: lineage_index_enabled default stays FALSE with an
  explicit rationale (post-sync Superset detail-call cost; flip after live
  indexer stability proof); consumers treat disabled index as empty read-model.

Verification: lineage + api vitest = 236 passed; vite build OK; eslint clean
for changed code (pre-existing ruff/require-each-key warnings untouched).
2026-08-07 15:27:24 +07:00
31d45a9d64 feat(039): REST scenario binding + verification pipeline views (T054-T057)
Close the 039 REST-binding and pipeline-view gaps found in the audit: the
scenario API client was never imported and pipeline views were not bound to
any page.

T054/T056 - dashboard-testing.ts gains compileScenario/validateScenario/
  resolveScenario (requestApi POST); WorkspaceModel.compileFromRest /
  validateFromRest / resolveFromRest give an agent-free REST preview path;
  DashboardScenarioWorkspaceModel.rest.test.ts (3 tests).
T055 - capture/vlm/disposition REST surface already landed on backend (038
  T057/T058); EvidencePanel autonomous binding deferred to follow.
T057 - DashboardDetailModel.loadVerificationRuns() + VerificationHistoryList
  bound on /dashboards/[id], consuming 037 T081 GET /verification/history;
  DashboardDetailModel.test.ts = 67 passed.

T058 (verify action) intentionally left open: VerificationRunRequest needs a
repository_id which dashboard metadata does not expose, and there is no
PREPROD deployment page in the frontend. Documented as a blocker in tasks.md.

Verification: 79 vitest passed (REST + detail model + api), vite build OK;
eslint clean for changed code (pre-existing URLSearchParams lint on old line
left untouched).
2026-08-07 15:15:01 +07:00
410afdf40e feat(037): verification pipeline automation + GET read-API (T080-T081)
Close the 037 pipeline-automation and read-API gaps found in the audit:
deploy/release hooks did not create VerificationRun, and GET endpoints for
history/detail were absent even though 039 UI and client call them.

T080 - _release_routes.py: create_release now fires best-effort
  _trigger_release_verification -> VerificationRun with trigger=release_create
  (metric+structure); verification scheduling failures never roll back the
  release transaction.
T081 - verification.py: add GET /verification/history (dashboard_id +
  environment_id filters, newest-first) and GET /verification/{run_id}
  (404 RUN_NOT_FOUND); reuse _record_to_response.
  - verification_run.py + alembic migration p2q3r4s5t6u7: nullable indexed
    dashboard_id populated from structure/visual/metric category_params.
  - verification_service.py: _derive_dashboard_id helper.

Verification: release routes (32) + verification API (8) + persistence (21)
= 53 passed; ruff clean for changed code (pre-existing RUF012/UP017 on old
lines left untouched).
2026-08-07 14:21:25 +07:00
d1e15904e1 feat(038): wire real VLM submit + real capture bytes (T057-T059)
Close the 038 MVP runtime gaps found in the audit: VLM analysis previously
returned empty findings with no provider call, and capture registered a
synthetic sha256 derived from run/step ids instead of real image bytes.

T057 - vlm.py: replace _default_submit stub with real submit_screenshot that
  resolves a multimodal provider via LLMProviderService (decrypted key,
  multimodal-required gate) and calls Plugin.Service.LLMClient.get_json_completion
  with the masked screenshot; analyze_screenshot is now async.
T058 - capture.py: dispatch_capture now REQUIRES real capture_bytes/masked_bytes
  and computes sha256 from the actual image bytes (synthetic hashes forbidden);
  scenario API accepts base64 capture/masked bytes.
T059 - test_scenario_vlm_e2e.py: capture -> VLM -> disposition end-to-end with
  real bytes (masked bytes reach the provider; digest matches sha256 of bytes).

Verification: tests/services/dashboard_testing/scenario/ = 103 passed, ruff clean
(existing B008 on pre-existing draft-pack route lines untouched).
2026-08-07 14:08:04 +07:00
a1b20bf2cf feat(040): wire RunnerPool into run_load_run — real load execution (T075-T079)
Close the 040 MVP runtime gap: run_load_run previously only slept through
RAMP->STEADY->DRAIN->COMPLETED with zero Superset requests. Now it actually
executes load-test chart queries through the 037 client.

- executor.py: 037-native adapter (execute_superset_chart via
  execute_dashboard_query_envelope; build_execution_items with stable ids)
- persistence.py: write_load_executions batched LoadExecution persistence
- plugin load_testing.py: run_load_run -> _resolve_execution_scope ->
  _run_bounded_pool (RunnerPool + shared client_registry.get_semaphore +
  CircuitBreaker -> circuit_breaker_abort)
- test_executor_runtime.py: 5 tests proving executor delegation, item
  determinism, real-execution persistence (2 items -> 2 LoadExecution rows,
  run -> COMPLETED), and lifecycle-only degradation on missing env

Verification: tests/services/load_testing/ = 76 passed, ruff clean.
2026-08-07 13:37:46 +07:00