- Validate and repair OpenAPI YAML contracts for 038, 043, and 046
- Canonicalize 038 JSON schema and fixtures around scenario_key,
content_hash, and logical_step_id; validate all fixtures with jsonschema
- Regenerate 038 validation evidence for compiler scope only
- Add reconcile_contracts.py for repeatable OpenAPI/JSON/fixture checks
- Replace raw CreateScenario payload with server-owned handles and document
transactional outbox/materialization saga for Registry-to-git persistence
- Record reconciliation outcome in REVIEW-042-047-CLOSURE.md
Scenario compile returned 422 VALIDATION_ERROR because the LLM sent
human-readable case names (smoke, data_integrity, filter_propagation) as
selected_case_ids, but the compiler requires registered catalog ids
(B01-B09, C01-C07, T01-T03) and raised KeyError on unknown ones.
- tools_038._compile_objective: resolve case ids through _resolve_case_ids,
which (1) passes through registered ids case-insensitively, (2) maps
human-readable synonyms to closest catalog cases, (3) drops unresolvable
tokens — a free-form name can never reach the compiler as a KeyError.
- CompileScenarioInput.objective_json description now enumerates the valid
catalog id ranges and gives an example so the LLM stops inventing names.
- Tests: human-readable mapping, unknown-id dropping, dedupe/first-registered
order (test_tools_038_parse.py, 21 passed).
Verification: ruff clean, 21 tests pass.
Runtime migrations failed with 'Multiple head revisions are present' because
037 T081 (p2q3r4s5t6u7 -> verification_runs.dashboard_id) and a concurrent
session-activity change (a1b2c3d4e5f7) both branched from o1p2q3r4s5t6.
The failed 'upgrade head' left dashboard_id unapplied, causing
'column verification_runs.dashboard_id does not exist' on
GET /verification/history.
Add a no-op merge revision (015281bd7759) collapsing both into a single head
so 'upgrade head' applies the verification_runs.dashboard_id column.
Verified: ScriptDirectory.get_heads() == ['015281bd7759'].
QA review of the 036-041 closure range returned FAIL with 3 criticals, all
confirmed. Fixes:
C1 - breaker dead: on_result=persist_batch is now wired into RunnerPool
(breaker.record() fed per result); added test_breaker_abort_persists_partials
proving CIRCUIT_BREAKER_ABORT reachability + partial persistence.
C2 - index-based result mapping corrupted data under concurrency: results now
map by execution_id to their source item; test uses two distinct payloads
and asserts chart->digest pairing (previously masked by identical fixtures).
C3 - double-acquire of the shared client semaphore (deadlock invariant):
RunnerPool no longer manually acquires the client semaphore; capacity is
enforced by worker count, the client bounds total concurrency.
C4 - duplicated ScenarioGraph.Vlm.Analyze region: outer region renamed
ScenarioGraph.Vlm [TYPE Module].
H1 - _default_submit stub removed: analyze_screenshot requires submit=; no
silent empty-findings fallback.
M1 - test_capture_dispatch.py region closed.
M3 - capture.py raw_sha256 bypass removed: digest always derived from real
capture_bytes (no caller-supplied hash).
Verification: load_testing (77) + scenario (103) = 180 passed; ruff clean;
all region pairs balanced.
- 038/039/040/041 validation.md: update PASS status to reflect completed
runtime closure (T057-T059, T054-T057, T075-T079, T045-T048); regenerate
039/040/041 digest tables; 039 T058 documented as the sole open task
- frontend/src/routes/dashboards/[id]/+page.svelte: include the 039 T057
VerificationHistoryList binding (was created in the 039 commit but the
page-level wiring was left unstaged)
All closure tasks across 036-041 are now complete except 039 T058 (blocked:
no repository_id in dashboard metadata; no PREPROD deployment page).
Close the 039 REST-binding and pipeline-view gaps found in the audit: the
scenario API client was never imported and pipeline views were not bound to
any page.
T054/T056 - dashboard-testing.ts gains compileScenario/validateScenario/
resolveScenario (requestApi POST); WorkspaceModel.compileFromRest /
validateFromRest / resolveFromRest give an agent-free REST preview path;
DashboardScenarioWorkspaceModel.rest.test.ts (3 tests).
T055 - capture/vlm/disposition REST surface already landed on backend (038
T057/T058); EvidencePanel autonomous binding deferred to follow.
T057 - DashboardDetailModel.loadVerificationRuns() + VerificationHistoryList
bound on /dashboards/[id], consuming 037 T081 GET /verification/history;
DashboardDetailModel.test.ts = 67 passed.
T058 (verify action) intentionally left open: VerificationRunRequest needs a
repository_id which dashboard metadata does not expose, and there is no
PREPROD deployment page in the frontend. Documented as a blocker in tasks.md.
Verification: 79 vitest passed (REST + detail model + api), vite build OK;
eslint clean for changed code (pre-existing URLSearchParams lint on old line
left untouched).
Close the 037 pipeline-automation and read-API gaps found in the audit:
deploy/release hooks did not create VerificationRun, and GET endpoints for
history/detail were absent even though 039 UI and client call them.
T080 - _release_routes.py: create_release now fires best-effort
_trigger_release_verification -> VerificationRun with trigger=release_create
(metric+structure); verification scheduling failures never roll back the
release transaction.
T081 - verification.py: add GET /verification/history (dashboard_id +
environment_id filters, newest-first) and GET /verification/{run_id}
(404 RUN_NOT_FOUND); reuse _record_to_response.
- verification_run.py + alembic migration p2q3r4s5t6u7: nullable indexed
dashboard_id populated from structure/visual/metric category_params.
- verification_service.py: _derive_dashboard_id helper.
Verification: release routes (32) + verification API (8) + persistence (21)
= 53 passed; ruff clean for changed code (pre-existing RUF012/UP017 on old
lines left untouched).
Close the 038 MVP runtime gaps found in the audit: VLM analysis previously
returned empty findings with no provider call, and capture registered a
synthetic sha256 derived from run/step ids instead of real image bytes.
T057 - vlm.py: replace _default_submit stub with real submit_screenshot that
resolves a multimodal provider via LLMProviderService (decrypted key,
multimodal-required gate) and calls Plugin.Service.LLMClient.get_json_completion
with the masked screenshot; analyze_screenshot is now async.
T058 - capture.py: dispatch_capture now REQUIRES real capture_bytes/masked_bytes
and computes sha256 from the actual image bytes (synthetic hashes forbidden);
scenario API accepts base64 capture/masked bytes.
T059 - test_scenario_vlm_e2e.py: capture -> VLM -> disposition end-to-end with
real bytes (masked bytes reach the provider; digest matches sha256 of bytes).
Verification: tests/services/dashboard_testing/scenario/ = 103 passed, ruff clean
(existing B008 on pre-existing draft-pack route lines untouched).
Address external review findings on the 036-041 package:
- 041 LIN-FR-016/Q2: sync with research R2 — SQL-expression parsing uses
the in-repo sqlparse-based extractor (sqlglot rejected as new dep);
exact-confidence bounded to authoritative column/metric refs; note that
sqlparse is intentionally non-validating (no SQL AST guarantees)
- 041 LIN-FR-002/data-model: snapshot pinning documented as an optimistic
consistency token, not a historical edge-set store (mismatch -> stale
notice, no edge-set restore)
- 041 research: second R9 renamed R10 (collision with R9 labels/metrics)
- 041 tasks T017: '11 matrix rows' -> '12 data-rows' (actual matrix count)
- 041 plan: storage counts 5 -> 6 new tables + 1 additive FK (matches
data-model); decision-memory R1-R8 -> R1-R10
- 041 spec Status: Draft -> Ready for Implementation; CHK021 reworded
(cycles impossible by construction per LIN-FR-015)
- 038/039/040/041 validation.md: regenerate digest tables; withdraw
038 'IMPLEMENTATION COMPLETE' claim and clarify each PASS certifies
spec/contract validity only while runtime closure tasks stay open
Fact-check the dashboard-testing spec packages against actual code and
amend the full speckit document set (spec, plan, research, tasks,
traceability, quickstart) so documented status matches reality:
- 037: discrete metric tools work, but deploy hooks do not create
VerificationRun and GET read-API endpoints are missing (T080-T081)
- 038: compiler/validator work; VLM _default_submit and capture dispatch
remain runtime stubs that never call LLMClient/ScreenshotService
(T057-T059)
- 039: UI components exist, but api/dashboard-testing.ts is unbound and
pipeline views are not wired to pages; depends on 037 read-API
(T054-T058)
- 040: run_load_run never invokes RunnerPool, so load runs execute zero
Superset requests; must wire 037 executor (T075-T079)
- 041: backend index works, but no /lineage frontend and lineage_index
stays opt-in (T045-T048)
- 036: confirmed operational, relations to reused plugin modules fixed
19 open closure tasks total; region pairs balanced.
start_maintenance accepted any environment_id, creating a stuck PENDING
event that never transitioned for unknown environments. Add synchronous
404 guard (mirrors preview_dashboards), inject config_manager via Depends,
and cover with a regression test proving no event row is created.
Also fix mock_task_manager to await broadcast_maintenance_event (AsyncMock),
aligning the fixture with the production route's awaited call.
Add auto_end flag to maintenance_events: when set with an end_time, a
60s APScheduler scan dispatches the end task automatically. The scan
survives restarts and is de-duped by task_id; end_time alone stays
informational. Includes alembic migration, route/schema wiring, Svelte
checkbox with validation, examples, and backend + frontend tests.
The start endpoint declared 409 in OpenAPI responses but returned the
already_active idempotency hit as a plain 200. Now returns HTTP 409
Conflict with the declared MaintenanceAlreadyActiveResponse body
{maintenance_id, status: 'already_active'}.
Consumers updated to treat 409 already_active as idempotent success:
- bash example: 409 case in api_call
- python example: 409 branch in start_maintenance
- frontend form: info toast instead of error
New test: TestStartIdempotency verifies 409 + body + no new task
dispatched (naive datetimes to match SQLite tz-stripping).
Two defects fixed:
- ensure_banner_chart created an orphan 'Maintenance Banner' markdown chart
(polluted Charts menu and dashboard exports). The banner is now a native
MARKDOWN element in position_json; chart_id is a synthetic layout key.
- insert_banner_markdown_at_top blindly targeted GRID_ID; on ROOT->TABS
dashboards (FI-0085) GRID_ID is orphaned and the banner never rendered.
The layout is now normalized to ROOT->GRID_ID->[ROW-banner, ...] with
recursive parents update, matching the proven-working manual example.
Review-driven hardening:
- liveness check verifies reachability from ROOT (children graph), so
dashboards corrupted by the old bug self-heal on the next start.
- insert removes stale ROW-banner-*/MARKDOWN-banner-* keys (single banner).
- decision memory (@RATIONALE/@REJECTED) added to both modules.
- new ops script scripts/cleanup_maintenance_banner_charts.py (dry-run by
default) deletes already-created bogus banner charts in prod.
sqlparse raises SQLParseError above MAX_GROUPING_TOKENS=10000 tokens
(~25KB of typical SQL). The try/except fallback already handled it, but paid
~1s per oversized SQL for a parse doomed to fail. Add _SQLPARSE_SKIP_THRESHOLD
(30k chars) to bypass sqlparse for oversized text (~15x faster, 1.2s->0.08s for
a 212KB SQL) while keeping literal filtering for SQL under the threshold.
Tests: oversized-SQL skip-threshold behavior.