Fact-check the dashboard-testing spec packages against actual code and amend the full speckit document set (spec, plan, research, tasks, traceability, quickstart) so documented status matches reality: - 037: discrete metric tools work, but deploy hooks do not create VerificationRun and GET read-API endpoints are missing (T080-T081) - 038: compiler/validator work; VLM _default_submit and capture dispatch remain runtime stubs that never call LLMClient/ScreenshotService (T057-T059) - 039: UI components exist, but api/dashboard-testing.ts is unbound and pipeline views are not wired to pages; depends on 037 read-API (T054-T058) - 040: run_load_run never invokes RunnerPool, so load runs execute zero Superset requests; must wire 037 executor (T075-T079) - 041: backend index works, but no /lineage frontend and lineage_index stays opt-in (T045-T048) - 036: confirmed operational, relations to reused plugin modules fixed 19 open closure tasks total; region pairs balanced.
12 KiB
12 KiB
#region DashboardLoadTesting.Research [C:3] [TYPE ADR] [SEMANTICS research,load-testing,workers,cache,gil] @BRIEF Phase 0 decisions for load execution: worker runtime, environment capacity, timing model, cache provenance, bounded result processing, persistence, and 041 integration.
Feature: 040-dashboard-load-testing | Date: 2026-07-22
R1 — Async Worker Runtime
- Decision: One
LoadRunis registered as one TaskManager task. The load plugin starts N asyncio workers reading a per-run FIFOasyncio.Queue[LoadExecutionSpec]. An environment-scopedLoadCapacityRegistryowns the load semaphore shared by all active runs. Workers check cancellation and circuit-breaker state before taking the next item; in-flight requests drain within a bounded timeout. - Rationale: The workload is 95–99% upstream I/O. Async workers represent execution concurrency directly, expose queue wait, and support deterministic drain. TaskManager already creates/tracks one asyncio task per plugin run and provides cancellation/events.
- Alternatives Considered: (a) HTTP connection count as concurrency — rejected: indirect, hides queue wait and circuit-breaker boundaries. (b) TaskManager task per execution — rejected: floods Task Center/Reports and destroys run-level lifecycle. (c) threads/process workers — rejected for v1: mostly I/O; GIL mitigation is bounded result processing (R5).
- Impact: New C5
LoadTesting.RunnerPool; no APScheduler or AsyncJobRunner for manual starts. Scheduled load callbacks may dispatch one TaskManager run through existing scheduler boundaries later.
R2 — Effective Capacity and Existing Client Semaphore Gap
- Decision: Effective run concurrency =
min(profile_requested, env.load_test_max_concurrent, max(1, env.connection_pool_size - env.load_test_reserved_slots), 25). Defaults: PROD=5; DEV/PREPROD=10; absolute ceiling=25;load_test_reserved_slots=5. Worker acquires load semaphore first, then request travels through the shared client semaphore; one global acquisition order prevents deadlock. - Repository Finding:
Environmentalready hasstage,is_production,connection_pool_size(default 20), andconnection_pool_timeout.client_registry.get_client()creates a semaphore but currently constructsAsyncAPIClient(...)without passingsemaphore=semaphore; therefore the documented per-env limit is not active. - Rationale: Load traffic must not occupy all shared Superset capacity and freeze ordinary UI/API operations. Capacity derives from the existing environment pool rather than an unrelated number.
- Alternatives Considered: Separate httpx client/pool — rejected: loses shared auth/CSRF cookies; explicitly rejected by
AsyncAPIClientcontract. Acquiring shared semaphore manually in load runner — rejected: once registry wiring is fixed, this double-acquires the same semaphore and can deadlock. - Impact: Foundational prerequisite: pass registry semaphore into
AsyncAPIClient; regression test proves ordinary calls and load calls share the same limit. Add config fieldsload_test_max_concurrent(optional derived default) andload_test_reserved_slots=5.
R3 — Timing Model and Percentiles
- Decision: Record three non-overlapping monotonic timings per execution:
queue_wait_ms(enqueued → worker takes),resource_wait_ms(worker takes → upstream request admitted),upstream_latency_ms(admitted → response/error).end_to_end_msis their sum. Dashboard performance percentiles useupstream_latency_ms; orchestration health reports queue/resource wait separately. Percentiles use nearest-rank over successful samples; errors/timeouts have dedicated rates and are not converted into synthetic latency values. - Rationale: Combining waits with HTTP latency would measure our orchestrator under contention rather than Superset. Circuit breaker p99 must reflect upstream degradation; queue saturation has its own trigger/diagnostic.
- Alternatives Considered: End-to-end p99 only — rejected: cannot attribute queue pressure vs Superset slowdown. Treat timeout as latency=timeout — rejected: biases percentile and double-counts error semantics.
- Impact:
LatencySummarystores sample_count, p50/p90/p95/p99 per timing dimension; breaker latency threshold evaluates upstream p99 only after minimum sample count.
R4 — Authoritative Cache State from Superset
- Decision: Parse
/api/v1/chart/dataJSON body fields per query:is_cached,cache_key,cached_dttm,queried_dttm,cache_timeout. Mapping: hit iffis_cached=true; bypassed iff request.force=true; disabled iffcache_timeout=-1; miss iffis_cached=null && cached_dttm=null && force=false; unknown for absent/inconsistent fields. Persist raw fields pluscache_state_source="superset_response". - Evidence: Audited
/home/busya/dev/superset/superset/common/query_context_processor.py(returns all fields),charts/data/api.py(serializes query results),charts/schemas.py(response schema),common/utils/query_cache_manager.py(hit semantics), and Superset integration tests (force/source execution producesis_cached=None). - Rationale: This is an authoritative execution signal from Superset, unlike latency inference.
- Alternatives Considered: HTTP headers — rejected: chart-data uses body metadata. Probe inference from latency delta — rejected: DB buffer cache, connection reuse, and network jitter confound it.
- Impact: 037 chart-data adapter must preserve these fields rather than normalize them away.
R5 — GIL and Bounded Result Processing
- Decision: Load path computes SHA-256 while consuming the raw response bytes, extracts row count/cache metadata, and retains only a bounded first-N-row sample for display. It does not run full 037 value normalization. Consistency compares digest for identical
(chart, filters_hash, role, time_range)coordinates. Default sample cap: 100 rows and 256 KiB serialized sample. - Rationale: At cap 25 and 0.5–3s upstream latency, workers are predominantly I/O-bound. The relevant GIL risk is large table JSON parsing/normalization (10k+ rows, 200–500ms continuous hold). Bounded processing keeps event-loop distortion low and ensures reported latency belongs to Superset.
- Alternatives Considered: ProcessPool normalization — deferred; IPC cost distorts small responses. Full 037 normalization — rejected for load path; belongs to correctness verification. Hashing post-parsed normalized JSON — rejected: incurs full parse and may hide byte-level divergence.
- Impact: If fixture instrumentation shows event-loop lag >10% of upstream latency at ceiling load, emit
<ESCALATION>for a process/streaming parser ADR; do not silently add ProcessPool.
R6 — Persistence and Event Granularity
- Decision: Persist
LoadProfile,LoadRun,LoadExecution, andConsistencyFindingin PostgreSQL. Workers buffer execution records and flush bounded batches (≤50 or ≤1s), while aggregate progress events are emitted at ≤4 Hz. Store digest/count/sample/cache/timing, never full large payload. Partial results survive stop/breaker/failure. - Rationale: Per-execution commit/event at concurrency 25 creates database/WebSocket amplification. Batching preserves auditability without turning the orchestrator DB into the bottleneck.
- Alternatives Considered: Full response blobs — rejected: unbounded storage/GIL cost and duplicates Superset. In-memory results only — rejected: reconnect and cross-run comparison require durability.
- Impact: C5 repository contract owns atomic batch insert + run aggregate update.
R7 — Circuit Breaker Semantics
- Decision: Evaluate on a rolling completed-execution window with minimum 20 samples. Defaults: error-rate threshold 25%; upstream p99 multiplier 3× the selected comparison run/profile reference; if no reference exists, latency breaker is armed only with an explicit absolute threshold. Breach closes queue intake, transitions to draining, preserves partial results, records triggering window/metric.
- Rationale: p99 on tiny samples is unstable; implicit latency reference would create false aborts. Error breaker works from the first minimum window.
- Alternatives Considered: Break on first timeout — rejected: transient failures are expected load data. Use end-to-end p99 — rejected by R3 attribution model.
- Impact: UI must show breaker
inactive|armed|trippedand why latency breaker may be inactive.
R8 — Variation Matrix Expansion
- Decision: Closed axes: filters, viewport, role, time_range. Normalize axis ordering and values, then cartesian-expand if total ≤ profile cap (default 500 combinations). Above cap, deterministic seeded reservoir sampling without materializing the full product. Each coordinate gets a stable variation id from canonical JSON hash.
- Rationale: Determinism supports reproducible comparisons; bounded sampling prevents combinatorial memory/request explosion.
- Alternatives Considered: Always cartesian — rejected: unbounded expansion. Random sampling without seed — rejected: cross-run comparison becomes invalid.
- Impact: Preview shows theoretical combination count, selected count, seed, chart count, total requests.
R9 — 041 Blast-Radius Read Model
- Decision: Profile validation calls 041 dependents endpoint with pinned fingerprint.
LoadRunpersists that fingerprint and the resolvedBlastRadiusReport. Start rejects a changed fingerprint with a refresh/reconfirm response, especially before PROD approval. Mid-run index updates do not mutate the run. - Rationale: Configure-time and gate-time impact must name the same dependent dashboards; otherwise user approval is not bound to actual impact.
- Impact: 040 cannot be implementation-complete before 041 LIN-FR-013 read contract exists; fixtures unblock independent domain tests.
R10 — PROD Gate and RBAC
- Decision: Stage source of truth is
Environment.stage(PROD) withis_productiontreated as backward-compatible alias; either marks PROD. PREPROD/DEV requiredashboard:loadtest:execute; PROD additionally requiresdashboard:loadtest:prodand 036 approval gate reason. Gate hash binds profile revision, effective cap, estimated request count, environment id, and 041 fingerprint. - Rationale: Existing configuration already exposes stage; dual interpretation avoids silently under-classifying legacy
is_production=trueenvironments. - Impact: Start-after-approval revalidates all hash inputs; mutation/stale gate dispatches zero requests.
Resolved Risks
| Risk | Mitigation |
|---|---|
| Load starves ordinary Superset calls | Effective cap reserves 5 shared slots; registry semaphore wiring fixed (R2) |
| Double semaphore acquisition deadlock | Single order: load semaphore then request-internal shared semaphore; never manually acquire shared semaphore |
| Event loop distortion on large tables | Raw digest + bounded sample; no full normalization (R5) |
| Cache state misclassification | Authoritative response body + force provenance (R4) |
| Per-execution DB/event amplification | Batches ≤50/1s; progress ≤4 Hz (R6) |
| Stale blast-radius approval | Fingerprint bound into gate hash; start revalidates (R9/R10) |
R11 — MVP Runtime Gap (audit 2026-08-07)
- Decision:
run_load_runMUST instantiateRunnerPool(capacity=effective_concurrency, executor=037-query-executor, env_semaphore=shared, breaker=run_breaker, on_result=persist)and drain queued executions through it. The current hold/sleep body (фазовый переход без запросов) is a stub and violates LOAD-FR-001. - Rationale: RunnerPool уже реализован и покрыт тестами; отсутствует только wiring в production-путь. Подключение 037
SupersetClient.ChartData.Executeчерез адаптерservices/load_testing/executor.pyсохраняет единый источник metric truth и bounded-normalization (LOAD-FR-018). - Alternatives considered: in-process pool без shared semaphore (rejected — R2), прямой httpx в load-пути (rejected — LOAD-FR-001/037 continuity), запуск load как 036 AgentRun (rejected — LOAD spec).
- Impact: T075–T079 в tasks.md Phase 9; quickstart шаги 3–7 не считаются пройденными до реального исполнения executor'а.
#endregion DashboardLoadTesting.Research