Files
busya ac95beb1a0 docs(specs): record 036-041 MVP runtime gaps as open closure phases
Fact-check the dashboard-testing spec packages against actual code and
amend the full speckit document set (spec, plan, research, tasks,
traceability, quickstart) so documented status matches reality:

- 037: discrete metric tools work, but deploy hooks do not create
  VerificationRun and GET read-API endpoints are missing (T080-T081)
- 038: compiler/validator work; VLM _default_submit and capture dispatch
  remain runtime stubs that never call LLMClient/ScreenshotService
  (T057-T059)
- 039: UI components exist, but api/dashboard-testing.ts is unbound and
  pipeline views are not wired to pages; depends on 037 read-API
  (T054-T058)
- 040: run_load_run never invokes RunnerPool, so load runs execute zero
  Superset requests; must wire 037 executor (T075-T079)
- 041: backend index works, but no /lineage frontend and lineage_index
  stays opt-in (T045-T048)
- 036: confirmed operational, relations to reused plugin modules fixed

19 open closure tasks total; region pairs balanced.
2026-08-07 12:56:40 +07:00

12 KiB
Raw Permalink Blame History

#region DashboardLoadTesting.Research [C:3] [TYPE ADR] [SEMANTICS research,load-testing,workers,cache,gil] @BRIEF Phase 0 decisions for load execution: worker runtime, environment capacity, timing model, cache provenance, bounded result processing, persistence, and 041 integration.

Feature: 040-dashboard-load-testing | Date: 2026-07-22

R1 — Async Worker Runtime

  • Decision: One LoadRun is registered as one TaskManager task. The load plugin starts N asyncio workers reading a per-run FIFO asyncio.Queue[LoadExecutionSpec]. An environment-scoped LoadCapacityRegistry owns the load semaphore shared by all active runs. Workers check cancellation and circuit-breaker state before taking the next item; in-flight requests drain within a bounded timeout.
  • Rationale: The workload is 95–99% upstream I/O. Async workers represent execution concurrency directly, expose queue wait, and support deterministic drain. TaskManager already creates/tracks one asyncio task per plugin run and provides cancellation/events.
  • Alternatives Considered: (a) HTTP connection count as concurrency — rejected: indirect, hides queue wait and circuit-breaker boundaries. (b) TaskManager task per execution — rejected: floods Task Center/Reports and destroys run-level lifecycle. (c) threads/process workers — rejected for v1: mostly I/O; GIL mitigation is bounded result processing (R5).
  • Impact: New C5 LoadTesting.RunnerPool; no APScheduler or AsyncJobRunner for manual starts. Scheduled load callbacks may dispatch one TaskManager run through existing scheduler boundaries later.

R2 — Effective Capacity and Existing Client Semaphore Gap

  • Decision: Effective run concurrency = min(profile_requested, env.load_test_max_concurrent, max(1, env.connection_pool_size - env.load_test_reserved_slots), 25). Defaults: PROD=5; DEV/PREPROD=10; absolute ceiling=25; load_test_reserved_slots=5. Worker acquires load semaphore first, then request travels through the shared client semaphore; one global acquisition order prevents deadlock.
  • Repository Finding: Environment already has stage, is_production, connection_pool_size (default 20), and connection_pool_timeout. client_registry.get_client() creates a semaphore but currently constructs AsyncAPIClient(...) without passing semaphore=semaphore; therefore the documented per-env limit is not active.
  • Rationale: Load traffic must not occupy all shared Superset capacity and freeze ordinary UI/API operations. Capacity derives from the existing environment pool rather than an unrelated number.
  • Alternatives Considered: Separate httpx client/pool — rejected: loses shared auth/CSRF cookies; explicitly rejected by AsyncAPIClient contract. Acquiring shared semaphore manually in load runner — rejected: once registry wiring is fixed, this double-acquires the same semaphore and can deadlock.
  • Impact: Foundational prerequisite: pass registry semaphore into AsyncAPIClient; regression test proves ordinary calls and load calls share the same limit. Add config fields load_test_max_concurrent (optional derived default) and load_test_reserved_slots=5.

R3 — Timing Model and Percentiles

  • Decision: Record three non-overlapping monotonic timings per execution: queue_wait_ms (enqueued → worker takes), resource_wait_ms (worker takes → upstream request admitted), upstream_latency_ms (admitted → response/error). end_to_end_ms is their sum. Dashboard performance percentiles use upstream_latency_ms; orchestration health reports queue/resource wait separately. Percentiles use nearest-rank over successful samples; errors/timeouts have dedicated rates and are not converted into synthetic latency values.
  • Rationale: Combining waits with HTTP latency would measure our orchestrator under contention rather than Superset. Circuit breaker p99 must reflect upstream degradation; queue saturation has its own trigger/diagnostic.
  • Alternatives Considered: End-to-end p99 only — rejected: cannot attribute queue pressure vs Superset slowdown. Treat timeout as latency=timeout — rejected: biases percentile and double-counts error semantics.
  • Impact: LatencySummary stores sample_count, p50/p90/p95/p99 per timing dimension; breaker latency threshold evaluates upstream p99 only after minimum sample count.

R4 — Authoritative Cache State from Superset

  • Decision: Parse /api/v1/chart/data JSON body fields per query: is_cached, cache_key, cached_dttm, queried_dttm, cache_timeout. Mapping: hit iff is_cached=true; bypassed iff request.force=true; disabled iff cache_timeout=-1; miss iff is_cached=null && cached_dttm=null && force=false; unknown for absent/inconsistent fields. Persist raw fields plus cache_state_source="superset_response".
  • Evidence: Audited /home/busya/dev/superset/superset/common/query_context_processor.py (returns all fields), charts/data/api.py (serializes query results), charts/schemas.py (response schema), common/utils/query_cache_manager.py (hit semantics), and Superset integration tests (force/source execution produces is_cached=None).
  • Rationale: This is an authoritative execution signal from Superset, unlike latency inference.
  • Alternatives Considered: HTTP headers — rejected: chart-data uses body metadata. Probe inference from latency delta — rejected: DB buffer cache, connection reuse, and network jitter confound it.
  • Impact: 037 chart-data adapter must preserve these fields rather than normalize them away.

R5 — GIL and Bounded Result Processing

  • Decision: Load path computes SHA-256 while consuming the raw response bytes, extracts row count/cache metadata, and retains only a bounded first-N-row sample for display. It does not run full 037 value normalization. Consistency compares digest for identical (chart, filters_hash, role, time_range) coordinates. Default sample cap: 100 rows and 256 KiB serialized sample.
  • Rationale: At cap 25 and 0.5–3s upstream latency, workers are predominantly I/O-bound. The relevant GIL risk is large table JSON parsing/normalization (10k+ rows, 200–500ms continuous hold). Bounded processing keeps event-loop distortion low and ensures reported latency belongs to Superset.
  • Alternatives Considered: ProcessPool normalization — deferred; IPC cost distorts small responses. Full 037 normalization — rejected for load path; belongs to correctness verification. Hashing post-parsed normalized JSON — rejected: incurs full parse and may hide byte-level divergence.
  • Impact: If fixture instrumentation shows event-loop lag >10% of upstream latency at ceiling load, emit <ESCALATION> for a process/streaming parser ADR; do not silently add ProcessPool.

R6 — Persistence and Event Granularity

  • Decision: Persist LoadProfile, LoadRun, LoadExecution, and ConsistencyFinding in PostgreSQL. Workers buffer execution records and flush bounded batches (≤50 or ≤1s), while aggregate progress events are emitted at ≤4 Hz. Store digest/count/sample/cache/timing, never full large payload. Partial results survive stop/breaker/failure.
  • Rationale: Per-execution commit/event at concurrency 25 creates database/WebSocket amplification. Batching preserves auditability without turning the orchestrator DB into the bottleneck.
  • Alternatives Considered: Full response blobs — rejected: unbounded storage/GIL cost and duplicates Superset. In-memory results only — rejected: reconnect and cross-run comparison require durability.
  • Impact: C5 repository contract owns atomic batch insert + run aggregate update.

R7 — Circuit Breaker Semantics

  • Decision: Evaluate on a rolling completed-execution window with minimum 20 samples. Defaults: error-rate threshold 25%; upstream p99 multiplier 3× the selected comparison run/profile reference; if no reference exists, latency breaker is armed only with an explicit absolute threshold. Breach closes queue intake, transitions to draining, preserves partial results, records triggering window/metric.
  • Rationale: p99 on tiny samples is unstable; implicit latency reference would create false aborts. Error breaker works from the first minimum window.
  • Alternatives Considered: Break on first timeout — rejected: transient failures are expected load data. Use end-to-end p99 — rejected by R3 attribution model.
  • Impact: UI must show breaker inactive|armed|tripped and why latency breaker may be inactive.

R8 — Variation Matrix Expansion

  • Decision: Closed axes: filters, viewport, role, time_range. Normalize axis ordering and values, then cartesian-expand if total ≤ profile cap (default 500 combinations). Above cap, deterministic seeded reservoir sampling without materializing the full product. Each coordinate gets a stable variation id from canonical JSON hash.
  • Rationale: Determinism supports reproducible comparisons; bounded sampling prevents combinatorial memory/request explosion.
  • Alternatives Considered: Always cartesian — rejected: unbounded expansion. Random sampling without seed — rejected: cross-run comparison becomes invalid.
  • Impact: Preview shows theoretical combination count, selected count, seed, chart count, total requests.

R9 — 041 Blast-Radius Read Model

  • Decision: Profile validation calls 041 dependents endpoint with pinned fingerprint. LoadRun persists that fingerprint and the resolved BlastRadiusReport. Start rejects a changed fingerprint with a refresh/reconfirm response, especially before PROD approval. Mid-run index updates do not mutate the run.
  • Rationale: Configure-time and gate-time impact must name the same dependent dashboards; otherwise user approval is not bound to actual impact.
  • Impact: 040 cannot be implementation-complete before 041 LIN-FR-013 read contract exists; fixtures unblock independent domain tests.

R10 — PROD Gate and RBAC

  • Decision: Stage source of truth is Environment.stage (PROD) with is_production treated as backward-compatible alias; either marks PROD. PREPROD/DEV require dashboard:loadtest:execute; PROD additionally requires dashboard:loadtest:prod and 036 approval gate reason. Gate hash binds profile revision, effective cap, estimated request count, environment id, and 041 fingerprint.
  • Rationale: Existing configuration already exposes stage; dual interpretation avoids silently under-classifying legacy is_production=true environments.
  • Impact: Start-after-approval revalidates all hash inputs; mutation/stale gate dispatches zero requests.

Resolved Risks

Risk Mitigation
Load starves ordinary Superset calls Effective cap reserves 5 shared slots; registry semaphore wiring fixed (R2)
Double semaphore acquisition deadlock Single order: load semaphore then request-internal shared semaphore; never manually acquire shared semaphore
Event loop distortion on large tables Raw digest + bounded sample; no full normalization (R5)
Cache state misclassification Authoritative response body + force provenance (R4)
Per-execution DB/event amplification Batches ≤50/1s; progress ≤4 Hz (R6)
Stale blast-radius approval Fingerprint bound into gate hash; start revalidates (R9/R10)

R11 — MVP Runtime Gap (audit 2026-08-07)

  • Decision: run_load_run MUST instantiate RunnerPool(capacity=effective_concurrency, executor=037-query-executor, env_semaphore=shared, breaker=run_breaker, on_result=persist) and drain queued executions through it. The current hold/sleep body (фазовый переход без запросов) is a stub and violates LOAD-FR-001.
  • Rationale: RunnerPool уже реализован и покрыт тестами; отсутствует только wiring в production-путь. Подключение 037 SupersetClient.ChartData.Execute через адаптер services/load_testing/executor.py сохраняет единый источник metric truth и bounded-normalization (LOAD-FR-018).
  • Alternatives considered: in-process pool без shared semaphore (rejected — R2), прямой httpx в load-пути (rejected — LOAD-FR-001/037 continuity), запуск load как 036 AgentRun (rejected — LOAD spec).
  • Impact: T075–T079 в tasks.md Phase 9; quickstart шаги 3–7 не считаются пройденными до реального исполнения executor'а.

#endregion DashboardLoadTesting.Research