docs(reports): add ss-prod agentic E2E gap analysis and task notes
Add 2026-09-08 production gap/spec-coverage/refresh-plan and dashboard scenario E2E plan/report, axiom-mcp feedback, and 2026-09-11 data-team maintenance dev-test run notes.
This commit is contained in:
110
docs/reports/axiom-mcp-feedback-2026-09-10.md
Normal file
110
docs/reports/axiom-mcp-feedback-2026-09-10.md
Normal file
@@ -0,0 +1,110 @@
|
||||
# AXIOM MCP — обратная связь потребителя: ортогональные оценки полезности и удобства
|
||||
|
||||
**Дата оценки:** 2026-09-10 (записано 2026-09-11)
|
||||
**Оценщик:** implementation/orchestration-агент сессии «agentic-runtime» (Front 6 Track A — INV_7-сплиты и семантическая курация; Front 6 Track B — live-chain).
|
||||
**Объект:** AXIOM MCP server, реально использованное подмножество — `axiom_search` (ops `search_contracts`, `workspace_health`, `rebuild`) и `axiom_audit` (`detect_missing_contracts`).
|
||||
**Объём выборки:** ~11 вызовов за сессию (5 + 2 + 2 + 2), 2 связанных workstream, один репозиторий (`/home/busya/dev/ss-tools`).
|
||||
**Версия сервера:** в ответах не экспонировалась; зафиксировать не удалось (это само по себе — находка, ось 3/7).
|
||||
**Оговорка:** оценки отражают наблюдения одной сессии на одном кодовом срезе; это не бенчмарк. Категории ортогональны (независимы); шкала 1–5.
|
||||
|
||||
---
|
||||
|
||||
## 1. Журнал использования (что вызывалось, зачем, с каким результатом)
|
||||
|
||||
| Op | Кол-во | Задача в сессии | Наблюдавшийся результат |
|
||||
|---|---|---|---|
|
||||
| `search_contracts` | 5 | Найти точные контракты-цели для 3 orphan-связей: `BaselineEngine.Catalog.Materialization`, `McpServer.ScenarioTools.register_draft_pack`, `ScenarioGraph.Models.CanonicalBytes`; подтвердить существование модулей-замен | Возвращал `contract_id` + `contract_type` + `file:line`. Позволил точно исправить `@RELATION` (T3) без ручного разбора анкеров |
|
||||
| `workspace_health` | 2 | Измерить эффект курации: unresolved/orphan до и после | `unresolved_relations 409 → 406` (ровно −3, совпало с числом правок); orphan-счётчик вырос 2809 → 2818 (новые модули) |
|
||||
| `rebuild` | 2 | Пост-правковый гейт после 6-модульного сплита и INV_1-обёрток | `0 parse warnings`; `edges 5352 → 5366 → 5369`; `contract_count ~10956 → 10961`; число изменённых файлов |
|
||||
| `detect_missing_contracts` | 2 | INV_1-аудит по `execution/` до/после обёртки naked-функций | Per-file таблица «contracted/naked»: 36 → 32 naked; мои 4 функции закрылись, остаток — pre-existing чужих модулей |
|
||||
|
||||
Итоговое состояние после сессии (для контекста): полная backend-суита `11527 passed / 0 failed`; индекс пересобран, 0 parse warnings.
|
||||
|
||||
---
|
||||
|
||||
## 2. Ортогональные категориальные оценки
|
||||
|
||||
| # | Ось (ортогональная) | Оценка | Ключевое основание |
|
||||
|---|---|---:|---|
|
||||
| 1 | Навигация: поиск контрактов и ID | **5/5** | Точные ID+`file:line` за один запрос; основа недорогого T3 вместо ручной археологии |
|
||||
| 2 | Достоверность сигналов | **5/5** | Стабильные ID между запросами; численный эффект правок совпал точно (−3 unresolved) |
|
||||
| 3 | Freshness / наблюдаемость поколения индекса | **2/5** | Нет `generation_id`/`built_at`/delta; инжектированный `root.map` заявил `@CONTRACTS 10561`, `rebuild` — ~10956–10961; неотличимо от учётной семантики |
|
||||
| 4 | Эргономика mutation-loop | **4/5** | `rebuild` быстрый, отдаёт `0 warnings` — хороший пост-правковый гейт; нет композита и contract-diff |
|
||||
| 5 | Точность аудитов/гардрейлов | **3/5** | Полезен, но флагует легитимные функции внутри `[TYPE Module]`-регионов как naked → косметические правки; `workspace_health` не перечисляет unresolved |
|
||||
| 6 | Экономика ответов / signal-to-noise | **3/5** | `include_body=false` по умолчанию полезен; крупные ответы вытеснялись из контекста → перезапросы; нет «только ID+file:line» режима |
|
||||
| 7 | Ясность поверхности / discoverability | **2/5** | `mcp_instructions` описывают task-shaped ops (`semantic_discovery`, `contract_patch`…), которых нет среди фактических инструментов (`axiom_search`/`axiom_audit` + enum) |
|
||||
| 8 | Интеграция с workflow docs/nav | **3/5** | `rebuild` обновляет индекс, но не `docs/api/nav|html`; после сплита tracked-доки устарели; единой regen-операции нет |
|
||||
|
||||
**Сводно:** сильные стороны — аналитика и измеримость (оси 1–2, 4 частично). Слабые — наблюдаемость состояния и онбординг (3, 7), плюс шероховатости аудита/экономики (5, 6).
|
||||
|
||||
---
|
||||
|
||||
## 3. Детализация по осям
|
||||
|
||||
### Ось 1 — Навигация (5/5)
|
||||
- **Что сработало:** `search_contracts` по подстроке («Materialization», «CanonicalBytes», «register_draft_pack») мгновенно давал `contract_id`, `contract_type`, `file:line`. Именно так были подтверждены точные цели связей: `BaselineEngine.Candidates.Materialization` (модуль), `McpServer.ScenarioTools` (block-владелец инструмента), `ScenarioGraph.Models.DashboardTestScenario.CanonicalBytes` (функция).
|
||||
- **Ценность:** без сервера — ручной проход `#region`/`@RELATION` по десяткам файлов; с сервером — 3 точных запроса.
|
||||
|
||||
### Ось 2 — Достоверность (5/5)
|
||||
- Идентификаторы и типы контрактов стабильны при повторе; `detect_missing_contracts` per-file таблицы совпали с фактическим состоянием кода, который я читал вручную.
|
||||
- Численный контроль: 3 исправленные связи → `unresolved_relations` ровно −3. Это дало проверяемый feedback-loop для курации.
|
||||
|
||||
### Ось 3 — Freshness / поколение индекса (2/5)
|
||||
- **Проблема:** ответы не содержат признака «насколько индекс свежий». Инжектированный в стартовый контекст `docs/api/nav/root.map` заявлял `# @CONTRACTS 10561 @MODULES 196 @TESTS 49%`, тогда как `rebuild` сообщал суммарный `contract_count` ~10956–10961. Я **не смог однозначно отличить** staleness от учётной семантики (`@TESTS 49%` collapse — возможно, `@CONTRACTS` считает только не-тестовые контракты).
|
||||
- AGENTS.md сам предупреждает: «Это снепшот на момент старта сессии: после крупных мутаций перечитать…» — но API не сигнализирует, что я читаю устаревшее.
|
||||
- **Эффект:** приходилось переспрашивать/пересобирать, чтобы быть уверенным в актуальности; часть времени — на верификацию состояния, а не на задачу.
|
||||
|
||||
### Ось 4 — Эргономика mutation-loop (4/5)
|
||||
- `rebuild` быстрый (инкрементальный) и сразу даёт жёсткий гейт `0 parse warnings` — идеально после правок анкеров.
|
||||
- **Чего не хватило:** (а) композитной «проверь после правки» (rebuild + parse-warnings + naked-delta + orphan-delta + unresolved-delta одним вызовом); (б) contract-diff «до/после» — какие `contract_id` появились/изменились/исчезли, а не только «changed files: 22».
|
||||
|
||||
### Ось 5 — Точность аудитов (3/5)
|
||||
- **Полезно:** `detect_missing_contracts` корректно и наглядно показал INV_1-долг по `execution/` (36 naked), а после моих обёрток — 32, причём мои модули стали 0.
|
||||
- **Проблема-1 (ложноположительный класс):** функция `_advance_run`, лежащая в регионе `#region ScenarioExecution.Runner.Walker [TYPE Module]`, была помечена как naked. Формально регион оборачивает функцию, но типизирован как Module → чекер требует смены типа. Мне пришлось сделать **чисто косметическую** правку `[TYPE Module] → [TYPE Function]`, чтобы удовлетворить аудит. Это конфликтует с духом INV_9 (не добавлять синтетику ради зелёного аудита) и съедает доверие к сигналу.
|
||||
- **Проблема-2:** `workspace_health` сообщает только счётчик unresolved (406), без перечня «owner_file:line → target». Три конкретные связи я нашёл вручную, кросс-сверяя результаты `search_contracts`.
|
||||
|
||||
### Ось 6 — Экономика ответов / SNR (3/5)
|
||||
- `include_body=false` по умолчанию — правильный выбор, тела контрактов не заливают контекст.
|
||||
- Но даже без тел ответы `search_contracts` крупные (много контрактов + метаданные + relations), и в моей среде часть tool-результатов вытеснялась («content cleared») → повторные запросы с уточнением. Компактный режим «только `contract_id` + `file:line` + краткий purpose» (или cap/`fields=[...]`) заметно снизил бы расход.
|
||||
|
||||
### Ось 7 — Ясность поверхности / discoverability (2/5)
|
||||
- **Мismatch инструкции и реальности:** `mcp_instructions` (в стартовом контексте) описывают task-shaped поверхность — `semantic_discovery`, `semantic_context`, `semantic_validation`, `contract_patch`, `contract_refactor`, `contract_metadata`, `testing_support`, `runtime_evidence`, `workspace_artifact/path/policy/command/checkpoint`, `security_workflow`. Фактически доступны инструменты `axiom_search` / `axiom_audit` / `axiom_docs` с enum-операциями. Мне пришлось вручную строить соответствие «что мне нужно → какая operation».
|
||||
- Рекомендованные стартовые ресурсы `axiom://project/tools` и `axiom://protocols/tool-selection` автоматически не подсовываются; я их не читал (мой недосмотр, но и сигнал: нет «push» в нужный момент).
|
||||
- Версия сервера/протокола в ответах не видна — нельзя приложить к отчёту.
|
||||
|
||||
### Ось 8 — Интеграция с workflow docs/nav (3/5)
|
||||
- `rebuild` перестраивает DuckDB-индекс, но **не** `docs/api/nav/**` и `docs/api/html/**`. После моего 6-модульного сплита (runner→walker/start_run/…) trackнутые nav-доки остались бы устаревшими без отдельного `make docs-nav`.
|
||||
- Двойная поверхность doc-gen (CLI `doc-gen --nav/--html` vs MCP `axiom_docs emit_nav_graph`) без единой «пересобрать всё» операции повышает риск рассинхрона индекс↔доки.
|
||||
|
||||
---
|
||||
|
||||
## 4. Топ-предложения по улучшению (конкретные)
|
||||
|
||||
1. **Метаданные свежести в каждом ответе:** `index_generation`, `built_at`, `contract_count`, `changed_since`. Отдельная op `status`. Снимает неоднозначность «10561 vs 10956» и предупреждает о работе по устаревшему индексу.
|
||||
2. **`workspace_health` с перечнем:** для каждого unresolved — `owner_contract_id`, `owner_file:line`, `target_id`, `target_exists: bool`. Сейчас счётчик без адресов заставляет искать вручную.
|
||||
3. **Композит `verify_after_edit(paths[])`** → `{parse_warnings[], naked_delta, orphan_delta, unresolved_delta, new_contracts[]}` — один вызов вместо 3–4.
|
||||
4. **`rebuild` с contract-diff:** added/changed/removed `contract_id` (а не только changed files). Это превращает гейт из «индекс пересобран» в «вот что изменилось семантически».
|
||||
5. **Согласовать `mcp_instructions` с фактическими tool-именами** (или перечислить оба слоя явно), плюс уточнить правило аудита для функций внутри `[TYPE Module]`-регионов — сейчас оно провоцирует косметические правки типа региона.
|
||||
6. **Единая связка index ↔ nav:** либо `rebuild` с флагом `--docs`, либо явная одна команда, чтобы исключить рассинхрон.
|
||||
|
||||
---
|
||||
|
||||
## 5. Методические замечания (self + tool-selection)
|
||||
|
||||
- **Не использовал `impact_analysis`** перед 6-модульным сплитом `runner.py` — а стоило: blast-radius переноса контрактов и `local_context` на seed-контракт дали бы страховку. Описание инструмента не подталкивает к pre-refactor проверке; в `tool-selection`-гайде стоило бы явно прописать «перед переносом контрактов — `impact_analysis` + `local_context`».
|
||||
- **Не прочитал `axiom://project/tools` и `axiom://protocols/tool-selection`** (мой недосмотр). При этом сами по себе они не всплывают в потоке — discoverability можно улучшить.
|
||||
- **Оценки 3/5 по freshness и 2/5 по ясности поверхности** — не про качество анализа, а про наблюдаемость/онбординг: сервер делает правильные вещи, но недостаточно сообщает о своём состоянии и не совпадает с собственным описанием поверхности.
|
||||
|
||||
---
|
||||
|
||||
## 6. Покрытие: что НЕ использовалось (честность выборки)
|
||||
|
||||
Не вызывались: `ast_search`, `read_outline`, `describe_tags`, `local_context`, `task_context`, `hybrid_query`, `trace_related_tests`, `scaffold_tests`, `map_trace_to_contracts`, `read_events`, `summarize`/`diff`/`rollback_preview` (checkpoints), `policy`, `status`, `server_metrics`, `axiom_docs` (все ops), `axiom_audit` ops кроме `detect_missing_contracts` (`audit_contracts`, `audit_belief_protocol`, `audit_belief_runtime`, `diff_contract_semantics`, `impact_analysis`, `scan`). Причины: не было задачи в этом срезе; часть требовала бы отдельного контекста; часть не всплыла из описания под мою задачу. Оценки выше — только по использованному подмножеству.
|
||||
|
||||
---
|
||||
|
||||
## 7. Рекомендация
|
||||
|
||||
**Использовать снова — да**, в первую очередь для: (а) навигации по контрактам и точных ID (ось 1 — реально незаменимо), (б) численного контроля семантической курации (ось 2), (в) пост-правкового гейта `rebuild` (ось 4).
|
||||
|
||||
**Пока не полагаться** на: freshness-сигналы (ось 3) и на самодостаточность описания поверхности (ось 7) — их нужно либо доделать, либо компенсировать процедурно (явный `status`/`rebuild` перед доверием к индексу, сверка с `docs/api/nav` после сплитов).
|
||||
239
docs/reports/ss-prod-agentic-e2e-baseline-gap-2026-09-08.md
Normal file
239
docs/reports/ss-prod-agentic-e2e-baseline-gap-2026-09-08.md
Normal file
@@ -0,0 +1,239 @@
|
||||
# SS-Prod Agentic Dashboard E2E — baseline gap audit (2026-09-08)
|
||||
|
||||
#region Report.SsProdAgenticE2EBaselineGap [C:4] [TYPE Verification] [SEMANTICS baseline,visual,metric,performance,scenario,agent-evaluation]
|
||||
|
||||
## Executive verdict
|
||||
|
||||
Baseline support is **substantial but not ScenarioRun-ready**.
|
||||
|
||||
- **A. Visual screenshot baseline:** 037 defines and implements release-pinned `VisualBaselineEntry`, exact/perceptual comparison, authoritative staleness fingerprints, durable expected-image bytes and gated publication. The compare path is real and unit-verified. Candidate capture/review, lifecycle and ScenarioRun ownership are incomplete or drifted.
|
||||
- **B. Numerical/dashboard metric baseline:** this is the strongest path. 037 implements authoritative Superset-native capture, stores raw response bytes, normalized raw/canonical values, release/filter/query provenance, tolerances, immutability detection and a three-stage approval gate. However catalog supersession/rebaseline/invalidation and ScenarioRun pinning are not closed.
|
||||
- **C. Execution performance/quality baseline:** **no approved baseline entity exists or is specified**. 044/045 persist timestamps and compare run outputs; 047 specifies rolling health/flakiness/trends. These are derived projections, not versioned approved performance baselines, and current analytics materially under-implements its own context contract.
|
||||
|
||||
Across **18 audited baseline contract units**: **COMPLETE 5, PARTIAL 5, MISSING 3, DRIFTED 5**. The prior minimum estimate needs a non-overlapping delta of **2 new modules**, **6 existing modules modified**, **580–1,000 production LOC**, **900–1,450 test LOC**, and **2.7–4.7 engineer-weeks** to make existing 037 baselines authoritative inputs to ScenarioRun. No extra screenshot comparator or metric engine is needed.
|
||||
|
||||
## Scope and evidence
|
||||
|
||||
Read-only audit of specs 037, 038, 044, 045 and 047, their data models/contracts/tasks/traceability, and production schemas/models/routes/services/tests. No production request or mutation was made.
|
||||
|
||||
Local verification:
|
||||
|
||||
- `102 passed` in `2.97s`: baseline catalog, visual comparison/lifecycle/executor, metric async compare and approval lifecycle edge suites.
|
||||
- API/persistence/MCP parity subset did not complete a test within a bounded 25-second run and was terminated. It is **not** reported as passing or failing.
|
||||
- Existing ss-prod scenario evidence did not execute baseline candidate publication or real visual comparison. “Implemented” below means source + passing local service tests unless explicitly stated otherwise.
|
||||
|
||||
Classification is the same as the parent spec-coverage report: `COMPLETE`, `PARTIAL`, `IMPLIED`, `MISSING`, `DRIFTED`. `tasks.md` is not standalone proof.
|
||||
|
||||
## A. Visual screenshot baseline
|
||||
|
||||
### Normative contract
|
||||
|
||||
`VisualBaseline` is explicitly a distinct expected visual state, never an LLM verdict (`037/spec.md:94-109`; `contracts/modules.md:183-215`):
|
||||
|
||||
- identity/context: `baseline_id`, release version + commit, dashboard, normalized filters, tab, optional ROI;
|
||||
- expected evidence: `expected_image_sha256`, optional opaque `expected_image_content_ref`, capture time;
|
||||
- staleness: query/dataset/filter/layout fingerprints;
|
||||
- compare policy: exact digest or perceptual `ssim_min` + `pixel_diff_threshold`;
|
||||
- lifecycle: reviewed screenshot → draft candidate → 036 approval gate → approved catalog entry;
|
||||
- outcomes: pass/fail/inconclusive/`stale_visual_baseline`/immutability violation;
|
||||
- inheritance and status: approved/superseded/retired under the same release rules as metric entries.
|
||||
|
||||
Normative anchors: `037/data-model.md:125-131`, `contracts/baseline-catalog.schema.json:66-135`, `contracts/dashboard-testing.openapi.yaml:746-780`, `contracts/modules.md:183-215`, `tasks.md:66-76`.
|
||||
|
||||
### Runtime implementation
|
||||
|
||||
- DTO and YAML round-trip exist: `schemas/dashboard_testing/catalog.py:108-157`, `baseline_catalog.py:61-126,240-290,297-378`.
|
||||
- Expected bytes are resolved from opaque DraftStorage refs, not a caller path: `visual_executor_async.py:38-159`.
|
||||
- Actual screenshot and expected screenshot must be from different AgentRuns; release/environment/catalog are server-resolved and fingerprints are recomputed: `visual_executor_async.py:204-323`.
|
||||
- Exact/SSIM, pixel threshold, layout/query/dataset/filter staleness and closed-period immutability exist in `visual_baseline.py:35-379` and `visual_ssim.py`.
|
||||
- Only approved entries are selected; superseded/retired are skipped: `catalog_queries.py:56-94`.
|
||||
|
||||
### Gaps and classification
|
||||
|
||||
- `VisualBaselineEntry` entity/schema: **COMPLETE**.
|
||||
- exact/perceptual compare + staleness: **COMPLETE**.
|
||||
- capture/review candidate: **DRIFTED**. Spec requires a reviewed screenshot and `ReviewDisposition` (`contracts/modules.md:202-212`), but REST `CandidateRequest` accepts caller-carried `approval`, digest/content ref/fingerprints (`schemas/dashboard_testing/candidates.py:48-114`). Service verifies artifact ownership/kind/digest (`candidates.py:97-110`) but does not verify a durable human review disposition. There is no authoritative visual capture endpoint analogous to metric `/baseline-candidates/capture`.
|
||||
- approve/materialize: **PARTIAL**. Durable gate and atomic YAML materialization exist, but the 037 module contract says git commit/push (`contracts/modules.md:132-142`) while `approvals.py:290-385`/`materialization.py:94-187` only write the working-tree YAML + DB state. MCP `consume_baseline_approval` is a fallback envelope, not REST parity (`mcp_server/tools_review.py:232-238`).
|
||||
- invalidate/rebaseline/retention lifecycle: **MISSING**. Status enum exists, but no explicit transition API/service defines supersede/retire/invalidate/rebaseline, byte-retention proof or deletion audit.
|
||||
|
||||
### Raw values and provenance
|
||||
|
||||
The approved visual entry retains baseline image bytes by opaque content ref, SHA-256, release identity, dashboard/filter/tab/ROI, four staleness fingerprints, actor/AgentRun and capture time (`catalog.py:119-149`). It does **not** normatively retain MIME, viewport, pixel dimensions, capture method/readiness profile or screenshot provider/version. Model/prompt provenance is intentionally not part of deterministic visual truth: 037 explicitly rejects LLM/VLM as visual baseline authority (`contracts/modules.md:197-199`). Those fields belong on `AgentEvaluation`, not `VisualBaseline`.
|
||||
|
||||
## B. Numerical/dashboard metric baseline
|
||||
|
||||
### Normative contract
|
||||
|
||||
`BaselineEntry` is a release-pinned expected `NormalizedValue` stored in repository YAML (`037/data-model.md:58-123`):
|
||||
|
||||
- chart or dataset + result key + normalized filters/time-range and filter hash;
|
||||
- redacted bounded raw value, canonical value, display/format and source metadata;
|
||||
- full Superset response SHA-256 and capture timestamp;
|
||||
- exact, absolute, relative, range and row-set comparison policies;
|
||||
- optional closed-period immutability block;
|
||||
- approved/superseded/retired state, provenance and release inheritance;
|
||||
- non-pass results preserve actual/expected/diff/policy/warnings and emit investigation signals.
|
||||
|
||||
API lifecycle is declared in `037/contracts/dashboard-testing.openapi.yaml:59-203`: list, create/capture candidate, request gate, decide gate, consume gate, compare, and VerificationRun create/read.
|
||||
|
||||
### Runtime implementation
|
||||
|
||||
- `/query-model`, `/filters/normalize`, `/queries/execute`, `/comparisons`, `/baselines` are RBAC-protected routes (`api/routes/dashboard_testing/core.py:65-206`).
|
||||
- Authoritative metric capture derives environment/repository/dashboard from approved release, executes a fresh Superset query model, hashes raw response bytes, stores the raw bytes in DraftStorage and registers `capture_execution` metadata (`candidate_capture.py:48-255`).
|
||||
- Direct metric candidate creation requires the server-issued capture artifact; coordinates, canonical expected value and stored-byte digest are reverified (`candidate_provenance.py:26-165`).
|
||||
- Gate request hash binds candidate content/path/operation/release/commit/period; decide and consume revalidate it (`candidate_guards.py:26-97`; `approvals.py:46-225,255-385`).
|
||||
- Catalog write uses per-catalog locking, atomic replacement and version-guarded compensation (`materialization.py:94-187`). Closed periods reject silent reclosure and compare response digests (`candidate_guards.py:186-249`; `immutability.py`).
|
||||
- Metric verification resolves an approved catalog entry and real deployment environment, executes a Superset-native envelope, preserves the full response hash and applies deterministic comparison (`metric_executor_async.py:232-367`).
|
||||
|
||||
### Gaps and classification
|
||||
|
||||
- metric entity/schema: **COMPLETE**.
|
||||
- authoritative raw capture/provenance: **COMPLETE**.
|
||||
- compare/tolerance/closed-period immutability: **COMPLETE**.
|
||||
- approval/publication: **DRIFTED**. REST gate/materialization is implemented, but no actual git commit/push occurs despite `BaselineEngine.Release.ApproveBaseline` contract; MCP consume is not parity. OpenAPI `BaselineEntry` requires `fingerprints`/`approval` (`contracts/dashboard-testing.openapi.yaml:676-707`) while the runtime metric DTO does not have those fields (`schemas/dashboard_testing/catalog.py:50-80`). Candidate request/response shapes also diverge from the checked-in OpenAPI.
|
||||
- version/inheritance: **PARTIAL**. Release version/commit and inheritance code exist, but baseline identity is not a first-class `BaselineRevision`; catalog version is only file content/Git state. Multiple approved entries for one coordinate are not prohibited. Duplicate `baseline_id` is silently skipped without verifying identical content (`baseline_catalog.py:357-362`), and lookup returns the first matching approved entry (`catalog_queries.py:37-49`).
|
||||
- invalidate/rebaseline: **MISSING**. There is no CAS transition that supersedes the prior approved coordinate, creates its successor and leaves an explicit audit edge.
|
||||
|
||||
### Raw values and provenance
|
||||
|
||||
The capture artifact retains the **full raw Superset response bytes** plus digest, byte size, environment/dashboard/chart/dataset/result, normalized filters, normalized raw/canonical value, AgentRun and release (`candidate_capture.py:103-149,210-245`). The approved catalog retains the redacted `raw_value`, canonical/display/format/source, expected value, response hash, capture time, release, chart/dataset/result/filter and actor/AgentRun (`schemas/dashboard_testing/results.py:17-31`; `catalog.py:50-72`).
|
||||
|
||||
Limitations:
|
||||
|
||||
- full raw bytes remain owned by an AgentRun DraftArtifact; `BaselineEntry` has no capture-artifact/content ref or retention guarantee tying that raw object to the approved entry;
|
||||
- query provenance is partly an unstructured `NormalizedValue.source` string plus response hash, rather than a typed query-envelope/version record;
|
||||
- time range is preserved only insofar as it is present in normalized filters;
|
||||
- screenshot digest/provider/model/prompt are not applicable to metric baseline truth.
|
||||
|
||||
## C. Execution performance/quality baseline
|
||||
|
||||
There is no normative `PerformanceBaseline`, `QualityBaseline`, approved threshold set, candidate/approval lifecycle or versioned SLO comparison.
|
||||
|
||||
What exists:
|
||||
|
||||
- `ScenarioRun`/`ScenarioStepRun` persist created/started/finished timestamps, attempts, status/outcome and artifacts (`models/scenario_run.py:34-87`); duration is derived.
|
||||
- 045 specifies run-to-run per-step latency/value deltas (`045/data-model.md:19-25`). Runtime `execution/comparison.py:28-73` compares status and output hashes but does **not** calculate latency/value deltas.
|
||||
- 047 specifies rolling 30-run flakiness and contextual 30-day quality ratios, not approved baselines (`047/data-model.md:30-40`; SCAN-FR-004/005/007/010/013 in `spec.md:77-87`).
|
||||
- Runtime analytics reduce context to `scenario_id:environment_id` and only compute usable runs/transitions/success rate (`analytics/flakiness.py:9-45`), violating the specified 044 SHA-256 `AnalyticsContextKey` and omitting baseline family, release/dashboard/lineage/principal, agent disagreement, low confidence and model instability.
|
||||
|
||||
Classification:
|
||||
|
||||
- persisted timing inputs: **PARTIAL**;
|
||||
- rolling quality/flakiness analytics: **DRIFTED**;
|
||||
- approved/versioned execution-performance baseline: **MISSING / not currently specified**.
|
||||
|
||||
An approved performance baseline is not required by current 037/044/047 specifications. If the product wants one, it requires a new normative feature decision rather than overloading rolling health thresholds.
|
||||
|
||||
## ScenarioRun binding and agent-evaluation integration
|
||||
|
||||
### Required behavior
|
||||
|
||||
038 provides typed `baseline.{baseline_id}` roots and says runtime preflight resolves baselines (`038/data-model.md:115-131`). 044 requires RunnerPlan and ScenarioRun to pin baselines, and result provenance to carry baseline revision (`044/data-model.md:9,66,236-238`). 047 grouping requires the exact baseline family inside `AnalyticsContextKey` (`047/data-model.md:32-40`).
|
||||
|
||||
For a visual/semantic step the safe order is:
|
||||
|
||||
```text
|
||||
approved 037 baseline entry + pinned release/catalog digest
|
||||
-> deterministic exact/SSIM or metric comparison
|
||||
-> immutable ComparisonResult evidence
|
||||
-> optional AgentEvaluation(input manifest includes screenshot + baseline + comparison refs)
|
||||
-> DecisionPolicy(evaluation + deterministic comparison)
|
||||
-> StepOutcome
|
||||
-> 047 analytics keyed by baseline family + model/prompt when applicable
|
||||
```
|
||||
|
||||
The LLM may explain or find semantic differences, but cannot replace deterministic baseline comparison or publish/rebaseline an expected value.
|
||||
|
||||
### Actual ScenarioRun gaps
|
||||
|
||||
- `derive_runner_plan()` copies an unvalidated `graph["baselines"]` blob (`runner_plan.py:125-141`); it does not resolve 037 catalog/release/approved status/digest.
|
||||
- launch `baseline_set` is written only to `target_snapshot` (`runner.py:234-250`) and is absent from `_request_hash()` inputs (`runner.py:212-215`). An idempotency replay with a changed baseline set can return the existing run.
|
||||
- generic assertion consumes embedded `expected` and maps an incomplete comparison policy (`executors.py:115-160`); it does not call the 037 catalog resolver or retain baseline ID/release/source hash.
|
||||
- result provenance derives one `baseline_revision` heuristically from the first runner-plan mapping (`execution/result.py:51-66`).
|
||||
- `AgentEvaluation`/DecisionPolicy are not implemented, so baseline/comparison refs cannot yet enter an immutable evaluation manifest.
|
||||
- runtime analytics invent a context key instead of using the specified run/step `AnalyticsContextKey`.
|
||||
|
||||
Classifications:
|
||||
|
||||
- ScenarioRun baseline resolve/pin/idempotency: **DRIFTED**;
|
||||
- baseline → AgentEvaluation/DecisionPolicy contract: **PARTIAL** normatively, runtime missing;
|
||||
- 045 baseline/result/run comparison projection: **PARTIAL**;
|
||||
- 047 baseline-family analytics binding: **DRIFTED**.
|
||||
|
||||
## Immutability, silent overwrite, CAS and RBAC
|
||||
|
||||
Strengths:
|
||||
|
||||
- approved entries are release version + Git commit pinned;
|
||||
- request hash freezes candidate content, path, release, commit and closure intent across request/decide/consume;
|
||||
- one-shot gate FSM and `dashboard:testing APPROVE` protect REST decisions/consume;
|
||||
- per-file locking, atomic replace and version-guarded compensation prevent stale rollback from erasing a concurrent write;
|
||||
- closed-period source-response changes are critical immutability violations, never auto-updates;
|
||||
- only `approved` entries participate in lookup; actual visual evidence must be independent of baseline capture.
|
||||
|
||||
Remaining silent-state risks:
|
||||
|
||||
- no catalog revision/CAS token is exposed to a publication client;
|
||||
- no coordinate uniqueness or explicit current-baseline pointer exists;
|
||||
- same `baseline_id` with different content is silently skipped rather than conflict-rejected;
|
||||
- no first-class supersede/retire/rebaseline transition/audit relation exists;
|
||||
- contract says git commit/push, implementation only writes working-tree YAML;
|
||||
- visual `approval` metadata is caller-carried before the real gate;
|
||||
- MCP consume advertises a tool but delegates to fallback instead of performing REST-equivalent materialization;
|
||||
- ScenarioRun idempotency does not bind `baseline_set`.
|
||||
|
||||
## Orthogonal count matrix
|
||||
|
||||
| Classification | Contract units | Count |
|
||||
|---|---|---:|
|
||||
| COMPLETE | visual entity/schema; visual compare/staleness; metric entity/schema; metric authoritative capture/provenance; metric compare/tolerance/immutability | 5 |
|
||||
| PARTIAL | visual approve/materialize; metric version/inheritance; persisted timing inputs; baseline→AgentEvaluation/DecisionPolicy contract; 045 result/run comparison projection | 5 |
|
||||
| MISSING | visual invalidate/rebaseline/retention lifecycle; metric invalidate/rebaseline lifecycle; approved performance/quality baseline | 3 |
|
||||
| DRIFTED | visual capture/review; metric approval/publication; 047 quality analytics; ScenarioRun baseline pin/idempotency; 047 baseline-family context binding | 5 |
|
||||
| **Total** | | **18** |
|
||||
|
||||
## Delta to previous production estimate
|
||||
|
||||
Module definition is unchanged: one cohesive production source file/component owning one runtime contract. The following is **additional**, because the previous seven new modules covered browser/screenshot bytes/AgentEvaluation/DecisionPolicy/UI but did not include 037 catalog resolution.
|
||||
|
||||
### Required minimum delta
|
||||
|
||||
| Work package | New modules | Existing modules modified | Prod LOC | Test LOC | Eng-weeks |
|
||||
|---|---:|---:|---:|---:|---:|
|
||||
| `ScenarioBaselineResolver` — approved entry/release/catalog digest resolution, baseline family, request/plan pinning | 1 | 3 | 280–480 | 400–650 | 1.3–2.2 |
|
||||
| `ScenarioVisualBaselineAdapter` — ScenarioArtifact-owner bridge to existing 037 exact/SSIM comparator and immutable ComparisonResult | 1 | 2 | 180–320 | 300–500 | 0.8–1.5 |
|
||||
| Result/analytics binding — carry exact baseline IDs/digests through result, AgentEvaluation manifest and 047 context | 0 | 1 | 120–200 | 200–300 | 0.6–1.0 |
|
||||
| **Required delta** | **2** | **6** | **580–1,000** | **900–1,450** | **2.7–4.7** |
|
||||
|
||||
No migration is intrinsically required for this minimum if the immutable pin stays in existing RunnerPlan/JSON snapshots. Add a migration only if baseline pin/query fields are normalized into columns.
|
||||
|
||||
### Hardened lifecycle delta
|
||||
|
||||
Beyond the required minimum, a first-class `BaselineCatalogRevision/CAS` owner plus explicit supersede/rebaseline audit would add **1 new module**, modify **4 existing modules**, add **300–550 prod LOC**, **500–850 test LOC**, and **1.5–2.7 engineer-weeks**. This is not double-counted with the previous H3 retention or H5 analytics packages; it covers only catalog concurrency/current-entry semantics.
|
||||
|
||||
An approved performance baseline would be a separate optional feature, not part of the current estimate: approximately **2 new modules**, **350–650 prod LOC**, **550–900 test LOC**, one policy-storage migration and **1.8–3.2 engineer-weeks**, after a new spec defines whether the authority is SLO policy, release baseline or rolling statistical envelope.
|
||||
|
||||
## Required spec amendments
|
||||
|
||||
### P0 — before ScenarioRun baseline implementation
|
||||
|
||||
1. **037 owner:** add `BaselineRevision/CatalogRevision` identity, coordinate uniqueness, current-entry/supersedes relation, CAS conflict, explicit invalidate/retire/rebaseline lifecycle and audit. Duplicate ID with different content must be conflict, never silent success.
|
||||
2. **037 owner:** define authoritative visual candidate capture: either a new visual capture operation or a strict ScenarioArtifact input contract. Require server-verified owner, bytes/MIME/digest/capture profile and durable `ReviewDisposition`; remove caller-provided approval as evidence.
|
||||
3. **037 owner:** reconcile `baseline-catalog.schema.json`, OpenAPI and Pydantic fields/policies/candidate payloads; state whether publication commits/pushes Git or only materializes a working tree.
|
||||
4. **044 owner:** add `ScenarioBaselineResolver` PRE/POST/INVARIANT and API schema. Pin baseline set/version, release+commit, entry IDs, catalog digest, source/image digest and baseline family in request hash, RunnerPlan, ScenarioRun and result. Missing/stale/ambiguous entries must fail before I/O.
|
||||
5. **038 + 044 owners:** define canonical baseline-ref resolution and deterministic visual/metric comparison step before optional AgentEvaluation. AgentEvaluation input manifest and DecisionPolicy truth table must distinguish baseline comparison failure, VLM disagreement, low confidence and missing/corrupt evidence.
|
||||
6. **045 + 047 owners:** expose pinned baseline/comparison/evaluation provenance in typed result DTO and consume the exact 044 `AnalyticsContextKey`; prohibit `scenario_id:environment_id` substitute grouping.
|
||||
|
||||
### P1 — minimum hardening
|
||||
|
||||
1. **037 + 046 owners:** define raw capture/screenshot retention, retirement/tombstone, digest-on-read, audit export and proof that approved baselines do not outlive required bytes accidentally.
|
||||
2. **037 + 050 owners:** make MCP capture/request/decide/consume/verification behavior and errors identical to REST; eliminate the consume fallback discrepancy.
|
||||
3. **045 owner:** specify and implement baseline-aware run comparison: actual/expected/delta, baseline revision/entry change, per-step latency delta and non-comparability reasons.
|
||||
4. **047 owner:** implement the complete baseline-family/release/dashboard/lineage/principal context and separate deterministic flakiness from model/prompt instability.
|
||||
5. **New spec only if required:** define `ExecutionPerformanceBaseline` authority, capture window, approval/version, thresholds, drift/rebaseline and retention. Do not infer it from rolling 047 health.
|
||||
|
||||
## Final decision
|
||||
|
||||
037 supplies reusable, tested baseline engines; rebuilding them would be wrong. Production agentic ScenarioRun still cannot claim a baseline-backed visual or metric verdict because it neither resolves nor cryptographically pins the approved catalog selection, its idempotency identity omits `baseline_set`, and its evaluation/result/analytics chain does not carry the baseline provenance. Close the six P0 amendments and implement the two-module resolver/adapter delta before adding baseline-backed canary acceptance.
|
||||
|
||||
#endregion Report.SsProdAgenticE2EBaselineGap
|
||||
339
docs/reports/ss-prod-agentic-e2e-production-gap-2026-09-08.md
Normal file
339
docs/reports/ss-prod-agentic-e2e-production-gap-2026-09-08.md
Normal file
@@ -0,0 +1,339 @@
|
||||
# SS-Prod Agentic Dashboard E2E Production Gap — 2026-09-08
|
||||
|
||||
#region Report.SsProdAgenticE2EGap [C:4] [TYPE Verification] [SEMANTICS scenario,browser,screenshot,llm,vision,production-gap]
|
||||
|
||||
## Executive verdict
|
||||
|
||||
**NO-GO for a production ScenarioRun that walks a dashboard in a browser, captures screenshots, asks an LLM/VLM to read them, and produces an authoritative verdict.** The repository is not greenfield: the deterministic ScenarioRun engine, durable browser/screenshot adapters, `ScreenshotService`, multimodal `LLMClient`, LLM provider CRUD, and typed advisory VLM findings already exist. The missing work is the production bridge and completion of several contracts, not a replacement screenshot or LLM stack.
|
||||
|
||||
The shortest production-ready vertical slice is estimated at:
|
||||
|
||||
- **7 new production modules**, **24 existing production modules modified**;
|
||||
- **2,000–3,430 production LOC** and **2,980–4,730 test LOC**;
|
||||
- **1 database migration**, **4 configuration/deployment units**;
|
||||
- **9.6–15.4 engineer-weeks** for one experienced backend or **6–10 calendar weeks** with two engineers and review/operations support.
|
||||
|
||||
A hardened production-ready implementation, cumulative over the minimum slice, is estimated at:
|
||||
|
||||
- **13 new production modules**, **48 existing production modules modified**;
|
||||
- **3,950–6,930 production LOC** and **6,480–10,330 test LOC**;
|
||||
- **3 database migrations**, **11 configuration/deployment units**;
|
||||
- **19.1–31.7 engineer-weeks** (roughly **10–18 calendar weeks** for a two-engineer team, including staged rollout).
|
||||
|
||||
These are bottom-up engineering ranges, not story-point conversions. They exclude rewriting the already present 8,057-line execution/screenshot/LLM surface, model training, browser-farm procurement, and repair of unrelated product features.
|
||||
|
||||
The immediate runtime blocker reported by ss-prod — `provider_loop=not_started`, browser/screenshot `unregistered`, bindings `0` — has **two independent roots**:
|
||||
|
||||
1. **Deployment/configuration:** `settings.scenario_live_execution_bindings` defaults to an empty list and ss-prod has zero configured bindings. Therefore startup has no lawful environment/query/principal tuple to register.
|
||||
2. **Application wiring defect:** the provider loop implements `start()`/`stop()`, but the FastAPI lifespan bootstraps providers without calling either lifecycle method. Even a valid binding would reach an adapter whose first guard returns `*_LOOP_UNAVAILABLE`.
|
||||
|
||||
Beyond that blocker, simply adding config is insufficient: the browser transport supports only `open_dashboard`, `wait_for_state`, `refresh`, `row_edit`, and `bulk_edit`, while the registry exposes filter, pagination, XLSX download, and cross-dashboard navigation actions; `vlm_analyze` is registered as an ordinary assertion; no immutable `AgentEvaluation` record/provider or `DecisionPolicy` exists; and the ScenarioRun UI has no screenshot/evaluation evidence viewer.
|
||||
|
||||
## Scope, method, and definitions
|
||||
|
||||
This is a read-only code/spec/runtime gap analysis. It reuses the evidence in [the production E2E report](ss-prod-dashboard-scenario-e2e-report-2026-09-08.md) and the follow-up contract in [the E2E plan](ss-prod-dashboard-scenario-e2e-plan-2026-09-08.md). No new production mutation was performed.
|
||||
|
||||
Evidence sources were independent:
|
||||
|
||||
- normative specs 017, 036, 038, 039, 042, 044, 045, 046, and 047;
|
||||
- semantic graph (`docs/api/nav/root.map`, `ScenarioExecution.map`, `ScreenshotService.map`, `LLMClient.map`, model/migration maps);
|
||||
- backend/frontend source, migrations, Docker build configuration, and tests;
|
||||
- already captured ss-prod REST/MCP/UI evidence from `2026-09-08T05:46:45Z`–`06:20Z`.
|
||||
|
||||
**Module definition:** one cohesive production source file or UI component that owns a distinct runtime contract. A migration, a deployment manifest/config record, a test fixture, and a test file are counted separately, not as production modules. “Modified modules” is a deduplicated set assigned to exactly one work package below, so totals do not double-count cross-package files. LOC means non-generated source added or materially rewritten; comments/contracts are included, generated OpenAPI/nav output is excluded.
|
||||
|
||||
Status vocabulary:
|
||||
|
||||
- `PASS`: implemented and evidenced in ss-prod for the relevant contract.
|
||||
- `DEFECTIVE`: code/surface exists but violates or cannot satisfy the contract.
|
||||
- `PRESENT_UNWIRED`: reusable implementation exists but is not connected to the ScenarioRun production path.
|
||||
- `MISSING`: required runtime contract has no production implementation.
|
||||
- `DEPLOYMENT_ONLY`: production code exists and remaining gap is trusted config/image/provider rollout.
|
||||
- `NOT_SPECIFIED`: desired behavior has no stable normative contract; product decision is required before implementation.
|
||||
|
||||
## Current end-to-end implementation trace
|
||||
|
||||
```text
|
||||
038 compile/validate
|
||||
-> immutable 042 ScenarioRevision
|
||||
-> 044 derive_runner_plan (pinned ActionRegistry descriptors)
|
||||
-> queued worker / deterministic DAG walker
|
||||
-> typed executor registry
|
||||
browser -> LiveCompositionRoot -> BrowserProvider -> Playwright transport
|
||||
screenshot -> LiveCompositionRoot -> ScreenshotProvider -> ScreenshotService
|
||||
assertion/transform/xlsx/report/artifact -> local executors
|
||||
-> ScenarioStepRun + ScenarioArtifact + StepOutcome
|
||||
-> result/events/analytics/automation
|
||||
|
||||
Existing but separate:
|
||||
ValidationPolicy/ValidationRun -> DashboardValidationPlugin Path A
|
||||
-> ScreenshotService -> JPEGs -> LLMClient multimodal -> ValidationRecord
|
||||
|
||||
Existing advisory preview path:
|
||||
scenario/vlm.py -> LLMProviderService -> LLMClient.get_json_completion
|
||||
-> VlmFinding[]
|
||||
|
||||
Missing ScenarioRun link:
|
||||
ScenarioArtifact screenshot -> AgentEvaluationProvider -> immutable AgentEvaluation
|
||||
-> deterministic DecisionPolicy -> authoritative StepOutcome
|
||||
```
|
||||
|
||||
### Scenario/revision/run orchestration
|
||||
|
||||
- `runner_plan.py:74-142` derives a deterministic plan from the current immutable revision, pins exact action descriptors, rejects bootstrap-only empty revisions, topologically orders the DAG, and marks human graphs manual-only.
|
||||
- `dispatch.py:32-80` gates dependencies and dispatches only exact `{tool, action}` descriptors; human is lifecycle control.
|
||||
- `runner.py:893-1109` creates durable step rows, claims leases before I/O, records real artifact refs/digests, blocks descendants, and aggregates outcomes.
|
||||
- `lifecycle.py:262-306,416+` implements durable cancel and bounded retries; `provider_operations.py` and migrations `0014`/`0015` provide capacity leases and browser operation receipts/reconciliation.
|
||||
- ss-prod evidenced create/reload/update, PROD approval, waiting-human resolution, result/events/compare, cancel and retry/resume conflict guards. That part is real and persisted, but the only completed run was correctly `inconclusive`, not visual PASS.
|
||||
- There is schema drift inside 038: `TOOLS` and the runtime executor registry include `sql_evidence` and `transform` (`templates/__init__.py:16-18`, `executors.py:449-464`), while `ScenarioStep.tool` omits both (`scenario/models.py:137-159`). The passing compiler tests do not prove those two tool values can survive the canonical Pydantic model.
|
||||
|
||||
### Browser execution provider
|
||||
|
||||
- `live_composition.py:193-279` constructs existing `SupersetClient`, `ScreenshotService`, screenshot provider, Playwright browser provider, and browser reconciler for each enabled trusted binding.
|
||||
- `providers/browser.py:191-398` validates binding/principal/action/mutation contracts, acquires capacity, persists operation receipts for mutations, stores screenshot evidence, and fails closed.
|
||||
- `providers/browser_transport.py:29-31,83-148` launches an isolated Playwright context using the server-owned screenshot login flow and closes context/browser on every path.
|
||||
- The transport is functionally incomplete: `_READ_ONLY_ACTIONS` contains only `open_dashboard`, `wait_for_state`, `refresh`; the registry at `scenario/templates/__init__.py:22-45` also promises `apply_filters`, `text_filter`, `table_filter`, `pagination`, `download_xlsx`, and `navigate_dashboard`. Unsupported actions fail before I/O at `browser_transport.py:91-92`.
|
||||
- The compiler compounds this: each apparently multi-stage template maps to one action tuple (`templates/__init__.py:153-174`), and `_compile_graph` emits one step per selected case (`compiler.py:181-203`), so names such as `browser_apply_observe_assert` do not build a browser→observe→assert chain.
|
||||
- The production plugin catalog containing `llm_dashboard_validation` proves the older validation plugin is installed, not that 044 providers are registered: plugin identity is declared at `plugins/llm_analysis/plugin.py:165-187`, whereas 044 registration is exclusively binding-driven in `live_composition.py`.
|
||||
|
||||
### Screenshot capture, storage, and artifacts
|
||||
|
||||
- `plugins/llm_analysis/_screenshot.py:21-37` is the reusable facade. `_screenshot_session.py:25-183` owns isolated Chromium launch, Superset UI login, dashboard navigation, chart readiness, and scroll warm-up.
|
||||
- `_screenshot_capture.py:30-200` recursively switches tabs to depth 3, captures CDP PNG with Playwright fallback, uses 1920×1200, and closes the browser. `capture_dashboard` at lines 204-252 produces JPEG inputs and WebP archive output.
|
||||
- `_screenshot_media.py:22-49,58-84,91-100` resizes JPEG to 1024px at quality 60, archives WebP at quality 80, and cleans intermediates. This materially implements spec 017 FR-040/041/059.
|
||||
- `execution/providers/screenshot.py:76-214` is already the 044 adapter the older prose at spec 044:186-188 says was missing: it performs capacity admission, calls `ScreenshotService`, enforces byte limits, hashes real JPEG bytes, stores opaque refs, and returns durable evidence. The spec/traceability text is stale relative to code.
|
||||
- The backend container can install Chromium only when `INSTALL_PLAYWRIGHT_BROWSERS=1` (`docker/backend.Dockerfile:33`); the isolated frontend E2E compose explicitly uses `0`. The prior ss-prod readiness stopped at `unregistered`, so it did not prove a Chromium executable is present in that deployed backend image.
|
||||
- `execution/artifacts.py` and `models/scenario_artifact.py:27-61` supply generic ScenarioRun-owned immutable metadata. But no ScenarioRun artifact-content GET route was found; the only draft-content route is AgentRun-specific (`api/routes/agent_runs.py:205`). The ScenarioRun UI exposes ref strings but cannot render the bytes.
|
||||
- `CaptureSpec.mask_selectors` exists (`scenario/models.py:74-87`) and the 036 bridge can register caller-supplied masked bytes (`scenario/capture.py:28-82`), but neither `ScreenshotService` nor the 044 screenshot provider applies DOM masks. Newer specs 036 and 038 make masking optional for local-only providers; spec 017 still requires it for its validation-task security posture. This is a spec-policy conflict, not a reason to claim masking is implemented.
|
||||
|
||||
### LLM settings, selection, image payload, and parsing
|
||||
|
||||
- `services/llm_provider.py:62-179,216-274` provides encrypted provider CRUD/decryption; `api/routes/llm.py` and MCP list/status expose redacted configuration. ss-prod reported zero providers and `ready=false`, so provider creation/selection is a deployment prerequisite.
|
||||
- The mature validation-task path constructs correct JPEG `image_url` payloads (`_llm_client_analysis.py:166-178`), chunks by provider `max_images`, calls chunks concurrently, and merges typed status/issues (`191-296`). `get_json_completion` uses JSON mode with fallback and five retry attempts (`_llm_client_core.py:255-329`).
|
||||
- `DashboardValidationPlugin._execute_path_a` is a real end-to-end consumer: capture at `plugin.py:455-474`, multimodal analysis at `479-526`, JPEG cleanup and persisted WebP/result at `528-603`.
|
||||
- That plugin **cannot be reused as the ScenarioRun orchestrator**. It persists `ValidationRecord`, directly treats the LLM `status` as validation truth, carries no immutable 038 AgentEvaluationSpec, no ScenarioRun provider receipt/capacity identity, and no DecisionPolicy. Reusing it wholesale would violate SCEX-FR-003/015/025.
|
||||
- The narrower `scenario/vlm.py:82-147` correctly resolves a configured multimodal provider and reuses `LLMClient`, but it labels every byte stream `image/png` at line 139 although the screenshot provider stores JPEG bytes. Its response normalizer returns `[]` for every unrecognized object (`151-163`), making malformed responses indistinguishable from no findings. Its header promises redacted raw-response persistence (`3-6`), but no persistence/redaction call exists in the module.
|
||||
- The validation-task payload limiter only reduces quality to 30 (`_llm_client_analysis.py:236-244`). It does not implement spec 017 FR-056’s second 800px reduction and final Path B fallback after a re-estimate.
|
||||
- `DashboardValidationPlugin` logs decrypted API-key length at `plugin.py:266-269`, contradicting `LLMProviderService`’s explicit no-length logging contract at `llm_provider.py:216-226`. No secret value is printed, but this should be removed during security hardening.
|
||||
|
||||
### Assertions, AgentEvaluation, verdict, and UI
|
||||
|
||||
- Deterministic comparison exists in `execution/executors.py` and result aggregation exists in `execution/result.py:24-155`.
|
||||
- `vlm_analyze` is registered as tool `assertion` (`templates/__init__.py:42`), so the default registry resolves it to the ordinary deterministic `assertion()` function (`executors.py:449-464`). It never invokes `scenario/vlm.py` during ScenarioRun execution.
|
||||
- No production `AgentEvaluation` SQLAlchemy model, migration, provider, executor, or deterministic `DecisionPolicy` mapper was found. This matches the explicit reconciliation note in `specs/038.../verification-program.md:58-76` and `specs/044.../traceability.md:17`.
|
||||
- The spec task list incorrectly marks T039 complete, while traceability and production code say absent. That checkbox is not implementation evidence.
|
||||
- `ScenarioResultView.svelte:11-40` renders counts, failures, and provenance only. `frontend/src/types/scenario-run.ts:74-85` has no evidence or agent-evaluation DTO. The required `EvidencePanel` and `VlmFindingReviewCard` named by spec 039:111-163 do not exist in `frontend/src`; the older validation-task report has screenshot thumbnails but is a different run model.
|
||||
|
||||
### Retry, cancel, recovery, analytics, and automation
|
||||
|
||||
- General runner retry closure, step leases, crash recovery, terminal signals, and durable cancellation are present and have substantial unit coverage. ss-prod exercised manual cancellation and terminal guards.
|
||||
- SCEX-FR-020 is only partial: `cancel_run` records/drains DB state (`lifecycle.py:262-306`) but does not dispatch provider cancellation by `operation_id`. Browser receipts/reconciliation exist; equivalent screenshot/LLM operation receipts and cancellation are absent.
|
||||
- Automation and analytics exist, but the production E2E found separate product defects: REST can persist disabled automation that MCP rejects (`DEF-02`), and an accepted investigation remains queued (`DEF-04`). Six automation reads are anonymous (`SEC-01`). These do not cause the immediate browser/VLM blocker, but must be closed for unattended hardened production.
|
||||
|
||||
## Spec → code → ss-prod orthogonal matrix
|
||||
|
||||
| Requirement | Spec reference | Code contract/module | ss-prod evidence | Status | Gap conclusion |
|
||||
|---|---|---|---|---|---|
|
||||
| First-class pinned run, deterministic DAG | 044:30-46; SCEX-FR-001..003 | `runner_plan.py:74-142`; `dispatch.py:32-80`; `runner.py:893-1109` | durable runs, reload, approval, result/events | PASS | Core orchestration exists. |
|
||||
| Human checkpoint lifecycle | 044:50-74; FR-004..007 | `lifecycle.py`; `runner.py:1032-1052` | waiting-human→inconclusive; cancel guards | PASS | Not a visual automated result. |
|
||||
| Startup live composition | 044:155-177; FR-022/023 | `app.py:111-120,224-231`; `live_composition.py:193-279`; `preflight.py:79-128` | loop `not_started`, bindings 0, providers unregistered | DEFECTIVE | Missing loop lifecycle plus missing deployment binding. |
|
||||
| Trusted binding/config | 044:170-177 | `config_models.py:240-269`, default `[]` | registered bindings 0 | DEPLOYMENT_ONLY | Configure exact snapshot/query/principal after code fix. |
|
||||
| Browser Playwright provider | 044:41-46, FR-002/007/016 | `providers/browser.py`; `browser_transport.py` | unregistered; no live browser step executed | PRESENT_UNWIRED | Provider exists, but action coverage is incomplete. |
|
||||
| Full browser action registry | 044 data-model:123; 038 ActionRegistry | registry `templates/__init__.py:22-45`; transport `29-31,91-92` | compiler fell back to manual B01 | DEFECTIVE | Registry/transport/compiler capability mismatch. |
|
||||
| Screenshot capture by existing stack | 017 FR-040/041/059; 044 FR-009 | `ScreenshotService` session/capture/media mixins | screenshot provider unregistered | PRESENT_UNWIRED | Reuse; do not build another capture engine. |
|
||||
| Durable screenshot evidence | 044 FR-018/019 | `providers/screenshot.py:76-214`; `artifacts.py`; `ScenarioArtifact` | evidence storage ready, no capture run | PRESENT_UNWIRED | Adapter exists; content delivery/UI is missing. |
|
||||
| Capture masking | 017 FR-029c vs 036 FR-009 superseded / 038 optional | `CaptureSpec`, `capture.py`; no masking in ScreenshotService | not exercised | NOT_SPECIFIED | Decide local-only policy; external providers require implementation. |
|
||||
| LLM provider configuration/selection | 017 FR-047/050; 038 FR-011 | `LLMProviderService`; LLM APIs/UI | provider count 0, not ready | DEPLOYMENT_ONLY | Add approved active multimodal provider; no new client needed. |
|
||||
| Image payload + multimodal call | 017 FR-040/056; 038 FR-011 | `_llm_client_analysis.py`; `scenario/vlm.py` | blocked by zero provider | PRESENT_UNWIRED | Reuse client; fix MIME, parser, fallback, provenance. |
|
||||
| Immutable AgentEvaluation | 038 verification-program:22-34; 044 FR-015/025 | none in `backend/src/models` or executor registry | no API/event/result record | MISSING | New model, migration, provider, DTO/event path required. |
|
||||
| Deterministic DecisionPolicy | same; 044 SC-009 | no production mapper | no authoritative visual verdict | MISSING | Must mediate model result; bare LLM status forbidden. |
|
||||
| Compiler emits visual execution chain | 038 FR-011/011c; 044 Story 2 | one-action templates and compiler one-step loop | manual-only compiled graph; DEF-01/03 | DEFECTIVE | Authoring must emit screenshot→evaluation→policy/assert graph. |
|
||||
| Scenario evidence/evaluation UI | 039 FR-011..013; 045 FR-005/012/013 | result view/types only; legacy validation report separate | no screenshot result UI | MISSING | Add authenticated artifact content and typed panels. |
|
||||
| Retry/timeout/crash recovery | 044 Story 4, FR-006/007/021 | `lifecycle.py`, leases, runner recovery | cancel/guards proven; live I/O unproven | PRESENT_UNWIRED | Local lifecycle exists; live provider retry not canaried. |
|
||||
| Operation-aware provider cancel | 044 FR-020 | browser receipt/reconcile; no lifecycle callback | no in-flight provider case | DEFECTIVE | DB drain is not provider cancellation acknowledgement. |
|
||||
| Capacity/concurrency | 044 FR-019/024 | `capacity.py`, migration 0014; browser/screenshot use it | no live provider capacity evidence | PRESENT_UNWIRED | Shared cross-run enforcement not production-proven. |
|
||||
| Result/artifact integrity | 044 FR-018/027 | digest/ref validation, result snapshot | persistence/result pass; no live screenshot | PRESENT_UNWIRED | Metadata path passes; byte retrieval and vision chain absent. |
|
||||
| Analytics/investigation | 044 FR-011; 047 | terminal signals and analytics services | DEF-04 accepted case still queued | DEFECTIVE | Independent lifecycle defect. |
|
||||
| Automation | 046 FR-001..018 | schedule/rule/policy services | DEF-02 REST/MCP eligibility split | DEFECTIVE | Blocks unattended trust, not manual visual slice. |
|
||||
| Read authorization | 046/enterprise security posture | automation routes | SEC-01 six anonymous reads | DEFECTIVE | Close before hardened production. |
|
||||
|
||||
## Runtime blocker root-cause tree
|
||||
|
||||
```text
|
||||
No live agentic visual ScenarioRun on ss-prod
|
||||
├─ A. No provider registration (observed)
|
||||
│ └─ settings.scenario_live_execution_bindings = [] (default and ss-prod)
|
||||
│ └─ bootstrap loop registers no Superset/browser/screenshot tuple
|
||||
├─ B. Provider runtime not started (observed + code defect)
|
||||
│ ├─ ProviderEventLoop.start/stop exist
|
||||
│ └─ FastAPI lifespan calls bootstrap but never start/stop
|
||||
├─ C. No LLM execution capability (observed deployment gap)
|
||||
│ └─ zero configured/active multimodal LLM providers
|
||||
├─ D. No executable visual program (code/product defect)
|
||||
│ ├─ compiler emitted manual-only B01 on current capabilities
|
||||
│ ├─ proposal/save path has DEF-01/DEF-03
|
||||
│ ├─ templates are one action, not browser→screenshot→evaluation chains
|
||||
│ └─ vlm_analyze routes to assertion, not an AgentEvaluation executor
|
||||
└─ E. No authoritative visual decision path (missing code)
|
||||
├─ no AgentEvaluation record/provider
|
||||
├─ no DecisionPolicy mapper
|
||||
└─ no ScenarioRun evidence/evaluation API/UI
|
||||
```
|
||||
|
||||
Fixing A alone changes readiness from `unregistered` only if bootstrap dependencies work, but B still rejects all I/O. Fixing A+B enables open-dashboard and screenshot capture, but not a legitimate LLM-derived ScenarioResult because D+E remain.
|
||||
|
||||
## Reuse decision: screenshotservice and LLM reader
|
||||
|
||||
**Reuse both; add a bounded adapter/provider, not a second stack.**
|
||||
|
||||
The minimum design is:
|
||||
|
||||
1. Keep `ScreenshotService` as the Playwright/login/tab/CDP/media implementation.
|
||||
2. Keep `providers/screenshot.py` as the 044 capture/capacity/durable-evidence adapter.
|
||||
3. Add an authenticated `ScenarioArtifactContentService` that verifies owner/run/step/ref/digest and returns bytes/MIME without exposing paths.
|
||||
4. Add `AgentEvaluationProvider` that accepts only an immutable 038 evaluation spec, retrieves only allowlisted screenshot refs, resolves the pinned multimodal provider/model/prompt, and calls the existing `LLMClient`/hardened `scenario.vlm` logic under capacity/deadline limits.
|
||||
5. Persist immutable `AgentEvaluation` plus raw-response artifact metadata under retention/redaction policy.
|
||||
6. Run a versioned pure `DecisionPolicy` mapper; only its result becomes `StepOutcome`.
|
||||
|
||||
Directly calling `DashboardValidationPlugin._execute_path_a` from ScenarioRun is rejected: it has the wrong persistence model and authority boundary. Copying its screenshot or LLM client is also rejected by SCEX-FR-009.
|
||||
|
||||
## Estimate A — minimum production-ready vertical slice
|
||||
|
||||
Definition of done: one approved, immutable, read-only scenario can open dashboard 11 on a non-destructive ss-prod canary account, navigate declared tabs/apply one native filter, capture durable screenshots, submit them to a pinned local multimodal provider, persist typed evaluation/provenance, map via deterministic policy, display evidence, retry a safe failed step, cancel safely, and clean up transient resources. No browser data mutation, arbitrary agent orchestration, distributed browser farm, or high-load SLO is included.
|
||||
|
||||
| WP | Non-overlapping deliverable | New modules | Existing modules modified | Prod LOC | Test LOC | Migration | Config/deploy units | Eng-weeks |
|
||||
|---|---|---:|---:|---:|---:|---:|---:|---:|
|
||||
| M1 | Provider loop start/stop, readiness order, trusted binding startup | 0 | 3 | 100–180 | 180–280 | 0 | 1 | 0.7–1.2 |
|
||||
| M2 | Compiler/model/action contract emits real browser→screenshot→evaluation DAG; close DEF-01/03 slice | 0 | 4 | 250–450 | 350–550 | 0 | 0 | 1.2–2.0 |
|
||||
| M3 | Browser action driver for tabs/native filter plus registry/transport/reconciler parity | 1 | 3 | 450–750 | 550–850 | 0 | 0 | 2.0–3.0 |
|
||||
| M4 | Scenario screenshot artifact content/MIME/ownership bridge and authenticated read route | 1 | 4 | 250–450 | 350–550 | 0 | 1 | 1.0–1.7 |
|
||||
| M5 | Immutable AgentEvaluation model/provider, DecisionPolicy, registry/events/result integration | 3 | 6 | 650–1,050 | 750–1,150 | 1 | 1 | 2.5–4.0 |
|
||||
| M6 | Scenario EvidencePanel + AgentEvaluationCard + typed monitor/result DTOs | 2 | 4 | 300–550 | 350–600 | 0 | 0 | 1.2–2.0 |
|
||||
| M7 | Real canary fixtures, local-provider contract test, preprod/prod smoke and rollback runbook | 0 | 0 | 0 | 450–750 | 0 | 1 | 1.0–1.5 |
|
||||
| **Total** | | **7** | **24** | **2,000–3,430** | **2,980–4,730** | **1** | **4** | **9.6–15.4** |
|
||||
|
||||
Expected new modules are: browser action driver; scenario artifact-content service; AgentEvaluation model; AgentEvaluation provider; DecisionPolicy mapper; EvidencePanel; AgentEvaluationCard. The migration creates immutable evaluation/policy linkage and indexes. The four deploy units are provider-loop lifecycle/readiness rollout, exact ss-prod live binding, approved multimodal provider configuration, and Playwright/canary image + runbook validation.
|
||||
|
||||
Assumptions:
|
||||
|
||||
- existing `ScreenshotService`, LLM provider CRUD, LLM client, ScenarioRun runner, scheduler, capacity tables, and artifact storage are retained;
|
||||
- minimum browser traversal means open, tab navigation, readiness/wait, refresh, and one native-filter action—not every mutation/export action in the registry;
|
||||
- LLM provider is local/approved and OpenAI-compatible; model procurement/evaluation is external;
|
||||
- DEF-02/04 and SEC-01 are excluded from manual vertical-slice critical path but included in hardened scope;
|
||||
- a product decision accepts newer 036/038 optional masking for an entirely local trust perimeter. If external vision endpoints are permitted, add the H3 security package before launch.
|
||||
|
||||
## Estimate B — hardened production-ready
|
||||
|
||||
Definition of done adds provider-aware cancellation, safe reconciliation and idempotency, multi-worker concurrency, security/redaction/retention, provider compatibility and payload fallback, automation/analytics consistency, observability/SLOs, load/chaos testing, and staged rollback. Counts below are **additional** to Estimate A; cumulative totals are the executive numbers.
|
||||
|
||||
| WP | Non-overlapping hardening deliverable | New modules | Existing modules modified | Prod LOC | Test LOC | Migration | Config/deploy units | Eng-weeks |
|
||||
|---|---|---:|---:|---:|---:|---:|---:|---:|
|
||||
| H1 | Provider cancellation broker, screenshot/LLM receipts, unknown-effect reconciliation | 1 | 3 | 350–600 | 500–800 | 0 | 0 | 1.5–2.5 |
|
||||
| H2 | Distributed capacity/leases, idempotent worker failover, concurrency fairness | 2 | 4 | 500–850 | 650–1,000 | 1 | 2 | 2.0–3.5 |
|
||||
| H3 | DOM masking/redaction policy, encrypted raw-response retention, access audit | 2 | 5 | 450–800 | 500–800 | 1 | 2 | 2.0–3.0 |
|
||||
| H4 | MIME/schema parser hardening, 1024→800 re-estimate, Path B fallback, provider matrix | 0 | 4 | 250–450 | 350–600 | 0 | 1 | 1.0–1.8 |
|
||||
| H5 | Close DEF-02/04/SEC-01; automation/analytics cross-surface invariants | 0 | 6 | 250–500 | 500–800 | 0 | 2 | 1.5–2.5 |
|
||||
| H6 | SLO metrics/traces, load/chaos/browser-version canaries and rollback automation | 1 | 2 | 150–300 | 1,000–1,600 | 0 | 4 | 1.5–3.0 |
|
||||
| **Additional** | | **6** | **24** | **1,950–3,500** | **3,500–5,600** | **2** | **7** | **9.5–16.3** |
|
||||
| **Cumulative A+B** | | **13** | **48** | **3,950–6,930** | **6,480–10,330** | **3** | **11** | **19.1–31.7** |
|
||||
|
||||
The hardened estimate is intentionally not “minimum plus 30%”: each additional package closes a distinct falsifiable production property.
|
||||
|
||||
## Critical path and dependency graph
|
||||
|
||||
```text
|
||||
M1 loop lifecycle ─┬─> exact live binding + Playwright readiness ─┐
|
||||
│ │
|
||||
M2 executable DAG ┴─> M3 browser action parity ─> M4 screenshots/artifact bytes
|
||||
│
|
||||
LLM provider deployment ───────────────────────────────────┤
|
||||
v
|
||||
M5 AgentEvaluation + DecisionPolicy
|
||||
│
|
||||
M6 evidence/result UI │
|
||||
v
|
||||
M7 real canary acceptance
|
||||
|
||||
After minimum acceptance:
|
||||
H1 cancellation/reconcile ─┬─> H2 distributed concurrency
|
||||
H3 security/retention ─────┤
|
||||
H4 provider compatibility ─┼─> H6 load/chaos/SLO rollout
|
||||
H5 automation/analytics ───┘
|
||||
```
|
||||
|
||||
Critical path is **M1 → M2/M3 → M4 → M5 → M7**. M6 may proceed after DTOs stabilize but must finish before operator-facing production release. An approved local multimodal provider and exact deployment binding are external prerequisites parallel to M1–M4.
|
||||
|
||||
## Top risks
|
||||
|
||||
1. **False authoritative verdict (critical):** wiring existing ValidationRecord status directly into ScenarioResult would violate the DecisionPolicy boundary.
|
||||
2. **Browser registry drift (high):** graph validation accepts actions the live transport rejects. A registry-wide provider contract test is mandatory.
|
||||
3. **Credential/RLS mismatch (high):** browser login identity, Superset API identity, and pinned principal fingerprint may observe different data unless canary evidence proves equivalence.
|
||||
4. **Artifact split-brain (high):** AgentRun `DraftArtifact`, ScenarioArtifact metadata, filesystem/WebP archive, and LLM JPEG temp files have different ownership/lifetime models.
|
||||
5. **Provider loop lifecycle/race (high):** ss-prod never starts it; the local unit test also timed out starting it. Startup ordering and shutdown drain need an isolated fix/proof.
|
||||
6. **Malformed or oversized LLM output (high):** current parser can silently turn invalid envelopes into empty findings; incomplete payload fallback can exceed provider limits.
|
||||
7. **Unsafe retries (high):** browser mutations and late LLM responses require operation receipts and deterministic winning-attempt rules; minimum slice must remain read-only.
|
||||
8. **Spec drift (medium):** 044 prose says screenshot provider absent although code exists; tasks mark AgentEvaluation complete although traceability/code say missing; masking requirements conflict.
|
||||
9. **Unattended control-plane inconsistency (medium):** DEF-02/04 and SEC-01 undermine automation/triage trust even after the visual path works.
|
||||
|
||||
## Acceptance criteria
|
||||
|
||||
### Minimum vertical slice
|
||||
|
||||
- `/api/ready` reports provider loop, evidence storage, browser, and screenshot `ready`, and exactly the intended binding fingerprint; no credentials/cookies appear.
|
||||
- A registry contract test enumerates every enabled browser action and proves supported or explicitly disabled before a revision becomes runnable.
|
||||
- One immutable scenario revision contains ordered browser, screenshot, agent-evaluation, and policy-bound assertion steps with exact refs and provider/model/prompt hashes.
|
||||
- A real isolated browser session logs into ss-prod with the deployment principal, opens dashboard 11, traverses declared tabs/filter state, and closes all resources.
|
||||
- Screenshot bytes are stored under opaque ScenarioRun-owned refs with correct MIME, byte length, SHA-256, capture metadata, and authenticated retrieval.
|
||||
- The LLM sees only allowlisted evidence; malformed/empty/oversized responses produce typed non-pass, never PASS.
|
||||
- `AgentEvaluation` is append-only and records pinned provider/model/prompt/evidence/input hashes, confidence/findings, raw-response artifact ref, timings, and attempt.
|
||||
- A versioned deterministic DecisionPolicy is the sole mapper to StepOutcome; ScenarioResult never consumes a bare model verdict.
|
||||
- Safe retry creates a new attempt/evaluation and retires only the active projection; cancellation leaves no active lease or late winning response.
|
||||
- UI reload shows screenshot, evaluation provenance, findings, policy-derived outcome, and evidence download without raw paths.
|
||||
- Cleanup proves no canary browser mutation, no leaked context/process/temp file, and no unrelated production record change.
|
||||
|
||||
### Hardened production
|
||||
|
||||
- Provider cancel returns `stopped|completed|unknown`; `unknown` blocks retry/PASS until reconciliation.
|
||||
- Capacity is atomic across all worker processes and workload classes; fairness and overload queue behavior meet declared SLOs.
|
||||
- Browser/LLM versions and binding/policy fingerprints are revalidated immediately before I/O; drift returns typed `*_CHANGED` without I/O.
|
||||
- DOM masking/redaction/retention policy matches the approved local/external provider posture; raw outputs and screenshots expire as declared.
|
||||
- 5-, 15-, and 50-tab dashboards, provider 401/403/429/5xx, browser crash, process restart, storage failure, late response, and concurrent duplicate triggers pass chaos/load acceptance.
|
||||
- Automation REST/MCP eligibility is identical, disposed cases leave the queue, and all operational reads require the intended permission.
|
||||
- Metrics cover queue time, browser/login/capture latency, LLM latency/tokens/cost, retry/cancel/reconcile, artifact bytes/retention, verdict distribution, and reason codes without secrets.
|
||||
|
||||
## Exact next implementation sequence
|
||||
|
||||
1. Reconcile specs first: mark 044 T039 open, update stale screenshot-provider prose, define the canonical `agent_evaluation` tool/action, and decide local-only masking policy.
|
||||
2. Fix ProviderEventLoop lifecycle in FastAPI startup/shutdown and add a failing lifespan test that requires `ready` before composition dispatch.
|
||||
3. Add one reviewed ss-prod live binding in non-enabled/canary form; validate environment/query/principal/RLS fingerprints and Playwright executable in the deployed image.
|
||||
4. Make compiler/templates emit a true multi-step visual program and reject any action unsupported by the selected provider capability fingerprint. Close the relevant DEF-01/03 authoring path.
|
||||
5. Complete only the read-only browser actions needed by the first canary (open, tab navigation, filter, wait/refresh); defer PROD mutations.
|
||||
6. Add ScenarioArtifact authenticated content retrieval and reconcile JPEG/WebP/MIME/digest/retention ownership.
|
||||
7. Implement immutable AgentEvaluation persistence, provider, capacity/receipt boundary, and strict response schema; reuse `LLMClient` and screenshot bytes.
|
||||
8. Implement/version DecisionPolicy and route `vlm_analyze` through it; add disagreement/low-confidence/missing-evidence tests.
|
||||
9. Add ScenarioRun evidence/evaluation DTOs and UI panels; do not reuse ValidationRun IDs or raw filesystem paths.
|
||||
10. Configure an approved local multimodal provider and run the M7 canary through browser→screenshot→LLM→policy→result with reload and cleanup proof.
|
||||
11. Only after the canary passes, implement H1–H5 and execute H6 staged load/chaos rollout.
|
||||
|
||||
## Verification performed
|
||||
|
||||
At `2026-09-08T06:58Z`–`07:01Z` UTC:
|
||||
|
||||
- `46 passed` in `1.35s`: compiler, scenario models, and typed executor unit suites.
|
||||
- `10 passed` in `1.22s`: VLM parsing/submission seam and fake-byte capture→VLM→disposition suites.
|
||||
- A hardcoded registry/transport check found **9 registered browser actions, only 3 supported overlaps** (`open_dashboard`, `row_edit`, `bulk_edit`), **6 registered-but-unsupported** actions, and 2 transport-only actions (`wait_for_state`, `refresh`) that no graph can select.
|
||||
- Hardcoded `ScenarioStep` fixtures for tools `transform` and `sql_evidence` both raised Pydantic `ValidationError`, confirming the schema/registry drift rather than inferring it from source names.
|
||||
- Provider runtime test failed before its first assertion: `ProviderEventLoop.start()` raised `PROVIDER_LOOP_START_TIMEOUT` after 5 seconds. Browser/screenshot provider suites repeated setup errors because their fixtures start the same runtime; the broad run was stopped rather than spending minutes repeating the identical failure.
|
||||
- No integration tests, real local browser, external LLM call, or PostgreSQL migration check was run in this gap-analysis turn. The prior ss-prod run remains the only real environment evidence and proved provider unavailability, not browser/VLM success.
|
||||
|
||||
The passing VLM tests use mocks/fake bytes and prove schema/seam behavior only; they do not close SCEX-FR-025. The provider-runtime failure may be sensitive to this sandbox’s thread/event-loop restrictions, so it is recorded as a local verification failure rather than independently declared a production defect. The independent production fact remains `provider_loop=not_started`, and source inspection proves the missing lifespan call.
|
||||
|
||||
## Limitations and remaining evidence
|
||||
|
||||
- No valid ss-prod LLM provider existed, so model compatibility, token use, latency, and real image comprehension remain unmeasured.
|
||||
- No exact live binding existed, so Superset browser principal/RLS equivalence and Playwright image availability were not exercised.
|
||||
- The production scenario was manual-only; there is no real ScenarioRun screenshot artifact, AgentEvaluation, or visual verdict to inspect.
|
||||
- This estimate assumes the existing APIs/models can be evolved without a backward-incompatible five-way VerificationProgram migration. Adopting the deferred full 038 IR decomposition would add approximately 3–6 modules, 1,000–2,000 production LOC, 1,200–2,500 test LOC, one migration, and 4–8 engineer-weeks beyond the ranges above.
|
||||
- Product must decide whether “full browser pass” includes XLSX download, pagination, cross-dashboard navigation, and non-PROD row/bulk mutations. The minimum estimate includes only read-only traversal plus one filter. Enabling every currently registered browser action adds approximately 500–900 production LOC, 700–1,100 test LOC, and 2–4 engineer-weeks.
|
||||
|
||||
#endregion Report.SsProdAgenticE2EGap
|
||||
170
docs/reports/ss-prod-agentic-e2e-spec-coverage-2026-09-08.md
Normal file
170
docs/reports/ss-prod-agentic-e2e-spec-coverage-2026-09-08.md
Normal file
@@ -0,0 +1,170 @@
|
||||
# SS-Prod Agentic Dashboard E2E — нормативное покрытие спецификациями (2026-09-08)
|
||||
|
||||
#region Report.SsProdAgenticE2ESpecCoverage [C:4] [TYPE Verification] [SEMANTICS scenario,browser,screenshot,llm,spec-coverage,traceability]
|
||||
|
||||
## Executive verdict
|
||||
|
||||
**Спецификации не покрывают весь production-контракт целиком.** Они подробно и проверяемо задают deterministic ScenarioRun, общий provider protocol, browser action semantics, screenshot metadata/digest, immutable `AgentEvaluation` и запрет прямого LLM-вердикта. Однако критический вертикальный шов от сохранённых screenshot bytes до авторизованного ScenarioRun UI и authoritative DecisionPolicy остаётся неполным. Поэтому наличие большого числа отмеченных задач не означает готовность реализовывать все семь gap-модулей без дополнительных решений.
|
||||
|
||||
Аудит охватывает **12 релевантных spec-наборов**, **19 ортогональных contract units** и **57 файлов в `contracts/`**. Итоговая классификация contract units: **COMPLETE 2, PARTIAL 4, IMPLIED 2, MISSING 1, DRIFTED 10**. Из семи новых модулей minimum slice только `AgentEvaluation` и `AgentEvaluationProvider` прямо названы как runtime сущности; browser driver, artifact-content service и итоговая UI card — архитектурный glue, который нельзя считать специфицированным только по `tasks.md`.
|
||||
|
||||
Нормативная готовность к реализации: **NO-GO до закрытия P0 amendments** для artifact content, DecisionPolicy truth table, provider lifecycle binding и единого ScenarioRun evidence/evaluation DTO. Это не отменяет уже реализованный reusable `ScreenshotService` и multimodal `LLMClient`: новый capture/LLM stack спецификациями прямо запрещён.
|
||||
|
||||
## Scope, method, vocabulary
|
||||
|
||||
Это read-only сопоставление spec → contract → acceptance → code/runtime evidence. Production вызовы и мутации не выполнялись. Runtime evidence взято из `ss-prod-agentic-e2e-production-gap-2026-09-08.md` и предшествующего фактического E2E отчёта; локальные утверждения повторно проверены `rg`, `find`, `nl` по реальным файлам.
|
||||
|
||||
Категории взаимоисключающие и относятся к **совокупному нормативному пакету**, а не только наличию имени:
|
||||
|
||||
- `COMPLETE` — явно определены responsibility, I/O/schema, errors/invariants, acceptance и непротиворечивая traceability;
|
||||
- `PARTIAL` — контракт назван, но один из обязательных элементов (обычно schema/errors/policy table) отсутствует;
|
||||
- `IMPLIED` — необходимость следует из соседних контрактов, но самостоятельного module contract нет;
|
||||
- `MISSING` — требуемая production boundary не имеет нормативного контракта;
|
||||
- `DRIFTED` — нормативные файлы противоречат друг другу либо заявленный implementation trace расходится с кодом/runtime.
|
||||
|
||||
`tasks.md` используется только как исторический сигнал. Checkbox без входов, выходов, ошибок, инвариантов и acceptance не повышает категорию.
|
||||
|
||||
## Inventory релевантных specs
|
||||
|
||||
Всего найдено: 12 `spec.md`, 11 `plan.md`, 11 `data-model.md`, 12 `tasks.md`, 10 `traceability.md`, 57 contract files и 14 checklist files.
|
||||
|
||||
| Spec | Role в agentic E2E | Status/version marker | Нормативные artifacts | Audit note |
|
||||
|---|---|---|---|---|
|
||||
| `017-llm-analysis-plugin` | существующий Playwright screenshot + multimodal LLM pipeline | `Implemented`, 2026-01-28 (`spec.md:4-6`) | spec/plan/data-model/tasks; `contracts/llm-api.yaml`, `modules.md`; 4 checklists | Полный старый ValidationTask pipeline, не ScenarioRun authority. |
|
||||
| `036-agent-test-stabilization` | evidence bridge, AgentRun artifacts, HITL | `Ready for Implementation`, 2026-07-07 (`spec.md:17`) | полный набор + 7 contract files | Bridge строго `CaptureSpec + AgentRun -> DraftArtifact[]` (`contracts/evidence.md:21-38`), поэтому не заменяет ScenarioRun content boundary. |
|
||||
| `037-superset-baseline-engine` | query model/baselines/comparison reuse | `Ready for Implementation`, 2026-07-07 (`spec.md:15`) | полный набор + schema/OpenAPI/modules | Supporting deterministic evidence contract. |
|
||||
| `038-dashboard-scenario-model` | canonical graph, ActionRegistry, capture/VLM/AgentEvaluation specs | `Reworked`, 2026-07-31 (`spec.md:16`) | полный набор; 7 contract files | В одном contract одновременно есть future normative IR и reconciliation «deferred» (`contracts/verification-program.md:58-76`). |
|
||||
| `039-dashboard-scenario-ui` | authoring workspace, EvidencePanel/VLM review | `Ready for Implementation`, 2026-07-07 (`spec.md:17`) | полный набор; UX/module contracts | Spec заявляет реализованный AgentWorkspace (`spec.md:187`), но он демонтирован и привязан к AgentRun, не ScenarioRun. |
|
||||
| `042-dashboard-scenario-registry` | immutable scenario/revision/activation | `Partially implemented`, audit pending (`spec.md:18`) | полный набор + OpenAPI/modules | Обязательная prerequisite boundary. |
|
||||
| `043-dashboard-scenario-editor` | edit/CAS/revalidation | `Partially implemented`, audit pending (`spec.md:16`) | полный набор + OpenAPI/modules | Обязательная revision-update boundary. |
|
||||
| `044-dashboard-scenario-execution` | runner, providers, artifacts, cancellation, evaluation | `Not production-complete` (`spec.md:2,19`) | полный набор + OpenAPI/modules | Главный normative spec; честно оставляет production composition/evaluation открытыми (`spec.md:319-327`). |
|
||||
| `045-dashboard-run-monitor` | run status/result/evidence/evaluation UI | `Partially implemented`, audit pending (`spec.md:17`) | полный набор + OpenAPI/modules/UX | Требует typed evaluation UI, но implementation DTO этого не содержит. |
|
||||
| `046-dashboard-scenario-automation` | schedules/triggers/policies/MCP parity | `Not production-complete` (`spec.md:17`) | полный набор + OpenAPI/modules | Контракт CRUD/enable-disable и deterministic scheduling явный (`spec.md:105-119`). |
|
||||
| `047-dashboard-scenario-analytics` | signals, queue/case, trends | `Not production-complete` (`spec.md:17`) | полный набор + OpenAPI/modules | Собственный audit фиксирует незакрытые evidence/case semantics (`spec.md:114-128`). |
|
||||
| `050-mcp-interface` | внешнее authoring/run/automation surface | `Ready for Implementation`, 2026-08-24 (`spec.md:18-20`) | только spec/tasks; **нет** plan/data-model/contracts/checklist/traceability | Supporting spec с сильными FR, но недостаточным contract package; status отстаёт от task claims/runtime. |
|
||||
|
||||
Primary runtime specs: 038, 044, 045, 046, 047. Reused implementation specs: 017, 036, 037. Authoring/persistence/UI prerequisites: 039, 042, 043. External client boundary: 050.
|
||||
|
||||
## Trace matrix: 7 minimum-gap modules
|
||||
|
||||
В колонках `API/schema`, `Semantics` и `Acceptance` указано наличие именно нормативного содержания, а не checkbox.
|
||||
|
||||
| Minimum module из gap report | Явно назван? | API/schema | PRE/POST/INVARIANT или UX_STATE | Acceptance | Implementation trace | Category |
|
||||
|---|---|---|---|---|---|---|
|
||||
| **Browser action driver** | Нет; назван более широкий `BrowserProvider` | Action catalog есть в 038 ActionRegistry и 044 `BrowserProviderActionContract` (`044/contracts/modules.md:308-347`) | Сильные browser invariants (`:282-298`) | Полный action/isolation/cancel canary profile (`:362-371`) | Transport реализует лишь 5 actions и расходится именами (`browser_transport.py:29-31,91-118`) | **IMPLIED** — внутренний adapter/driver decomposition требует собственного contract и mapping table. |
|
||||
| **ScenarioArtifact content service** | Нет | Есть только metadata `EvidenceReceipt` (`044/contracts/openapi.yaml:222-250`); content GET/range/auth/stream schema отсутствует | Generic owner/digest заданы (`044/data-model.md:77-79,218-223`), выдача bytes не задана | Нет test profile на authorization, MIME sniffing, digest-on-read, retention/tombstone | ScenarioRun хранит ref, но content route найден только для AgentRun; UI не может получить bytes | **MISSING**. |
|
||||
| **AgentEvaluation model** | Да (`038/contracts/verification-program.md:22-26`; `044/data-model.md:58-62`) | Full + summary schemas (`044/contracts/openapi.yaml:239-255+`) | Immutable/bounded/no-mutation contract явный | SCEX-SC-009 (`044/spec.md:220-221`) | Production model/migration отсутствуют, что честно отмечено 038 reconciliation | **COMPLETE** нормативно; implementation gap не является spec gap. |
|
||||
| **AgentEvaluationProvider** | Да (`044/spec.md:149-151`; `contracts/modules.md:266-270`) | Provider protocol + evaluation schema есть | Явные bounded inputs, authority limits и output ownership | Common provider suite + SC-009 | `tasks.md:136` помечает T039 `[x]`, но `traceability.md:17` — `[ ]`, а code содержит лишь capacity key | **DRIFTED** — сам contract достаточен, implementation trace недостоверна. |
|
||||
| **DecisionPolicy mapper** | Да (`038/contracts/verification-program.md:26`) | Поля политики названы; отдельной versioned schema/registry API нет | Запрет bare verdict и deterministic ownership явны | SCEX-SC-009 | Mapper отсутствует | **PARTIAL** — нет исчерпывающей truth table, precedence, missing/contradictory evidence semantics и error taxonomy. |
|
||||
| **EvidencePanel** | Да (`039/contracts/modules.md:108-120`) | DTO обещан, но contract зависит от `AgentRuns.DraftList`, не ScenarioRun artifact/evaluation API | Полные UX_STATE/feedback/recovery/test/invariant | UX tests перечислены | `039/spec.md:187` заявляет реализацию; компонент отсутствует, а продуктовая surface была демонтирована | **DRIFTED** — owner/API boundary устарела. |
|
||||
| **AgentEvaluationCard** | Нет под этим именем | Два соседних, неэквивалентных прообраза: `VlmFindingReviewCard` (`039/spec.md:142-163`) и `AgentEvaluationPanel` (`045/data-model.md:33`) | UX требования распределены между authoring disposition и immutable run result | 045 FR-012/013 задают display acceptance (`045/spec.md:126-127`) | ScenarioResult DTO/card отсутствуют | **IMPLIED** — требуется канонизировать один read-only ScenarioRun component contract. |
|
||||
|
||||
Вывод по семи: прямо и достаточно специфицирован **1** (`AgentEvaluation`); прямо специфицирован, но traceability drifted **1** (`AgentEvaluationProvider`); прямо, но неполно **1** (`DecisionPolicy`); явно назван, но устарел **1** (`EvidencePanel`); inferred glue **2** (browser driver, evaluation card); полностью отсутствует boundary **1** (artifact content service).
|
||||
|
||||
## Trace matrix: modified subsystems
|
||||
|
||||
| Contract unit | Главный normative anchor | Code/runtime evidence | Category | Что не закрыто |
|
||||
|---|---|---|---|---|
|
||||
| Deterministic scenario/revision/run orchestration | SCEX-FR-001..007 (`044/spec.md:107-114`), RunnerPlan/StepOutcome schemas | `runner_plan.py`, `dispatch.py`, `runner.py`; ss-prod persistence/resume/cancel evidence | **COMPLETE** | Live provider result не доказан, но core contract определён и реализован. |
|
||||
| Live composition + provider event loop | `044/spec.md:155-177`; ProviderRuntime PRE/POST (`contracts/modules.md:199-226`) | `app.py:224-231` bootstraps without `provider_runtime.start()` although `provider_runtime.py:62-93` supplies it; ss-prod `not_started` | **DRIFTED** | Startup/shutdown binding violates explicit exactly-one-loop contract. |
|
||||
| Compiler + ActionRegistry | `038/data-model.md:50-79`; 044 FR-002/016 | Model omits `sql_evidence`/`transform` (`scenario/models.py:140-159`) while registry/executors include them; one template emits one action | **DRIFTED** | Canonical schema, template expansion and runtime registry are not one catalog. |
|
||||
| Browser provider/transport parity | Browser per-action policy (`044/contracts/modules.md:324-347`) | Registry promises filters/pagination/download/navigation; transport accepts only open/wait/refresh/row_edit/bulk_edit | **DRIFTED** | Exact action names, I/O schemas and capability readiness do not match. |
|
||||
| ScreenshotService + Scenario provider | 017 FR-039..041/059 (`017/spec.md:118-123`); 044 FR-009/018 | Capture/media stack and 044 provider exist, yet `044/spec.md:319-320` still says no adapter deployed | **DRIFTED** | Spec mixes code absence, deployment absence and stale implementation claim; masking policy also conflicts with 036 local-only residual risk. |
|
||||
| Multimodal LLM reader | 017 FR-045/047/050/056 (`017/spec.md:126-136,155`); 038 VlmAnalysisSpec | Existing client builds images; `scenario/vlm.py:128-163` hardcodes PNG and turns malformed response into `[]` | **PARTIAL** | Scenario evaluation request/response, MIME propagation, parser failure, provenance/raw-response handling and payload fallback lack one contract. |
|
||||
| Generic artifact/evidence ownership | SCEX-FR-018 (`044/spec.md:129-132`); EvidenceReceipt schema | Provider writes digest/ref metadata; no content delivery contract | **PARTIAL** | Producer receipt is strong; consumer bytes/auth/retention read semantics are absent. |
|
||||
| Cancellation/retry/reconcile | SCEX-FR-020/021 (`044/spec.md:135-140`); ProviderOperations (`contracts/modules.md:392-403`) | Runner drains DB state; browser receipt exists; lifecycle does not invoke provider cancel by operation_id | **DRIFTED** | Exact spec is stronger than runtime; screenshot/LLM receipts and cancellation are absent. |
|
||||
| Scenario result/evidence/evaluation UI | RUNMON-FR-012/013 (`045/spec.md:126-127`); `AgentEvaluationPanel` (`045/data-model.md:33`) | `ScenarioResultView.svelte:11-40` renders counts/failures/provenance; DTO `scenario-run.ts:74-85` has no evidence/evaluation | **DRIFTED** | Typed content URL, evidence manifest, evaluation and policy outcome projections. |
|
||||
| Automation | SCAUTO-FR-010/015/017/018 (`046/spec.md:111-119`) | ss-prod found REST/MCP eligibility divergence (DEF-02) | **DRIFTED** | Same service contract exists but transports do not enforce identical eligibility. |
|
||||
| Analytics/investigation | 047 factual audit (`047/spec.md:114-128`) | ss-prod accepted investigation remained queued (DEF-04) | **DRIFTED** | Case transition/evidence snapshot does not satisfy declared lifecycle. |
|
||||
| MCP external surface | MCPX-FR-022..029 (`050/spec.md:187-194`) | Tools exist, but spec package has no OpenAPI/tool schemas/modules/checklist/traceability; status still “Ready” | **PARTIAL** | Inputs/outputs/errors/idempotency/permission parity are prose-only or code-derived, not stable normative artifacts. |
|
||||
|
||||
### Count reconciliation
|
||||
|
||||
The 19 mutually exclusive units are the seven minimum modules plus twelve modified subsystems above:
|
||||
|
||||
| COMPLETE | PARTIAL | IMPLIED | MISSING | DRIFTED | Total |
|
||||
|---:|---:|---:|---:|---:|---:|
|
||||
| 2 | 4 | 2 | 1 | 10 | 19 |
|
||||
|
||||
## Contracts present in specs but absent or divergent at runtime
|
||||
|
||||
1. **Browser action parity:** 044 specifies 15 actions from `open_dashboard` through `refresh` (`contracts/modules.md:324-347`); runtime transport accepts five and uses `row_edit` where the spec says `edit_row`. This is code drift, not missing product intent.
|
||||
2. **Provider binding/loop:** exactly one startup-owned long-lived loop is explicit (`contracts/modules.md:204-215`), but app startup only registers composition. ss-prod `provider_loop=not_started` matches the code path.
|
||||
3. **Screenshot evidence bytes/MIME:** receipt requires non-zero byte length/content type/SHA (`contracts/openapi.yaml:222-238`), and screenshot provider creates it; no normative or runtime ScenarioRun content read exists. `scenario/vlm.py:139` labels every payload PNG although the capture pipeline sends/stores JPEG.
|
||||
4. **Multimodal evaluation:** immutable AgentEvaluation/summary schemas exist, but no SQLAlchemy model/migration/provider/executor. `vlm_analyze` remains an ordinary assertion action and cannot lawfully produce SCEX-SC-009 authority.
|
||||
5. **DecisionPolicy:** policy is named and exclusive, but mapping rules are not an executable truth table. Runtime has no mapper.
|
||||
6. **Cancellation/retry:** operation-aware cancel/reconcile is explicit (`contracts/modules.md:192-196,392-403`); runtime cancellation stops DB scheduling without provider cancellation acknowledgement.
|
||||
7. **Result UI:** 045 requires evidence and evaluation separated from StepOutcome; actual DTO/view expose only aggregate counts, failures and generic provenance.
|
||||
8. **Automation:** 046 requires identical REST/MCP validation outcomes; ss-prod DEF-02 demonstrates divergence.
|
||||
9. **Analytics:** 047 requires resolved case lifecycle/evidence; its own factual audit and ss-prod DEF-04 show incomplete transition semantics.
|
||||
|
||||
## Code gaps versus deployment/config-only gaps
|
||||
|
||||
### Code/spec gaps
|
||||
|
||||
- new AgentEvaluation persistence/provider/DTO/event flow and deterministic DecisionPolicy;
|
||||
- authenticated ScenarioArtifact content read/stream contract and implementation;
|
||||
- browser action driver parity and compiler chain generation;
|
||||
- provider loop startup/shutdown wiring and operation-aware cancellation;
|
||||
- MIME-correct VLM payload, strict response validation and immutable raw-response provenance;
|
||||
- ScenarioRun evidence/evaluation UI projections;
|
||||
- automation transport parity and analytics case-transition repair.
|
||||
|
||||
### Deployment/config-only after code closure
|
||||
|
||||
- populate trusted `settings.scenario_live_execution_bindings[]` with server-owned environment/query/principal snapshots;
|
||||
- ship/verify Chromium in the backend execution image;
|
||||
- configure an approved active multimodal LLM provider and model/prompt versions;
|
||||
- provision durable evidence storage/retention and run PREPROD browser/screenshot/VLM canaries.
|
||||
|
||||
Current ss-prod browser/screenshot `unregistered`, bindings `0`, and LLM providers `0` are deployment evidence, but `provider_loop=not_started` is an application wiring defect. Configuration alone cannot close action parity, artifact content, evaluation or policy gaps.
|
||||
|
||||
## Critical normative gaps
|
||||
|
||||
### P0 — must be amended before implementation
|
||||
|
||||
1. **ScenarioArtifact content contract:** add authenticated `GET/HEAD` schema, owner authorization, accepted MIME, length/digest verification, retention/tombstone behavior, streaming/range limits, redacted errors and UI-safe URL/ref semantics.
|
||||
2. **DecisionPolicy contract:** versioned schema plus total truth table for hard deterministic failure, model pass/fail/inconclusive, confidence thresholds, disagreement, missing/corrupt evidence, provider/parser errors and precedence. Prove bare LLM verdict cannot set result.
|
||||
3. **Provider lifecycle composition:** specify exact startup/shutdown owner, ordering, idempotency, readiness transition and failure behavior; bind `start/stop` to FastAPI lifespan acceptance.
|
||||
4. **Canonical evaluation chain:** amend 038/044 with an explicit screenshot → owned EvidenceReceipt → AgentEvaluationProvider → immutable AgentEvaluation → DecisionPolicy → StepOutcome graph and descriptor/ref schemas.
|
||||
5. **Unified ScenarioRun result API:** content/evaluation/event DTO consumed by 045; remove dependence on AgentRun DraftList.
|
||||
|
||||
### P1 — required for minimum production release
|
||||
|
||||
- canonical browser action mapping table across 038 schema, templates, provider and transport, including exact `edit_row` naming and output schemas;
|
||||
- operation-aware cancel/reconcile contracts for screenshot and LLM providers, not browser only;
|
||||
- strict multimodal MIME/parser/error/provenance/payload-limit contract;
|
||||
- reconcile `EvidencePanel`, `VlmFindingReviewCard`, `AgentEvaluationPanel` into authoring-review versus immutable run-result components;
|
||||
- align 046 REST/MCP eligibility and 047 case transition acceptance with executable contract tests.
|
||||
|
||||
### P2 — hardening/document integrity
|
||||
|
||||
- resolve 017 mandatory masking versus 036/038 local-provider optional masking policy;
|
||||
- update stale status, task and traceability claims in 038/039/044/050;
|
||||
- give 050 a plan, data model, tool schemas, module contracts, security checklist and traceability matrix;
|
||||
- specify cost/token telemetry, retention deletion proof, load/concurrency SLOs and browser/VLM canary rollback.
|
||||
|
||||
## Spec artifacts to add or amend, in order
|
||||
|
||||
1. **038 amendment:** make current ActionRegistry-backed graph versus future five-way VerificationProgram an explicit version decision; add `contracts/agent-evaluation.schema.json`, `decision-policy.schema.json`, total policy examples and canonical visual-chain fixture.
|
||||
2. **044 amendment:** add `contracts/artifact-content.openapi.yaml` (or routes to current OpenAPI), `ScenarioArtifactContentService`, `AgentEvaluationStore/Provider`, `DecisionPolicyEvaluator`, and lifecycle startup/shutdown contracts with PRE/POST/INVARIANT and error taxonomy.
|
||||
3. **044 acceptance:** add contract suites for all browser actions, JPEG/PNG MIME integrity, digest-on-read, corrupt/missing bytes, malformed VLM response, cancellation at each external-I/O phase, restart/reconcile and no-bare-verdict invariant.
|
||||
4. **045 amendment:** define ScenarioResult evidence/evaluation DTOs and separate `EvidenceViewer`/`AgentEvaluationCard` UX_STATE, including loading/expired/forbidden/corrupt/inconclusive states.
|
||||
5. **039 reconciliation:** either archive AgentWorkspace evidence claims as superseded or re-scope authoring review to typed ScenarioRun DTOs; remove false implementation assertion at `spec.md:187`.
|
||||
6. **017/036 boundary note:** state that capture/client/redaction are reusable primitives, while ScenarioRun ownership/content/evaluation are 044 responsibilities; choose one masking policy by trust-boundary class.
|
||||
7. **046/047 amendments:** encode the observed REST/MCP eligibility and case-state invariants as shared service contract tests.
|
||||
8. **050 completion:** add missing plan/data-model/contracts/checklists/traceability; keep MCP authoring separate from runtime reasoning authority.
|
||||
9. Reconcile every status/checkbox only after linked executable tests pass; never infer completion from `tasks.md`.
|
||||
|
||||
## Verification notes and limitations
|
||||
|
||||
- This audit did not run a new ss-prod cycle and made no production mutation.
|
||||
- No application test suite was rerun because the request was normative coverage, not implementation verification. Existing test names/claims are treated as trace links only unless supported by code/runtime evidence from the prior report.
|
||||
- Absence checks were bounded to the relevant spec and production roots. In particular, no ScenarioRun artifact content path was found in 044 OpenAPI; no production `AgentEvaluation`, `AgentEvaluationProvider` or `DecisionPolicy` implementation was found; and the only runtime hit for `agent_evaluation` was a capacity key/frontend readiness field.
|
||||
- Specs 017 and 036 describe reusable services with different owning entities and trust assumptions. Their existence is not evidence that the ScenarioRun chain is normatively closed.
|
||||
|
||||
## Final decision
|
||||
|
||||
The specs cover the **architecture and safety intent**, but not every module/contract needed for a production browser+screenshot+LLM ScenarioRun. Implementation may safely reuse the present capture and LLM stack, and can implement the already-complete deterministic runner/AgentEvaluation data intent. Before coding the missing vertical bridge, the five P0 amendments above must establish byte delivery, policy authority, lifecycle ownership, chain identity and UI DTOs. Otherwise engineers would have to invent security- and verdict-critical behavior in code.
|
||||
|
||||
#endregion Report.SsProdAgenticE2ESpecCoverage
|
||||
@@ -0,0 +1,25 @@
|
||||
# Agentic E2E specification refresh
|
||||
|
||||
## Scope and decisions
|
||||
|
||||
Update the twelve related specification packages: 017, 036, 037, 038, 039, 042, 043, 044, 045, 046, 047, 050. The scope follows the user's request in the context of the production E2E, spec coverage, and baseline audits. Unrelated repository features are outside this refresh.
|
||||
|
||||
The implementation worker owns the normative files and refresh report. The architect reviews completeness and coordinates independent verification after edits settle.
|
||||
|
||||
Normative requirements must distinguish target behavior from observed implementation. Adding contracts does not close production acceptance gates. Existing successful checks remain historical evidence; unsupported completion claims must be reconciled.
|
||||
|
||||
Reuse ScreenshotService, LLMClient, and baseline comparison engines. Specify the missing integration boundaries explicitly. An approved execution-performance baseline remains an optional future feature, not a silently added requirement.
|
||||
|
||||
## User constraint: external agent interaction only
|
||||
|
||||
Frontend must not expose agent prompts/chat, agent-assisted editing, agent proposal generation or an agent workspace. External MCP clients own agent interaction. Ordinary manual editing, human approvals and read-only run evidence/evaluation remain product UI responsibilities. Existing agent editing panels are implementation drift requiring a removal task; this specification refresh does not itself change frontend runtime code. Supersede contradictory active requirements rather than leaving competing instructions in appendices.
|
||||
|
||||
## Acceptance
|
||||
|
||||
- All twelve packages link to their current responsibilities and implementation gaps.
|
||||
- P0/P1/P2 findings have normative owners, schemas where applicable, errors/invariants and falsifiable acceptance.
|
||||
- Baseline identity is pinned across request idempotency, plan, run, result and analytics.
|
||||
- Artifact delivery, evaluation policy, provider lifecycle and REST/MCP parity have concrete contracts.
|
||||
- Plan/tasks/traceability/status agree with remaining implementation work.
|
||||
- Available static validators, schema parsing and local-reference checks run; limitations are reported.
|
||||
- Runtime source and unrelated existing user changes remain untouched.
|
||||
@@ -0,0 +1,46 @@
|
||||
# SS-Prod Dashboard Scenario E2E — plan and evidence ledger
|
||||
|
||||
## Objective
|
||||
|
||||
Execute and evidence the complete lifecycle for creating and operating a dashboard testing scenario on `ss-prod`, including every exposed read-only and mutating tool, then produce an orthogonal report.
|
||||
|
||||
## Safety and scope
|
||||
|
||||
- Use uniquely named test entities carrying the marker `codex-e2e-20260908`.
|
||||
- Mutations are authorized by the user only inside the dashboard-testing scenario workflow.
|
||||
- Record identifiers before every destructive transition.
|
||||
- Prefer reversible mutations; delete only test entities created by this run.
|
||||
- Do not alter unrelated production scenarios, dashboards, credentials, schedules, or users.
|
||||
|
||||
## Orthogonal verification matrix
|
||||
|
||||
1. Lifecycle: discover prerequisites → create → read/list → update → execute → inspect results/analytics → automation/schedule if exposed → delete/cleanup.
|
||||
2. Tool surface: enumerate all scenario-related tools and invoke each with a valid fixture; separately exercise mutation tools.
|
||||
3. State machine: idle/loading/success/error/empty and persistence after reload where observable.
|
||||
4. Data integrity: request/response identity, persisted fields, run linkage, timestamps/statuses, absence after cleanup.
|
||||
5. Failure boundaries: missing field, invalid type/value, missing resource or external failure where safely reproducible.
|
||||
6. Isolation: prove no unrelated production object was mutated.
|
||||
|
||||
## Evidence ledger
|
||||
|
||||
The verification worker will append exact endpoints/tools, entity IDs, timestamps, observed results, cleanup state, and defects/gaps here or in a sibling final report.
|
||||
|
||||
## Decisions
|
||||
|
||||
- @RATIONALE Production mutations are constrained to uniquely marked disposable entities so destructive coverage does not widen into unrelated production state.
|
||||
- @REJECTED Reusing or deleting an existing scenario: ownership and cleanup cannot be proven.
|
||||
- @REJECTED Treating UI visibility as proof of persistence: reload/API evidence is required where available.
|
||||
|
||||
## Follow-up: production-readiness gap analysis
|
||||
|
||||
Compare the executed ss-prod behavior with repository specifications and implementation for the complete agentic loop:
|
||||
|
||||
`scenario orchestration -> browser provider -> screenshots -> LLM visual reading -> assertions/verdict -> retry/recovery -> artifacts/analytics`.
|
||||
|
||||
Required evidence:
|
||||
|
||||
- Trace each specified capability to concrete production modules and registered runtime providers.
|
||||
- Inspect the existing screenshot service and LLM image-reading path rather than estimating them as greenfield work.
|
||||
- Classify gaps as missing, present-but-unwired, incomplete, defective, or deployment/configuration-only.
|
||||
- Estimate additional modules and source/test code using explicit scope assumptions and non-overlapping work packages.
|
||||
- Separate minimum production-ready scope from reliability/scale hardening.
|
||||
206
docs/reports/ss-prod-dashboard-scenario-e2e-report-2026-09-08.md
Normal file
206
docs/reports/ss-prod-dashboard-scenario-e2e-report-2026-09-08.md
Normal file
@@ -0,0 +1,206 @@
|
||||
# SS-Prod Dashboard Scenario E2E Report — 2026-09-08
|
||||
|
||||
#region Report.SsProdDashboardScenarioE2E [C:4] [TYPE Verification] [SEMANTICS scenario,e2e,ss-prod,cleanup]
|
||||
|
||||
## Executive verdict
|
||||
|
||||
**PARTIAL PASS / execution-environment BLOCKED.** The authenticated ss-tools lifecycle completed durable create, list/detail/reload, two persisted updates, status transitions, PROD-gated approval, human-checkpoint completion, results/events/compare, analytics investigation, automation CRUD, cancel/retry/resume guards, and cleanup. No foreign production object was mutated.
|
||||
|
||||
The available compiler emitted a manual-only scenario. It completed honestly as `inconclusive`; no visual PASS was fabricated. Automated browser/screenshot/VLM execution remains blocked because the provider loop is not started, browser and screenshot providers are unregistered, and no LLM provider is configured.
|
||||
|
||||
Exact accounting:
|
||||
|
||||
- MCP scenario/dashboard-testing catalog: **35 relevant tools; PASS 28, FAIL 1, BLOCKED 5, N/A 1**.
|
||||
- Live REST/OpenAPI catalog: **61 relevant operations; PASS 58, FAIL 0, BLOCKED 3, N/A 0**. Expected 4xx negative/state guards count PASS only when rejection was the asserted contract.
|
||||
- Created: 2 AgentRuns, 3 scenarios (including clone), 5 revisions, 2 ScenarioRuns, 2 authoring workspaces, editor drafts/proposal, 1 policy, 1 schedule, 1 trigger rule, and 1 investigation case.
|
||||
- Cleanup: policy/schedule/rule deleted; all 3 scenarios archived; runs terminal (`inconclusive`, `cancelled`); investigation case accepted. Audit records remain because no delete API exists.
|
||||
|
||||
## Environment and evidence
|
||||
|
||||
| Item | Evidence |
|
||||
|---|---|
|
||||
| Window | Initial auth audit `2026-09-08T03:14:41Z`; authenticated lifecycle `05:46:45Z`–`06:20Z` (UTC; local UTC+03:00) |
|
||||
| ss-tools | UI `http://127.0.0.1:5173`; backend `http://127.0.0.1:8000`; UI login succeeded; `/api/auth/me` → `200`, Admin principal |
|
||||
| Target | `environment_id=ss-prod`, Dashboard `11`; query model `200`: 10 charts, 1 dataset, 7 filters; Superset title `Sales Dashboard`, published slug `sales` |
|
||||
| Readiness | `/api/ready` → `200 ready`; evidence storage ready; provider loop `not_started`; browser/screenshot `unregistered`; bindings `0` |
|
||||
| Plugins/health | `/api/plugins` → 11 records; health summary `200` empty; encryption health `200 healthy` |
|
||||
| LLM | MCP provider list → 0; LLM status not ready |
|
||||
| MCP | Authenticated `/mcp/`, protocol `2025-06-18`, server `ss-tools 2.2.0`; live `tools/list` exactly 61 tools |
|
||||
| Inventory sources | Live OpenAPI + authenticated MCP `tools/list` + backend routes/registrations + frontend API/models |
|
||||
| Secrets | Password, bearer, cookies, token and MCP session values are absent from report/evidence |
|
||||
|
||||
The earlier credential pair was tested in both normal ss-tools flows (rendered login and direct backend login) and returned `401`, not `403` or CSRF/session failure. The later authorized credential succeeded in those flows. Browser-standard storage checks found the authenticated token only after success; it was used in-memory and cleared after the run.
|
||||
|
||||
## Test invariants and hardcoded fixtures
|
||||
|
||||
| Invariant | Fixture / expectation | Result |
|
||||
|---|---|---|
|
||||
| `INV-TARGET` | `ss-prod`, dashboard `11` resolves to published Sales Dashboard | PASS |
|
||||
| `INV-MARKER` | Names/comments/idempotency keys use `codex-e2e-20260908` | PASS |
|
||||
| `INV-OWNERSHIP` | Mutate/delete/archive only IDs issued here | PASS |
|
||||
| `INV-PERSISTENCE` | Scenario survives list/detail/revision checkout/reload | PASS |
|
||||
| `INV-PINNING` | Run pins scenario, revision, content hash, `ss-prod` | PASS |
|
||||
| `INV-NO-FABRICATION` | Manual visual assertion cannot be reported passed without evidence | PASS: `inconclusive` |
|
||||
| `INV-PROD-GATE` | PROD run starts pending approval | PASS |
|
||||
| `INV-CLEANUP` | Automation absent; scenarios archived; runs terminal | PASS with audit residuals |
|
||||
|
||||
Hardcoded fixtures: marker `codex-e2e-20260908`; target `ss-prod`; dashboard `11`; case `B01`; step `phase-5-B01-human_checkpoint-1`; disabled cron `0 0 1 1 *` UTC; missing provider `codex-e2e-20260908-missing`; invalid trigger `invalid_codex_e2e`.
|
||||
|
||||
## Inventory: MCP tools
|
||||
|
||||
All were reconciled against authenticated `tools/list` and source. `PASS-negative` is an expected typed rejection without foreign mutation.
|
||||
|
||||
| Tool(s) | Result | Evidence |
|
||||
|---|---|---|
|
||||
| `list_environments`, `get_health_summary`, `search_dashboards` | PASS | 2 envs; empty health; one Sales match |
|
||||
| `list_llm_providers`, `get_llm_status` | PASS | 0 providers; not ready |
|
||||
| `create_agent_run`, `get_agent_run` | PASS | `70491960-a40b-4798-94a8-8be886de8d2f`; prior surface concern resolved |
|
||||
| `inspect_dashboard_context` | PASS | 10 charts, 1 dataset, 7 filters; browser capability false |
|
||||
| `inspect_scenario`, `validate_scenario`, `scenario_resolve` | PASS | manual B01 graph; valid; resolver returned hash `32c7b39c…` |
|
||||
| `generate_draft_pack`, `register_draft_pack` | PASS | pack `42acb0cf-8ddd-4cef-8de6-56f06e89bf74`, context verified |
|
||||
| `bootstrap_authoring_scenario` | PASS | scenario `d5491d2c-9509-4bcd-8212-0cfa6bd381fb`; rev `0c52f215-cd3e-4562-8b29-6119188cfac6`; workspace `2b0e907f-7220-4e27-a966-eb14d96d13f7` |
|
||||
| `create_authoring_session`, `propose_test_plan` | PASS | workspace `99be7861-aa9f-46ad-ae5e-946b8009cc7e`; CAS advanced |
|
||||
| `start_exploration` | BLOCKED | callable/persisted; `sandbox_unavailable` |
|
||||
| `get_exploration_result` | PASS | same request/operation and blocker |
|
||||
| `propose_graph_revision`, `get_graph_diff` | PASS | proposal `3ac514b8-b6bf-47be-b2e3-9380d5cf7712` |
|
||||
| `promote_to_scenario` | FAIL | server proposal rejected unsafe; `DEF-01` |
|
||||
| `request_save` | BLOCKED | workspace `validation_blocked`, not awaiting review |
|
||||
| `activate_revision` | BLOCKED | workspace draft, not candidate |
|
||||
| `start_scenario_run` | PASS | run `09b9e548-4cdb-4998-983b-f7501a8ef801`, pending approval |
|
||||
| `get_task_status` | PASS-negative | run UUID is not task; `not_found` |
|
||||
| `list_checkpoints`, `decide_checkpoint` | PASS | checkpoint listed; duplicate decision safely `not_found` |
|
||||
| `list_pending_approvals` | PASS | empty invocation approvals; run gate is separate REST gate |
|
||||
| `decide_approval` | N/A | no owned MCP invocation approval; no foreign decision |
|
||||
| `get_scenario_automation_policy`, `get_scenario_automation_metrics` | PASS | scenario-specific not-found/default metrics |
|
||||
| `upsert_scenario_automation_policy` | PASS | disabled, PROD gate retained; `5c6e15ee-64ef-4a5a-8cf4-20ccbaf652ec` |
|
||||
| `list_scenario_schedules`, `list_scenario_trigger_rules` | PASS | filtered reads |
|
||||
| `upsert_scenario_schedule`, `upsert_scenario_trigger_rule` | BLOCKED | `AUTOMATION_INELIGIBLE_HUMAN_STEP`; REST differs (`DEF-02`) |
|
||||
| `delete_scenario_schedule`, `delete_scenario_trigger_rule` | PASS | exact owned IDs deleted |
|
||||
|
||||
**28 PASS / 1 FAIL / 5 BLOCKED / 1 N/A = 35.**
|
||||
|
||||
## Inventory: REST/UI surface
|
||||
|
||||
Every live OpenAPI operation was invoked. UI actions map to these endpoints; login, registry list and persisted reload were also observed in UI. Visual execution is blocked separately.
|
||||
|
||||
| Group | Operations and result |
|
||||
|---|---|
|
||||
| Compile/capture (7) | compile, validate, resolve, draft-pack, disposition PASS; capture BLOCKED (`VALIDATION_ERROR`, no capture provider); VLM BLOCKED (`STALE_PROMPT`, no provider) |
|
||||
| Registry (12) | list/create/revision-list/checkout/diff/transition/clone/archive/restore/health/detail PASS; raw revision create PASS-negative (`422` missing snapshot) |
|
||||
| Editor (6) | load/apply/save/agent-propose/proposal-save PASS; revalidate PASS-negative (`409` after archive; earlier missing-query `422`) |
|
||||
| Runs (12) | create/list/compare/detail/events/cancel/approval/human/result/history PASS; retry/resume PASS-negative (`409` on cancelled run) |
|
||||
| Analytics (7) | queue/open/detail/disposition/health/trends/recurring PASS; queue defect below |
|
||||
| Automation writes (11) | schedule/rule/policy create/update/delete and event dispatch PASS; direct trigger BLOCKED (manual-only) |
|
||||
| Automation reads (6) | schedules/rules/policies/notifications/metrics/retention PASS functionally; anonymous exposure `SEC-01` |
|
||||
|
||||
**58 PASS / 0 FAIL / 3 BLOCKED / 0 N/A = 61.**
|
||||
|
||||
## Orthogonal matrix
|
||||
|
||||
| Dimension | Result | Evidence |
|
||||
|---|---|---|
|
||||
| Prerequisites | PASS | auth, ready, env/dashboard/query-model/plugin/provider reads |
|
||||
| Create | PASS | REST and MCP create pipelines; clone |
|
||||
| Persisted reload | PASS | marker list/detail/revisions/checkout/editor load |
|
||||
| Update/config/status | PASS | revs `37f68e67-…`, `097070ad-…`; READY→DISABLED→READY |
|
||||
| Execution | PARTIAL/BLOCKED | gated manual run inconclusive; providers absent |
|
||||
| Status/results/logs | PASS | detail, SSE, terminal result, compare |
|
||||
| Analytics | PASS with defect | case opened/accepted; queue row remained |
|
||||
| Automation | PASS with defect | disabled REST CRUD/delete; MCP rejects same config |
|
||||
| Mutations | PASS | clone, transitions, archive/restore, approve, decide, cancel; guards |
|
||||
| Negative boundaries | PASS | missing fields, invalid values, missing resource/task, state conflicts |
|
||||
| Isolation/cleanup | PASS | exact owned IDs only |
|
||||
|
||||
## Chronological lifecycle
|
||||
|
||||
1. `05:46:45Z`: authenticated readiness/provider/plugin audit; Dashboard 11 model resolved.
|
||||
2. REST AgentRun `c953464e-78bc-4f89-ab75-840e71b941b5` created after expected invalid-intent/objective 422s.
|
||||
3. Compiler produced manual B01. Initial draft registration returned `CONTEXT_QUERY_MODEL_REQUIRED`; echoing the authoritative query model and recomputing canonical hash made pack `40203635-a1af-475d-9003-89a41b8ec966` eligible.
|
||||
4. Scenario A `99e0a888-c7e6-459c-b977-e51755f2b19c`, rev `135b7c5d-39e7-4801-8db0-ebd05f6e057f` persisted and reloaded.
|
||||
5. Authoring plan/exploration/proposal/diff ran. Sandbox unavailable; promotion rejected its server-derived graph.
|
||||
6. MCP AgentRun `70491960-…`, pack `42acb0cf-…`, and scenario B/current revision/workspace persisted.
|
||||
7. Run `09b9e548-4cdb-4998-983b-f7501a8ef801` started pending approval, advanced to waiting-human, then finished `inconclusive/completed` at `06:08:08Z`.
|
||||
8. Editor apply/save and agent-propose/save produced two candidate revisions. State traversed READY→DISABLED→READY. Clone `64cf8888-ecda-45e9-a78f-dee0c5f2cef0` created. Scenario A archive→restore exercised.
|
||||
9. Disabled policy/schedule/rule created/updated. Invalid trigger → `422`; event dispatch → `202`, zero runs; direct trigger rejected manual-only.
|
||||
10. Run `b269e15a-ec38-48d6-9da3-2540d85bfe7c` created pending approval, cancelled, and remained terminal; retry/resume guarded with 409. Compare → 200.
|
||||
11. Owned analytics case `53598184-6e70-401c-b748-0f36cf39c112` opened and accepted, version 2.
|
||||
12. Schedule/rule/policy deleted; all three scenarios archived; zero automation configs verified.
|
||||
|
||||
## Mutation ledger
|
||||
|
||||
| Entity | ID(s) | Final state |
|
||||
|---|---|---|
|
||||
| AgentRuns | `c953464e-78bc-4f89-ab75-840e71b941b5`; `70491960-a40b-4798-94a8-8be886de8d2f` | audit records; no delete |
|
||||
| Scenario A / rev | `99e0a888-c7e6-459c-b977-e51755f2b19c` / `135b7c5d-39e7-4801-8db0-ebd05f6e057f` | ARCHIVED |
|
||||
| Scenario B / current rev | `d5491d2c-9509-4bcd-8212-0cfa6bd381fb` / `0c52f215-cd3e-4562-8b29-6119188cfac6` | ARCHIVED |
|
||||
| Editor revs | `37f68e67-8aa1-4ba9-89bd-03d0383ec1f8`; `097070ad-6950-4222-aae9-615e2f939514` | under archived scenario |
|
||||
| Clone / rev | `64cf8888-ecda-45e9-a78f-dee0c5f2cef0` / `5400204f-b824-451b-9067-46ae2a704130` | ARCHIVED |
|
||||
| Runs | `09b9e548-4cdb-4998-983b-f7501a8ef801`; `b269e15a-ec38-48d6-9da3-2540d85bfe7c` | inconclusive; cancelled |
|
||||
| Policy | `5c6e15ee-64ef-4a5a-8cf4-20ccbaf652ec` | DELETE 204; list 0 |
|
||||
| Schedule | `3eb4e020-7fd6-4326-b589-56882b08af36` | MCP delete; list 0 |
|
||||
| Trigger rule | `62a0911f-8173-49b6-92b5-42eef8cb168d` | MCP delete; list 0 |
|
||||
| Investigation | `53598184-6e70-401c-b748-0f36cf39c112` | accepted v2; no delete |
|
||||
|
||||
## Negative tests
|
||||
|
||||
| Edge | Observed | Result |
|
||||
|---|---|---|
|
||||
| Invalid AgentRun intent | `422`; required build intent | PASS |
|
||||
| Missing objective | `422`; goal/case IDs required | PASS |
|
||||
| Missing context authority | `422 CONTEXT_QUERY_MODEL_REQUIRED` | PASS contract; `DEF-03` chain mismatch |
|
||||
| Invalid exploration fields | schema reject; valid request sandbox-unavailable | PASS/BLOCKED |
|
||||
| Duplicate checkpoint decision | safe `not_found` | PASS |
|
||||
| Missing task/resource | task `not_found`; missing Superset dashboard `404` | PASS |
|
||||
| Invalid trigger | `422 INVALID_TRIGGER` | PASS |
|
||||
| Missing idempotency header | direct trigger `422` | PASS |
|
||||
| Cancelled retry/resume | `409 RETRY_CONFLICT` / `RESUME_CONFLICT` | PASS |
|
||||
| Providerless capture/VLM | `422 VALIDATION_ERROR` / `STALE_PROMPT` | BLOCKED |
|
||||
|
||||
## Defects and blockers
|
||||
|
||||
### `DEF-01` High — server-derived proposal cannot promote
|
||||
|
||||
`propose_graph_revision` reported a valid proposal, but `promote_to_scenario` blocked it as unsafe SQL/executable/path content. The diff unexpectedly added `assertions`/`dependencies`, changed `parameters []` to `{}`, and did not remove the requested step. MCP request-save/activation therefore cannot complete.
|
||||
|
||||
### `DEF-02` High — MCP and REST disagree on automation eligibility
|
||||
|
||||
MCP schedule/rule upserts correctly returned `AUTOMATION_INELIGIBLE_HUMAN_STEP`; REST POSTs persisted the same scenario/revision as disabled configs (`201`). Direct REST trigger later rejected it. Invalid configs can enter through one public surface; both were deleted.
|
||||
|
||||
### `DEF-03` Medium — compiler output is not directly draft-registerable
|
||||
|
||||
Compiled graph lacked authoritative `dashboard_context.query`; registration failed until the separately returned query model was echoed and the hash recomputed. The compiler/inspect composition contract is incomplete or undocumented.
|
||||
|
||||
### `DEF-04` Medium — accepted investigation remains queued
|
||||
|
||||
Case `53598184-…` was accepted at version 2, but subsequent queue read still returned its source row as `queued`, allowing duplicate triage/reopen.
|
||||
|
||||
### `SEC-01` Medium — six automation reads are anonymous
|
||||
|
||||
Unauthenticated schedules, trigger-rules, policies, notifications, metrics and retention GETs returned `200`; adjacent writes returned `401`. Initial collections were empty: no secrets, identities, scenario/env IDs, cron, policy or notification body appeared. Only zero aggregates and retention tiers were exposed. Populated endpoints could disclose operational metadata; no broader exploitation occurred.
|
||||
|
||||
### Environment blockers, not defects
|
||||
|
||||
- browser/screenshot providers unregistered; provider loop not started; bindings zero;
|
||||
- no LLM/VLM provider;
|
||||
- exploration sandbox unavailable;
|
||||
- compiled scenario manual-only, hence automation legitimately ineligible.
|
||||
|
||||
Authenticated `tools/list` includes `create_agent_run` and `get_agent_run`; the earlier provisional surface concern is resolved.
|
||||
|
||||
## Cleanup proof, limitations, next actions
|
||||
|
||||
Final reads: schedules `[]`, trigger rules `[]`, policies `[]`; each exact scenario detail has `entry.lifecycle_status=ARCHIVED`; scenario B has exactly two owned terminal runs (cancelled and inconclusive). No dashboard, dataset, environment, foreign scenario, user, approval or foreign case changed.
|
||||
|
||||
Residuals are audit state, not missed deletion: registry has archive/restore but no scenario hard-delete; AgentRuns, revisions, workspaces/drafts/proposals, ScenarioRuns and cases expose no deletion. `DEF-04` leaves one owned queue row incorrectly queued.
|
||||
|
||||
Exact next actions:
|
||||
|
||||
1. Register/start browser and screenshot providers/provider loop; configure approved LLM only for VLM.
|
||||
2. Fix `DEF-01`; repeat promotion→save→activation on a fresh marker.
|
||||
3. Apply automation eligibility identically to REST/MCP and add cross-surface tests.
|
||||
4. Make compiler output satisfy draft context authority or provide/document composition.
|
||||
5. Dequeue atomically when a case opens/disposes; prove fingerprint absence.
|
||||
6. Decide and enforce authenticated VIEW permission for the six anonymous reads.
|
||||
7. Rerun with a real provider-backed read-only step; require evidence-backed terminal PASS.
|
||||
|
||||
#endregion Report.SsProdDashboardScenarioE2E
|
||||
94
docs/tasks/2026-09-11-data-team-maintenance-dev-test-run.md
Normal file
94
docs/tasks/2026-09-11-data-team-maintenance-dev-test-run.md
Normal file
@@ -0,0 +1,94 @@
|
||||
# Задача: тестовый прогон maintenance (banner) на DEV и интеграция скриптов
|
||||
|
||||
**Команда:** Data Engineering
|
||||
**Приоритет:** P2
|
||||
**Окружение:** только DEV (PROD не трогать)
|
||||
**Дата постановки:** 2026-09-11
|
||||
**Спека:** `specs/031-maintenance-banner/`
|
||||
|
||||
## Контекст
|
||||
|
||||
В superset-tools реализован функционал maintenance banner: внешний инструмент (ETL, cron,
|
||||
CI/CD, Airflow) через API запускает «плашку» «Ведутся технические работы» на дашбордах
|
||||
Superset и снимает её. Нужно прогнать функционал на DEV и встроить скрипты запуска/снятия
|
||||
плашек в инструменты команды.
|
||||
|
||||
## Цель
|
||||
|
||||
1. Добавить в инструменты команды (Airflow DAG / CI job / cron / runbook) скрипты запуска
|
||||
и снятия плашек, работающие против DEV-окружения.
|
||||
2. Прогнать функциональный сценарий на DEV и подтвердить корректность баннеров.
|
||||
|
||||
## 1. Подготовка доступа
|
||||
|
||||
- Создать API-ключ в superset-tools с permissions: `maintenance:start`, `maintenance:end`,
|
||||
`maintenance:end_all` (Admin → API Keys).
|
||||
- Узнать точный id DEV-окружения: `GET {BASE_URL}/api/environments`. Во всех командах
|
||||
ниже `{DEV_ENV_ID}` — это реальный id из ответа (в примерах фигурирует `ss-dev` —
|
||||
сверить с конфигом стенда).
|
||||
- Определить `{BASE_URL}` DEV-стенда (например, `http://127.0.0.1:8000` для локального
|
||||
или адрес dev-стенда команды).
|
||||
- Ключ хранить в секретах (CI/CD secrets, vault, `.env` вне git). **Не передавать ключ
|
||||
аргументом командной строки.**
|
||||
|
||||
## 2. Включить скрипты в инструменты команды
|
||||
|
||||
Взять готовые примеры и адаптировать под свой пайплайн:
|
||||
|
||||
| Файл | Назначение |
|
||||
|---|---|
|
||||
| `examples/maintenance/maintenance-api-bash.sh` | bash/curl — для shell, cron, CI |
|
||||
| `examples/maintenance/maintenance-api-python.py` | Python/requests — для ETL/Airflow (`start_maintenance` / `end_maintenance` / `end_all_maintenance` можно импортировать) |
|
||||
| `examples/maintenance/README.md` | описание API, параметров и ошибок |
|
||||
|
||||
Требования к интеграции:
|
||||
|
||||
- скрипты добавлены в репозиторий команды (например `dags/`, `ci/`, `scripts/`);
|
||||
- окружение (`{DEV_ENV_ID}`) и `{BASE_URL}` параметризованы, по умолчанию — DEV;
|
||||
- API-ключ читается из переменной окружения/секрет-хранилища;
|
||||
- ненулевой код возврата на ошибках, чтобы пайплайн не «проглатывал» сбой.
|
||||
|
||||
## 3. Тест-сценарии на DEV
|
||||
|
||||
| # | Сценарий | Команда (bash) | Ожидание |
|
||||
|---|---|---|---|
|
||||
| 1 | Старт maintenance на известной таблице | `./maintenance-api-bash.sh start schema.table 4 {DEV_ENV_ID}` | `202`, `task_id` + `maintenance_id`; плашка на затронутых дашбордах |
|
||||
| 2 | Проверка событий и баннеров | `GET /api/maintenance/events`, `GET /api/maintenance/dashboard-banners` | событие `active`, баннер виден в Superset |
|
||||
| 3 | Ручное завершение | `./maintenance-api-bash.sh end {maintenance_id}` | `202`; при повторе `already_completed` (идемпотентно); плашка снята |
|
||||
| 4 | Идемпотентность старта | повтор сценария 1 с тем же окном | `409 {status:"already_active"}`, тот же `maintenance_id` |
|
||||
| 5 | Авто-завершение | `./maintenance-api-bash.sh start schema.table 4 {DEV_ENV_ID} "ETL" --auto-end` | плашка снимается автоматически в `end_time` |
|
||||
| 6 | Аварийное снятие всех | `./maintenance-api-bash.sh end-all {DEV_ENV_ID}` | все активные события завершены, плашки сняты |
|
||||
| 7 | Негатив: неизвестное окружение | `start schema.table 4 unknown-env` | `404` |
|
||||
| 8 | Негатив: ключ без права | ключ без `maintenance:start` | `403` |
|
||||
|
||||
Проверить в UI superset-tools (Admin, вкладка Maintenance) список активных/завершённых
|
||||
событий и настройки; баннер — в Superset на дашборде, который использует указанную таблицу.
|
||||
|
||||
## 4. Артефакты
|
||||
|
||||
- PR/ссылка на скрипты, добавленные в инструменты команды.
|
||||
- Короткий отчёт: команды, HTTP-коды, `task_id` / `maintenance_id` по каждому сценарию.
|
||||
- Скриншоты плашки «до» и «после» снятия на DEV.
|
||||
- Подтверждение, что PROD не затронут.
|
||||
|
||||
## Критерии приёмки
|
||||
|
||||
- [ ] Скрипты запуска/снятия плашек есть в инструментах команды, параметризованы по окружению.
|
||||
- [ ] На DEV плашка появляется на дашбордах, затронутых указанными таблицами.
|
||||
- [ ] `end` и `end-all` снимают плашку; повторные вызовы идемпотентны.
|
||||
- [ ] `--auto-end` снимает плашку по истечении окна.
|
||||
- [ ] Ошибки `403` / `404` / `409` обрабатываются скриптом, пайплайн не падает молча.
|
||||
- [ ] Ключ не передаётся в командной строке; используется секрет-хранилище.
|
||||
|
||||
## Ограничения и безопасность
|
||||
|
||||
- Тестировать только на DEV. **Fan-out без `environment_id` стартует во всех PROD-окружениях
|
||||
— на этом прогоне не использовать.**
|
||||
- `end-all` снимает плашки со всех дашбордов — применять осознанно.
|
||||
- Не коммитить API-ключ и `.env` со секретами.
|
||||
|
||||
## Ссылки
|
||||
|
||||
- `examples/maintenance/README.md`
|
||||
- `specs/031-maintenance-banner/spec.md`, `quickstart.md`
|
||||
- API: `{BASE_URL}/api/maintenance/*` (все мутации возвращают `202` + `task_id`)
|
||||
Reference in New Issue
Block a user