- skill §11 rewritten: one invariant (byte-identical prefix) + explicit
preserve/invalidate lists + discipline (load self-orchestration once;
persona/toolFilter/model fixed for a worker's whole life)
- worker skills: "disposable context" -> "long-lived context"; role lines
now say "leaf, long-lived, refined in place via send_message"
Workers are continuable long-lived children: spawn once, refine in place
via send_message as the feature evolves; a worker's own session persists
and compacts independently. Re-spawn only when the context is poisoned.
- skill §3: delegation tree rewritten — "refine, don't re-spawn"
- skill §5: "refine, don't re-spawn" coordination rule
- skill §7: stalled worker → send_message (keep context), fresh spawn
only for poisoned context
- contracts: Self.Worker.Implement is "long-lived, refined in place";
add long-lived invariant
Drop the "never read raw files" rule — it was a false economy. The
architect's context is large and auto-compacting, so reading is cheap and
necessary for decomposition and verification. The real boundary is
EXECUTION (edits/builds/tests) stays with workers, not READING.
- skill §1/§6/§10: "read freely" replaces "never read raw files";
prefer read_outline/search_contracts to locate, read/grep/glob to understand
- @RATIONALE reframed: reading is cheap; only DECISIONS must not live
solely in context (they go to files)
- contract: invariant becomes "reads freely for decisions, delegates execution"
Session-log review found the orchestrator diverging from its own design.
Encode every divergence as a rule + invariant:
- no polling: get_goal/list_agents are state tools, not completion checks
(settlement/report are the signals) — was 101 get_goal + 98 list_agents
- no raw reads: structure via read_outline/search_contracts; file CONTENT
is delegated (was 77 read vs 1 read_outline)
- <RESULT> enforcement: a worker result with no envelope is blocked
(was 4/19 envelopes)
- closed role taxonomy: only Implement/Verify/Curate (was ad-hoc
"code reviewer"/"auditor"/"adversarial")
- fork semantics: fork inherits MY context, not a stalled worker's;
no "takeover" via fork
- interrupt only to redirect; let workers finish (was 7 interrupts)
- decision memory written by me to files (was zero file writes)
- skill hygiene: load self-orchestration once; never load worker skills
- cross-workspace: point Axiom at the target or delegate all reading
Child subagents join the parent's preset composition, so by default they
inherited the orchestrator persona AND the delegation tools — and drifted
into orchestrating instead of working (observed: a "worker" called
list_agents/get_goal/send_message and spawned its own grandchildren).
Fixes:
- preset: subagent/subagent_fork now carry a role-agnostic worker `persona`,
`toolFilter.deny` for all delegation/coordination tools, and `maxDepth: 1`
(children cannot spawn grandchildren) — hard enforcement at the boundary
- skill: every worker prompt opens with a mandatory role-reset block
- contracts: Self.Worker.Implement/Verify gain a LEAF invariant (no subagents,
no delegation tools)
- ENCRYPTION_KEY: smoke test generates a fresh key inline; templates use a placeholder
- drop RUSAL_ROOT.cer (client-specific public cert) + gitignore it
Remove dead schema left behind by removed features:
- dataset-review family (dataset_review_sessions, dataset_profiles, and
related children) from c3ad0afc — its non-cascading FKs broke environment
deletion with ForeignKeyViolation
- connection_configs from 74e64622
Both are unreachable from the app (no models register them).
- bash: propagate api_call failures (exit 1 on 400/401/403/404/network),
write diagnostics to stderr, escape message for JSON safety, help without
API key
- python: argparse options after subcommand (parents), single error message
per failure, network errors without traceback, idempotent already_completed
- move scripts to examples/maintenance/ with README instructions
- backend: correct stale envelope-shape comment in maintenance schemas
- Validate and repair OpenAPI YAML contracts for 038, 043, and 046
- Canonicalize 038 JSON schema and fixtures around scenario_key,
content_hash, and logical_step_id; validate all fixtures with jsonschema
- Regenerate 038 validation evidence for compiler scope only
- Add reconcile_contracts.py for repeatable OpenAPI/JSON/fixture checks
- Replace raw CreateScenario payload with server-owned handles and document
transactional outbox/materialization saga for Registry-to-git persistence
- Record reconciliation outcome in REVIEW-042-047-CLOSURE.md
Scenario compile returned 422 VALIDATION_ERROR because the LLM sent
human-readable case names (smoke, data_integrity, filter_propagation) as
selected_case_ids, but the compiler requires registered catalog ids
(B01-B09, C01-C07, T01-T03) and raised KeyError on unknown ones.
- tools_038._compile_objective: resolve case ids through _resolve_case_ids,
which (1) passes through registered ids case-insensitively, (2) maps
human-readable synonyms to closest catalog cases, (3) drops unresolvable
tokens — a free-form name can never reach the compiler as a KeyError.
- CompileScenarioInput.objective_json description now enumerates the valid
catalog id ranges and gives an example so the LLM stops inventing names.
- Tests: human-readable mapping, unknown-id dropping, dedupe/first-registered
order (test_tools_038_parse.py, 21 passed).
Verification: ruff clean, 21 tests pass.
Runtime migrations failed with 'Multiple head revisions are present' because
037 T081 (p2q3r4s5t6u7 -> verification_runs.dashboard_id) and a concurrent
session-activity change (a1b2c3d4e5f7) both branched from o1p2q3r4s5t6.
The failed 'upgrade head' left dashboard_id unapplied, causing
'column verification_runs.dashboard_id does not exist' on
GET /verification/history.
Add a no-op merge revision (015281bd7759) collapsing both into a single head
so 'upgrade head' applies the verification_runs.dashboard_id column.
Verified: ScriptDirectory.get_heads() == ['015281bd7759'].
QA review of the 036-041 closure range returned FAIL with 3 criticals, all
confirmed. Fixes:
C1 - breaker dead: on_result=persist_batch is now wired into RunnerPool
(breaker.record() fed per result); added test_breaker_abort_persists_partials
proving CIRCUIT_BREAKER_ABORT reachability + partial persistence.
C2 - index-based result mapping corrupted data under concurrency: results now
map by execution_id to their source item; test uses two distinct payloads
and asserts chart->digest pairing (previously masked by identical fixtures).
C3 - double-acquire of the shared client semaphore (deadlock invariant):
RunnerPool no longer manually acquires the client semaphore; capacity is
enforced by worker count, the client bounds total concurrency.
C4 - duplicated ScenarioGraph.Vlm.Analyze region: outer region renamed
ScenarioGraph.Vlm [TYPE Module].
H1 - _default_submit stub removed: analyze_screenshot requires submit=; no
silent empty-findings fallback.
M1 - test_capture_dispatch.py region closed.
M3 - capture.py raw_sha256 bypass removed: digest always derived from real
capture_bytes (no caller-supplied hash).
Verification: load_testing (77) + scenario (103) = 180 passed; ruff clean;
all region pairs balanced.
- 038/039/040/041 validation.md: update PASS status to reflect completed
runtime closure (T057-T059, T054-T057, T075-T079, T045-T048); regenerate
039/040/041 digest tables; 039 T058 documented as the sole open task
- frontend/src/routes/dashboards/[id]/+page.svelte: include the 039 T057
VerificationHistoryList binding (was created in the 039 commit but the
page-level wiring was left unstaged)
All closure tasks across 036-041 are now complete except 039 T058 (blocked:
no repository_id in dashboard metadata; no PREPROD deployment page).
Close the 039 REST-binding and pipeline-view gaps found in the audit: the
scenario API client was never imported and pipeline views were not bound to
any page.
T054/T056 - dashboard-testing.ts gains compileScenario/validateScenario/
resolveScenario (requestApi POST); WorkspaceModel.compileFromRest /
validateFromRest / resolveFromRest give an agent-free REST preview path;
DashboardScenarioWorkspaceModel.rest.test.ts (3 tests).
T055 - capture/vlm/disposition REST surface already landed on backend (038
T057/T058); EvidencePanel autonomous binding deferred to follow.
T057 - DashboardDetailModel.loadVerificationRuns() + VerificationHistoryList
bound on /dashboards/[id], consuming 037 T081 GET /verification/history;
DashboardDetailModel.test.ts = 67 passed.
T058 (verify action) intentionally left open: VerificationRunRequest needs a
repository_id which dashboard metadata does not expose, and there is no
PREPROD deployment page in the frontend. Documented as a blocker in tasks.md.
Verification: 79 vitest passed (REST + detail model + api), vite build OK;
eslint clean for changed code (pre-existing URLSearchParams lint on old line
left untouched).
Close the 037 pipeline-automation and read-API gaps found in the audit:
deploy/release hooks did not create VerificationRun, and GET endpoints for
history/detail were absent even though 039 UI and client call them.
T080 - _release_routes.py: create_release now fires best-effort
_trigger_release_verification -> VerificationRun with trigger=release_create
(metric+structure); verification scheduling failures never roll back the
release transaction.
T081 - verification.py: add GET /verification/history (dashboard_id +
environment_id filters, newest-first) and GET /verification/{run_id}
(404 RUN_NOT_FOUND); reuse _record_to_response.
- verification_run.py + alembic migration p2q3r4s5t6u7: nullable indexed
dashboard_id populated from structure/visual/metric category_params.
- verification_service.py: _derive_dashboard_id helper.
Verification: release routes (32) + verification API (8) + persistence (21)
= 53 passed; ruff clean for changed code (pre-existing RUF012/UP017 on old
lines left untouched).
Close the 038 MVP runtime gaps found in the audit: VLM analysis previously
returned empty findings with no provider call, and capture registered a
synthetic sha256 derived from run/step ids instead of real image bytes.
T057 - vlm.py: replace _default_submit stub with real submit_screenshot that
resolves a multimodal provider via LLMProviderService (decrypted key,
multimodal-required gate) and calls Plugin.Service.LLMClient.get_json_completion
with the masked screenshot; analyze_screenshot is now async.
T058 - capture.py: dispatch_capture now REQUIRES real capture_bytes/masked_bytes
and computes sha256 from the actual image bytes (synthetic hashes forbidden);
scenario API accepts base64 capture/masked bytes.
T059 - test_scenario_vlm_e2e.py: capture -> VLM -> disposition end-to-end with
real bytes (masked bytes reach the provider; digest matches sha256 of bytes).
Verification: tests/services/dashboard_testing/scenario/ = 103 passed, ruff clean
(existing B008 on pre-existing draft-pack route lines untouched).