Systematic rename of all semantic anchors (#region, [DEF], @RELATION)
across 1400+ files — backend Python, frontend Svelte/TS, specs, docs:
- Flat anchors become Namespace.Module.Entity
- @RELATION references updated to match new anchor paths
- Zero business logic changes
Enhance the reliability and observability of the task execution engine
by introducing retry mechanisms, idempotency, and structured progress
tracking.
- Implement centralized retry logic with exponential backoff support
in `JobLifecycle`.
- Add `retry_task` API endpoint and `TaskManager` method for manual
task restarts.
- Introduce task idempotency using `_idempotency_key` to prevent
duplicate executions.
- Add `retry_count`, `max_retries`, `last_error`, and `progress` fields
to the `Task` model and ensure persistence via `TaskPersistenceService`.
- Upgrade `SchedulerService` to use differential synchronization with
the persistent `SQLAlchemyJobStore` for better job durability.
- Implement structured heartbeat logging to support real-time progress
updates.
- Update project documentation and ADRs to reflect the new plugin
runtime and task resilience patterns.
- Add comprehensive unit and integration tests for the new task
lifecycle features.
Create AsyncJobRunner — centralized bridge between APScheduler
(BackgroundScheduler, sync thread pool) and async coroutines.
Fixes:
- P0: execute_run() called without await from APScheduler thread,
causing coroutine to be silently discarded (root cause: no
translation history)
- P0: get_async_job_runner() deadlock when called from APScheduler
thread pool without running event loop
- P1: ID mismatch in disable_schedule/delete_schedule routes
(job_id passed instead of schedule_id)
- P1: asyncio.run() in APScheduler callbacks incompatible with
running event loop
- Delete unused llm_analysis/scheduler.py (not used in production)
Changes:
core/async_job_runner.py — new: AsyncJobRunner class
core/scheduler.py — use runner.run()/run_later()
translate/scheduler.py — use runner.run() for execute_run
mapping_service.py — remove unused BackgroundScheduler
dependencies.py — add get_async_job_runner() DI
app.py — init runner in lifespan
api/routes/migration.py — use runner.run()
_schedule_routes.py — fix ID mismatch
plugins/migration.py — use runner.run()
llm_analysis/scheduler.py — delete (unused)
tests: 151 new/updated tests, all passing
Translations:
- FR-045: new-key-only execution mode — translate only rows with unseen keys
- Compare source row keys against TranslationRecord.source_data from last succeeded run
- Baseline expired (>90 days) fallback to full mode with baseline_expired event
- run_noop early return when zero new keys (skip LLM + SQL)
Scheduler reliability:
- Stale PENDING protection: runs older than 1h auto-marked FAILED, no longer block schedule
- load_schedules() reloads active translation schedules from DB on restart
- add_translation_job/remove_translation_job register/unregister with APScheduler
- execution_mode column on translation_schedules with additive DB migration
Automation view:
- GET /settings/automation/translation-schedules endpoint
- Translation schedules displayed on Automation page below validation policies
Trace propagation:
- seed_trace_id() in all background entry points:
TaskManager._run_task/_flusher_loop, SchedulerService._trigger_backup,
websocket_endpoint, IdMappingService, all standalone scripts
Migrations:
- dictionary_entries: origin_run_id, origin_row_key, origin_user_id
- translation_schedules: execution_mode
Tests:
- 3 new test modules (22 tests): core_scheduler, executor_filter, scheduler_execution+guard
- Fix pre-existing translate test isolation (conftest.py)
- Fix test_list_runs_filter_status parameter