donto box — workers, jobs & capacity (live catalogue, 2026-06-29)
A verified-live inventory of every worker, job, daemon and queue on the donto box, the
hardware it runs on, and a safe-concurrency budget (what can run together without
overloading it). Built from a 6-way parallel live sweep (systemctl / docker / Temporal / psql /
coordinator API), not from the docs — several doc entries were stale and are corrected inline.
Raw data: donto-resources/infra/worker-catalogue-raw-2026-06-29.json.
1. The box
| Host | OVH VPS (migrated off GCloud) · public IP 15.235.185.42 |
| CPU | 8 vCPU (Intel Haswell, 1 thread/core) |
| RAM | 22 GiB (≈16 GiB available, ~9 GiB buff/cache) |
| Swap | 55 GiB (8 GiB /swapfile + 48 GiB /swapfile2) — both ~fully used |
| GPU | none (virtual Cirrus GD-5446 only). All GPU embedding is off-box (donto.org/help contributors) |
Disk / |
193 GB, 62% used (75 GB free) |
Disk /mnt/donto-data |
984 GB, 34% used (623 GB free) — pgdata, blobs, backups, workspace |
| Substrate DB | donto-pg (Postgres 16 + pgvector, docker, host :5432), ~42.4M donto_statement rows |
Correction to the docs: CLAUDE.md still says 4 vCPU / 16 GB / GCloud
34.135.161.87. The live box is 8 vCPU / 22 GB / OVH15.235.185.42, with much larger disks.
This box is multi-tenant. Besides donto it runs: the Omega Discord stack (bot/web/jobs/mongo/postgres/deployer), MobTranslate (web/tts/pg), the full tpmjs.com stack (web/playground/tutorial + 2 Deno sandboxes), and the observability stack (OpenObserve/OTel/Umami). That is the always-on baseline every heavy job competes with.
2. ⚠️ The #1 capacity problem right now: a runaway health-view poll
The single biggest thing overloading the box is not a legitimate worker. A donto-role host
service polls dontosrv's substrate-health endpoint — select row_to_json(donto_v_substrate_health.*) —
on a long-lived connection. That view full-scans the 42M-row substrate (~20+ minutes each), and
because the donto role has no statement_timeout, repeated polls pile up: at sweep time
there were 9 concurrent copies, the oldest 24 minutes, driving load to 25–33 on 8 vCPU with
~18 GB swap in use.
This single pile is the cause of: the embedding contributor's submit timeouts (Cloudflare 504/524), slow builds, and the test-runner worker-start timeouts seen all session.
Relieved (cancelled the 18 stuck read-only queries — safe, no data touched), but it recurs because the poller keeps issuing it. Recommended durable fix (needs a decision — touches the substrate):
- Make
donto_v_substrate_healthcheap — back it withreltuples+ the existing matviews instead of livecount(DISTINCT …)and aSUMover the 6.3M-rowdonto_contradiction_pressurematview (the same reltuples trick already used by the admin overview). Migration0147_substrate_health_faststarted this; it needs finishing for today's row counts/load. - Throttle / cache the poller (it does not need sub-minute fresh exact counts).
- Belt-and-suspenders: a generous
statement_timeoutfor that read path (pg_dump sets its own=0, so backups are unaffected; the align daemon already self-caps).
3. Worker / job catalogue (by function)
All VERIFIED active/running unless marked. 30 donto systemd units, 0 failed.
3.1 Substrate & core APIs
| Service | Port | Role |
|---|---|---|
donto-pg (docker) |
5432 | THE substrate DB (PG16+pgvector, ~42.4M statements) |
dontosrv |
7879 | Rust substrate HTTP API (the GOAL service; all consumers read through it) |
donto-memory-api |
7900 | Memory consumer API (/memorize /recall /search /jobs) |
donto-api |
8000 | Genealogy extraction/ingest API (doc says ~4001 — stale) |
temporal + temporal-ui (docker) |
7233 / 8233 | Workflow/queue engine + UI |
3.2 Extraction (Rust + Temporal fleet — replaced donto-agent-beam)
The doc's headline
donto-agent-beam.serviceis gone/inactive. Extraction is now a Temporal Rust pipeline:
| Service | Queue | Status |
|---|---|---|
donto-extract-temporal-rust-scheduler |
leases donto_extract_queue → per-lane queues |
active, but throughput ~0 (recurring "connection not ready" restarts; MAX_STARTS_PER_TICK=4) |
donto-extract-temporal-rust@hyades |
extract-hyades |
active, idle — the bulk lane (free holo3.1) |
donto-extract-temporal-rust@codex |
extract-codex |
active, idle — best-anchored |
donto-extract-temporal-rust@glm |
extract-glm |
active, idle — lane capped |
donto-extract-temporal-rust@openrouter |
extract-openrouter |
active, idle — do-not-use (per-token) |
donto-extract-citer |
7942 | always-on evidence anchoring (bge-small) |
donto-api-worker |
Temporal donto-extraction |
active, idle (genealogy path) |
3.3 Embedding fabric (bge-small-en-v1.5 / 384-dim)
| Component | Port | Status |
|---|---|---|
donto-embed-coordinator |
7930 | active — work queue donto_embed_queue is DRAINED (4.6M done · 233 stuck-leased · 0 pending) |
| Remote workers — NUC8i7HNK pool (gpu1/gpu2 + w1–w11) | — | only gpu2 leasing; failing every submit (504/524) due to the box overload → 0 productive now |
donto-memory-embed-768 |
7912 | bge-base 768-dim query vectors (episodic recall arm) — different model, query-time only |
embed_memory.py |
— | memory_chunk backfill, ~2–4 chunk/s CPU-bound, not running |
| Coverage | — | predicate ~100% (1.08M) · entity ~3.0M · claim ~1.49M · memory_chunk ~459K |
3.4 Alignment
| Component | Status |
|---|---|
donto-align-daemon |
active, 600 s tick: embed-gap fill → propose → GLM-adjudicate → ~12 h closure/identity rebuild. Backlogs: predicate-candidate 77,492, identity 1,684. total_accepted=0 / total_rebuilds=0 since the 09:49 restart (last real rebuild 2026-06-22); GLM-rate-limited. CPU-light itself; heavy embed offloaded to the fabric. |
| Query-time folding | donto_predicate_alignment (98,882) · donto_predicate_closure (~1.1M ≈ predicate count). Note: the doc's donto_match_aligned table does not exist. |
3.5 Memory
| Service | Port | Role |
|---|---|---|
donto-memory-api |
7900 | API |
donto-memory-worker |
— | sleep-path reconsolidation worker (the live successor to the old Temporal memory-worker, which is gone) |
donto-memory-rerank |
7903 | cross-encoder reranker (hybrid recall) |
donto-memory-embed-768 |
7912 | query vectors |
Queue donto_x_memory_reconsolidation_queue |
— | 44,526 rows queued, not draining (its Temporal worker is inactive) |
3.6 Web · Bot · Observability · Research
| Function | Services |
|---|---|
| Web | donto-home:3100 · donto-memory-web:3101 · donto-genealogy-web:3102 · donto-admin-web:3103 · dontopedia-web:3000 (doc says 8000 — stale) · donto-debug:3002 |
| Bot (Omega) | omega-tellus-bridge:7940 · omega-pov-view:7941 · omega-vm-omega-bot (Discord, GLM-backed) · omega-vm-omega-jobs (Temporal omega-ai, profile-analysis ~30 min) |
| Observability | otel-collector (4317/4318) · openobserve:5080 · umami:7960 · donto-logs:7950 · donto-scanner-api:7920 · temporal-ui:8233 |
| Research | donto-research-worker (Temporal donto-research, SearXNG+GLM) · donto-searxng:8889 |
| Backups/timers | donto-db-backup (cron.daily, verified — fresh 17 GB dump today, KEEP=5) · donto-resource-sweep.timer (hourly :17) |
4. Inference lanes — the real constraint
Extraction/adjudication/research throughput is gated by inference capacity, not the box. Live state:
| Lane | Model | Cost | Status now |
|---|---|---|---|
| Hyades | holo3.1 (256K ctx) + AEON-7/Qwen3.6-27B (128K) |
free / no cap (user's gateway) | UP — the workhorse with headroom. Route bulk here. |
| codex | gpt-5.3-codex-spark (ChatGPT Pro) |
flat (hidden weekly cap) | UP, leftover — best-anchored |
| GLM | z.ai coding | flat (weekly cap) | EXHAUSTED (error 1310) until 2026-07-01 16:43Z → degrades omega bot, research, tribunal, align adjudication |
| Cerebras | — | flat / sold-out | DEAD (payment_required) |
| OpenRouter | — | per-token | reachable but OFF-LIMITS (budget) |
Net usable inference right now = Hyades + codex. Everything that depends on GLM is degraded until 2026-07-01.
5. Queues & backlogs (the work waiting)
| Queue | Depth | Worker | State |
|---|---|---|---|
donto_extract_queue |
352,175 pending / 10,051 done (1.33M claims) | extract scheduler + lanes | STALLED — scheduler not draining; growth ~0 since Jun-27 |
donto_x_memory_reconsolidation_queue |
44,526 | (Temporal memory-worker) | not draining (worker inactive) |
donto_embed_queue |
0 pending / 4.6M done / 233 stuck-leased | coordinator + remote | drained (embed CPU is free) |
| align predicate-candidate | 77,492 | align daemon | draining slowly (GLM-capped) |
6. Safe-concurrency budget
Baseline (idle, healthy): Postgres + Temporal(java) + OTel + OpenObserve + ~15 web/python services ≈ load 4–8 at rest on 8 vCPU. RAM is comfortable (16 GB available); swap is heavily used but cold.
The CPU-bound contenders (each wants a core under load): donto-align-daemon (bge-small),
donto-extract-citer (bge-small), donto-memory-rerank (cross-encoder), embed_memory.py backfill,
local embed services, and any Next.js build. The extraction lane workers are network/inference-bound
(the CPU cost is the citer+ingest, not the model), so extraction throughput is capped by the lane,
not the box.
Empirical (measured this session):
- one
next build→ load +3–6; four parallel builds → load ~25 (serialize them). - the test runner (vitest) can't boot its workers if load > ~12 (60 s startup timeout).
- the runaway
substrate_healthpoll alone → load +20.
Rules of thumb:
- Fix §2 first. With the health-view pileup gone, the legit fleet sits at ~load 4–8 and everything else has room.
- Serialize heavy CPU jobs. Run at most one of {a build, an
embed_memorybackfill, a CPU-heavy extraction batch} at a time. Two+ pushes load > 15 and starves the daemons. - Extraction scales on the lane, not the box. You can run many
extract-hyades/extract-codextasks concurrently against the inference lanes with little box load — as long as the citer + ingest keep up (those are the CPU cost). Hyades parallelizes (~5 concurrent cleanly). - Embedding backlog is currently zero, so embed CPU is free — but a big new backfill is single-process CPU-bound (~2–4 chunk/s); budget it as one heavy job and prefer the off-box GPU contributors for volume.
- Keep ~2 cores free for Postgres + Temporal at all times; they back every worker.
7. Issues found — prioritized action list
- ⚠️
substrate_healthpoll overload (§2) — relieved, recurs. Make the view cheap + throttle the poller. Biggest single win. - ⚠️ Duplicate Postgres on the same data dir.
donto-pg-pre-pgvectoris a second live postmaster bind-mounted to the same/mnt/donto-data/pgdataasdonto-pg(verified identicalPG_VERSION). Two postmasters on one data dir is a corruption hazard — stop/remove it. (No published port, so nothing reaches it — but I3 "never lose data" says fix this.) dontopedia-workercrash-looping. Container "Up" but fatal-looping onTEMPORAL_ADDRESS=10.128.0.3:7233(a dead GCloud IP). Repoint totemporal:7233.- Extraction stalled — 352K backlog, scheduler not draining ("connection not ready" restarts;
MAX_STARTS_PER_TICK=4). Route to Hyades+codex once the box is healthy. - Memory reconsolidation backlog 44,526 not draining (worker inactive).
- GLM lane exhausted until 2026-07-01 — omega bot, research, tribunal, align-adjudication all degraded.
- 7 of 8 matviews drift (no refresh job) — incl.
donto_contradiction_pressure(1.7 GB),subject_count(1.26 GB),predicate_idf. Onlydonto_mv_context_statsauto-refreshes. donto_align_daemon_tickhistory stopped writing 2026-06-22 (status table still ticks).omega-vm-caddy-1stuck "Created" (port-80 conflict with host caddy) — harmless leftover.