The substrate
contradiction-preserving · evidence-first

Omega × donto — making the integration even better (2026-06-07)

A grounded review of how the Omega Discord bot actually uses donto, from the live logs (bot stdout since the dontoQuery deploy ~01:00Z), the tool source (recallMemories.ts, dontoQuery.ts), the system prompt, and the substrate search impl (search.rs) — plus live probes of /search. Every finding below is evidence-backed, with the fix.


TL;DR

Omega is using donto on every relevant turn — but only through one path: the recallMemories tool (memory /recall + substrate /search). The headline finding:

The new 12-lens dontoQuery reasoning tool has been invoked exactly 0 times. It is deployed, BM25-selected, and loaded as a core tool on every turn — yet the model never calls it. It answers everything with recallMemories.

That's not a bug in dontoQuery; it's a tool-overlap and prompting problem (root cause below). The three levers that make Omega materially better at donto:

  1. Disambiguate the two tools so the model knows when to reason with the graph vs fetch text — this alone unlocks dontoQuery.
  2. Teach it to investigate, not one-shot — every turn today is a single recallMemories call then an answer; no search → entity_facts → cross_source → contradictions chaining (the "take it further" goal).
  3. Fix substrate recall qualityplainto_tsquery ANDs words (phrases collapse to zero), ts_rank scores are near-degenerate, and the semantic embedding fabric that already exists isn't in Omega's path.

Method & data

  • Logs: docker logs omega-vm-omega-bot-1 — the bot prints every tool call's args + result and a per-turn tool-selection summary. Window: since the dontoQuery deploy (~4h of live traffic) + the full retained buffer for the "ever called?" check.
  • Source: apps/omega/packages/agent/src/tools/{recallMemories,dontoQuery}.ts, lib/systemPrompt.ts, toolRegistry/metadata.ts; substrate crates/donto-memory/src/api/routes/search.rs.
  • Live probes: POST https://memories.apexpots.com/search with single vs multi-word queries.

What's already working

  • Donto is in the answer loop. recallMemories is a CORE tool and the BM25 selector reliably surfaces it (and dontoQuery) for knowledge questions — e.g. Top tools: memorize, dontoQuery, … on a "what do you remember about X" turn.
  • Cross-source reach. recallMemories(scope:'all') fans out to memory /recall and substrate /search in parallel (Promise.allSettled) and tags each row by source — so a single call already pulls from ctx:genes/*, ctx:discord/*, memory, and the 40.9M-statement substrate.
  • It grounds. Real turns show it pulling genuinely relevant rows (e.g. lablancheex:patrice-francois-lablanche hasFullName "Patrice François Lablanche" from the val-mauritius corpus) and answering from them. The user's read — "working great with donto" — is fair for the recall path.

Finding 1 — dontoQuery is deployed but dormant (0 invocations)

Evidence. Across the full retained log: Tool called: dontoQuery = 0; dontoQuery tool-usage report blocks = 0. In the last 6h every knowledge turn resolved to recallMemories (5 calls), even though the BM25 selector listed dontoQuery among the top tools and it was loaded as 1 of the 8 core tools every turn.

Root cause — tool-purpose overlap. recallMemories's own description says:

"This reaches the ENTIRE donto knowledge graph, not just chat history: … the whole substrate — genealogy research, the ctx:genes/ corpora, and ~39M statements across every context."*

From the model's vantage, recallMemories already does "search donto." dontoQuery's value (graph traversal, contradictions, identity, semantic neighbours, bitemporal) is real but abstract, and the only thing telling the model to prefer it is one subtle system-prompt line ("recall fetches text; dontoQuery lets you THINK with the graph"). A subtle distinction loses to an established habit + an overlapping, confident description every time.

Consequence. All of dontoQuery's differentiated power — contradictions, identity, predicate_alignment, semantic_neighbors, cross_source, bitemporal, the raw_sql escape hatch — is sitting unused. Omega is doing keyword recall, not analysis.


Finding 2 — single-shot lookup, not investigation

Every observed turn is one recallMemories call (occasionally two near-identical ones, e.g. "two split problem" then "split problem") followed immediately by the answer. There is no multi-step chase — no "got an entity, now pull its facts, now see who else attests it, now check contradictions." The user's explicit goal was for Omega to "think like that with the information and take it further when needed." Today it takes one keyword swing and stops.

Two near-duplicate recall calls in a turn (just dropping a stop-word) is the model trying to dig but lacking a real traversal tool in-hand — exactly the gap dontoQuery fills, if it were used.


Finding 3 — substrate recall quality

The substrate arm Omega leans on is /search (FTS), and it has three measurable weaknesses:

  1. Phrases AND to zero. search.rs uses plainto_tsquery('simple', $1), which ANDs every lexeme. Live probe:

    • "love"20 rows · "infinite love"6 · "love verify real"3 · "love death meaning"0. The more natural-language the query, the fewer hits — the opposite of what you want. Omega partly works around this by querying single words, but that's fragile and loses precision.
  2. Degenerate ranking. Recall results show the same score (0.07599088549613953) repeated across completely unrelated rows. ts_rank('simple', <humanized IRI-segment projection>) over short, structurally-similar text can't discriminate — so the "relevance score" is mostly noise, and the top-20 is effectively arbitrary among matches.

  3. No semantic arm in Omega's path, and noisy contexts. The bot container sets no DONTO_*/embed env, so it uses the FTS defaults — the hybrid vector recall (bge-small + RRF) that the LongMemEval work proved load-bearing (FTS-only → hybrid lifted hit@10 0.85→0.98) is not reaching Omega's substrate search. Result: abstract queries surface low-signal rows ("hate"ex:hate-speech-content-N sameAs … chains; "rotors" → car-maintenance chunks) and test/quarantine/smoketest contexts (ctx:genealogy/smoketest, ctx:_quarantine/*) leak into answers.


Recommendations (prioritized)

P0 — Unlock dontoQuery by disambiguating the two tools

This is the single highest-leverage change; it's pure prompting, no infra.

  • Rewrite recallMemories's description to own ONE job: "what was said/remembered" (chat history + memorized text). Drop the "reaches the ENTIRE knowledge graph / ~39M statements" claim — that overlap is exactly what starves dontoQuery.
  • Rewrite dontoQuery's description to own "what is true / who is who / how things relate" (entities, facts, relationships, contradictions, identity). Already good; make the boundary explicit: "Use this, not recallMemories, for any question about a person/place/thing/claim or how they connect."
  • Add explicit IF→THEN routing + a worked example to the system prompt, e.g.:
    • "who is X / what do we know about X / is X the same as Y / what conflicts about X / what's connected to X"dontoQuery (start search, then chain entity_factscross_sourcecontradictionsidentity/semantic_neighbors).
    • "what did I/we say about X / what happened in this channel / remind me"recallMemories.
    • One concrete 3-step dontoQuery transcript in the prompt so the model has a pattern to imitate.

P0 — Fix /search query semantics

  • Switch plainto_tsquerywebsearch_to_tsquery (handles phrases, OR, quotes) or OR-join lexemes with a phrase-boost, so natural-language queries stop collapsing to zero.
  • Exclude test/quarantine contexts from substrate results by default (context NOT LIKE 'ctx:%test%' AND context NOT LIKE 'ctx:_quarantine%' AND context NOT LIKE '%smoketest%').

P1 — Put the semantic arm in Omega's path

  • Point Omega's recall at the hybrid recall (the embedding fabric already exists: pgvector + bge-small, RRF) — either set the embed-service env on the bot so /recall uses hybrid, or add a vector arm to /search. This is the proven precision lever (LongMemEval: hit@10 0.85→0.98).
  • Add a lightweight rerank + de-dup + time-decay pass on the merged rows (OMEGA/Mastra-style cross-encoder rerank is the gap vs the leaders) so the top-k Omega reads is actually the best-k, not 20 tied rows.

P1 — Make dontoQuery output induce chaining

dontoQuery returns raw rows. Wrap each result with a tiny summary + suggested next lens (e.g. after search: "3 candidate entities; call entity_facts on ex:… to go deeper"; after entity_facts: "2 contradictions detected — try contradictions"). Models follow affordances; this turns one-shot lookups into investigations without relying on prompt willpower alone.

P1 — Nudge investigation in the system prompt

Add a short directive: "For knowledge questions, don't answer from the first lookup. Resolve the entity, pull its facts, check who else attests it and what conflicts, then answer — cite the contexts." Pair with the worked example from P0.

P2 — Observability & hygiene

  • Add an Omega-usage panel to admin.donto.org (mirror the embedding tracker): per-tool call counts, dontoQuery lens distribution, multi-step rate, empty-result rate. Right now adoption is invisible — we only found the 0-calls fact by grepping stdout.
  • De-duplicate near-identical recall rows before returning (the hasContentSnippet/listsSnippet pairs and sameAs chains flood the top-k).
  • Revisit lens_name: likely_identity_v1 + scope:'all' defaults once the two tools are split.

Exact change sketches

recallMemories.ts description (narrow it):

"Recall what was said or remembered — Omega's chat history and memorized notes for this holder/session. Use for 'what did I say about X', 'remind me', 'what happened in this channel'. For facts about a person/place/thing or how they relate, use dontoQuery instead."

System prompt routing block (replace the subtle distinction):

KNOWLEDGE ROUTING
- "who is X" / "what do we know about X" / "is X the same as Y" / "what conflicts about X"
  / "what's connected to X"  → dontoQuery. Investigate: search → entity_facts → cross_source
  → contradictions → identity/semantic_neighbors. Don't stop at the first lookup.
- "what did I/we say about X" / "what happened here" / "remind me"  → recallMemories.
Example: user "who is Caroline Brown?" → dontoQuery search "caroline brown" → entity_facts
  ex:caroline-rose-brown → cross_source → contradictions → answer, citing contexts.

search.rs (one-line query fix):

-- was: plainto_tsquery('simple', $1)
WITH q AS (SELECT websearch_to_tsquery('simple', $1) AS tsq)
... AND context NOT LIKE 'ctx:%test%' AND context NOT LIKE 'ctx:_quarantine%'

How we'll know it worked (measurement)

Re-pull the bot logs after the changes and track:

  • dontoQuery call-rate on knowledge turns (target: from 0% → the majority of "who/what/how-related" questions).
  • Multi-step rate — turns with ≥2 chained donto calls of different lenses (target: >0, today 0).
  • Empty/degenerate-result rate on /search (phrases returning 0; tied scores).
  • Grounded-answer spot-check — does the reply cite the contexts/entities it pulled?

Appendix — evidence

  • 0 dontoQuery calls: docker logs … | grep -cE "Tool called: dontoQuery|[0-9]+/[0-9]+: dontoQuery" → 0 (full retained log).
  • Overlap: recallMemories.ts:42"reaches the ENTIRE donto knowledge graph … ~39M statements across every context."
  • plainto AND: search.rs:43,148 plainto_tsquery('simple',$1); live "love"=20 → "love death meaning"=0.
  • Degenerate score: recall rows repeat "score": 0.07599088549613953 across unrelated subjects.
  • No semantic arm for the bot: omega-bot container sets no DONTO_*/EMBED* env → FTS defaults.
  • Single-shot: every observed turn = 1 (occasionally 2 near-duplicate) recallMemories calls, then answer.

Generated 2026-06-07 from the live Omega bot + donto substrate. The P0 changes are prompting-only (no infra) and should flip dontoQuery from dormant to default; P1 adds the proven semantic-recall precision lever. Re-run the §measurement checks after deploying to confirm adoption.