# Donto Canon — The Operating System for Contested Reality
### iteration 6 · CANONICAL · 2026-06-10

> Authored as the project's north star, synthesizing five prior iterations — the company vision,
> the Lens Engine, the Claim Substrate, Generative Abundance, and the operator's Epistemic-OS
> direction (`DONTO-EPISTEMIC-OS.md`, the proximate inspiration for this document) — against what
> the live system has actually proven. Where this document disagrees with its inspirations, this
> document wins.

---

## 1. The thesis

**Donto turns agent work into verifiable epistemic state** — claims with receipts, disagreements
with structure, beliefs with time — **and uses reality's feedback to decide what to ask next.**

The shorthand stands: *donto is not a knowledge graph; it is the operating system for contested
reality.* But the load-bearing word is neither "knowledge" nor "reality" — it is **trail**. When
an agent (or a society of agents) works through donto, it leaves behind something no chat log,
vector store, or classical KG can produce: an inspectable, contestable, time-travelable record of
what was claimed, on what evidence, against what disagreement, under which lens, and with what
standing — that *improves* as reality answers back.

That trail is the product. Everything else — the calculus, the lenses, the frames, the standards
bridges — exists to make the trail trustworthy, queryable, and *steering*.

## 2. Why this bet (and not the adjacent ones)

- **Not "the best ontology."** Ontology is a commodity claim and a brittle one; the system's own
  thesis (predicates are freely minted, alignment is query-time, identity is a hypothesis) is a
  *refutation* of the one-true-schema project. Donto is where ontologies compete and earn
  standing — it must never become one.
- **Not "the best agent memory."** We benchmarked this honestly: on LongMemEval, a strong reader
  with full context matches any memory system; on BEAM-10M we beat the author's published score
  in a partial state — but the score conflates reader×memory and the categories we win
  (event-ordering, temporal, abstention) are exactly the *trail* properties (bitemporality,
  evidence-first honesty), not raw recall. Memory is a consumer of the substrate, and a proving
  ground. It is not the destination.
- **The trail has no incumbent.** Chat logs are unverifiable. Vector stores forget why. Classical
  KGs corrupt on conflict. Provenance standards (PROV, nanopubs) describe single assertions, not
  living disagreement under time and lens. The system that holds *contested, evidence-anchored,
  time-indexed agent output at firehose scale* — and ranks what to do next — does not exist
  elsewhere. It half-exists here, today, with 41.7M claims.

## 3. The mechanism: three loops that must compound

The inspiration document organized the future into five capability phases. That is the wrong
shape: this project's recurring failure mode is **capability built ahead of usage** (127 tables;
2,433 argument edges against 41.7M statements; claim frames at zero rows). The canon organizes by
*loops*, and a capability activates **only when a loop pulls it** — pull, never push.

### Loop A — HOLD (exists; keep scaling)
*Observe → extract → anchor → claim → contest.*
The abundance engine: free-typed extraction, the always-on citer, paraconsistent holding,
bitemporal belief. This loop works — it ingested a 13.3M-token benchmark corpus, two literature
corpora, a genealogy research program, and a Discord society. Its health metrics are coverage
metrics: % of claims evidence-anchored (3.8% today — the single most important number to move),
contradiction edges on contested subjects, valid-time coverage.

### Loop B — JUDGE (half exists; make standing real)
*Align → identify → review → stand.*
Alignment and identity are live and proven (the closure folds at query time; identity resolves
per hypothesis). What's thin is the *standing* of a claim — and here the canon deliberately cuts
the inspiration's nine-component standing vector down to what can ship and be trusted:

> **standing v1 = ⟨maturity, corroboration, contradiction-pressure, recency⟩** — every component
> computable from tables that exist today (flags, evidence_link counts, paraconsistency_density,
> tx_time). Ship it as one SQL function and one API field, then let v2 earn components (identity
> stability, source reliability, downstream utility) when a loop demonstrates the need.

Review becomes citable state (`donto_review_decision` goes load-bearing on the genealogy
adjudications that are *already happening* in prose), and lenses become first-class objects —
the one piece of the inspiration adopted whole, because it is correct and everything needed
already exists except the registry table.

### Loop C — STEER (the frontier; donto's actual differentiator)
*Rank what would disambiguate → act → measure → re-rank.*
This is "let reality prune the graph" made operational — and the canon's sharpest disagreement
with its inspiration is about where it starts. Not materials science. **It starts where the
contested corpus already lives: the genealogy program.** Every research loop this project runs
already ends with a hand-written "decisive next action" list (order this certificate, open that
register, test this filiation). That ranking — *which evidence acquisition most reduces
uncertainty on the contested claims* — is currently produced by an agent ad hoc. Formalize it:

> `donto.suggest_next_evidence(scope, lens)` → ranked actions, each tied to the contradiction or
> identity hypothesis it would resolve, with the expected standing shift.

The killer demo is then run on materials we hold: take a real contested identity (the Kitty
disambiguation; Val's apex couple; Otto Davis's birth year), have donto enumerate the
contradiction structure, rank the next evidence, dispatch agents to fetch what is fetchable,
ingest, and **show the standing shift** — before/after, with receipts. The materials-science
version of the demo is the same machinery pointed at a new corpus later; the genealogy version is
achievable this month, with zero new budget.

## 4. The constraint set is part of the vision

This system runs on one 4-core box, zero per-token spend, rotated subscriptions, and agent labor.
The inspiration treats scale as a given; the canon treats the constraints as *design inputs*:

- **Self-hosting first.** Donto's heaviest users are the agents already on this box — Omega, the
  extraction fleet, the research loops, this assistant. The shortest path to "epistemic OS for
  agents" is to make *our own agents* bind to the substrate for everything they produce: research
  notes as claims, rule-outs as negative claims, decisions as reviewed state. Dogfooding is not a
  growth hack here; it is the proof that the trail is worth leaving.
- **Proof discipline is the brand.** What this project demonstrably does better than its peers is
  honest verification: no-shortcuts reconciliation, controlled negatives, benchmark honesty
  (publishing the caveats next to the wins). The canon elevates that from culture to product:
  every headline claim donto makes about itself ships with the query that proves it.
- **Distribution through contribution.** The embed-fleet pattern (donto.org/help — strangers'
  machines doing verifiable work through a dumb-worker protocol) is the template for how the
  commons grows: small, safe, inspectable units of contribution. Review, evidence-fetching, and
  adjudication can all eventually take this shape.

## 5. What is adopted from the inspirations, and what is corrected

| From | Adopted | Corrected |
|---|---|---|
| Epistemic OS (iter 5) | The OS-for-contested-reality frame; lenses as first-class; dormant-schema-as-roadmap; "schema deep, protocol simple"; MCP toolset; loss-reported standards bridges; the Observatory surfaces | Category-phases → compounding loops; 9-part standing → standing v1; materials-science demo → genealogy demo now; A2A/markets/tournaments → horizon, entered only when a loop pulls them |
| Abundance (iter 4) | Emit-free/defer-joining; abundance as signature not problem; measurement as steering wheel | Abundance is Loop A, not the whole story — holding everything is table stakes; *judging and steering* are the moat |
| Lens Engine (iter 2) | Discovery at lens intersections; the verifier as the moat | Discovery is Loop C and it starts with evidence-ranking, not analogy-mining |
| Company vision (iter 1) | Open-core posture; one regulated vertical eventually | GTM deferred until the trail demos itself |

## 6. The execution spine (already in motion)

- **Phase 0 — version everything: DONE** (2026-06-10). Every operational system under git; the
  Infrastructure Register is law.
- **Phase 1 — formal legibility: drafted.** The Calculus, Lens Spec, Protocol + JSON schemas,
  Activation PRD, Standards Map — in `donto/docs/`, grounded in the live system, every construct
  marked exists/TO-BUILD. These are infrastructure for all three loops and proceed regardless of
  vision iteration.
- **Next, in order of loop-pull:**
  1. **standing v1** (one function, one API field) + lens registry (`donto_lens`) — Loop B's spine.
  2. **Review goes load-bearing** on real genealogy adjudications; releases go load-bearing on
     the Val report + BEAM outputs (`donto_dataset_release`).
  3. **`suggest_next_evidence` v1** + the genealogy steering demo — Loop C's first heartbeat.
  4. **The MCP epistemic toolset** in loop-pull order: assert/attach_evidence/explain_claim/
     find_contradictions/suggest_next_evidence first; the rest as consumers appear.
  5. **The Observatory** (Substrate Observatory → Evidence Explorer → Alignment Lab → Nebula
     Studio), built on the design system, each surface shipping when its loop has state worth
     showing.
- **Horizon (entered when pulled, never before):** science frames (a measurement is a claim with
  a unit lens), standards import/export (an export is a lens with a loss report), A2A capability
  cards, hypothesis tournaments, cross-domain analogy.

## 7. Success, honestly measured

1. **Anchoring coverage** — % of live claims evidence-anchored (3.8% → 25% is the first mountain).
2. **Contest density where it matters** — argument edges + review decisions on the subjects we
   actually dispute, not global averages.
3. **Standing in every answer** — every API/Observatory answer carries standing v1 + "why" one
   join away.
4. **Steering throughput** — disambiguations per week that `suggest_next_evidence` initiated and
   reality settled (the loop-C heartbeat; today: zero, by hand: ~weekly).
5. **Benchmark honesty maintained** — BEAM/LongMemEval numbers published with caveats, deltas
   attributed to architecture (the vector arm, valid-time) not reader strength.
6. **Self-hosting depth** — fraction of our own agents' research output that lands as claims
   rather than prose.

## 8. What we will not do

No one-true-schema. No merge-on-conflict. No silent deletion (I3 is forever). No per-token spend.
No capability without a consuming loop. No benchmark number without its caveats. No claim about
donto that donto cannot itself evidence.

---

*The sentence to keep:* **agents bind to donto to leave a trail worth trusting — and donto tells
them where to look next.**
