The substrate
contradiction-preserving · evidence-first

donto — The Engine for Contested Reality

Every modern knowledge system has the same reflex: when two facts disagree, kill one of them. A vector database dedups near-identical chunks. A knowledge graph picks a winner per (subject, predicate) slot, or invalidates the old value when a new one arrives. A retrieval pipeline ranks and returns the top result. The disagreement vanishes, and the system reports a single confident answer.

That reflex is fine when reality is settled. It is catastrophic when reality is contested — when the sources genuinely conflict and the conflict is the most important thing in the record. donto is built on the opposite reflex. It is bitemporal, paraconsistent, and evidence-first: it holds incompatible claims forever as legal state, anchors each to the exact span of source text that asserts it, links divergent claims with typed argument edges, and re-ranks them by reality over time instead of deleting on conflict. Nothing collapses. Everything is held, with its receipt attached.

This report describes the first working slice of that engine — the Tribunal — running live on the Australian frontier-violence record, with an adversarial credibility audit of that slice and no softening. The verdict is mixed, and the mix is itself the most useful result.


The contradiction problem

Consider a single historical event for which the surviving sources record death tolls of 1, 2, 3, 4, 5, and 12. A normal store forces a choice. Dedup keeps the most frequent figure. A KG keeps the last-written one. A "trusted source" heuristic keeps whichever came from the highest-ranked publication. In every case the system now knows one number — and it is lying, because the record knows six.

The collapse is not a storage detail; it is an epistemic decision made silently, at write time, by machinery that has no idea what the dispute is about. Worse, it is irreversible: once the other five tolls are gone, no downstream query, no later evidence, and no reviewer can recover what the record actually contained. The contradiction — the signal a historian would care about most — is the first thing thrown away.

donto's wager is that for the age of LLM-generated abundance, where claims are cheap and conflicting claims are everywhere, holding the contradiction is the product. This is an architectural moat, not a feature: a system that collapses on write cannot retroactively un-collapse. donto never collapses, so it never has to.


The Tribunal on Rannes 1855

The proving ground is the frontier-massacre corpus (ctx:genealogy/frontier-massacres/*): roughly 56,000 evidence-anchored claims, of which 173 events carry eight or more numeric claims — enough quantitative disagreement to be worth adjudicating. Death tolls are an unusually clean contradiction substrate: they are numbers, so divergence is unambiguous, with none of the natural-language-inference fuzziness that plagues general claim-conflict detection.

A word on what this corpus is. It is a record of colonial violence against Aboriginal people. The Tribunal monitors contradictions in the historical record — which sources said what, and where they disagree. It does not, and must not, frame people as targets or events as scores. The subject under analysis is the documentary record and its internal conflicts, anchored to the words of the sources themselves. That framing is load-bearing, not decorative.

The flagship case is the Rannes 1855 event. donto holds, simultaneously and as legal state, death-toll claims of 1, 2, 3, 4, 5 — and a later 1913 oral recollection of 12 — each anchored to a span of source text. The toll of around five traces to contemporary 1855 reportage such as "Four out of the five constituting the native force fell victims," and "three of the troopers were done to death and two others crippled." The toll of 1 comes from a 1909 Capricornian account; the 12 from a 1913 Capricornian oral recollection. No reconciliation makes these one number. The substrate keeps all of them, each with its quote, its date, and its provenance. The contemporary 1-through-5 disagreement is currently wired as 102 typed rebuts edges in donto_argument (after the adjudication described below); the 1913 "12" is held as a claim but the adjudicator does not yet wire it as a contradiction — an honest residual, not a hidden one.

That is the thesis working exactly as designed: where every other system would have a single confident figure, donto has the whole contested field, and can show you why.


Three loops: JUDGE, OBSERVATORY, STEER

The Tribunal is structured as three compounding loops over the substrate, each reading what the previous one wrote.

JUDGE is the contradiction detector (tribunal_driver.py). For each event it runs one GLM tool-calling pass over the event's numeric claims, extracting every death-toll claim with its value, source, and reliability, and minting rebuts edges between divergent tolls. It is deliberately non-brittle — there is no hardcoded predicate list, no synonym table, no if/elif ladder over field names; the model reads the claims and decides. The writes are I3-safe (pure inserts, never overwrite) and idempotent. Across the 173-event slice, JUDGE surfaced five candidate contested events (of which three survived the adversarial audit described below).

OBSERVATORY measures the shape of the contradiction field, not just its presence. A Stage-0 cellular-sheaf analytic (donto analyze sheaf-h1 --seed arguments --scope-context) computes an H¹ "contradiction-pressure" signal and, crucially, finds cocyclesloop contradictions that pairwise edges cannot see (A rebuts B, B rebuts C, C rebuts A is a different beast than three independent disagreements). Scoped to the Tribunal's corrected edge set it ran over 33 statements and found 61 cocycles, and donto_standing.contradiction_pressure rose from 0 to ~41 on the Rannes tolls — every toll now carries pressure where before they carried none (the global, unscoped seed had crowded the frontier out of its node budget entirely; scoping is what lit it up). This pressure is one of four terms in the canonical standing-v1 vector: ⟨maturity, corroboration, contradiction_pressure, recency⟩, the weighting by which donto re-ranks contested claims by reality over time.

STEER closes the loop back to the world. donto_suggest_next_evidence reads the rebuts edges and ranks fetch_conflict_resolving_source actions per contested predicate — the system proposing what evidence to go find to resolve each contradiction it is holding. Honest caveat: STEER's priority ordering currently reads the older paraconsistency-density layer rather than the new sheaf pressure, so the ranking is presently flat even though the proposed actions are correct. Wiring STEER's priority onto the sheaf signal is a known, scoped follow-up.


The credibility audit (the honest part)

A headline demo that survives only friendly inspection is worthless. So the five contested events were put through an adversarial pass — an independent verifier and a skeptic per event, both reading the actual evidence spans — with an explicit kill criterion: if the auto-detected contradictions mislead more than they illuminate, the thesis fails at JUDGE, and we say so. Here is that verdict, faithfully:

Two of five contradictions are genuine; three are artifacts.

  • Rannes 1855 survives every attack. Even granting full reconciliation of the 2/3/4/5 "deaths-over-days" cluster, the 1909 toll of 1 and the 1913 oral toll of 12 cannot be reconciled with the ~5 cluster or with each other for one event. This is donto's thesis working as intended.
  • Beresford 1883 survives, with a qualifier. It is a documented contemporary retraction (5→1), held paraconsistently with an explicit correction edge — but both spans share a single dataset revision, so it is a secondary-sourced adjudicated dispute, not two provenance-independent primaries.
  • The other three collapse. Wyandotte 1871 is the clearest failure: 10 / 8 / 2 are a total and its own subcomponents (8 men + 2 women = 10) from one source, and the Tribunal wired rebuts edges pitting a total against its own parts. Dunk Island 1877 mistook 3 abducted for 3 killed — there is one death. Wilmot Lagoon 1855 is a scope artifact: the curator's own text says one "11" may belong to a different creek and reports zero killings at the lagoon, while "over 100" is a multi-site aggregate — different populations at different places, treated as one dispute.

On a strict 2-of-5 hit rate, the Tribunal as currently wired misleads more than it illuminates at the level of its headline output. The auditor's diagnosis is precise, and it is the no-brittle-logic / no-shortcuts failure made concrete: the contradiction detector is not toll-aware. Specifically, the numeric comparator (a) folds non-death predicates — wounded counts, victim totals, detachment strength, sheep, miles — into the toll contest; (b) pits subcomponents against totals and additive per-day increments against each other; (c) ignores scope, treating campaign aggregates and single-engagement counts as same-population disputes; and (d) lets a recency-blind "contemporary-firsthand wins" rule enshrine the wrong "5" at Beresford, where the later contemporary correction is authoritative.

The most important finding of the audit, though, is where the failure is not. The substrate passes; the Tribunal layer over it does not yet. donto genuinely holds the contradictions and even carries the disambiguating signal — it knows that "8" is tagged men, that "3" is tagged abducted, that "11" has unclear assignment. The descriptive layer is right. The failure is in JUDGE's rebuts-edge generation, which is not yet consuming that signal. As the auditor put it: "It illuminates Rannes; it manufactures three contradictions it should have suppressed."

So the headline — "where everyone else collapses to one toll, donto holds all of them, correctly weighted" — was, at the moment of the audit, defensible in principle and demonstrated on Rannes, but not yet defensible as a general claim. Stating that plainly is the point of building an engine for contested reality. What happened next is the more important half of the story.


The self-correction

The audit was not the end of the run; it was an input to it. Two fixes followed immediately, and crucially, neither destroyed anything.

First, JUDGE became evidence-aware. A second adjudication pass (tribunal_driver_v2.py) now reads each candidate number with its evidence span and decides — still non-brittle, the model judging, no hardcoded predicate list — which values are genuine death tolls for the same population at the same place, explicitly excluding wounded, abducted/captured, party sizes, replacements, livestock, distances, subtotals of another count, and different-scope figures. Re-run over the five events, it reached the verdict the auditor demanded: Wyandotte collapsed to "no genuine conflict" (it recognised 8 men + 2 women as components of the total 10), and Dunk Island collapsed to a single death (the "3" correctly reclassified as abducted). Rannes and Beresford were retained as genuine; Wilmot was retained as contested with an explicit scope caveat (two independent analyses disagree on it — itself an honest signal, not a hidden one).

Second — and this is the part a collapsing store cannot do — the three artifact events' edges were retracted, not deleted. Each wrong rebuts edge had its bitemporal transaction-time interval closed and its review state set to rejected; it vanishes from the current view but remains in history, queryable as "a contradiction the detector once proposed and later withdrew." The corrected Tribunal now holds three genuine contested events (Rannes, Wilmot-with-caveat, Beresford) over current edges, with the two clear artifacts (Wyandotte, Dunk Island) suppressed but not erased. The OBSERVATORY was recomputed over the corrected edges; contradiction pressure now flows only to genuine conflicts, and STEER's ranking was wired to read that same sheaf signal (so the highest-pressure real contradictions surface first).

This is the thesis turned on the system itself: because donto decided nothing irreversibly on write, it could be wrong about which conflicts mattered, be told it was wrong by its own adversarial pass, and correct itself — without ever having destroyed the record needed to do so. The headline now stands on three events instead of one, with its known residual blemishes (a misclassified "4" still in the Rannes set; Wilmot's scope genuinely unsettled) stated rather than hidden.


Limitations and what's next

Two of the three changes the audit demanded are now done; one remains.

  1. Make JUDGE predicate-aware — done (v1). The evidence-aware adjudication pass is now deaths-only, subset-aware, and scope-aware, and it consumes the men/abducted/scope signal the substrate already carried. Residual work: tighten the long tail (a misclassified "4" survives in the Rannes set; Wilmot's scope is genuinely unsettled and flagged as such), and run it over every event, not just the five audited.
  2. Make standing recency-correct — open. A later authoritative correction must be able to outrank an earlier firsthand figure (the Beresford lesson), rather than a flat "contemporary-firsthand wins." This is the next standing-layer change.
  3. Wire STEER onto the sheaf — done. suggest_next_evidence now reads the H¹ contradiction-pressure sheaf-first (the same signal donto_standing uses), so the highest-pressure unresolved contradictions surface first instead of a flat ranking.

Beyond those, the broader program runs the full ~2,800-row corpus through a hardened, predicate-aware Tribunal; sequences the analysis sheaf-first so OBSERVATORY pressure drives JUDGE and STEER rather than trailing them; and — non-negotiably — submits the methodology to historian review against the Newcastle Colonial Frontier Massacres map before any general claim is published. The kill criterion remains in force.


Why this is the moat

It would have been easy to ship the Rannes screenshot and call it a win. The reason not to is the same reason donto exists: a system whose value is holding contested reality honestly cannot earn trust by overclaiming. The audit found three manufactured contradictions, and reporting them is not an embarrassment — it is the engine being used on itself.

The architectural fact underneath all of this is simple and, for competitors, hard. A store that collapses on write has already decided, silently and irreversibly, which version of reality to keep. donto decides nothing on write. It holds every claim with its evidence span, links the conflicts, measures their shape, and re-ranks by reality as evidence accrues — and it can be wrong about which conflicts matter without ever having destroyed the record it would need to correct itself. That is the difference between a database that answers and an engine that adjudicates. The substrate is sound. Making the layer over it as honest as the substrate is the work that remains.