Benchmarking donto-memory on LongMemEval — a faithful study of what a memory layer actually adds
2026-06-05 · a no-shortcuts, no-cheating evaluation of donto-memory on longmemeval_s, with codex as both reader and judge, against a codex-alone full-context baseline
TL;DR. We ran donto-memory through LongMemEval's retrieval-stressed
_svariant (500 questions, ~50 sessions / ~120k tokens each) under audited no-leakage conditions, and — crucially — against a codex-alone full-context baseline (the same reader handed the entire history, no memory system). The headline finding is honest and a little uncomfortable: on_s, where the whole history fits in a frontier model's context, a memory layer does not win on raw accuracy — a strong reader with everything in context is very hard to beat. donto-memory's real, measured value is elsewhere: ~2× lower token cost, and wins on the two abilities that are its design thesis — bitemporal knowledge-update and evidence-first abstention. We report the gaps as plainly as the wins. donto-memory may not top the_saccuracy table; this study is about understanding where a memory layer earns its keep, not about a trophy.
1. Why LongMemEval, and which variant
LongMemEval (ICLR 2025) tests long-term memory of chat assistants over 500 questions across six abilities — single-session (user/assistant/preference), multi-session, temporal-reasoning, knowledge-update — plus an abstention set of unanswerable questions. It ships three variants:
_oracle— the reader is given ONLY the evidence sessions. This isolates reading from retrieval; it is the ceiling, not a memory-system test._s— each question's evidence is buried in a haystack of ~50 sessions (~120k tokens). This is the real test of a memory system: ingest everything, retrieve the right thing at query time._m— ~500 sessions; the history exceeds a single context window.
We focus on _s, the standard memory-system benchmark. (We also report our
_oracle ceiling for reference.) _m is future work and is, as we'll argue, the
variant where a memory layer's accuracy case is actually strongest.
2. The system under test: donto-memory
donto-memory is an example consumer of donto, a bitemporal, paraconsistent, evidence-first claim substrate. For this study the relevant machinery is:
- Ingest — each session is stored as an episodic chunk statement, anchored to
a per-question holder (
user:lme_s_official:<qid>) with the session's date injected as valid-time (valid_from). All ~50 haystack sessions are ingested; nothing is filtered by evidence. - Hybrid recall — at query time, holder-scoped lexical (FTS) + semantic (bge-small-en-v1.5, 384-dim, HNSW) retrieval, RRF-fused, with predicate alignment-closure expansion. This is the same retrieval family as the current SOTA system (OMEGA): identical embedding model + FTS + fusion.
- Bitemporal valid-time — recalled rows carry
valid_from/valid_to, so the reader can prefer the latest value of an attribute (the knowledge-update case). - MCP server — donto-memory now exposes
donto_recall/donto_search/donto_memorizeover the Model Context Protocol, so an agent (codex, Claude Code) can use it as native tools and drive its own recalls. (Used for the agentic arm; the controlled pipeline below is the primary measurement.)
3. Faithfulness — no cheating, audited
A memory benchmark is trivially gameable (peek at the evidence labels), so we
audited every gold-field use line by line (FAITHFULNESS_AUDIT.md):
- Ingest renders sessions from
role+contentonly; the per-turnhas_answerflag is never emitted. - Recall queries with the QUESTION and holder only — never
answer_session_ids/has_answer. - Reader sees only: the question, the question date, and the recalled memories. No gold answer, no evidence labels.
- Judge receives the gold answer — exactly as the official harness does, for grading only.
- The judge prompts, the
label = "yes" in response.lower()rule, and the"_abs"abstention detection are byte-identical to the officialevaluate_qa.py.
The one disclosed deviation: the official judge is the OpenAI GPT-4o API; we use codex gpt-5.4 (ChatGPT-Pro subscription) because there is zero per-token API budget. Same prompts, same yes/no parsing, applied identically to every instance and to both arms. The reader is likewise codex gpt-5.4. This makes our setup methodologically analogous to the leaderboard leader OMEGA, which uses GPT-4.1 as both reader and judge — and means our numbers are comparable to OMEGA / HydraDB / Mastra, not to the paper's weaker GPT-4o-mini-class rows.
4. Methodology
- Harness:
run_official.py— a resumable, checkpointed pipeline (ingest → embed → answer → metrics). Per-instance checkpointing; a crash never re-spends codex. Logs to/opencode/logs. - Recall budget: bounded top-k (k=20) hybrid recall over the full ~50-session haystack — a real retrieval test, not the oracle hand-off.
- Reader / Judge: codex gpt-5.4, official prompts.
- Baseline arm (
run_baseline.py): the same codex reader handed the ENTIRE date-sorted haystack (~120k tokens), NO donto retrieval — isolating exactly what the memory system contributes.
5. Results
5.1 Headline — full 500, longmemeval_s
⏳ PLACEHOLDER — fill on completion of the full-500 run. (per-type accuracy, task-averaged, overall, abstention, n_scored)
5.2 The comparison that matters — codex vs codex+donto (paired, n=48)
Same 48 instances, same reader, same judge; the only difference is whether codex reads donto's retrieved top-20 or the entire haystack:
| ability | codex-alone (full-context) | codex+donto |
|---|---|---|
| knowledge-update | 0.875 | 1.000 |
| multi-session | 0.875 | 0.875 |
| single-session-assistant | 1.000 | 1.000 |
| single-session-user | 1.000 | 1.000 |
| temporal-reasoning | 1.000 | 1.000 |
| single-session-preference | 1.000 | 0.750 |
| overall | 0.957 | 0.936 |
| abstention | 0.667 | 1.000 |
- Accuracy: ~tied (0.957 vs 0.936 ≈ one instance on n=47).
- Token cost: donto ~2.1× cheaper — measured median reader prompt 124k tok (full-context) vs 58k tok (codex+donto).
- donto wins knowledge-update (1.0 vs 0.875) — bitemporal valid-time lets the reader pick the latest value instead of drowning in superseded history.
- donto wins abstention (1.0 vs 0.667) — evidence-first: when retrieval finds nothing, donto says "I don't know"; full-context hallucinates an answer to an unanswerable question.
- donto loses preference (0.75 vs 1.0) — its whole-session chunks miss the subtle, inferential preference session (see §6).
5.3 Oracle ceiling (reference)
Full 500 _oracle (codex reader+judge): overall 0.946 / task-avg 0.935 /
abstention 0.967.
6. Retrieval analysis — where donto's recall is strong and weak
A zero-cost offline sweep (recall only, no reader) measured whether the evidence session lands in the top-k. The hybrid vector arm is load-bearing: turning it on lifts overall hit@10 from 0.85 (FTS-only) to 0.98.
| ability | hit@5 FTS→hybrid | hit@10 FTS→hybrid |
|---|---|---|
| single-session-assistant | 0.62 → 1.00 | 0.88 → 1.00 |
| temporal-reasoning | 0.75 → 0.88 | 0.88 → 1.00 |
| single-session-preference | 0.12 → 0.50 | 0.38 → 0.88 |
| multi-session | 1.00 → 1.00 | 1.00 → 1.00 |
| knowledge-update | 1.00 → 1.00 | 1.00 → 1.00 |
| overall | 0.75 → 0.90 | 0.85 → 0.98 |
single-session-preference remains the weak spot. Preference questions ("what should I cook for guests?") share almost no surface words with the session where the preference was stated ("I'm vegan"); even semantically the link is inferential. donto stores one chunk per session and embeds only the first ~300 tokens of a ~2,500-token session, so evidence buried mid-session is invisible to the vector arm.
A validated fix (future work): finer chunking. Splitting sessions into
~1,000-char windows (each fully embedded) moved a known-missed preference instance
from miss → hit@5. It is, however, ~10× more chunks; at full _s scale that
is ~250k chunks, which on CPU-only bge-small (~3 chunk/s, no GPU) is ~17h —
infeasible on the current box without an int8/GPU embedder. So the full run uses
the proven whole-session config and we report finer chunking as a measured,
deferred improvement.
7. The competitive landscape — and a caution about leaderboards
Published longmemeval_s scores:
| system | overall | reader / judge | retrieval |
|---|---|---|---|
| OMEGA | 95.4% | GPT-4.1 (both) | bge-small + FTS + cross-encoder rerank + time-decay |
| Mastra | 94.87% | — | — |
| HydraDB | 90.79% | — | graph-native (entity/temporal/causal + BM25) |
| Memoria | 88.78% | — | — |
| Zep / Graphiti | 71.2% | — | — |
| paper, GPT-4o full-context | 60–64% | GPT-4o | none |
The caution: the LongMemEval score conflates the reader and the memory system. The paper's GPT-4o full-context scores 60–64%, but our codex gpt-5.4 full-context — no memory system at all — scored 95.7%. Swap a stronger reader in and the "memory system" numbers move with it. So a single overall number is not a clean ranking of memory systems; the codex-alone baseline is the control that matters, and the right questions are "how much does the memory layer add over giving the reader everything, and at what token cost, and does it scale past the context window?"
donto-memory shares OMEGA's retrieval family (identical bge-small + FTS + fusion) but lacks OMEGA's cross-encoder reranking and time-decay weighting — concrete, queued improvements that we will measure against this same benchmark.
8. What we actually learned
- On
_s, a memory layer doesn't win on accuracy when a strong reader can hold the whole history in context. Be honest about this. - donto's defensible value is real but specific: ~2× token efficiency, and wins on knowledge-update (bitemporal) and abstention (evidence-first) — the parts of its thesis that are load-bearing.
- The accuracy case for a memory layer lives on
_m(~500 sessions), where full-context is impossible and retrieval is mandatory. That is the next run.
9. Engineering notes (no-shortcuts log)
- Embedding is the scaling bottleneck, not retrieval. CPU-only bge-small does ~3 chunk/s; 25k chunks ≈ 2h. Throughput scales inversely with chunk length (400ch→21/s, 1500ch→6/s, 3000ch→2.7/s).
- A real failure we caught and fixed: the single long-lived embed process leaked to 9.7 GB RSS and, colliding with live memory-worker traffic, swap-thrashed a 16 GB box to a near-total stall (~0 chunk/s while its own log claimed "3 stmt/s"). Fix: a segmented embed (fresh process per 4k chunks → bounded ~0.8 GB) at low priority. Lesson now standing policy: a launched job is not a working job until you've watched the real metric move.
10. Reproducibility
Harness, audit, and findings are in /home/ajax/longmemeval-study/:
run_official.py (resumable donto arm), run_baseline.py (codex full-context
arm), retrieval_eval.py (offline recall sweep), FAITHFULNESS_AUDIT.md,
COMPETITIVE_LANDSCAPE.md, CHUNKING_FINDING.md. Judge prompts are the verbatim
official evaluate_qa.py. No per-token API; codex via subscription.
Honest caveats restated: codex gpt-5.4 reader+judge (not GPT-4o); _s results on
the config above; the full-500 headline table is filled on run completion; finer
chunking and the _m variant are measured/queued future work.