Hyades/holo3.1: the empty tool_calls.arguments bug — a gateway diagnostic (2026-07-02)
A full technical characterization of the single biggest throughput tax on donto's extraction
fleet: roughly 1 in 3 structured calls to holo3.1 on hyades.gnostr.cloud returns HTTP 200
with a well-formed tool-call envelope whose function.arguments is an empty string — while
usage.completion_tokens shows 600–1,000+ tokens were generated and billed. The model does
the work; the payload never reaches the wire. This page is written for the Hyades operator:
it contains the exact request shape, raw captured response bodies, production-scale
measurements (1,064 calls / 204 runs in one day), every hypothesis we ruled out, and where in
the gateway we believe the bug lives.
TL;DR for the operator: your OpenAI-compat layer detects that holo3.1 made a tool call (it emits
finish_reason:"tool_calls"and atool_calls[0]envelope with a freshchatcmpl-tool-…id) but ~33% of the time theargumentsstring arrives empty, with completion tokens in the same range as successful calls. Identical serial requests succeed 5/8 and fail 3/8. Our best hypothesis is the tool-call parser between the raw model generation and the response serializer (e.g. a vLLM-style<tool_call>-tag parser that matches the opening tag but fails to extract the JSON body, then emits the envelope with an empty buffer instead of erroring). The raw generation for a failing request very likely contains the full facts payload — if you can log/replay raw generations for one of the failingchatcmpl-tool-*ids below, that either confirms or kills the hypothesis in one look.
1. Who is calling, and how
donto (an evidence-anchored claim-extraction substrate) runs a fleet of Rust workers that
extract structured facts from historical documents. Hyades/holo3.1 is our primary volume lane
(free, uncapped, ~6 concurrent deep extractions). Every call is an OpenAI-compatible
POST /v1/chat/completions with structured output via forced tool-calling — the model is
asked to return facts by calling a submit_facts function; the client never parses JSON out of
free text.
The exact request shape (sanitized only of the bearer token):
{
"model": "holo3.1",
"temperature": 0.1,
"max_completion_tokens": 64000,
"reasoning_budget": 8000,
"messages": [{ "role": "user", "content": "<extraction prompt + SOURCE, 1–30k chars>" }],
"tools": [{
"type": "function",
"function": {
"name": "submit_facts",
"description": "Submit the extracted facts.",
"parameters": {
"type": "object",
"properties": { "facts": { "type": "array", "items": {
"type": "object",
"properties": { "subject": {"type":"string"}, "predicate": {"type":"string"}, "object": {"type":"string"} },
"required": ["subject","predicate","object"] } } },
"required": ["facts"]
}
}
}],
"tool_choice": { "type": "function", "function": { "name": "submit_facts" } }
}
Request headers: Authorization: Bearer …, Content-Type: application/json,
X-Hyades-Thread: donto-agent/<uuid> (the server-side conversation memory — the broad pass
seeds a thread; gleaning nudges continue it), User-Agent: curl/8 -family (plain
python-urllib UAs are Cloudflare-1010-blocked; reqwest/curl pass).
Call pattern per document: 1 "broad" pass (full prompt + source) then 3–7 short "nudge" passes on the same thread (lens-directed gleaning). Fleet volume on 2026-07-02: 204 extraction runs, 1,064 structured HTTP calls to holo3.1 between 00:00 and ~07:00 UTC.
2. The failure, exactly
Two raw response bodies from a controlled probe (2026-07-02 ~07:35 UTC, serial calls,
identical request, ~2.6k-char source, fresh X-Hyades-Thread per call).
Failing call (probe call 3; 19s wall; this shape = 3 of 8 probe calls):
{
"id": "chatcmpl-hyades",
"object": "chat.completion",
"created": 1782975308,
"model": "holo3.1",
"choices": [{
"index": 0,
"message": {
"role": "assistant",
"tool_calls": [{
"id": "chatcmpl-tool-b00f234a8126becc",
"type": "function",
"function": { "name": "submit_facts", "arguments": "" }
}]
},
"finish_reason": "tool_calls"
}],
"usage": { "prompt_tokens": 996, "completion_tokens": 865, "total_tokens": 1861 }
}
Succeeding call (probe call 5; 14s wall; same request):
{
"id": "chatcmpl-hyades",
"object": "chat.completion",
"created": 1782975351,
"model": "holo3.1",
"choices": [{
"index": 0,
"message": {
"role": "assistant",
"tool_calls": [{
"id": "chatcmpl-tool-a0424187d99ada9b",
"type": "function",
"function": { "name": "submit_facts",
"arguments": "{\"facts\":[{\"subject\":\"Meridian\",\"predicate\":\"arrived from\",\"object\":\"Thursday Island\"}, …" }
}]
},
"finish_reason": "tool_calls"
}],
"usage": { "prompt_tokens": 996, "completion_tokens": 626, "total_tokens": 1622 }
}
Read those two usage blocks together — that is the whole bug in one line:
the failing call generated MORE completion tokens (865) than the succeeding one (626) and
delivered zero of them. The envelope is constructed (unique tool id, correct function name,
finish_reason:"tool_calls"), the arguments buffer is empty.
Full probe series (8 identical serial calls):
| call | wall | finish_reason | tool_calls | arguments chars |
completion_tokens |
|---|---|---|---|---|---|
| 1 | 25s | tool_calls | 1 | 1,098 | 898 |
| 2 | 22s | tool_calls | 1 | 985 | 911 |
| 3 | 19s | tool_calls | 1 | 0 | 865 |
| 4 | 22s | tool_calls | 1 | 1,085 | 953 |
| 5 | 14s | tool_calls | 1 | 1,165 | 626 |
| 6 | 20s | tool_calls | 1 | 0 | 1,004 |
| 7 | 18s | tool_calls | 1 | 1,106 | 1,060 |
| 8 | 9s | tool_calls | 1 | 0 | 602 |
3/8 empty at concurrency one — this is not a load phenomenon.
2b. The second variant: finish_reason:"stop" with nothing at all
In production logs there is a second, less common shape: finish_reason:"stop",
no tool_calls, content empty/absent (113 of 362 production empties). We did not reproduce
it in the serial probe; it may be thread-state-related (it shows up on retries within an
existing X-Hyades-Thread) or a different path through the harness. The dominant variant
(249/362) is the empty-arguments envelope above.
3. Production-scale measurements (2026-07-02, 00:00–07:00 UTC)
All numbers from per-run structured event logs (JSONL; every HTTP attempt, status, latency, tool_choice form, parse outcome). 204 runs / 1,064 calls.
Failure rate by tool_choice form — both forms fail; the "fallback" was worse:
| tool_choice | ok | empty | success rate |
|---|---|---|---|
{"type":"function","function":{"name":"submit_facts"}} (forced-named) |
542 | 239 | 69% |
"required" |
125 | 114 | 52% |
Failure rate by pass kind + input size — bigger inputs fail more:
| pass kind | ok | empty | empty rate |
|---|---|---|---|
| broad (full prompt + source) | 185 | 142 | 43% |
| nudge (short follow-up, same thread) | 515 | 222 | 30% |
Median source size on failing broad passes: 2,608 chars vs 1,240 chars on succeeding ones (~2×). Note the probe reproduced 37.5% empties at 2.6k chars — consistent.
Failure rate by hour (UTC) — flat; NOT diurnal, NOT load:
| hour | empty/calls | rate |
|---|---|---|
| 00 | 48/133 | 36% |
| 01 | 59/182 | 32% |
| 02 | 49/108 | 45% |
| 03 | 68/209 | 33% |
| 04 | 30/84 | 36% |
| 05 | 60/163 | 37% |
| 06 | 39/141 | 28% |
Our client concurrency changed 3 → 6 at ~06:20 UTC with no change in the empty rate (06:00 is actually the day's lowest).
Latency. Time-to-headers is ~13–16s for successes and empties alike. Body-read after headers: p50 ≈ 0s for both, but the tails diverge violently — p90 body-read 55s for successes vs 296s for empties (n=691/367). A minority of empties grind for minutes before delivering nothing; probe empties were fast (9–20s). So there appear to be two sub-populations: a fast parser-drop and a slow grind-then-drop.
Cost to us: with a mean ~79s per wasted call and 362 empties in ~7 hours, holo's empty responses consumed ~6.5 hours of worker-slot time in one day — the single largest inefficiency in the extraction pipeline.
4. What we ruled out
- Client-side parsing. We captured raw bodies (§2);
argumentsis literally""on the wire. - Tool schema validity. The identical schema succeeds 69% of the time, including calls in the same serial probe.
max_completion_tokensexhaustion. 64,000 allowed; failing calls used ~1,900 total.finish_reasonis not"length".reasoning_budgetstarvation. Budget 8,000; failing calls generated 602–1,004 completion tokens — nowhere near any cap.- Load/concurrency. Serial probe reproduces at 37.5%; production rate flat across hours and across our 3→6 concurrency raise.
- tool_choice form. Both
forced-namedand"required"fail (69%/52% success). - Cloudflare/transport. Clean HTTP 200s with complete, well-formed JSON bodies; correct UA in use.
- Prompt-injection weirdness in sources. The probe used a bland, fixed synthetic source; still 3/8.
- Randomness in our request. temperature 0.1, identical byte-for-byte probe bodies (fresh thread id only).
5. Where we think it lives (ranked hypotheses for the harness)
- Tool-call parser drop (most likely). The pipeline stage that turns raw model text into
the
tool_callsarray recognizes that a tool call started (hence the envelope + correct function name +finish_reason:"tool_calls") but fails to extract the JSON body — a vLLM-style tag parser (<tool_call>…</tool_call>/ hermes-style) that matches the opening tag but chokes mid-body (interleaved reasoning tokens? malformed/truncated tag? nested braces?) and then serializes an empty buffer instead of surfacing a parse error. The billed completion tokens are the model's real (probably fine) generation. Test: log the raw generation for a failingchatcmpl-tool-*id and eyeball it — if the facts JSON is present in the raw text, this is confirmed. - Reasoning/content channel interleave. holo3.1's hidden-reasoning stream and the tool
payload share a channel and the demux sometimes attributes the whole generation to
reasoning; the tool payload is "generated" (token count) but routed to a channel that is
discarded. The
finish:"stop"+ empty-everything variant (§2b) fits this better than #1. - Streaming/flush race in the response assembler. The arguments buffer is streamed internally and finalization races generation end; the long-grind tail (p90 296s empties) could be a stall in exactly this path. Less likely than #1 because the envelope is complete and consistent.
6. What would fix it (server-side wishlist, in order of value)
- Never emit a 200 with an empty
argumentsenvelope. If the tool-body parse fails, return an explicit error (finish_reason:"error", a 5xx, or an OpenAI-style"error":{...}object). Clients can retry a known failure immediately instead of discovering an empty payload after a full generation. - Log raw generations for tool-call requests (or keep a ring buffer keyed by the
chatcmpl-tool-*id) so parse-drops are observable and replayable. - Harden the tool parser against reasoning-interleave: strip/complete the reasoning
channel before scanning for tool tags; accept partial/streamed tag bodies; on failure,
fall back to emitting the raw text as
content(our client already handles content-channel answers — glm answers that way every call). - Bound the grind tail: if generation exceeds N seconds past
reasoning_budgetexhaustion with no emitted payload, cut and error. Empties that take 296s hurt double.
7. What we already did client-side (so the tax is bounded meanwhile)
- Dropped the dominated fallback rung. Our retry ladder used to step down from
forced-named to
"required"; measured per-roll success is 69% vs 52%, so retries now re-roll forced-named. Worst-case slot burn per pass halved (8 → 4 calls). - Fresh-thread engine retry for the residual all-rolls-fail tail (~1%).
- Content-channel fallback on every response — if the model answers in
contentinstead oftool_calls, we strict-parse it (this is how GLM answers 100% of calls); no regression risk from harness changes that route to content. - Full per-run event logs + conversation transcripts are now published for every run at
admin.donto.org/runs/<run-id>(pass timeline, per-call tool_choice/latency/status, raw model responses), so any harness-side change is verifiable against live traffic within minutes.
8. Reproduction recipe
Serial, ~2 minutes, reproduces at ~1-in-3 with any bearer key:
for i in 1 2 3 4 5 6 7 8; do
curl -sS -A 'curl/8' --max-time 300 \
-H "Authorization: Bearer $HYADES_KEY" -H "Content-Type: application/json" \
-H "X-Hyades-Thread: probe/$RANDOM" \
https://hyades.gnostr.cloud/v1/chat/completions -d '{
"model": "holo3.1", "temperature": 0.1,
"max_completion_tokens": 64000, "reasoning_budget": 8000,
"messages": [{"role":"user","content":"Extract every factual claim from this text as subject/predicate/object triples and return them by calling submit_facts: The schooner Meridian, under Captain J. Halloran, arrived from Thursday Island bearing forty tons of trepang consigned to Burns, Philp & Co. A crewman was lost overboard off Cape Melville; an inquest will be held before police magistrate H. M. Chester on Tuesday."}],
"tools": [{"type":"function","function":{"name":"submit_facts","description":"Submit the extracted facts.","parameters":{"type":"object","properties":{"facts":{"type":"array","items":{"type":"object","properties":{"subject":{"type":"string"},"predicate":{"type":"string"},"object":{"type":"string"}},"required":["subject","predicate","object"]}}},"required":["facts"]}}}],
"tool_choice": {"type":"function","function":{"name":"submit_facts"}}
}' | python3 -c 'import json,sys; d=json.load(sys.stdin); c=d["choices"][0]; tc=(c["message"].get("tool_calls") or [{}]); print("finish:",c.get("finish_reason"),"| args_chars:",len((tc[0].get("function") or {}).get("arguments") or ""),"| usage:",d.get("usage"))'
done
A failing iteration prints finish: tool_calls | args_chars: 0 | usage: {...completion_tokens: 600–1000...}.
9. Context: why this matters to donto
holo3.1 on Hyades is donto's volume workhorse — free, uncapped, and (measured separately, see
the donto model lab #1: holo3.1 report) capable of ≥48-way concurrency, 256K real context,
and deep multi-pass extraction with reasoning_budget: 8000. The lens-directed gleaning loop
now drives 6+ passes per document; at 8 concurrent slots the fleet moves ~1,500–2,000
statements/hour into the substrate. A flat 33% call tax — each costing 15–300s of a slot —
is the difference between the 2,800-document frontier-violence corpus taking days versus
weeks. Fixing the arguments-drop in the harness recovers roughly a third of the lane's
effective throughput with zero model or quality change.
Method notes: production numbers derive from per-run JSONL event logs written by the Rust extraction engine (every HTTP attempt/status/latency, tool_choice form, parse outcome, per-pass fact counts); the probe is 8 identical serial requests with fresh thread ids, raw bodies retained. Analysis scripts and raw probe bodies available on request. All times UTC, 2026-07-02. Contact: [email protected].