New tools vs production
Branch codex/agent-improvements @ 36803954 with prompt v29, against production (main + prompt v24). 100 real questions, 10 Oct 2026.
The new tools give better answers at about the same cost. A blinded judge preferred them in 50 of the 72 cases it could decide (69%). The gain comes from questions that name a source, need the full text, or follow up on an earlier turn. On plain questions the two arms are even.
Answers are a little better grounded, and serious errors drop from 6 to 2. The price is about 1 s more wait for first text and +5% cost per turn. Next: get the model to apply filters more consistently (it filtered in only 20 of 62 questions that named a source), and look at hadith takhrīj, where it trails.
| Production | New | Change | |
|---|---|---|---|
| Qualityblinded wins, both orders agree | 22 | 50 | 69% decided win rate95% CI 58–79% · p = 0.001 |
| Accuracyclaims fully supported | 72.8% | 76.6% | +3.8 pts95% CI +1.0 to +6.9 |
| Serious errorsanswers that would mislead | 6 | 2 | −4 answersCI −9 to +1 pts, not conclusive |
| Cost per turntokens + rerank | $0.049 | $0.051 | +5%paired CI includes zero |
| Latencyp50 to first text | 13.4 s | 14.1 s | +1.1 smedian paired diff · CI +0.1 to +1.9 |
Decided win rate for the new arm
Share of decided cases won by the new arm. Bars are 95% Wilson intervals; the dashed line is parity. A case is decided only when the judge picks the same winner in both answer orders. “Names a source” means a filter on book, author or period would help; “full text” means the answer needs passages around a hit (expansion); “plain” is neither.
Where it wins, where it loses
Wins
- Questions that name a source: 34–7. It stayed within the requested book or author in 79% of answers, against 61% for production, and never ignored the constraint (production ignored it 3 times).
- Follow-ups: 10–0. It keeps earlier tool results in its history, so later turns build on the passages it already found.
- Full-text questions: 18–7. Provided context was used in 65% of answers, against 50%. Production’s expand mostly fails: 12 of its 15 calls returned nothing.
Loses or no gain
- Plain questions: 12–11. With no source named and no expansion needed, the new arm is no better.
- Hadith work. It trails in takhrīj, 2–4 (n 9). Its one serious error that production avoided was a misidentified narrator.
- Shorter answers lose: 8–16. The judge leans towards longer answers. At similar lengths the new arm still wins 24–12 (p = 0.07).
- Slightly slower and costlier. First text arrives 1.1 s later. Cost is up 5%, mostly rerank, which now runs on every search (2.9 against 2.1 calls per turn).
- Filters underused. Applied in only 20 of the 62 questions that named a source. Arabic factual errors also rose slightly, from 8 to 10, though serious ones fell from 4 to 2.
Does Akhṣar al-Mukhtaṣarāt treat dog and pig impurity as heavy or light?
The new arm looked the book up and searched inside it. Production answered from other works: 4 claims contradicted the sources and 3 attributions were fabricated.
In a chain from al-Nasāʾī’s al-Sunan al-Kubrā, who is ʿAbdullāh ibn Muḥammad?
The new arm ran 8 searches without a catalog lookup and settled on an identification the sources did not support. The audit flagged it as serious.
Questions are paraphrased.
Method and caveats
- Selection. 100 questions picked by hand from production traces (1 Sep to 10 Oct 2026): 50 Arabic and 50 English (4 of them other Latin-script languages), 62 naming a source, 34 needing expansion, and 11 real multi-turn follow-ups, for 112 turns. Personal, product and chit-chat messages were excluded, and no user ids were kept.
- Blinded judging. Claude Opus compared the answers without labels, in both orders, and a case counts only when both orders agree (79% did). A separate Claude Opus auditor checked every claim against the passages that arm itself retrieved: 1,518 claims for the new arm, 1,450 for production.
- Cost model. Measured tokens at Gemini 3.7 Flash intro pricing ($0.75/M input, $3.75/M output including reasoning, cached input assumed at 10%), plus Cohere rerank at $0.0025 per call. Turbopuffer and catalog lookups are not priced. Prices double on 1 Jan 2027.
- Production fidelity. The production arm is the harness frozen from main, with shims restoring main’s behaviour where shared helpers have drifted: raw-query keyword search, version-only expand and main’s label wording. It runs prompt v24.
- Provider. Both arms called Vertex Gemini 3.7 Flash directly at the same thinking level, not through the Cloudflare gateway, so latency leaves out the gateway hop. Retries were off on both arms, and all 200 trials completed.
- Noise. Each arm answered each case once. Intervals are Wilson (win rates) or bootstrap (paired differences), and p-values come from a sign test. Treat slices under about 15 cases as hints.