QafAgent benchmark

New tools vs production

Branch codex/agent-improvements @ 36803954 with prompt v29, against production (main + prompt v24). 100 real questions, 10 Oct 2026.

Better than production

The new tools give better answers at about the same cost. A blinded judge preferred them in 50 of the 72 cases it could decide (69%). The gain comes from questions that name a source, need the full text, or follow up on an earlier turn. On plain questions the two arms are even.

Answers are a little better grounded, and serious errors drop from 6 to 2. The price is about 1 s more wait for first text and +5% cost per turn. Next: get the model to apply filters more consistently (it filtered in only 20 of 62 questions that named a source), and look at hadith takhrīj, where it trails.

ProductionNewChange
Qualityblinded wins, both orders agree225069% decided win rate95% CI 58–79% · p = 0.001
Accuracyclaims fully supported72.8%76.6%+3.8 pts95% CI +1.0 to +6.9
Serious errorsanswers that would mislead62−4 answersCI −9 to +1 pts, not conclusive
Cost per turntokens + rerank$0.049$0.051+5%paired CI includes zero
Latencyp50 to first text13.4 s14.1 s+1.1 smedian paired diff · CI +0.1 to +1.9

Decided win rate for the new arm

Share of decided cases won by the new arm. Bars are 95% Wilson intervals; the dashed line is parity. A case is decided only when the judge picks the same winner in both answer orders. “Names a source” means a filter on book, author or period would help; “full text” means the answer needs passages around a hit (expansion); “plain” is neither.

Where it wins, where it loses

Wins

  • Questions that name a source: 34–7. It stayed within the requested book or author in 79% of answers, against 61% for production, and never ignored the constraint (production ignored it 3 times).
  • Follow-ups: 10–0. It keeps earlier tool results in its history, so later turns build on the passages it already found.
  • Full-text questions: 18–7. Provided context was used in 65% of answers, against 50%. Production’s expand mostly fails: 12 of its 15 calls returned nothing.

Loses or no gain

  • Plain questions: 12–11. With no source named and no expansion needed, the new arm is no better.
  • Hadith work. It trails in takhrīj, 2–4 (n 9). Its one serious error that production avoided was a misidentified narrator.
  • Shorter answers lose: 8–16. The judge leans towards longer answers. At similar lengths the new arm still wins 24–12 (p = 0.07).
  • Slightly slower and costlier. First text arrives 1.1 s later. Cost is up 5%, mostly rerank, which now runs on every search (2.9 against 2.1 calls per turn).
  • Filters underused. Applied in only 20 of the 62 questions that named a source. Arabic factual errors also rose slightly, from 8 to 10, though serious ones fell from 4 to 2.
New wins · French · named book
Does Akhṣar al-Mukhtaṣarāt treat dog and pig impurity as heavy or light?

The new arm looked the book up and searched inside it. Production answered from other works: 4 claims contradicted the sources and 3 attributions were fabricated.

Production wins · Arabic · narrator
In a chain from al-Nasāʾī’s al-Sunan al-Kubrā, who is ʿAbdullāh ibn Muḥammad?

The new arm ran 8 searches without a catalog lookup and settled on an identification the sources did not support. The audit flagged it as serious.

Questions are paraphrased.

Method and caveats