Quran Gallery Blog · Articles
Why AI invents Islamic citations, and what it took to stop it
Give a language model a classical library and it will still fabricate references — just more convincingly. What we measured while grounding a Qur'an and hadith assistant in 8,594 works, and the change that finally made a fabricated citation structurally impossible.
10 min read3 sources
Author:The Quran Gallery team
The most requested thing in our chat app’s entire history is not a feature. It is a demand for evidence. Pulled verbatim from our own production logs:
“Can you cite the source and quote it verbatim?”
“Any other citations?”
“Any more evidence?”
And, from one user, a pasted line: AL-MUGHNI BY IBN QUDAMAH VOL.4, PG.405 —
asking the assistant to check whether the book says what someone claimed it says.
The assistant’s answers, back then, looked like scholarship. They cited things like “Al-Mas’udi, Muruj al-Dhahab” — a real author, a real book, no volume, no page, and no way for any reader to check whether the sentence attributed to it appears anywhere in it.
That is not a citation. It is the shape of a citation.
Questions about hadith authenticity alone are 10.8% of everything the app is asked. This is not a polish problem.
“Just give it the library” does not work
So we connected one. A complete Maktabah al-Shamela 4 install — 8,594 works, Lucene-indexed. Warm queries return in 2–60 ms. Retrieval speed was never the problem for a single day.
Retrieval ranking was.
Lucene ranks by term frequency. Term frequency is not scholarly authority, and in a corpus spanning eleven centuries the two come apart immediately:
- Search التيمم (dry ablution) across the whole library: 26,098 hits. The top result is الدر الفريد وبيت القصيد, a 710 AH poetry anthology. The word appears in verse.
- Search إنما الأعمال بالنيات — arguably the most famous hadith in Islam: 6,815 hits, topped by a work of legal theory that merely quotes it. Sahih al-Bukhari never appears in the results at all.
Hand a model those results and it will cite the poetry anthology. It has no way to know the top hit is wrong. It looks exactly like a search result, because it is one.
The fix is unglamorous: never search unscoped. Every query now carries an explicit scope — the eight primary hadith collections pinned to specific printed editions, or a set of category ids for a fiqh domain, or one named book and nothing else. Scoped, that same hadith query returns صحيح البخاري , page 1/6, with the chain of transmission intact.
Precision made recall worse
Exact-phrase search sounds like precisely the right tool for “is this the wording, and where?” For hadith, it is not.
Searching إنما الأعمال بالنيات as an exact phrase returned one result — and missed Sahih al-Bukhari. Transmitted wording varies between printings: بالنيات in one edition, بالنية in another. A token search, requiring all the words on one page in any order, found Bukhari, Ibn Majah, Abu Dawud and al-Nasa’i.
So both passes are needed, in that order, for opposite reasons. The phrase pass answers “where is this exact sentence”; the token pass answers “which pages carry these words.” Running only the stricter one turns a genuine hadith into a not-found — which, in this domain, is its own kind of false claim.
Real users do not type keywords
We pulled 2,832 real questions out of production and ran the retrieval layer against all of them. The dominant failure was strict AND: every token in the query had to appear somewhere on a single page.
صفة صلاة الوتر matches 406 pages. وكيفيتها matches 873 pages. All four words together: zero.
The library holds the answer. The shape of the question hid it. On 18 realistic sentence-shaped fiqh queries, 10 returned nothing. On 30 real production questions, 5 did.
We fixed this in the index layer rather than by trimming the user’s words, because Lucene can weigh every subset of the query at once while a heuristic has to guess which word to sacrifice — and it guesses badly. Dropping by prefix loses اليهودية off a question about the Khazars. Dropping the rarest term splits ابن سينا into ابن. Instead the search walks its own requirement down and stops at the most of the user’s own words that any page actually carries, then reports that it relaxed, so the answer can be labelled a near match instead of an exact one.
Zero-result rate on those 30 real questions: 5/30 → 0/30. A question typed as witr prhne ka traika — Roman Urdu, no diacritics, no Arabic at all — went from
nothing to nine sources led by al-Bidaya and al-Fatawa al-Hindiyya.
The sophisticated upgrade that mostly made things worse
The library ships with Arabic morphological analysis. Turning it on globally looked obviously correct and was not:
| Query | Plain | With morphology |
|---|---|---|
مس الذكر ينقض الوضوء | 566 hits, right fiqh manual on top, 17 ms | 4,362 hits, legal-theory works on top — worse |
حكم الوضوء من أكل لحم الجزور | 0 hits | 159 hits, the correct manuals — rescued, 467 ms |
Morphology expands recall and wrecks ranking. So it is not a setting; it is a fallback. Plain search first, morphology only to turn nothing into something.
Which printing?
Back to that pasted line: AL-MUGHNI BY IBN QUDAMAH VOL.4, PG.405.
Two traps here, and neither is visible until you probe the actual data.
A work is not an edition. al-Mughni exists in this library in several printings, and pagination is exactly what differs between them. Page 4/405 in discusses a sale conditioned on building or planting. Page 4/405 in discusses a worsening illness. Different text, same coordinate. Picking the highest-ranked edition is a coin flip with a citation attached — so every printing is opened and named, and when a page cannot be resolved that is reported as a failure rather than quietly substituted.
Printed pages are not database pages. Internal page ids do not map one-to-one onto printed page numbers. In Bukhari’s ط السلطانية the ratio runs roughly eight internal ids per printed page. A resolver stepping one id at a time drifts out of range and reports a page that genuinely exists as missing. Correcting by the slope the samples reveal, instead of by a fixed step, fixes it.
The negative result turned out to be the most valuable output of all this. “That wording is not in al-Mughni” is only sayable because the search was confined to al-Mughni.
The worst failure we produced
This is the one that changed the architecture.
Asked to verify a fabricated quotation attributed to al-Muwatta, the assistant answered NOT FOUND — correct — and then footnoted that answer with three genuine al-Muwatta pages.
Every citation was real. Every page number checked out. Audit any one of them and it passes. The conclusion was still unproven, and the three real citations made it read as more verified than a bare assertion would have.
The cause was not a bad prompt. It was a discarded result. The phrase search that actually established the absence returned nothing, so it was thrown away; the fallback token search returned neighbouring pages; and the model inferred the absence from their silence.
A failed search is evidence. Discarding it and keeping the fallback’s near-misses is how a correct answer ends up supported by the wrong reasons.
What the architecture became
Three ideas, in increasing order of how much they moved.
1. A ledger, not a prompt
Every passage retrieved during a turn is registered server-side under an id. The
model writes [S2]. It never writes a book title, an edition, or a page number —
those are resolved from the ledger afterwards and rendered into the citation
card.
Asking a model nicely not to invent page numbers is a prompt. Removing its ability to write one is an architecture. The fields are structurally unreachable to it.
2. Tool inputs can be refused; streamed text cannot be unsent
The answer body streams to the reader token by token. Nothing can un-send it, so anything applied to that prose afterwards is a ceiling on damage, not prevention — and we document it that way, because calling mitigation a guarantee is its own kind of overclaim.
But the citation list does not stream. It arrives as the arguments to a final tool call, and arguments can be rejected before a reader ever sees them. A citation whose id the ledger never issued is refused outright. The consequence is deliberate and slightly harsh: an ungrounded turn displays no source cards at all, rather than plausible-looking ones.
3. The turns with no evidence were the turns with no rules
This was the load-bearing realization, and we would never have reached it by reasoning — only by reading logs.
Our entire citation discipline was keyed on the presence of retrieved evidence. Which means the turns most capable of fabricating — a greeting, an outage, a question the library could not answer — were the only ones governed by nothing at all.
Worse, a turn that searched and found nothing fell through to the general instructions, which open with “Answer from general knowledge.” The model was not bypassing the evidence gate. It was obeying an instruction. Straight from the database: asked to locate a specific wording in a named book, it reported the failure honestly, then wrote “However, addressing your question as a matter of general knowledge:” and supplied the attribution anyway.
There are now three routes where there were two: grounded, ungrounded by design (greetings, pastoral questions), and searched-and-found-nothing — a different situation from never having searched, which has to say so out loud.
What still does not work
Publishing only the wins would be the same failure this whole post is about.
- Free prose remains a ceiling. Citation cards cannot be faked. Sentences in the answer body can still carry an attribution the retrieval did not earn. We scrub the obvious patterns and we are explicit that scrubbing is mitigation. The scrub itself broke first, instructively: matching روى and أخرج as attribution markers destroyed four out of five legitimate sentences, because they are ordinary Arabic verbs, before it was constrained to require an actual collection name.
- Narrative sourcing lags slot sourcing. On a four-school comparison, the Maliki position was supplied from an annotation sitting inside a Hanafi work. The per-school verification gate caught it — and the model routed around the gate in prose. Binding the narrative to the same eligibility rules as the verified slots is open work.
- Coverage is not uniform, and pretending otherwise would be the same sin. The hadith cross-reference index in this install covers exactly ten books. Of the eight primary collections, only two — al-Bukhari and al-Tirmidhi — have entries at all. Cross-collection takhrij can never resolve outside those ten, regardless of how relevant a page is.
- Not everything in the library is citable. A meaningful portion is paginated by the software rather than by a published edition, so those page numbers point at nothing you could pull off a shelf. Those works are excluded from citation entirely.
If you are building this for another corpus
Legal, medical, scriptural — the transferable parts:
- Ranking is not authority. Scope every query. A general index will confidently hand you a poem instead of a legal manual, and it will look exactly like a correct result.
- Make fabrication structurally impossible, not merely discouraged. If the model can emit a page number, eventually it will. Take the field away from it and resolve it server-side.
- Refuse at tool inputs, not in streamed prose. Anything you can only fix after it renders is mitigation. Design so the dangerous fields arrive somewhere refusable.
- Failed searches are results. Carry them forward. Discarding them is how correct conclusions acquire the wrong evidence.
- Govern the empty turns. The moment with no retrieved evidence is the moment with the most freedom, and it is usually the moment nobody wrote a rule for.
And one lesson that is not engineering at all: not one defect in this post was found by reading the code. Every single one surfaced by probing the live library with real questions and checking the result against the printed page. Throughout all of it, the system type-checked cleanly and its unit tests passed.
Grounded answers are live now at qurangallery.com/chat. If you find a citation that does not hold up, please tell us — being checkable is the entire point.
May Allah ﷻ protect us from speaking about His religion without knowledge.
Related
Journey Through Salah Salām: The Words That Release You From the Prayer The salām is not the prayer trailing off — it is the act that ends it, the counterpart to the takbīr that began it. What the words are, how far the head turns, who is actually being greeted, whether one salām or two, and the dhikr the sources place immediately after it. Articles Loving the Messenger, Loving Allah's Beloved The Qur'an does not accept a claim of love for Allah at face value — it appends a test, and that test runs straight through love for His Messenger. What that test actually asks of a believer, and what it promises in return, is not the same thing as feeling moved by his name. Journey Through Salah Before the Salām: Ṣalawāt on the Prophet, and a Duʿāʾ of Your Own The last stretch of the prayer, after the taḥiyyāt and before the salām, runs on a taught order: praise, then ṣalawāt on the Prophet, then whatever you yourself need to ask for. Every step of it is named in a hadith you can open, including the one place where the fixed script hands the words back to you.
