A new preprint tested whether tweaking conscionable 1 portion of a root changes AI hunt citations, pinch everything other kept the same. In the earthy numbers, the apical consequence sewage cited astir doubly arsenic often arsenic the fifth. When researchers Sriram Selvam and Anneswa Ghosh reversed the bid of matched sources, the effect was overmuch smaller and measured zero successful a follow-up test.
Posted to arXiv connected September 14, the insubstantial isn’t peer-reviewed. The study covers 1 GPT-5.4 hunt supplier that uses Exa arsenic its hunt provider, featuring offline replayed conversations and nary unrecorded webpage edits.
How The Test Worked
The researchers prompted the GPT-5.4 supplier to reply 130 communal questions by having it execute independent web searches. They recorded each connection and hunt consequence from the 129 questions it addressed. From these transcripts, they chose pairs of pages that appeared successful the aforesaid hunt outcomes and were some screened arsenic supporting the aforesaid fact. This screening aimed to find situations wherever either page could beryllium reasonably cited. When existent matches were identified, immoderate in installments differences were owed to really the exemplary apportioned nickname betwixt the 2 sources, some confirming the aforesaid fact.
That near 113 pairs. A later blinded quality cheque confirmed 103 of them arsenic genuine matches. The researchers replayed each saved speech 4 ways, placing 1 page supra aliases beneath the different and showing its matter either arsenic plain paragraphs aliases rewritten pinch headings and lists aliases a table. Only the last reply was generated again.
Both versions of the matter were generated by AI rewrites of the original page. Grok 4.3 created astir each of them, pinch GPT-5.4 utilized arsenic a fallback for 1 pair, and a abstracted Grok reappraisal checked that the facts matched. The wording varies betwixt the 2 versions, truthful the authors statement that the trial compares 2 rewrites but does not specifically isolate formatting differences.
Raw Position Gap Was Larger Than Swap Effects
In the first position of a hunt call, pages were cited 85.1% of the clip successful saved transcripts, compared to 42.8% for pages successful the 5th position. This creates a quality of 42.3 percent points.
Here, ‘position’ simply intends the bid of the 5 Exa results returned successful a azygous search, not wherever a page ranks connected Google aliases its position connected the unrecorded web.
The study highlights that hunt providers usually put much applicable pages astatine the top, truthful the earthy quality reflects some the position and the value of the pages. When the researchers moved the aforesaid page higher wrong its pair, the chance it was cited astatine each went up by 7.9 percent points. However, this uncovering wasn’t considered statistically important aft accounting for aggregate tests.
Another testing group pinch 56 pairs, wherever only the bid was switched, showed an estimate of 0.0 points, pinch a 95% assurance interval from -5.4 to +5.4.
The study explains that the earthy spread and the switch results measurement different aspects. Overall, it suggests that position did power citations successful immoderate cases, but averages from earthy position information aren’t reliable.
Structured Rewrites Got More Credit, Not Clearer Entry
Pages that were rewritten pinch headings and lists received an mean of 0.50 much citation markers per reply compared to the aforesaid pages written arsenic plain paragraphs, pinch a 95% assurance interval from 0.20 to 0.84. The answers successful the trial were heavy cited, pinch a median of 29 markers crossed six documents.
The full number of citations per reply didn’t rise, and the number connected the different page hardly changed. The authors spot this arsenic in installments being focused much connected the rewritten page.
The main trial the researchers conducted, which they planned earlier starting the experiment, was to spot if the page sewage cited astatine all. They recovered that utilizing system matter accrued that likelihood by 4.5 percent points, pinch a 95% interval from -1.4 to +10.4. The insubstantial points retired that this consequence isn’t conclusive and mentions that the study could reliably observe only effects of astir 8.5 points aliases more.
A much strict comparison, wherever each connection stayed the aforesaid but the layout was adjusted to 1 condemnation per database row, boosted citation rates crossed each 113 pairs. When they repeated the trial pinch a subset, the effect reversed.
In the chat conception of the paper, the authors shared these insights:
“This is an attribution-sensitivity warning, not an optimization tactic.”
Reruns Changed Citation Outcomes
The researchers tested 120 responses again utilizing the aforesaid inputs, and recovered that the determination to mention aliases not for the target page changed successful 15% of these cases, astir 1 successful seven.
The mean count effect remained accordant crossed these reruns. They estimate that astir 45% of the variety successful a azygous run’s effect is owed to exemplary randomness.
The authors urge rerunning citation tests aggregate times and sharing really accordant the results are crossed those runs.
Additionally, SparkToro reported successful January that ChatGPT and Google’s AI Overviews each produced the aforesaid marque database little than 1% of the clip erstwhile fixed the aforesaid punctual repeatedly.
Why This Matters
The earthy position spread successful this trial was overmuch larger than the mean effect observed erstwhile researchers swapped root order. An Ahrefs report from May showed pages cited by AI were astir 3 times much apt to see JSON-LD schema, but adding schema didn’t intelligibly summation citations.
This raises questions astir whether a relationship successful a vendor study aliases your search was ever tested by changing the variable, and a azygous reply is simply a anemic ground for labeling a citation arsenic won aliases lost.
The study can’t corroborate if reformatting a unrecorded page boosts citations, since rewrites only applied to matter already retrieved, excluding crawling, retrieval, and ranking processes.
Looking Ahead
The researchers re-ran the saved searches connected Grok 4.3, discovering that the system rewrites leaned the aforesaid way. However, little than half of Grok’s first replies followed the correct citation format.
The authors urge much investigation to trial each script a fewer times, research different hunt providers and models, and salary attraction to some really often citations hap and if a page is cited astatine all.
Featured Image: Accogliente Design/Shutterstock