[Submitted connected 31 Aug 2026]
View PDF HTML (experimental)
Abstract:Ambient AI scribes draught objective notes, and published audits find their ascendant correction is omission: accusation the brushwood established that the statement fails to record. The modular cheque is an LLM judge: a 2nd exemplary sounds the statement against the transcript and flags problems. We inquire whether judges observe omissions. Public corpora cannot proviso the reply key: their clinician reference notes and transcripts are materially discrepant. Our benchmark has 500 single-error statement pairs from audited truth sheets, 298 pinch a named truth surely absent and 202 added-or-altered controls. Across 8 judge designs, paired favoritism (the flawed statement beneath its cleanable twin, 0.5 a coin flip) sounds 0.79-0.94 connected added aliases altered contented and 0.50-0.63 connected omissions. On azygous notes, nary creation flags omissions reliably much often than cleanable notes. Wording changes, voting and GEPA punctual optimisation move the operating constituent without creating usable detection. Restructuring the task recovers it: database the facts the transcript establishes, past cheque the statement for each. Two methods scope it independently and waste and acquisition off: a per-fact pipeline, and a GEPA-evolved punctual doing the aforesaid successful 1 call. The pipeline's flags sanction the missing truth and its severity astatine 2.7% mendacious alarms. The azygous telephone detects much (36.9% against 24.6%, p=0.002) astatine 6.2% mendacious alarms and a tenth of the costs per note. A expert writer validated 70 items and, wherever the 2 routes disagree, sided pinch the pipeline connected 10 of 10 (p=0.002). A 2nd clinician, not an author, graded the severity rubric unsighted and agrees to wrong a grade. On existent vendor notes from a companion census nary benchmark period transfers, but the re-calibrated azygous telephone detects much than the champion of the 8 astatine half its false-alarm rate. Omissions whose truth is restated elsewhere conclusion some routes. We merchandise the benchmark, prompts and judgements.Submission history
From: Sebastian Fox [view email]
[v1] Mon, 31 Aug 2026 15:59:54 UTC (337 KB)
English (US) ·
Indonesian (ID) ·