Incidental information contaminates patient notes and disrupts clinical reasoning in large language models

Krithik Vishwanath, Brandon Ye, Anton Alyakin, John E. Markert, Aaron Hsieh, Michał Mańkowski, Eric K. Oermann

arXiv:2610.08585 · 2026-10-11 공개 · arXiv · PDF

large-language-models model-evaluation clinical-ai clinical-reasoning noise-robustness model-contamination ambient-documentation incidental-information

Abstract

Large language models (LLMs) are increasingly relied upon to support ambient documentation and clinical reasoning. Here we examine the impact of a failure mode shared between these two applications by assessing their sensitivity to information incidental to the patient encounter. In 576 patient-clinician dialogues, we found that frontier models inserted small-talk exchanges into 35% of notes, while mean quality scores changed by at most 0.20 points on five-point scales. In 3.7% of frontier notes, models misattributed the asides or used them clinically. In 57 mock recorded consultations, background speech from a separate patient encounter at -10 dB leaked into 48.2% of transcripts, with contamination detected in 5.3% of downstream notes generated by four open-weight models. We propose a dual encoding hypothesis of clinical reasoning and distraction in LLMs, with preliminary evidence that LLM components associated with disruption by incidental information also support clinical reasoning. These findings support evaluating resistance to incidental information before clinical use, with safeguards that prevent contamination while preserving clinical reasoning.

한국어 요약

한 줄 요약

대규모 언어 모델(LLM)이 임상 기록과 추론 과정에서 부수적 정보에 민감하게 반응하여 의료 오류를 유발할 수 있음을 밝힘.

핵심 기여도

핵심 아이디어

LLM이 임상 기록을 생성할 때, 환자와 무관한 대화나 배경 소음이 노트에 포함되어 의료 오류를 유발할 수 있음을 밝혔다. 특히, **Open-weight 모델**은 **35.0%** 대비 **50.1%**의 노트에 부수적 정보를 포함시켰으며, 이는 **25.7%**의 경우 환자에게 잘못 적용되었음.

이러한 현상을 설명하기 위해 **Dual Encoding Hypothesis**를 제안함. 즉, LLM 내 특정 어텐션 헤드는 임상 추론과 동시에 부수적 정보에 민감하게 반응하며, 이 헤드를 억제하면 정확도가 **24.4–37.8% 포인트** 감소함. 이는 부수적 정보에 대한 방어와 임상 추론 능력을 동시에 유지하는 것이 중요함을 시사함.

기술적 접근법

주요 결과

의의 및 한계

본 연구는 LLM이 임상 기록과 추론 과정에서 부수적 정보에 민감하게 반응할 수 있음을 처음으로 실증적으로 밝힘. 특히, **MedDistractNotes**와 **MedDistractAudio** 데이터셋은 부수적 정보 오염을 평가하는 기준을 제공하며, **CCHG**를 통한 메커니즘 실험은 오염 방지 전략 개발에 기여함.

그러나, 연구는 **제어된 환경**에서 수행되었으며, 실제 임상 상황의 **음성 복잡성**과 **대화 다양성**은 반영되지 않았음. 또한, **단일 LLM**에 의한 오염 판단과 **의료 전문가 검증**이 누락되었으며, **일반 지식 벤치마크**에서의 평가가 이루어지지 않았음.

실용적 활용

본 연구는 **Ambient Documentation 시스템**의 안전한 배포를 위해 **부수적 정보에 대한 민감도 평가**가 필수적임을 시사함. **의료 기록 시스템**에서 **발화자 식별**, **의료 관련성**, **참조 구간 검증**을 구조화하는 방식이 필요하며, **CCHG**를 활용한 메커니즘 개입이 오염 방지에 유용할 수 있음.