The Mask Is Not the Model: Auditing Prefix Invariance in Attention, State-Space, and Hybrid Sequence Models

Taebong Kim, Youngsik Hong, Minsik Kim, Sunyoung Choi, Jaewon Jang, Minseo Kim

arXiv:2608.22876 · 2026-08-26 공개 · arXiv · PDF

state-space-models attention-mask hybrid-models injected-faults zamba2 nemotron-h prefix-invariance causality-audit

Abstract

We formalize prefix invariance: representations at position t must not depend on future inputs. We give a lightweight audit, two forward passes, no training or gradients, that localizes exactly where causality breaks. Attention-mask inspection is incomplete: leaks can occur via scans or normalization despite correct masks. Across 192 injected-fault trials on eight checkpoints, mask inspection found none, while our audit localized all 192/192, also finding a defect in Zamba2 and Nemotron-H.

한국어 요약

한 줄 요약

전향 마스크 검사가 누설을 누락하는 이유를 규명하고, 새로운 인과성 검증 방법을 제시한다.

핵심 기여도

핵심 아이디어

기존의 어텐션 마스크 검사는 미래 정보 누설을 완전히 막지 못한다. 스캔이나 정규화 과정을 통해 누설이 발생할 수 있기 때문이다. 이 논문은 위치 t의 표현이 미래 입력에 의존하지 않아야 한다는 prefix invariance 원칙을 정식화하고, 이를 검증하기 위한 새로운 감사 프로토콜을 제시한다. 이 접근법은 훈련이나 그래디언트 없이 두 번의 순방향 패스만으로 인과성 위반을 정확히 특정할 수 있다.

기술적 접근법

주요 결과

의의 및 한계

실용적 활용

이 방법은 대규모 언어 모델의 인과성 검증, 모델 감사, 보안 검증 등에서 활용 가능하며, 특히 모델 신뢰도를 높이는 데 기여할 수 있다.