Post-Training Leaves Behavioral Shadows on Unrelated Decisions

Ziyang Zhang, Yubin Jing, Yuanhao Zeng, Yuyao Li, Haofan Wang, Yichen Gong

arXiv:2609.29233 · 2026-09-29 공개 · arXiv · PDF

language-models post-training prompt-engineering human-eval commonsense-reasoning model-updates capability-transfer taskless-distillation

Abstract

We find that language models can transfer capabilities through task-unrelated text. Post-training typically improves language models using task-specific data. Prior work on subliminal learning shows that information about these updates can pass through unrelated generations, but has largely focused on traits or preferences using extensive teacher outputs. We introduce Active Taskless Distillation (ATD), which achieves capability transfer using only a single word from the teacher per prompt. ATD probes the behavioral shadow of post-training by selecting prompts where the teacher and student's shared public ancestor is nearly indifferent between two ordinary words. A student initialized from this ancestor learns solely from the resulting prompt-word pairs, without target-task examples, teacher logits, or teacher parameters. In the primary coding experiment with Qwen2.5-1.5B, 5,664nses yield a 5.34 pp gain on HumanEval+ over an exact nuisance-matched control thadisrupts prompt-resperiments showtransfer in scientific knowledge, commonsense reasoning, and reading comprehensins across additional model generations, sizes, and families. Functional analyses show that the learned sid composable, andthat its strength tracks the teacher's update strength.

한국어 요약

한 줄 요약

Active Taskless Distillation(ATD)를 통해 단일 단어만으로 사후 학습 정보를 전달하고, HumanEval+에서 5.34pp 성능 향상.

핵심 기여도

핵심 아이디어

기존 연구는 사후 학습 정보가 무관한 생성물에도 퍼질 수 있음을 보여주었으나, 대부분은 긴 응답을 필요로 했다. 본 연구는 ATD를 제안하여, 사후 학습 정보가 단일 단어 선택에도 반영될 수 있음을 보여준다. ATD는 공통 공개 조상 모델을 기반으로, 두 단어 간 확률이 거의 동일한 프롬프트를 선정하고, 사후 학습된 교사 모델이 선택한 단일 단어를 학습 데이터로 사용한다. 이는 학습 대상이 사후 학습의 효과를 간접적으로 학습할 수 있도록 한다. 핵심 통찰은, 학습 신호가 단일 선택에 내재되어 있음에도 불구하고, 축적하면 전체 성능에 긍정적인 영향을 미친다는 점이다.

기술적 접근법

주요 결과

의의 및 한계

본 연구는 사후 학습의 영향이 타겟 태스크 외에도 무관한 입력에 퍼질 수 있음을 입증하며, 이는 모델의 비공개 학습 내용을 간접적으로 추적할 수 있음을 시사한다. ATD는 단일 단어만으로도 학습 정보를 전달할 수 있음을 보여주어, 모델 학습 과정의 투명성과 모델 간 지식 전달 가능성에 중요한 통찰을 제공한다. 그러나 모든 태스크와 모델에서 동일한 성능 향상을 보장하지는 않으며, 학습 신호의 효과는 모델 구조와 학습 조건에 따라 달라질 수 있다.

실용적 활용

ATD는 모델 학습 과정의 투명성 증대, 모델 간 지식 전달, 사후 학습 효과 분석 등 연구 및 산업 분야에서 활용 가능하다. 특히, 모델의 비공개 학습 내용을 추적하거나, 모델 간 지식 공유를 효율적으로 수행할 수 있는 기반 기술로 활용될 수 있다.