Spectral Feedback for Test-Time Alignment of Protein Diffusion Models

arXiv:2609.30456 · 2026-09-28 공개 · arXiv · PDF

model-agnostic reward-maximization spectral-feedback protein-diffusion test-time-alignment inverse-folding edit-sets stability-reward

Abstract

Reward maximization alignment methods for discrete diffusion models have primarily focused on steering the reverse process, either by influencing token logits or by selecting favorable sequences at intermediate steps. These approaches largely treat inference as a unidirectional process, lacking mechanisms for revisiting undesirable token selections. We introduce Spectral Feedback, an algorithm that selects edit-positions in a feedback loop, allowing the model to iteratively correct its own generations. This approach leverages the mask structure of discrete diffusion models by re-masking and re-sampling tokens, analogous to image editing methods that reintroduce noisy latents and re-run the reverse process. While prior alignment methods focus on what token labels to assign to maximize a target reward, we instead treat which tokens to revisit as the central alignment problem. Selecting edit-positions is challenging because edit effects are interdependent: the impact of modifying one token depends on which others are edited simultaneously. We define an edit-set as a set of token positions to re-mask and re-sample. Motivated by prior work on sparse interactions in biological systems, we find empirically that edit-set value functions for protein inverse folding admit sparse Fourier representations. This structure enables Spectral Feedback to efficiently learn and optimize the value functions for edit-position selection. Spectral Feedback is model-agnostic and can be applied to pretrained, test-time aligned, and fine-tuned diffusion models. For all of these models, the algorithm improves alignment performance without modifying the underlying generative process. Applied to inverse folding with a protein stability reward oracle, it achieves a 32.3% increase in stable proteins for a pretrained model, 24.8% for Best-of-10, and 5.8% for a state-of-the-art RL fine-tuned diffusion model.

한국어 요약

한 줄 요약

Spectral Feedback는 테스트 타임에 단백질 생성 모델의 생성물을 반복적으로 수정하여 안정성을 32.3% 향상시키는 알고리즘이다.

핵심 기여도

핵심 아이디어

기존의 테스트 타임 정렬 방법은 주로 **토큰 로짓 조정** 또는 **중간 단계에서 유리한 시퀀스 선택**에 집중했으며, 생성 과정을 일방향적이라고 가정했다. 이에 반해, Spectral Feedback은 **반복적인 수정**(feedback loop)을 도입하여, 생성된 시퀀스에서 **불필요한 토큰 선택을 재검토**하고 **재샘플링**함으로써 품질을 개선한다. 이는 이미지 편집에서 사용되는 **재노이징**(re-noising)과 유사한 방식으로, **재마스킹**(re-masking)과 **재샘플링**(re-sampling)을 반복한다.

핵심 통찰은 **편집 효과**(edit effects)가 **상호의존적**이라는 점이다. 즉, 한 토큰을 수정할 때 그 효과는 다른 토큰이 동시에 수정되는지에 따라 달라진다. 이를 해결하기 위해, **편집 집합**(edit-set)을 정의하고, 이 집합의 가치 함수를 **Fourier 기반 희소성**을 통해 근사한다. 이는 생물학적 시스템에서 흔히 관찰되는 **에피스타시스**(epistasis) 구조와 유사하며, 실제로 단백질 역접힘에서 **편집 집합의 가치 함수가 희소한 Fourier 표현**을 가진다는 것을 실험적으로 확인했다.

기술적 접근법

주요 결과

의의 및 한계

Spectral Feedback은 **기존 테스트 타임 정렬 방법**(token logit alignment, tree-search)과 달리, **모델 내부 구조를 수정하지 않고도 품질 향상**을 도모할 수 있다. 이는 **모델-에이전트**(model-agnostic) 특성 덕분이며, 다양한 단백질 생성 모델에 쉽게 적용 가능하다. 또한, **Fourier 기반 희소성**을 활용하여 **조합 최적화 문제**(combinatorial optimization)를 효율적으로 해결함으로써, **생물학적 시스템의 상호작용 구조**를 반영한 정렬 방법을 제시한다.

하지만, **scRMSD**(structure-based RMSD)가 일부 모델(예: DRAKES)에서 증가하는 경향이 있어, **구조적 정확도와 안정성 간의 트레이드오프**가 존재함을 시사한다. 또한, **고차 Fourier 항**(higher-order terms)이 포함된 경우, 계산 복잡도가 증가할 수 있다.

실용적 활용

Spectral Feedback은 **단백질 설계**(inverse folding), **생물학적 속성 최적화**(stability, β-sheet 등)에 적용 가능하다. 특히, **사전 학습된 단백질 생성 모델**을 테스트 타임에 정렬하여 **추가 학습 없이도 생성 품질을 향상**시키는 데 유용하다. **생물학 연구**, **의약품 개발**, **단백질 엔지니어링** 분야에서 실용적 활용 가능성이 높다.