diffusion-transformers denoising fid-metric image-synthesis structured-pathways adaptive-connectivity layer-guidance residual-connectivity
Abstract
Diffusion Transformers (DiTs) have established themselves as a scalable backbone for high-fidelity image synthesis. However, unlike U-Net based diffusion models that rely on rigid, hand-crafted skip connections, DiTs predominantly use a uniform residual stream that integrates all preceding layers as a monolithic state. In this work, we rethink residual connections in diffusion transformers and propose to transform them from passive summation into an active retrieval mechanism optimized for image denoising. First, we conduct a systematic analysis of DiT's internal representation, revealing a latent preference for early-layer feature reuse and symmetric layer guidance. Motivated by this, we introduce a structured connectivity design that explicitly integrates local residual connections with long-range pathways. Instead of static skip connections or dense all-layer routing, our method enables each transformer block to selectively ``attend'' to critical earlier representations, dynamically retrieving spatial and semantic cues through direct, differentiable cross-depth paths. Experiments show that our adaptive connectivity leads to faster convergence, with up to 1.73times fewer training iterations, and significant gains in FID and visual quality with less than 0.1% additional parameters, further improving a strong REPA-XL/2 model from 5.9 to 4.34 FID without guidance and reaching 1.39 FID with classifier-free guidance. Our findings suggest that adaptive cross-layer connectivity is a critical yet underexplored factor in diffusion transformers, and that incorporating structured information pathways provides a simple and effective direction for improving scalable generative models.
한국어 요약
한 줄 요약
Diffusion Transformers에서 구조화된 잔차 연결을 도입해 FID 1.39 달성 및 1.73× 빠른 수렴을 실현.
핵심 기여도
- **Mirror routing**이라는 새로운 잔차 연결 메커니즘을 제안, 기존의 단순 합산에서 활성화된 정보 검색으로 전환.
- **REPA-XL/2 모델**에서 FID 5.9 → 4.34 (0.35M steps), classifier-free guidance 시 1.39 FID 달성.
- **1.73× 적은 학습 반복 수**로 동일 수준 성능 달성.
- **0.1% 미만의 추가 파라미터**로 성능 향상.
핵심 아이디어
기존 Diffusion Transformers(DiT)는 단일 잔차 스트림을 사용해 모든 이전 레이어의 정보를 누적하는 방식을 채택했으나, 이는 레이어별 정보의 희석과 경사도 흐름의 제한을 초래한다. 본 연구는 DiT 내부 표현 분석을 통해 **초기 레이어 특징 재사용**과 **대칭 레이어 간 의존성**이 존재함을 발견하고, 이를 기반으로 **Mirror routing**이라는 구조화된 잔차 연결 방식을 제안한다. 이는 각 디코더 블록이 **단일 대칭 인코더 레이어**에 대해 **토큰 기반 가중치**로 정보를 동적으로 검색하는 방식이다. 이는 기존의 정적 스킵 커넥션과 밀집된 레이어 라우팅과 달리, **입력과 시간에 따라 유연하게 연결**되며, **미분 가능한 cross-depth 경로**를 통해 공간적·세멘틱 정보를 복구한다.
기술적 접근법
- **Mirror routing**: 각 디코더 레이어가 대칭 인코더 레이어의 정보를 **softmax 기반 가중치**로 선택적으로 참조.
- **Structured connectivity design**: 로컬 잔차 연결과 장거리 경로를 결합, **단일 레이어 대신 대칭 쌍**으로 연결.
- **REPA-XL/2 모델** 기반 실험, **0.35M 학습 스텝**에서 FID 4.34 달성.
- **Attention Residuals**와 비교해 **낮은 활성화 비용**으로 동일 성능.
- **DiT-S/2** 모델을 사용한 **ablation study** 수행.
주요 결과
- **REPA-XL/2 모델**에서 FID 5.9 → 4.34 (0.35M steps, +0.56 FID 개선), classifier-free guidance 시 **1.39 FID** 달성.
- **1.73× 적은 학습 반복 수**로 동일 수준 성능.
- **0.1% 미만의 추가 파라미터**로 성능 향상.
- **ImageNet 데이터셋**에서 DiT-S/2 기반 실험으로 FID 개선과 빠른 수렴 확인.
의의 및 한계
본 연구는 DiT에서 잔차 연결의 중요성을 재조명하고, **구조화된 정보 흐름**이 성능 향상에 기여함을 입증한다. Mirror routing은 기존의 단순 잔차 합산 방식을 **동적 정보 검색 메커니즘**으로 전환함으로써, **낮은 추가 비용**으로 높은 성능을 달성하는 효과적인 접근법을 제시한다. 그러나 Mirror routing은 **단일 대칭 레이어만 참조**하기 때문에, 더 복잡한 다중 경로 연결이 필요한 상황에서는 한계가 있을 수 있다. 또한, **모든 레이어가 대칭 구조를 가진 모델**에만 적용 가능하며, 비대칭 아키텍처에는 추가 연구가 필요하다.
실용적 활용
본 연구는 이미지 생성 모델에서 **빠른 수렴**과 **높은 생성 품질**을 요구하는 산업 분야(예: 콘텐츠 생성, 의료 영상, 자율주행)에 적용 가능하다. 특히, **파라미터 증가 없이 성능 향상**이 가능한 Mirror routing은 대규모 모델의 **비용 효율적 최적화**에 기여할 수 있다. 또한, **LLM과의 연계 가능성**도 제시하며, 다모달러 모델 설계에 활용될 수 있다.