In this work, we show that a single Transformer block, applied recurrently, can match the accuracy of a full-depth vision encoder at comparable inference FLOPs without intermediate feature distillation. reViT restores depth-specific transformations by representing the FFN at each recurrent depth as a convex combination of a small shared expert bank. A continuous normalized-depth coordinate programs this mixture, defining a resampleable trajectory through FFN parameter space. We evaluate this design in two regimes: supervised ImageNet-1k training and distillation from a DINOv2 teacher. Across both regimes, controlled adaptations identify weight-space merging as the strongest tested MoE family at a matching one-FFN budget, ahead of the token-dispatch and output-mixture alternatives. Trained from scratch, reViT-B/16 attains DeiT III accuracy with about 70\% fewer stored parameters. An 8-experts model distilled using only the teacher's output features retains nearly all of its DINOv2 teacher's linear-probe accuracy and transfers across classification, segmentation, and depth prediction. Elastic-depth training allows one checkpoint (trained model) to operate at multiple tested depths by resampling the same normalized coordinate interval. For fixed-depth deployment, the recurrent block can be materialized as a conventional dense graph, removing online routing and merging without changing the one-FFN-per-depth compute but expanding deployment storage.
한 줄 요약
reViT은 단일 Transformer 블록을 반복 적용하여 ViT의 전체 깊이 성능을 유지하면서 파라미터 수를 73% 줄인다.
핵심 기여도
- reViT은 ViT의 전체 깊이 스택을 단일 반복 Transformer 블록으로 대체함.
- FFN은 각 반복 깊이에서 공유된 8개의 전문가 은행의 볼록 결합으로 표현됨.
- reViT-B/16은 DeiT III와 유사한 정확도를 73% 적은 파라미터로 달성함.
- DINOv2 선생 모델로부터의 지도 학습에서 8개 전문가 모델이 선생의 선형 프로브 정확도를 거의 유지함.
핵심 아이디어
기존 ViT는 각 블록에 독립적인 파라미터를 사용하지만, reViT은 단일 Transformer 블록을 반복적으로 사용함으로써 파라미터 수를 줄인다. 이 모델은 각 반복 깊이에서 FFN을 공유된 전문가 은행의 볼록 결합으로 표현하며, 정규화된 깊이 좌표 $ s_t = t/(L-1) $를 조건으로 사용한다. 이는 FFN 파라미터 공간에서 연속적인 경로를 정의하며, 기존 ViT의 독립적인 FFN을 대체한다. reViT은 동일한 계산 예산(1개 FFN당)에서 토큰 디스패치 및 출력 혼합 방식보다 더 우수한 성능을 보인다.
기술적 접근법
- reViT은 ViT의 $ L $개 독립 블록을 단일 반복 pre-norm 모듈로 대체함.
- FFN은 공유된 $ E $개 전문가 은행의 볼록 결합으로 표현됨.
- 정규화된 깊이 좌표 $ s_t = t/(L-1) $를 조건으로 사용하여 FFN 혼합 계수를 결정함.
- 학습된 파라미터 경로는 FFN 공간을 연속적으로 탐색함.
- $ E = 4 $인 reViT-B/16은 DeiT III와 유사한 정확도를 73% 적은 파라미터로 달성함.
- $ E = 8 $인 모델은 DINOv2 선생 모델의 선형 프로브 정확도를 거의 유지함.
주요 결과
- reViT-B/16은 DeiT III와 유사한 ImageNet-1k 정확도를 73% 적은 파라미터로 달성함.
- $ E = 8 $인 모델은 DINOv2 선생 모델의 선형 프로브 정확도를 거의 유지함.
- reViT은 단일 체크포인트로 여러 깊이에서 작동 가능하며, 고정 깊이 배포 시 FFN을 고밀도 그래프로 변환 가능함.
의의 및 한계
reViT은 ViT의 파라미터 수를 줄이면서도 정확도를 유지하는 새로운 접근법을 제시한다. 특히, 단일 블록을 반복 적용하면서도 깊이별 계산을 복원할 수 있다는 점에서 혁신적이다. 그러나, 모델이 학습된 경로가 선생 모델과의 정렬에 의존한다는 한계가 있다. 또한, 전문가 은행 크기 $ E $가 증가할수록 성능 향상이 감소하며, 이는 파라미터 효율성과의 균형을 고려해야 한다는 점에서 주의가 필요하다.
실용적 활용
reViT은 이미지 분류, 세그멘테이션, 깊이 예측 등 다양한 비전 작업에 적용 가능하다. 특히, 파라미터 효율성이 높은 모바일 및 임베디드 시스템에서 유용하며, 단일 체크포인트로 다양한 깊이에서 작동하는 탄력적 배포가 가능하다.