- #1Measuring AI Ability to Complete Long Software Tasks
AI가 소프트웨어 작업을 완료하는 시간을 인간 기준으로 측정하는 새로운 지표를 제안하고, 2019년 이후 AI의 시간 범위가 7개월마다 2배씩 증가하고 있음을 밝힘.
- #2WARM: On the Benefits of Weight Averaged Reward Models
WARM은 가중치 평균을 통해 생성된 보상 모델로, 보상 해킹을 완화하고 정책 학습의 정확도를 79.4%까지 향상시킨다.
- #4Training Object Permanence in World Models
WROP 데이터셋을 활용한 16B 규모 PWM-WROP 모델이 물리적 추론 능력 향상에 기여함.
- #5DeltaWAM: Delta World Action Models for Bimanual Manipulation
DeltaWAM은 시각 델타와 동작을 결합하여 이전 WAM 모델의 계산 효율성을 개선한 이중 조작 모델이다.
- #6TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation
TokenFlow는 다중 모달 이해와 생성을 통합하는 이중 코드북 구조의 이미지 토크나이저로, 기존 VQ 기반 접근법의 한계를 극복한다.
- #7AlphaMath Almost Zero: process Supervision without process
AlphaMath는 MCTS와 value model을 결합해 수학적 추론 능력을 자율적으로 향상시키는 LLM 훈련 프레임워크다.
- #8Score identity Distillation: Exponentially Fast Distillation of Pretrained Diffusion Models for One-Step Generation
SiD는 데이터 없이 사전 학습된 디퓨전 모델을 단계 1 생성자로 빠르게 증류하는 새로운 방법이다.
- #9VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks
VLM2Vec은 대규모 다모달 임베딩 벤치마크 MMEB와 함께, 임베딩 성능을 10~20% 개선하는 새로운 대비 학습 프레임워크를 제시한다.
- #10Scalable Best-of-N Selection for Large Language Models via Self-Certainty
Self-certainty는 LLM의 내재적 확률 분포를 활용해 외부 보상 모델 없이도 Best-of-N 선택을 효율적으로 수행하는 새로운 메트릭이다.
- #11SEA-RAFT: Simple, Efficient, Accurate RAFT for Optical Flow
SEA-RAFT는 기존 RAFT에 비해 정확도와 효율성을 동시에 향상시킨 광학 흐름 추정 모델로, Spring 벤치마크에서 EPE 3.69, 1px 0.36 성능을 달성하며 2.3× 빠른 처리 속도를 보인다.
- #12DigiRL: Training In-The-Wild Device-Control Agents with Autonomous Reinforcement Learning
DigiRL은 Android-in-the-Wild 데이터셋에서 67.2% 성공률을 달성하며 기존 최고 성능 모델을 28.7% 개선한 자율 강화학습 기반 디지털 에이전트 학습 프레임워크이다.
- #13Token-Budget-Aware LLM Reasoning
TALE은 CoT 추론 과정에서 토큰 비용을 67%까지 줄이며 정확도를 3% 미만으로 유지하는 토큰 예산 인식 추론 프레임워크이다.
- #18Agent-Editing World Model: Rethinking World Modeling for LLM Agents
AEWM은 LLM 에이전트의 장기적 태스크 수행을 위해 환경 예측 대신 결정 효과와 에이전트 상태 편집을 모델링하는 새로운 월드 모델을 제안한다.
- #305Neural Spectral Capacity: Measuring and Designing Architectures from Network Specification Alone
Neural Spectral Capacity (NSC)는 Transformer 아키텍처의 스펙트럼 구조를 사전에 계산해 최적 설계를 CPU에서 2초 내 수행하는 NSC-DP 알고리즘을 제시한다.
- #307World Action Agent: Harnessing VLMs for Robot Manipulation via World Action Rehearsal
WAA는 VLM을 직접 조종하여 로봇 조작을 수행하는 시각 행동 작업공간을 통해 75.6% 성공률을 달성한 다중 에이전트 허네스를 제시한다.
- #309Learning to Discover Interesting Mathematics
- #311Rufus-Air: An Open LLM Post-Training Recipe
Rufus-Air는 GLM-4.5-Air-Base(106B-A12B)를 기반으로 8단계의 포스트-트레이닝 파이프라인을 통해 개선된 오픈 소스 LLM 레시피이다.
- #314Your Transformer Can Hold Two Thoughts at Once: Evidence of Linear Superposition in LLMs
Transformer는 선형 결합된 입력을 중첩된 출력으로 생성하며, 이를 통해 단일 추론 패스로 두 개의 일관된 텍스트를 생성할 수 있음.
- #315Qwen-Planner-Agent: A Closed-Loop AI-for-AI Framework for Real-World Mobile Planner Agents
Qwen-Planner-Agent는 모바일 환경에서 실행되는 AI 에이전트 개발을 위한 AI-for-AI 프레임워크를 통해 개발된, 모바일PA-Bench에서 최고 성능을 달성한 통합 모델-하버스 시스템이다.
- #320WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation
WanPE는 397B 파라미터를 가진 텍스트-비디오 생성을 위한 프롬프트 향상 모델로, 1.05M 개의 실제 비디오를 기반으로 영화 감독 수준의 촬영 계획을 학습한다.
- #321ViRDM: Taming Representation Distribution Matching for Few-Step Causal Video Generation
- #322RewardVerse: Rubric-Guided Policy Optimization for Video Reward Modeling
RewardVerse는 동적 평가 기준을 사용해 동영상 생성 모델의 보상 불안정성을 해결하는 Rubric-Guided Policy Optimization(RGPO) 알고리즘을 제안한다.
- #323EmbodiedSWE: Coding Agents for Long Horizon Dexterous Robotics
- #324OmniEcho: Spatial Audio Understanding for Embodied Agents
OmniEcho는 공간 음향 정보를 통합한 다중 모달 모델로, 실제 환경에서의 음향-시각 인식 및 탐색 성능을 향상시킨다.
- #325PACT: From Credit Assignment to Critic Alignment
- #327Capable yet Parsimonious: Extracting and Characterizing Hidden Chain-of-Thought in Frontier Models
- #328GeoPair: Geometry-Preserving Cross-Layer Factorization for Training-Free Transformer Compression
- #329Verifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms
- #330Parts-of-Speech as Emergent Categories in SAE Latent Space
- #331IterSynth: Rethinking Deep Search Agents via Role-Decoupled Iterative Synthesis