- #1AA-CLIP: Enhancing Zero-Shot Anomaly Detection via Anomaly-Aware CLIP
AA-CLIP은 CLIP의 이상 감지 능력을 향상시킨 두 단계적 접근법을 통해 제로샷 이상 탐지를 개선한 모델이다.
- #2Epona: Autoregressive Diffusion World Model for Autonomous Driving
Epona는 자율주행을 위한 고해상도, 장기 예측을 지원하는 오토회귀 확산 월드 모델로, 7.4% FVD 개선과 2분 이상 예측을 달성한다.
- #3World Action Modeling with Progressive Visual Planning
- #4MotorMind: Scaffolding General Vision Language Models for Zero-Shot Robot Manipulation
MotorMind는 VLM 기반의 Zero-Shot 로봇 조작 시스템으로, LIBERO-PRO에서 66.7% 성공률을 달성하며 xArm6 로봇에서도 95% 성공률을 보인다.
- #5AbstentionBench: Reasoning LLMs Fail on Unanswerable Questions
AbstentionBench는 LLM이 불확실한 질문에 대해 답변을 거절하는 능력을 평가하는 대규모 벤치마크로, 추론 훈련이 거절 능력을 오히려 악화시킨다는 점을 밝혀낸다.
- #6Spatial Memory Intelligence: Endowing World Models with Understanding-Driven Long-Term Memory
SMI는 장기 비디오 월드 모델에서 공간 기억 관리를 위한 첫 번째 통합 프레임워크로, 4가지 원자적 연산을 통해 기억 효율성과 생성 안정성을 향상시킨다.
- #7DivPrune: Diversity-based Visual Token Pruning for Large Multimodal Models
DivPrune는 시각 토큰의 다양성 최대화를 통해 LMM의 추론 속도와 메모리 사용량을 줄이는 토큰 프루닝 방법이다.
- #8SimLingo: Vision-Only Closed-Loop Autonomous Driving with Language-Action Alignment
SimLingo는 시각 정보만으로 자율주행과 언어-행동 정렬을 동시에 수행하는 첫 번째 모델로, CARLA 벤치마크에서 최고 성능을 달성했다.
- #9Difix3D+: Improving 3D Reconstructions with Single-Step Diffusion Models
Difix3D+는 단계적 3D 업데이트와 실시간 디퓨전 모델을 활용해 3D 재구성 품질을 2× 향상시키는 새로운 파이프라인이다.
- #10d1: Scaling Reasoning in Diffusion Large Language Models via Reinforcement Learning
d1은 마스킹된 확산 대형 언어 모델(dLLMs)에 강화학습(RL)을 통합하여 추론 성능을 향상시키는 2단계 프레임워크로, diffu-GRPO라는 새로운 정책 경사 알고리즘을 제안한다.
- #11The Impact of Reasoning Step Length on Large Language Models
Chain of Thought (CoT)의 추론 단계 길이가 LLM의 추론 성능에 큰 영향을 미친다는 것을 실험적으로 밝혔다.
- #12The Road Less Scheduled
스케줄 없이 학습률을 조정하는 Schedule-Free 알고리즘이 기존 스케줄 기반 방법을 능가한다.
- #13CHASE-SQL: Multi-Path Reasoning and Preference Optimized Candidate Selection in Text-to-SQL
CHASE-SQL은 Text-to-SQL 작업에서 질 높고 다양한 SQL 후보를 생성하고 정확하게 선택하는 새로운 에이전트 기반 프레임워크로, BIRD 데이터셋에서 73.01%의 실행 정확도를 달성했다.
- #273HelixWorld: A Real-time Interactive Audio-Visual World Model
- #274EditHero: A Benchmark for Long-Horizon Part-Level 3D Editing and Vibe Modeling
- #306From Retrieval to Typed Decisions: Calibrated System One Models from Biomedical Sentence Encoders
- #307Does Learning Protein Folding Generalize to Broader Reasoning?
- #308FrameMorrow: Future-guided Frame Selection with Prospective Tokens for Long-Horizon Video Generation
- #309Dream4ACT: A Shared Visual Action Interface for Multi-Embodiment Video-Action Modeling
Dream4ACT는 다양한 로봇 구조에 공통적으로 적용 가능한 시각적 행동 인터페이스를 도입하여 88.98% 성공률을 달성한 비디오-행동 월드 모델이다.
- #310Native Action-Prior Learning from Videos for World Action Models
- #311RealCompanion: Benchmarking Human Understanding from Reasoning over Longitudinal Real-World Conversations
- #312On-Policy Parameter Update Direction Underlies Generalization in LLM Post-Training
- #313WEFT: Scaling Tool-Use Post-Training for General-Purpose Agents
- #314PDE-JEPA: Predictive Representation Learning of Latent Dynamics Modeling for Parametric PDEs
- #315Pivot-SD: Efficient Self-Distillation for Masked Diffusion Language Models
Pivot-SD는 마스킹된 확산 언어 모델(dLM)의 핵심 토큰만 학습해 효율적인 자기-디스틸레이션을 구현한다.
- #316SimuVerity: Benchmarking Agents for Engineering-Grade Simulink Model Generation
SimuVerity는 10개 공학 분야에서 101개의 텍스트-실행가능 Simulink 모델 생성 작업을 평가하는 벤치마크로, 최고 시스템의 종합 점수는 42.86에 불과하다.
- #317HyperBrowseComp: A Multilingual and Multimodal Stress Test for Web-Browsing Agents
HyperBrowseComp는 13개 언어와 8개 모달성을 포함한 다국어·다모달 웹 탐색 벤치마크로, 423개의 수작업 질문을 통해 웹 에이전트의 정보 탐색 능력을 테스트한다.
- #318Science or Slop?: Benchmarking and Mitigating Scientific Slop in AI-Generated Papers
- #319LexReward: A Taxonomy-Driven Reward Framework for Legal Language Models
- #320Source Preference in the Wild: How LLM Agents Favor Items by Source, and How to Reduce It