- #1PhysBench: Benchmarking and Enhancing Vision-Language Models for Physical World Understanding
PhysBench 벤치마크와 PhysAgent 프레임워크를 제안하여 VLM의 물리 세계 이해 능력을 평가 및 향상시킨다.
- #2Hierarchical Diffusion Policy for Kinematics-Aware Multi-Task Robotic Manipulation
계층적 확산 정책 HDP는 로봇 운동학 제약을 고려한 다중 작업 조작 성능을 향상시킨다.
- #3WorldCrafter: Consistent Video World Model with Implicit 3D-aware Memory
- #4Think Like a World Model, Act Like a VLA: Distilling World-Model Representations into Compact Robot Policies
- #5RRSI: Regularized Recursive Self-Improvement of Agent Harnesses
- #6Train for the Worst, Plan for the Best: Understanding Token Ordering in Masked Diffusions
마스킹 확산 모델(MDM)의 토큰 순서 결정 전략이 추론 성능에 미치는 영향을 분석하고, 적응적 추론이 성능을 7%에서 90%까지 향상시킴을 보인다.
- #7A General Framework for Inference-time Scaling and Steering of Diffusion Models
Feynman-Kac (FK) steering을 통해 추론 시 확산 모델을 보상 함수로 제어하고, 훈련 없이도 샘플 품질과 제어력을 향상시킨다.
- #8WildTeaming at Scale: From In-the-Wild Jailbreaks to (Adversarially) Safer Language Models
WildTeaming은 실제 사용자 대화 로그를 분석해 5.7K개의 새로운 제이블레이크 전략을 발견하고, 이를 기반으로 262K개의 대규모 오픈소스 안전 훈련 데이터셋 WildJailbreak를 생성한 LLM 안전성 테스트 프레임워크이다.
- #9Plan-and-Act: Improving Planning of Agents for Long-Horizon Tasks
Plan-and-Act는 LLM 기반 에이전트의 장기적 계획 능력을 향상시키기 위해 Planner와 Executor 모듈을 분리하고, 합성 데이터 생성을 통해 성능을 개선한 프레임워크이다.
- #10Repeat After Me: Transformers are Better than State Space Models at Copying
트랜스포머는 GSSM 대비 컨텍스트 복사 능력에서 근본적 우위를 보인다.
- #11InfLoRA: Interference-Free Low-Rank Adaptation for Continual Learning
InfLoRA는 지속 학습에서 새로운 작업의 간섭을 제거하여 안정성과 유연성의 균형을 맞춘 저랭크 파인튜닝 방법이다.
- #12Neural Video Compression with Feature Modulation
DCVC-FM은 학습 가능한 양자화 스케일러와 주기적 템포럴 피처 리프레시 기법을 통해 NVC의 실용성을 크게 향상시킨다.
- #13XFeat: Accelerated Features for Lightweight Image Matching
XFeat는 저비용 CPU에서도 실시간으로 작동하는, 가속화된 로컬 특징 추출 및 매칭 CNN 모델로, 최대 5배 빠르면서 정확도를 유지한다.
- #304From Pretraining to Proficiency: Real-World Subtask RL for Long-Horizon Manipulation with Minimal Human Intervention
PARTS는 실세계에서 최소한의 인간 개입으로 장기적 조작 작업의 성능을 향상시키는 하위 작업 기반 강화 학습 프레임워크이다.
- #305GAVEL: Graph World Models for Verified and Efficient Long-Horizon LLM Task Planning
GAVEL은 그래프 월드 모델을 기반으로 LLM 기반 장기 로봇 계획의 신뢰성과 효율성을 향상시킨다.
- #306GameHorizon Suite: Multi-Horizon Data and Evaluation in Gameplay
- #307Grounded Action Model: 3D Grounding as a Foundation for Robotics
- #308HuRo: Robotizing Human Videos for Scalable VLA Pretraining
- #309Transferring the Intelligence of VLMs to Robotic Control
- #310IntBMoE: Integrating Block-Level Conditioning into Expert Composition for Full-Participation Mixture-of-Experts
IntBMoE는 참여, 실행, 메모리화를 독립적으로 조절하는 블록 기반 MoE로, AMap 추천 시스템에서 60ms 내 2.4% UVCTR 향상.
- #311OmniEdu: Open Foundation Models for Learning and Teaching
- #312CARE: Experience-Guided Atomic Corrective Execution for Vision-Language-Action Policies
- #313Helix-FNO: Spectral-Domain Operator Learning Coupled with a High-Fidelity Mechanistic Model for Fast Surrogate Simulation
Helix-FNO는 고정밀 기계 모델과 결합한 스펙트럼 도메인 오퍼레이터 학습을 통해 1000배 빠른 대체 시뮬레이션을 실현한 신경 오퍼레이터이다.
- #314Document Retrieval-Aware Chunking (D-RAC): Universal Retrieval-Aware Ingestion of Enterprise Documents via PDF Normalization and Multimodal Markdown Conversion
- #3151% of Tokens Can Be Enough: On Gradient Estimation in On-Policy Distillation
- #316CADWorld: Computer-Use Benchmark for Long-Horizon Computer-Aided Design
CADWorld는 FreeCAD 기반의 200개 엔지니어링 작업을 통해 컴퓨터 사용 에이전트의 장기적 CAD 작업 능력을 평가하는 벤치마크이다.
- #317VideoGen-Agent: Reinforcing Video Generation Agents
- #318OmniVChat: Synthesizing, Benchmarking, and Training for Native Audio-Visual Dialogue
OmniVChat은 사용자의 음성 및 영상 입력을 직접 처리해 대화를 생성하는 멀티모달 대화 시스템이다.
- #319Designer-RSI: Evolving Procedural Memory from User Traffic for Agentic Graphic Design
- #320onPanda: Efficient Annotation of On-Policy Alignment Data for LLMs and Agents via Token-Level Correction