- #1HAMSTER: Hierarchical Action Models For Open-World Robot Manipulation
- #2Embodied Agent Interface: Benchmarking LLMs for Embodied Decision Making
Embodied Agent Interface를 통해 LLM의 몸체화 결정 과제 성능을 체계적으로 평가하고, 핵심 모듈별 오류를 분석하여 활용 방향을 제시한다.
- #3Ovis-Embedding: Pushing the Frontiers of Universal Omni-Modal Embeddings
Ovis-Embedding은 텍스트, 이미지, 영상, 오디오를 통합한 공통 임베딩 공간을 생성하는 최신 옴니모달 임베딩 모델이다.
- #4InternW0: A Foundational Physical World Model for Efficient Real-World Interactions
- #5Transformers without Normalization
- #6WonderWorld: Interactive 3D Scene Generation from a Single Image
- #7GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities
- #8Quantifying the Persona Effect in LLM Simulations
- #9In-Context Edit: Enabling Instructional Image Editing with In-Context Generation in Large Scale Diffusion Transformer
- #10Taming Rectified Flow for Inversion and Editing
RF-Solver와 RF-Edit을 제안하여, Rectified Flow 기반 모델의 인버전 정확도와 편집 성능을 훈련 없이 향상시킨다.
- #11Sparse VideoGen: Accelerating Video Diffusion Transformers with Spatial-Temporal Sparsity
- #12Deep Compression Autoencoder for Efficient High-Resolution Diffusion Models
- #41Embedding Physics Priors in Robot Learning: A Survey
로봇 학습에 물리학 사전지식을 통합하는 방법을 체계적으로 분류하고, 주요 응용 분야와 개방적 문제를 조망한 서베이 논문.
- #297Lean Pool: An AI-Maintained Archive of Formalized Mathematics
Lean Pool은 AI 에이전트가 관리하는 형식화 수학 아카이브 시스템이다.
- #305HappyWorld-Bench
- #306JEV-as-a-Judge: Accept When Confident, Escalate When Unsure
JEV-as-a-Judge는 저비용으로 신뢰도 높은 판단을 제공하며, 불확실한 경우에만 강력한 모델로 전달하는 방식으로 99%의 정확도를 유지한다.
- #307The Tasteful Agent: Measuring and Improving Taste in Long-Horizon Tasks
Taste-Bench라는 새로운 벤치마크를 제안하여 LLM 에이전트의 장기적 판단력(‘taste’)을 측정하고 훈련한다.
- #308SpeakerMem-R1: Speaker-Centered Dual-Track Memory for Multi-Party Dialogue
- #309Uranus: Building the Next-Generation Simulation Infrastructure for Embodied AI
- #310PackLab: A Comprehensive Framework for Developing, Training, and Evaluating MLLMs in Robotic Bin Packing
- #311MemBodied: Recurrent Associative Memory for Vision-Language-Action Models
- #312Do Synthetic Personas Predict Real Audience Response? A Sim-to-Real Study Where a No-Persona Baseline Beats Persona-Based Copy Simulation
Upworthy A/B 테스트 데이터를 기반으로, 인공 인물(persona)을 사용한 마케팅 콘텐츠 시뮬레이션은 예측 정확도가 낮아, 단순 LLM 기반 순위 매기기가 더 효과적임을 밝힘.
- #313X-Planner: Event-Structured Task Planning for Embodied Intelligence
- #314Hunyuan-A13B Technical Report
- #315Spatial-Interactor: Learning Spatial Reasoning through Interaction with the Observable Physical World
- #316GAE: Learning a Geometry-Native Latent Space for 3D-Consistent World Generation
GAE는 3D 일관성 있는 세계 생성을 위한 기하학적 인코딩을 갖춘 압축된 잠재 공간을 학습하여 생성 및 인식 간의 공유 인터페이스를 제시한다.
- #317All-in-One Multilingual Scene Text Recognition with Script-aware Mixture-of-Experts
- #318The Past Frames the Future: Memory for Autoregressive Video Generation
- #319Just-in-Time Memory: Learning to Curate Task-Adaptive Memory for LLM Agents
- #320Bellman Policy Optimization