- #1Training Object Permanence in World Models
- #2ManiGaussian: Dynamic Gaussian Splatting for Multi-task Robotic Manipulation
ManiGaussian은 다중 작업 로봇 조작에서 장면 수준 시공간 역학을 학습하는 동적 가우시안 스플래팅 프레임워크를 제안한다.
- #3InternW0: A Foundational Physical World Model for Efficient Real-World Interactions
- #4DeltaWAM: Delta World Action Models for Bimanual Manipulation
- #5Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge
Meta-Rewarding 기법을 통해 LLM이 스스로 판단 능력을 향상시키고, 지시사항 수행 능력도 개선하는 자가진화 메커니즘을 제시한다.
- #6The Revolution of Multimodal Large Language Models: A Survey
다중모달 대형 언어 모델(MLLMs)의 최근 발전과 주요 연구 방향을 종합적으로 조사한 리뷰 논문.
- #7MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI
MMT-Bench는 31,325개의 다중 선택형 시각 질문을 포함한 종합적인 다모달 벤치마크로, LVLM의 다태스크 AGI 평가를 위한 새로운 기준을 제시한다.
- #8Improving LoRA in Privacy-preserving Federated Learning
FFA-LoRA는 LoRA의 통신 비용을 절반으로 줄이고, 프라이버시 보장 연방 학습 환경에서 더 안정적인 성능을 제공한다.
- #9Towards Measuring and Modeling “Culture” in LLMs: A Survey
90개 이상의 논문을 바탕으로 LLM에서 문화 표현과 포용성을 조사하고, 문화의 정의 부재와 탐색 방법의 한계를 지적한다.
- #10ProSA: Assessing and Understanding the Prompt Sensitivity of LLMs
ProSA는 LLM의 프롬프트 민감도를 평가하고 이해하기 위한 프레임워크로, 새로운 지표 PromptSensiScore(PSS)를 제안하고 해석력을 제공한다.
- #11InstructIR: High-Quality Image Restoration Following Human Instructions
InstructIR은 인간의 자연어 지시를 기반으로 다양한 이미지 복원 작업에서 최고 성능을 달성한 다중 작업 모델이다.
- #12Flowedit: Inversion-Free Text-Based Editing Using Pre-Trained Flow Models
FlowEdit는 기존 이미지 편집에서 사용되는 inversion 과정 없이, 텍스트 기반으로 편집 가능한 모델 구조를 제안한다.
- #13REPA-E: Unlocking VAE for End-to-End Tuning with Latent Diffusion Transformers
REPA-E는 VAE와 LDM을 함께 엔드투엔드로 훈련하여 생성 성능을 향상시키며, ImageNet 256×256에서 FID 1.26을 달성한다.
- #244Agent-Editing World Model: Rethinking World Modeling for LLM Agents
- #305HappyWorld-Bench
- #306Neural Spectral Capacity: Measuring and Designing Architectures from Network Specification Alone
- #307World Action Agent: Harnessing VLMs for Robot Manipulation via World Action Rehearsal
- #308Uranus: Building the Next-Generation Simulation Infrastructure for Embodied AI
- #309SpeakerMem-R1: Speaker-Centered Dual-Track Memory for Multi-Party Dialogue
- #310PackLab: A Comprehensive Framework for Developing, Training, and Evaluating MLLMs in Robotic Bin Packing
- #311MemBodied: Recurrent Associative Memory for Vision-Language-Action Models
- #312Qwen-Planner-Agent: A Closed-Loop AI-for-AI Framework for Real-World Mobile Planner Agents
- #313Hunyuan-A13B Technical Report
- #314Your Transformer Can Hold Two Thoughts at Once: Evidence of Linear Superposition in LLMs
- #315Spatial-Interactor: Learning Spatial Reasoning through Interaction with the Observable Physical World
- #316The Past Frames the Future: Memory for Autoregressive Video Generation
- #317Just-in-Time Memory: Learning to Curate Task-Adaptive Memory for LLM Agents
JitMem은 과제에 맞춤화된 메모리 추출을 위해 메모리 정제를 읽기 시점으로 미루는 방식으로, ALFWorld, WebShop, τ²-bench에서 기존 메모리 기반 에이전트 대비 16.2~16.3% 성공률 개선을 달성한다.
- #318Rufus-Air: An Open LLM Post-Training Recipe
- #319WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation
- #320RewardVerse: Rubric-Guided Policy Optimization for Video Reward Modeling