HotPaper.ai — 매일 자동 큐레이션되는 AI 논문 Top 25 + 한국어 요약

2026-09-26 featured 논문 30편

  1. #1Training Object Permanence in World Models

  2. #2ManiGaussian: Dynamic Gaussian Splatting for Multi-task Robotic Manipulation

    ManiGaussian은 다중 작업 로봇 조작에서 장면 수준 시공간 역학을 학습하는 동적 가우시안 스플래팅 프레임워크를 제안한다.

    Guanxing Lu, Shiyi Zhang, Ziwei Wang, Changliu Liu, Jiwen Lu

  3. #3InternW0: A Foundational Physical World Model for Efficient Real-World Interactions

  4. #4DeltaWAM: Delta World Action Models for Bimanual Manipulation

  5. #5Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge

    Meta-Rewarding 기법을 통해 LLM이 스스로 판단 능력을 향상시키고, 지시사항 수행 능력도 개선하는 자가진화 메커니즘을 제시한다.

    Tianhao Wu, Weizhe Yuan, Olga Golovneva, Jing Xu, Yuandong Tian

  6. #6The Revolution of Multimodal Large Language Models: A Survey

    다중모달 대형 언어 모델(MLLMs)의 최근 발전과 주요 연구 방향을 종합적으로 조사한 리뷰 논문.

    Davide Caffagni, Federico Cocchi, Luca Barsellotti, Nicholas Moratelli, Sara Sarto

  7. #7MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI

    MMT-Bench는 31,325개의 다중 선택형 시각 질문을 포함한 종합적인 다모달 벤치마크로, LVLM의 다태스크 AGI 평가를 위한 새로운 기준을 제시한다.

    Kaining Ying, Fanqing Meng, Jin Wang, Zhiqiang Li, Han Lin

  8. #8Improving LoRA in Privacy-preserving Federated Learning

    FFA-LoRA는 LoRA의 통신 비용을 절반으로 줄이고, 프라이버시 보장 연방 학습 환경에서 더 안정적인 성능을 제공한다.

    Youbang Sun, Zitao Li, Yaliang Li, Bolin Ding

  9. #9Towards Measuring and Modeling “Culture” in LLMs: A Survey

    90개 이상의 논문을 바탕으로 LLM에서 문화 표현과 포용성을 조사하고, 문화의 정의 부재와 탐색 방법의 한계를 지적한다.

    Farid Adilazuarda, Sagnik Mukherjee, Pradhyumna Lavania, Siddhant Singh, Ashutosh Dwivedi

  10. #10ProSA: Assessing and Understanding the Prompt Sensitivity of LLMs

    ProSA는 LLM의 프롬프트 민감도를 평가하고 이해하기 위한 프레임워크로, 새로운 지표 PromptSensiScore(PSS)를 제안하고 해석력을 제공한다.

    Jingming Zhuo, Songyang Zhang, Xinyu Fang, Haodong Duan, Dahua Lin

  11. #11InstructIR: High-Quality Image Restoration Following Human Instructions

    InstructIR은 인간의 자연어 지시를 기반으로 다양한 이미지 복원 작업에서 최고 성능을 달성한 다중 작업 모델이다.

    Marcos V. Conde, Gregor Geigle, R. Timofte

  12. #12Flowedit: Inversion-Free Text-Based Editing Using Pre-Trained Flow Models

    FlowEdit는 기존 이미지 편집에서 사용되는 inversion 과정 없이, 텍스트 기반으로 편집 가능한 모델 구조를 제안한다.

    Vladimir Kulikov, Matan Kleiner, Inbar Huberman-Spiegelglas, T. Michaeli

  13. #13REPA-E: Unlocking VAE for End-to-End Tuning with Latent Diffusion Transformers

    REPA-E는 VAE와 LDM을 함께 엔드투엔드로 훈련하여 생성 성능을 향상시키며, ImageNet 256×256에서 FID 1.26을 달성한다.

    Xingjian Leng, Jaskirat Singh, Yunzhong Hou, Zhenchang Xing, Saining Xie

  14. #244Agent-Editing World Model: Rethinking World Modeling for LLM Agents

  15. #305HappyWorld-Bench

  16. #306Neural Spectral Capacity: Measuring and Designing Architectures from Network Specification Alone

  17. #307World Action Agent: Harnessing VLMs for Robot Manipulation via World Action Rehearsal

  18. #308Uranus: Building the Next-Generation Simulation Infrastructure for Embodied AI

  19. #309SpeakerMem-R1: Speaker-Centered Dual-Track Memory for Multi-Party Dialogue

  20. #310PackLab: A Comprehensive Framework for Developing, Training, and Evaluating MLLMs in Robotic Bin Packing

  21. #311MemBodied: Recurrent Associative Memory for Vision-Language-Action Models

  22. #312Qwen-Planner-Agent: A Closed-Loop AI-for-AI Framework for Real-World Mobile Planner Agents

  23. #313Hunyuan-A13B Technical Report

  24. #314Your Transformer Can Hold Two Thoughts at Once: Evidence of Linear Superposition in LLMs

  25. #315Spatial-Interactor: Learning Spatial Reasoning through Interaction with the Observable Physical World

  26. #316The Past Frames the Future: Memory for Autoregressive Video Generation

  27. #317Just-in-Time Memory: Learning to Curate Task-Adaptive Memory for LLM Agents

    JitMem은 과제에 맞춤화된 메모리 추출을 위해 메모리 정제를 읽기 시점으로 미루는 방식으로, ALFWorld, WebShop, τ²-bench에서 기존 메모리 기반 에이전트 대비 16.2~16.3% 성공률 개선을 달성한다.

  28. #318Rufus-Air: An Open LLM Post-Training Recipe

  29. #319WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation

  30. #320RewardVerse: Rubric-Guided Policy Optimization for Video Reward Modeling

인기 키워드

reinforcement-learning (377) diffusion-models (216) llm (214) llm-agents (174) video-generation (153) llm-evaluation (149) transformer (143) vision-language (138) large-language-models (119) long-context (113) vlm (110) code-generation (102) long-horizon (98) benchmark-evaluation (97) benchmarking (92) language-models (90) foundation-models (80) retrieval-augmented (80) world-models (76) flow-matching (73)