- #1Memento 3: Model-Based Recursive Self-Improvement through Reflective Rulebooks
- #2Nemotron-CC: Transforming Common Crawl into a Refined Long-Horizon Pretraining Dataset
Nemotron-CC는 Common Crawl 데이터를 활용해 정확도와 데이터 양의 균형을 맞춘 6.3T 토큰의 훈련 데이터셋을 제시한다.
- #3Devils in Middle Layers of Large Vision-Language Models: Interpreting, Detecting and Mitigating Object Hallucinations via Attention Lens
LVLM에서 시각 정보 처리 단계를 분석하여 오브젝트 환영을 감지하고 완화하는 새로운 방법을 제안한다.
- #4Multi-Agent Egocentric World Model with Fine-Grained Embodied Interaction
- #5DreamTrue: Action-Faithful Robot World Model with Counterfactual Post-Training
- #6U-Space: Uncovering When and Why Uncertainty Arises in Language Models
- #7LiveBench: A Challenging, Contamination-Limited LLM Benchmark
LiveBench는 테스트 세트 오염과 LLM 평가의 편향을 방지하기 위해 설계된, 월별 업데이트와 객관적 평가를 기반으로 한 LLM 벤치마크이다.
- #8Investigating Cultural Alignment of Large Language Models
LLM의 문화적 정합성을 평가하고 언어와 사전학습 데이터의 영향을 분석하여, 문화적 다변성을 반영한 새로운 프롬프팅 방법을 제안한다.
- #9Teaching Large Language Models to Regress Accurate Image Quality Scores Using Score Distribution
DeQA-Score는 MLLM 기반의 정확한 이미지 품질 점수 회귀를 위해 분포 기반 소프트 라벨과 퓨이델리티 손실을 도입한 모델이다.
- #10The Unreasonable Effectiveness of Entropy Minimization in LLM Reasoning
엔트로피 최소화(EM)를 통해 라벨 없이도 LLM의 수학, 물리, 코딩 성능을 크게 향상시킬 수 있음을 보인 연구.
- #11AdaIR: Adaptive All-in-One Image Restoration via Frequency Mining and Modulation
AdaIR은 주파수 분해와 조절을 통해 다양한 이미지 복원 작업에서 최신 성능을 달성한 적응형 통합 복원 모델이다.
- #12MotionLCM: Real-time Controllable Motion Generation via Latent Consistency Model
MotionLCM은 텍스트와 제어 신호를 기반으로 실시간 인체 동작을 생성하는 새로운 모델로, 1단계 추론과 ControlNet을 통해 빠른 속도와 정밀 제어를 달성한다.
- #13Robust Prompt Optimization for Defending Language Models Against Jailbreaking Attacks
RPO는 적대적 프롬프트 공격에 대응하기 위해 최적화 기반 방어 알고리즘으로, GPT-4와 Llama-2에서 ASR을 각각 6%와 0%까지 감소시킨다.
- #14DuQuant: Distributing Outliers via Dual Transformation Makes Stronger Quantized LLMs
DuQuant은 회전 및 순열 변환을 활용해 4비트 정량화에서도 LLM 성능을 크게 향상시키는 새로운 정량화 방법이다.
- #305WorldGuide: Goal-Directed Video World Model for Procedural Task Execution
- #306A GPU-Parallel Framework for Heterogeneous Multi-Task Reinforcement Learning
- #307MiMo-V2.6: Scaling Reinforcement Learning Towards Self-Improvement
- #308Learn2Play Bench: How Well Do LLM Agents Learn from Experience in Unfamiliar Environments?
- #309OuroWorld: Bringing Any 3D World Alive as Diverse, Endlessly Looping 3D Cinemagraphs
- #310SPW-Nav: A Streaming Panoramic World Model for Language-Guided Navigation
- #311Foundations of Large Language Models
BERT 모델의 기본 아키텍처와 하이퍼파라미터 설정을 설명하며, LLM 개발의 기초를 다룬다.
- #312TestPrism: Rethinking Test Evaluation Beyond a Single Reference
- #313MC-Sparse: Deconstructing and Closing the Dense-Sparse Attention Gap in Diffusion Transformers
MC-Sparse는 토큰 수준의 정밀한 어텐션 선택과 타일 정렬 그룹핑을 결합한 훈련 없이 가능한 확산 트랜스포머 가속 프레임워크로, 1.80×~2.32×의 가속 성능을 보인다.
- #314Post-Training Frontier Text-to-Image Models by Composing Preference and Rubric Rewards
- #315Beyond Spatio-Temporal Priors: A Generalizable Approach for Dense Correspondence Matching
- #316Synthesis Through Simulation: Generating Coherent Enterprise Data via Scalable Agent-System Interaction
- #317LEGO: A Lifting-Free Approach for Exocentric-to-Egocentric Video Generation
- #318OneSearch-VL: Unified Multimodal Deep Research Agent for Image and Video
- #319SparseEngine: Sparse-First Inference Engine
- #320OmniCapBench: A Deep-Structured Evaluation Framework for Fine-Grained Audio-Visual Captioning