- #1To Believe or Not to Believe Your LLM
- #3HVI: A New Color Space for Low-light Image Enhancement
- #4Efficient Deformable ConvNets: Rethinking Dynamic and Sparse Operator for Vision Applications
- #5A Careful Examination of Large Language Model Performance on Grade School Arithmetic
- #6BLEnD: A Benchmark for LLMs on Everyday Knowledge in Diverse Cultures and Languages
- #7MimicMotion: High-Quality Human Motion Video Generation with Confidence-aware Pose Guidance
- #8HPSv3: Towards Wide-Spectrum Human Preference Score
- #9Reasoning with Latent Thoughts: On the Power of Looped Transformers
- #10Disentangling Length from Quality in Direct Preference Optimization
- #11LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images
- #304BI-Agent and BI-Bench: Towards Automating End-to-End Business Intelligence
- #306GAVEL: Graph World Models for Verified and Efficient Long-Horizon LLM Task Planning
- #307DeformSmith: Physics Harness-Guided Hierarchical Generation of Deformable Assets for Robot Manipulation
- #309OmniVBench: A Benchmark and Large-Scale Dataset for Omni Reference-to-Video Generation
- #311IntBMoE: Integrating Block-Level Conditioning into Expert Composition for Full-Participation Mixture-of-Experts
- #312Grounded Skill Synthesis from Code at Scale for Agentic Intelligence
- #313CodeMidas: Scaling Agentic Coding RL Environments from Code Itself
- #314EvoOntology: A Self-Evolving Ontology Layer for Data Agents
- #315From Pretraining to Proficiency: Real-World Subtask RL for Long-Horizon Manipulation with Minimal Human Intervention
- #316When AI Reviews Train AI Reviewers: Scientific-Judgment Collapse and Mitigation
- #317RecreationWorld: Scalable and Verifiable Environments for Hybrid Computer-Use Agents
- #326When2Think: Learning Difficulty-Aware Length Control for Efficient Hybrid Reasoning Models
When2Think는 문제 난이도에 따라 추론 깊이를 조절하여 효율적인 하이브리드 추론을 학습하는 포스트-트레이닝 프레임워크이다.
- #327What Does Privileged Information Add to On-Policy Self-Distillation?
On-policy self-distillation에서 privileged information의 추가 가치를 AMPLE-Math 데이터셋을 통해 분리 평가한 연구.
- #328UFO: Chain-of-Evaluation for Omni-Condition Alignment in Multi-Modal Image Generation
UFO는 다중 조건 일치 평가를 위한 체인형 평가 프레임워크로, 기존 방법 대비 15.25% 높은 인간 평가와의 상관성을 달성했다.
- #329OmniVChat: Synthesizing, Benchmarking, and Training for Native Audio-Visual Dialogue
- #330Paint-Anything: Unified Any-Color Control for Image Generation and Editing
- #331Sample Count Is Not Enough: Candidate-Generation Strategy Shapes the Energy and Performance of LLM Test-Time Scaling
테스트 타임 스케일링에서 후보 생성 방식이 성능과 에너지 소비에 큰 영향을 미친다.
- #332FAMOS: Feed-Forward 3D Articulation Modeling from Sparse Observations
FAMOS는 희소한 단일 뷰에서 3D 조인트 모델링을 수행하는 피드포워드 모델로, Multi-state Articulation Transformer와 observed articulation span loss를 통해 정확도를 향상시킨다.
- #333Srijika: OpenType-Layout-Reusing Font Restyling for Nine Indic Scripts
Srijika는 9개의 브라흐미 글꼴 스크립트에 대한 OpenType 레이아웃 재사용 기반 글꼴 리스타일링 시스템이다.
- #334Learning Foresight without Explicit Trajectories for 3D Diffusion Policies