- #1Are VLMs Ready for Autonomous Driving? An Empirical Study from the Reliability, Data, and Metric Perspectives
12개 VLM을 17개 설정에서 평가한 DriveBench를 통해 자율주행에서의 VLM 신뢰도 문제를 실증적으로 분석하고, Robust Agentic Utilization(RAU)을 제안한다.
- #2Stop Regressing: Training Value Functions via Classification for Scalable Deep RL
가치 함수 학습에서 회귀 대신 분류 손실을 사용하면 다양한 도메인에서 성능과 확장성이 크게 향상된다.
- #3Adaptive Attacks Break Defenses Against Indirect Prompt Injection Attacks on LLM Agents
LLM 에이전트의 간접 프롬프트 주입 공격 방어책은 적응형 공격에 50% 이상의 성공률로 무너진다.
- #4Benchmarking and Enhancing Skill-Level Memory for Partially Observable Robotic Manipulation
- #5Reliable Conflictive Multi-View Learning
ECML은 신뢰도를 함께 제공하는 방식으로, 충돌이 있는 다중 뷰 데이터를 처리하는 새로운 학습 문제 RCML을 해결한다.
- #6InterEvolve: Test-Time Evolution of Reward Programs for Humanoid Loco-Manipulation
- #7Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond
Light-R1은 공개 데이터와 모델을 활용해 재현 가능한 방식으로 장기 추론 모델을 학습하는 오픈소스 툴킷으로, 14B 모델에서 AIME24 점수 74.0을 달성했다.
- #8VoiceCraft: Zero-Shot Speech Editing and Text-to-Speech in the Wild
VoiceCraft는 실생활 음성 편집 및 제로샷 TTS에서 최고 성능을 보이는 신경 코드 언어 모델이다.
- #9Stealing Part of a Production Language Model
생산 환경 언어 모델에서 임베딩 프로젝션 계층을 추출하는 최초의 모델 도난 공격이 제시되었다.
- #10Diffusion-Based Planning for Autonomous Driving with Flexible Guidance
Diffusion Planner는 rule-based 정제 없이 다중 모달 운전 행동을 모델링하고, nuPlan 및 200시간 배송 차량 데이터셋에서 최고 성능을 달성한 transformer 기반 확산 모델 기반 경로 계획 알고리즘이다.
- #11Gaussian Shading: Provable Performance-Lossless Image Watermarking for Diffusion Models
Gaussian Shading은 디퓨전 모델에 성능 저하 없이 훈련 없이 워터마킹을 적용하는 기법으로, 256비트 용량의 강력한 워터마킹을 달성한다.
- #12Political Compass or Spinning Arrow? Towards More Meaningful Evaluations for Values and Opinions in Large Language Models
LLM의 가치관 평가에서 강제 선택형 평가가 현실과 동떨어진 결과를 유발한다는 점을 밝히고, 보다 현실적인 평가 방법을 제안한다.
- #296World Observer: Joint Actor-Observer Generation for Persistent World Modeling
- #304OneStreamer: Unifying Perception, Memory, and Proactive Response in Streaming Video Interaction
- #305RLE-Bench: A Qualifying Exam for Coding Agents as Robot Learning Engineers
- #306Robust Online Aero-Engine Blade Defect Detection via Dual-Alignment Test-Time Adaptation
- #307A Missing Piece for Trustworthy AI Reviewers: From Benchmarking Rhetorical Robustness to SciCore Review
- #308E-MoE: Enhanced Mixture-of-Experts for Non-Factorized Diffusion Language Models
- #309Make Sparse Rewards Count: Density-Aware Reward Aggregation for Multi-Reward RL
- #310EgoTools: Towards Tool-Centric Reasoning in Real-World Egocentric Videos
- #311Decentralized Master-Mind: Joint Action Refinement through Iterative Intent Denoising in Multi-Agent Pathfinding
- #312LOCI: Spatial Linear Memory for Streaming World Models
- #313Devils in Question Relay: Source-Conditioned Relay Steering to Mitigate Hallucinations in Audio-visual Large Language Models
- #314Prompt2Skill: Unsupervised Skill Optimization From Natural Language Instructions
- #315X-Tree: Tokenizing Reusable Experience for Efficient Agent Generalization
- #316OTRetarget: Joint Robot and Object Motion Retargeting via Optimal Transport
- #317Fewer Tokens, Better Action: GPT-6 Astra Robot Agents with 14% Higher Success Rate but 65% Fewer Tokens
- #318On-Policy or Off-Policy Learning? A Systematic Study of Distillation Dynamics
- #319Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows
- #320CorrGRPO: Correlation-Normalized GRPO for Multi-Reward Learning