- #3DoRA: Weight-Decomposed Low-Rank Adaptation
DoRA는 LoRA의 학습 능력을 향상시키며 추론 오버헤드 없이 정확도를 개선하는 가중치 분해 기반 파라미터 효율적 미세조정 방법이다.
- #4SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
SpatialVLM은 20억 개의 3D 공간 추론 VQA 데이터로 학습한 VLM을 통해 공간 추론 능력을 향상시킨다.
- #5OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
OSWorld는 실제 컴퓨터 환경에서 다중 모달 에이전트를 평가하는 첫 번째 벤치마크이다.
- #6Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs
Cambrian-1은 시각 중심의 다모달 LLM을 제시하며, 20개 이상의 시각 인코더를 기반으로 SVA 모듈을 통해 성능을 향상시킨다.
- #7Are We on the Right Way for Evaluating Large Vision-Language Models?
MMStar는 1,500개의 인간 검수된 시각-언어 평가 샘플로 LVLM의 실제 멀티모달 능력을 정확히 평가하는 새로운 벤치마크이다.
- #8Refusal in Language Models Is Mediated by a Single Direction
대규모 언어 모델의 거부 행동이 단일 방향으로 매개됨을 밝혀내고, 이를 기반으로 간단한 가중치 수정으로 안전 메커니즘을 무력화하는 방법을 제시한다.
- #9OpenHands: An Open Platform for AI Software Developers as Generalist Agents
OpenHands는 개발자처럼 행동하는 AI 에이전트를 개발하고 평가할 수 있는 오픈 플랫폼이다.
- #10LGM: Large Multi-View Gaussian Model for High-Resolution 3D Content Creation
LGM은 텍스트나 단일 뷰 이미지로부터 5초 이내에 512 해상도의 고해상도 3D 가우시안 모델을 생성하는 새로운 프레임워크이다.
- #11Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs
최근 멀티모달 LLMs는 CLIP 기반 시각 인코더의 한계로 인해 기본적인 시각 패턴을 인식하지 못하는 문제가 있음을 밝혀냈다.
- #12Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads
Medusa는 병렬 토큰 예측을 통해 LLM 추론 속도를 2.3~3.6× 가속하는 프레임워크이다.
- #315BM25 Wins at Scale: A Scaling Study of Retrieval-Augmented Generation Paradigms
BM25는 1000만 토큰 이상의 대규모 문서에서 RAG 성능에서 File-System Agent를 20점 이상 앞선다.
- #316Flux-OPD: On-Policy Distillation with Evolving Contexts
Flux-OPD는 역 KL 분해를 통해 진화하는 컨텍스트를 활용한 오퍼레이션 폴리시 디스틸레이션(OPD) 패러다임으로, 개방형 도메인에서의 작업 선호도를 효과적으로 학습한다.
- #318Beyond Borrowed Histories: Person-Aligned User Simulation for Interactive Role-Playing Evaluation
PALATE는 사용자 맞춤형 시뮬레이션을 통해 대화형 역할극 에이전트를 평가하는 새로운 벤치마크를 제시한다.
- #319RefCaptioner: Multi-Reference Image-Grounded Video Captioning
RefCaptioner는 다중 참조 이미지 기반 영상 캡션 생성을 위한 새로운 프레임워크이다.
- #320MindForge: Teaching Small Language Models Whole-Life-Cycle Software Engineering via Source-Free Program Synthesis
MindForge는 소스코드 없이 프로그램을 처음부터 완성하는 소프트웨어 엔지니어링 훈련 파이프라인으로, Qwen3.6-27B의 ProgramBench 성능을 37.98%에서 49.51%로 향상시킨다.
- #321See2Think: Do Multimodal Models Really Use Intermediate Visual States?
See2Think은 멀티모달 모델이 중간 시각 상태를 얼마나 효과적으로 사용하는지 평가하는 통합 프레임워크이다.
- #322Is Deep Research Reliable? Misleading Knowledge Induces False Conclusions
MisKnow-Agent라는 프레임워크를 통해 생성된 오도적 지식이 Deep Research 에이전트의 최종 보고서에 오류를 유도하는 실패 모드를 분석하고, 이를 완화하는 방어 전략을 제시한다.
- #323SpatialCLI: Learning to Reason With Spatial Tools, Then Without Them
SpatialCLI는 VLM이 공간 도구를 활용한 후 이를 내재화하여 공간 추론 성능을 향상시키는 3단계 프레임워크이다.
- #324β-OPSD: Deriving with Policy Optimization, Training with Self-Distillation
β-OPSD는 정책 최적화 기반의 자가 교사 학습(SEL)을 효율적으로 근사하는 새로운 자가 교사 학습 알고리즘으로, 수학적 추론 성능을 향상시킨다.
- #325SpecFirst: Behavioral Specification Elicitation as a First-Class Step in Agent-Based Program Synthesis from Scratch
- #326Chimera: Designing and Chinchilla-Scaling Hybrid Visual Diffusion Transformers
- #327πR^2: Reactive Real-time Flow Policies
πR²는 대규모 사전 학습된 백본을 유지하면서도 실시간 반응성을 확보한 흐름 정책 프레임워크이다.
- #328Grading the Narrators: An Isnad-Rijal Framework for Claim-Level Provenance in Multi-Agent Knowledge Systems
- #329Can Large Language Models Execute Parent Orders?
LLM 기반 PACE 프레임워크가 주문 실행 비용을 0.65 bps 개선하며 인간 투자자와 다른 행동 패턴을 보인다.
- #330MemHarness: Memory Is Reconstructed, Not Replayed
- #331INTACT: Isomorphic Intent-to-Action Learning for Search-Free World Models
- #332Explorative Modeling: Unlocking a Third Pretraining Axis and End-to-End Generation
- #334LEDGERMIND: Provenance-Constrained Multimodal Agentic Reasoning with a Structured Evidence Ledger
- #335Echoverse: Deep, Evolving Environments for Training Computer-Use Agents at Scale
- #337Multi-Head Attention Residuals