- #1PhiZero: A World Model Built Around Physical Language
PhiZero는 물리적 세계의 진화를 추론-렌더링 방식으로 모델링하는, 물리적 언어 기반의 월드 모델이다.
- #2MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?
MME-RealWorld는 13,366개의 고해상도 이미지로 구성된 최대 규모의 수작업 라벨 벤치마크로, 28개 MLLM이 60% 미만의 정확도를 기록하며 모델의 실제 세계 인식 능력 한계를 드러낸다.
- #3Optimizing Instructions and Demonstrations for Multi-Stage Language Model Programs
MIPRO 알고리즘은 다단계 언어 모델 프로그램에서 지시문과 샘플을 최적화하여 최대 13%의 정확도 향상을 달성한다.
- #4Σ-Mem: An Online Reliability Memory for LLM-based Multi-Agent Systems
Σ-Mem은 LLM 기반 다중 에이전트 시스템에서 신뢰도를 기반으로 동적으로 조정하는 온라인 메모리 시스템이다.
- #5HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
HarmBench는 자동 레드팀 테스트와 LLM의 안정적 거부를 평가하기 위한 표준 평가 프레임워크로, 18개 공격 방법과 33개 모델을 대규모 비교한 결과를 제시한다.
- #6OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems
OlympiadBench는 수학·물리 올림피아드 수준의 이중 언어·다중 모달 과학 문제를 포함한 8,952개 문제로 구성된 AGI 연구를 위한 도전적인 벤치마크이다.
- #7SimPO: Simple Preference Optimization with a Reference-Free Reward
SimPO는 DPO 대비 6.4~7.5 포인트 개선된 성능을 보이는 reference-free reward 기반 오프라인 선호 최적화 알고리즘이다.
- #8Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction
Visual AutoRegressive (VAR) 모델은 이미지 생성에서 확률적 확장법과 제로샷 일반화를 달성한 새로운 AR 프레임워크이다.
- #9The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale
FineWeb은 15조 토큰 규모의 웹 기반 LLM 학습 데이터셋으로, 교육적 텍스트를 필터링한 FineWeb-Edu(1.3조 토큰)를 포함하며, 공개 데이터셋 대비 성능 향상이 입증됨.
- #10Grounding Image Matching in 3D with MASt3R
MASt3R은 3D 이미지 매칭 성능을 향상시키기 위해 DUSt3R에 새로운 헤드와 빠른 매칭 알고리즘을 결합한 방법이다.
- #11YOLO-World: Real-Time Open-Vocabulary Object Detection
YOLO-World는 CLIP 기반 텍스트 인코더와 RepVL-PAN을 활용한 실시간 오픈-바보카 보물 감지 모델로, LVIS에서 35.4 AP, 52.0 FPS를 달성한다.
- #12DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models
DeepSeekMoE는 전문가 분화를 극대화한 MoE 아키텍처로, 2B 모델에서도 GShard 2.9B와 유사한 성능을 달성한다.
- #73ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Engine
ACE-Data-0은 실내 환경에서 사람-객체 상호작용을 다중 센서로 동기화 기록한 대규모 데이터셋이다.
- #304AskChem: Claim-Centered Infrastructure for Chemistry Literature Synthesis
AskChem은 문서 중심 검색 대신, 과학적 주장(claim)을 중심으로 화학 문헌을 통합하는 인프라를 제시한다.
- #305Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents
Qwen-UI-Agent는 실제 모바일 기기와 크로스 플랫폼 환경에서 뛰어난 성능을 보이는 GUI 에이전트로, 92.2%의 MobileWorld-Real 정확도를 달성했다.
- #306Metis: Memory Foundation Model
Metis는 기존의 외부 메모리 모듈을 대체하는 첫 번째 내재 메모리 기반 모델로, 메모리 상태와 절차를 모델 내부에 통합하여 효율적이고 유연한 메모리 처리를 가능하게 한다.
- #308Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering
Frontis-MA1은 MLE 분야에서 재귀적 자기 개선을 위한 AI4AI 모델로, OpenMLE 스택과 연계해 Medal Average 60.61%를 달성한다.
- #309ShadowDancer: Teaching Video World Models Any Action by Learning Unified Dynamics Representations from a Video and Its Shadow
ShadowDancer는 동영상과 그림자 쌍을 통해 동작을 학습하여 다양한 동역학에 대한 프레임 수준의 정밀 제어를 가능하게 하는 새로운 인터페이스를 제시한다.
- #311DistillAlign: Coordinating Mode Covering and Mode Seeking in Autoregressive Video Distillation
DistillAlign은 DMD와 CD를 결합하여 초기화와 최종 정제 단계의 분포 일치를 동시에 달성하는 비디오 디스틸레이션 방법이다.
- #312VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System
VideoCoCo는 실행 가능한 코드를 이용해 물리적으로 일관된 동영상을 생성하는 이중 엔진 시스템을 제안한다.
- #313Memory Decoder at Scale: A Pretrained, Parametric Long-Term Memory
Memory Decoder at Scale는 6.9B 파라미터 규모로 확장된 장기 기억 모듈을 제안하여 언어 모델 성능을 파라미터 효율적으로 향상시킨다.
- #314Beacon: Knowing When and How to Perform Agentic Visual Reasoning
Beacon은 도구 사용의 적절성과 효과성을 동시에 향상시킨 새로운 에이전트형 시각 추론 모델로, 강화 학습 기반의 Necessity-Aware Adaptive Reward와 Hint-Guided Capability Expansion을 핵심으로 한다.
- #315BM25 Wins at Scale: A Scaling Study of Retrieval-Augmented Generation Paradigms
BM25는 1000만 토큰 이상의 대규모 문서에서 RAG 성능에서 File-System Agent를 20점 이상 앞선다.
- #316Flux-OPD: On-Policy Distillation with Evolving Contexts
- #317MPIE-Bench: Benchmarking Anatomically Plausible Multi-Person Interaction Editing
MPIE-Bench는 다인물 상호작용 편집의 해부학적 일관성을 평가하기 위한 2,500개 샘플의 벤치마크와 MPIE-Eval이라는 새로운 평가 프로토콜을 제안한다.
- #318Beyond Borrowed Histories: Person-Aligned User Simulation for Interactive Role-Playing Evaluation
PALATE는 사용자 맞춤형 시뮬레이션을 통해 대화형 역할극 에이전트를 평가하는 새로운 벤치마크를 제시한다.
- #319RefCaptioner: Multi-Reference Image-Grounded Video Captioning
- #320See2Think: Do Multimodal Models Really Use Intermediate Visual States?
- #321MindForge: Teaching Small Language Models Whole-Life-Cycle Software Engineering via Source-Free Program Synthesis
- #322Is Deep Research Reliable? Misleading Knowledge Induces False Conclusions