- #1AREX-2: Advancing Self-Improving Agents through Long-Horizon Reflective Tasks
- #2ControlNet++: Improving Conditional Controls with Efficient Consistency Feedback
- #3Ferret-UI: Grounded Mobile UI Understanding with Multimodal LLMs
Ferret-UI는 모바일 UI 화면을 정밀하게 이해하고 지시에 따라 행동할 수 있는 다중모달 대형 언어 모델로, GPT-4V보다 기본 UI 작업에서 우수한 성능을 보인다.
- #43DGStream: On-the-Fly Training of 3D Gaussians for Efficient Streaming of Photo-Realistic Free-Viewpoint Videos
3DGStream은 실시간 렌더링과 초당 200프레임(FPS)을 지원하는 3D 가우시안 기반의 실내/실외 동적 장면의 실시간 프리뷰(Free-Viewpoint Video) 스트리밍 기법이다.
- #5SpatialTracker: Tracking Any 2D Pixels in 3D Space
SpatialTracker는 3D 공간에서 2D 픽셀 추적을 수행하여 복잡한 시나리오에서도 정확한 장거리 픽셀 추적을 달성하는 새로운 방법이다.
- #6Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
Ada-KV는 주의력 헤드별로 적응적 예산을 할당하여 KV 캐시 제거 성능을 향상시키는 최초의 전략이다.
- #7COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability
- #8Masked Diffusion Models are Secretly Time-Agnostic Masked Models and Exploit Inaccurate Categorical Sampling
마스킹된 확산 모델(MDMs)은 시간 변수와 무관한 마스킹 모델과 동등하며, 범주형 샘플링의 수치적 불안정성으로 인해 성능 평가가 왜곡될 수 있음.
- #9Trial and Error: Exploration-Based Trajectory Optimization for LLM Agents
- #10MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents
- #11Text-IF: Leveraging Semantic Text Guidance for Degradation-Aware and Interactive Image Fusion
Text-IF는 텍스트 유도를 활용해 저품질 이미지 복원과 상호작용 기반 퓨전을 통합한 새로운 이미지 퓨전 모델이다.
- #165Mid-Harness: Scaling Actions Between Model and Harness for Terminal Agents
- #287A Data-Free Physics-Informed Neural Operator for Level-Set Interface Advection
- #292LoopVL: Recurrent Visual Intelligence
- #295EVOKE: Eliciting World Knowledge in Agents for Transferable Decision-Making
- #306UniEvo-VL: An On-policy Self-Distillation Training Recipe for Multimodal Model Self-improvement
- #307False Frontiers: Diagnosing and Mitigating Co-Cheating in Self-Evolving Search Agents
- #308ATLAS: Aligned Transport of Latent Structure for Reliable World Model Planning
- #309MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution
- #310Rethinking Latent Visual Reasoning: Grounding Latent Reasoning in Visual Evidence
- #311LANTERN: Illuminating Hidden Mathematical Knowledge in Language Models
- #312Breaking Babel: A Self-Evolving Multi-Agent System for Long-Form Subtitle Translation
- #313Unmask the State: When Does State Adaptation Matter for Masked Diffusion Language Models
- #314It's Not What the Image Shows: Irrelevant Context Destabilises VLM Judges Without Informing Them
- #315DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes
- #316TERRA: Terrain-Aware Reconstruction, Retargeting and Control for Musculoskeletal Locomotion
- #317OSWorld-Science: A Benchmark of Computer Use Agents for Learning and Using Scientific Software
- #318Label-free steering: Compressing test-time reinforcement learning into bias-only subspaces
- #319FlashDiffusion: Fused Tiled Kernel Spectral Decomposition
- #320Automated Evaluation of Multi-Turn Dialogues in In-Car Conversational Assistants