- #3Ego2Act: Evaluating Goal-Directed Manipulation in Egocentric Video Generation
Ego2Act는 목표 지향적 제1인칭 영상 생성 모델을 평가하기 위한 2,640개 영상 기반 벤치마크로, Seedance-2.0 모델이 64.0점에 그친 것을 보여준다.
- #4Streaming Long Video Understanding with Large Language Models
VideoStreaming은 긴 동영상의 효율적인 이해를 위해 Memory-Propagated Streaming Encoding과 Adaptive Memory Selection을 결합한 VLLM 모델이다.
- #5Transcoders Find Interpretable LLM Feature Circuits
Transcoders를 활용한 MLP 서브레이어 분석으로 GPT2-small의 "greater-than circuit"을 포함한 해석 가능한 LLM 회로를 발견한다.
- #6Poseidon: Efficient Foundation Models for PDEs
Poseidon은 PDE 해 연산자 학습을 위한 다중 스케일 오퍼레이터 트랜스포머 기반의 효율적인 펀다멘탈 모델로, 15개의 다운스트림 태스크에서 뛰어난 성능을 보인다.
- #7TabArena: A Living Benchmark for Machine Learning on Tabular Data
TabArena는 표 형태 데이터 기계 학습 모델을 평가하는 첫 번째 지속적으로 유지되는 라이빙 벤치마크 시스템이다.
- #8Large Language Models Can Learn Temporal Reasoning
TG-LLM을 제안하여 언어 기반 시공간 추론 성능을 향상시킨다.
- #9LLaVA-3D: A Simple Yet Effective Pathway to Empowering LMMs with 3D Capabilities
LLaVA-3D는 3D 포지션 임베딩과 3D 패치를 활용해 2D LMM인 LLaVA를 3D 시나리오로 확장한 모델로, 3D 객체 인식 및 추론 성능을 향상시키며 3.5배 빠른 수렴 속도를 보인다.
- #10Audio Flamingo 2: An Audio-Language Model with Long-Audio Understanding and Expert Reasoning Abilities
Audio Flamingo 2는 3B 파라미터 소형 언어 모델로, 20개 이상의 벤치마크에서 최고 성능을 달성한 다기능 오디오-언어 모델이다.
- #11HALC: Object Hallucination Reduction via Adaptive Focal-Contrast Decoding
HALC는 LVLM에서 발생하는 오브젝트 환각(OH)을 줄이기 위해 설계된 플러그 앤 플레이 디코딩 알고리즘이다.
- #12OGBench: Benchmarking Offline Goal-Conditioned RL
OGBench는 오프라인 목표 조건화 강화학습 연구를 위한 8개 환경, 85개 데이터셋, 6개 알고리즘 구현을 포함한 새로운 벤치마크이다.
- #13VideoPhy: Evaluating Physical Commonsense for Video Generation
VideoPhy는 생성된 동영상이 물리적 상식을 얼마나 잘 반영하는지 평가하는 벤치마크로, CogVideoX-5B가 39.6%만 정확히 생성함을 보여준다.
- #308Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It
ран크-8 LoRA를 통한 미세 조정이 프리트레인드 트랜스포머의 참조 추적 능력을 극적으로 향상시킨다.
- #314Video Generation Models: A Survey of Post-Training and Alignment
비디오 생성 모델에서 사전 학습 후 정렬(post-training and alignment)을 체계적으로 분류하고, 주요 방법론과 도전 과제를 정리한 최초의 종합적 서베이.
- #317Persona Dosing: Calibrated Activation Steering for Graded Trait Control
PersonaDose는 FLAS 기반 컨트롤러를 활용해 인격 특성의 강도를 행동 단위로 조절하는 기법으로, Llama-3.1-8B 등 3개 모델에서 평균 17.8~33.2 포인트의 표현 강도 향상을 달성했다.
- #319Honeycomb: Constant-Size Scene Memory Representation for Video World Models
Honeycomb은 고정 크기의 HexMemory를 사용해 장기 비디오 생성 시 장면 일관성을 유지하는 비디오 월드 모델이다.
- #320Architect-Ant: Editable Automatic Furnishing of Architectural Floor Plans
- #322SILSA: Sliding-Window Slice Latents for Topology-Preserving High-Resolution 3D Generation
SILSA는 고해상도 3D 생성에서 위상 일관성을 유지하면서 생성 비용을 58.5% 절감하는 슬라이딩-윈도우 슬라이스 잠재량 기반 프레임워크이다.
- #323Decoding Looped Transformers Better for (Almost) Free
- #324Multimodal Flow: Unified Flow Modeling of Language and Vision in Embedding Spaces
- #326PhysVista: Benchmarking Physical Intelligence in VLMs via a Perception-Reasoning-Assessment Loop
- #327Omni-Embed-Mini: Binding Modalities Without Forgetting via Dense Distillation
- #328Better Supervision Is Nearby: Neighborhood On-Policy Self-Distillation
- #329Smaller Models, Better Rejects: Preference Distillation Scaling
- #330Where-OPD: Spatially Guided On-Policy Self-Distillation of MLLMs with Synthetic Scenes
- #331DataMagic: Authoring Data Videos through Declarative Multi-Agent Orchestration
- #332OpenTumorBoard: A Real-World Benchmark of Multidisciplinary Tumor Board Discussion Trajectories
- #333AgSpec: Pushing the Limits of Retrieval-Based Speculative Decoding in Coding Agent Pipelines
- #334SemanTok: Predictable Semantic Tokens for Efficient Autoregressive Video Generation
- #335Generalization Is Stability, Not Accuracy: Multi-Axis Evaluation of LLMs
- #336Removing the NEEDLE in the Haystack: Backdoor Removal in LLMs via Weight Orthogonalisation