OmniScientist: An Omni-Modal Omni-Discipline AI Scientist

Bobo Li, Hao Fei, Tianjie Ju, Mong-Li Lee, Wynne Hsu

arXiv:2608.13558 · 2026-08-16 공개 · arXiv · PDF

autonomous-agents scientific-discovery omni-modal ai-scientist heterogeneous-data evidence-grounded research-workflow multidisciplinary

Abstract

Recent advances in foundation models have enabled AI scientists to automate increasingly complete research workflows, from hypothesis generation and code execution to manuscript preparation. Yet workflow coverage alone does not provide access to the full evidence on which scientific discovery depends. Existing systems typically reason over text, code, labels, or precomputed summaries, leaving scientifically decisive spatial, temporal, cross-channel, and procedural relations unavailable to the agent. We introduce OmniScientist, an end-to-end, omni-modal AI scientist that conducts multidisciplinary research directly from heterogeneous raw evidence. A perception layer and 3 autonomous agents for ideation, experiment, and writeup operate within a deterministic pipeline, allowing observations to shape research questions, experimental decisions, and final claims throughout the research lifecycle. By running idea, rigour, and claim checks in code, the system enforces novelty screening, statistical validity, execution provenance, and numerical traceability. We evaluate OmniScientist on 36 real-data cases spanning 5 discipline families, 4 families of scientific evidence, and modalities including images, signals, audio, video, 3-D structures, trajectories, tables, formulae, and graphs. The system completes the full path from raw data to a compiled manuscript in all 36 cases and achieves a mean overall paper score of 6.3 with the reference reasoning backbone. In paired comparisons against a blind variant that receives only precomputed scalar features, direct perception improves all 7 evaluation dimensions and wins 85% of head-to-head judgments. These results show that lifecycle-wide perception is essential for evidence-grounded scientific discovery and provides a practical path toward broadly capable AI scientists.

한국어 요약

한 줄 요약

OmniScientist는 다학문적 연구를 이종 원천 데이터에서 직접 수행하는 종단간 다모달 AI 과학자 시스템이다.

핵심 기여도

핵심 아이디어

기존 AI 과학자 시스템은 텍스트, 코드, 요약 정보를 기반으로 작동하며, 공간적, 시간적, 채널 간, 절차적 관계를 누락시킨다. OmniScientist는 이종 원천 데이터를 직접 감지하고 이를 연구 전 과정에 반영함으로써 과학적 발견의 근거를 확장한다. 핵심 아이디어는 **감지-이유-행동**(ReAct) 루프를 통해 원천 데이터가 연구 질문, 실험 설계, 결과 해석, 최종 주장에 직접적으로 영향을 미치도록 하는 것이다.

**감지 레이어**(perception layer)는 이미지, 신호, 3D 구조, 테이블, 그래프 등 다양한 모달을 처리하며, **아이디어**(ideation), **실험**(experiment), **논문 작성**(writeup) 에이전트는 각각 연구 과정을 자율적으로 진행한다. **아이디어 체크**(idea check), **엄격성 체크**(rigour check), **주장 체크**(claim check)는 코드 내에서 실행되어 **신뢰성**(novelty screening), **통계적 유효성**(statistical validity), **실행 추적**(execution provenance), **수치 추적 가능성**(numerical traceability)을 보장한다.

기술적 접근법

주요 결과

의의 및 한계

OmniScientist는 과학적 발견의 근거를 확장하고, AI 과학자 시스템의 **증거 기반**(evidence-grounded) 연구를 가능하게 한다. 기존 시스템은 워크플로우는 완전하지만, **증거는 불완전**(workflow-complete but evidence-incomplete)하다는 한계를 극복한다. 이 시스템은 **이종 데이터를 직접 인식**하고, **연구 전 과정에 반영**함으로써 과학적 주장의 신뢰성을 높인다.

한계로는, **모델 백본**(reasoning backbone)의 성능이 최종 결과에 큰 영향을 미친다는 점이 언급된다. 또한, **새로운 학문 분야 추가 시**는 **새로운 명세 파일**(specification file)만 추가하면 되지만, **기본 파이프라인 변경 없이** 동작한다는 점에서 유연성은 높으나, **모델의 일반화 능력**이 여전히 중요한 과제이다.

실용적 활용

OmniScientist는 이종 원천 데이터(이미지, 신호, 3D 구조, 테이블 등)를 처리할 수 있어, 생물학, 물리학, 환경 과학, 공학, 사회과학 등 다양한 학문 분야에서 연구를 자동화할 수 있다. 특히, 실험 설계, 데이터 분석, 논문 작성 등 연구 워크플로우의