Poseidon: Efficient Foundation Models for PDEs

Maximilian Herde, Bogdan Raoni'c, Tobias Rohner, R. Käppeli, Roberto Molinaro, Emmanuel de B'ezenac, Siddhartha Mishra

arXiv:2405.19101 · 2026-07-27 공개 · arXiv · PDF

foundation-models open-source generalization downstream-tasks pde-solution fluid-dynamics pretraining-strategy operator-transformer

Abstract

We introduce Poseidon, a foundation model for learning the solution operators of PDEs. It is based on a multiscale operator transformer, with time-conditioned layer norms that enable continuous-in-time evaluations. A novel training strategy leveraging the semi-group property of time-dependent PDEs to allow for significant scaling-up of the training data is also proposed. Poseidon is pretrained on a diverse, large scale dataset for the governing equations of fluid dynamics. It is then evaluated on a suite of 15 challenging downstream tasks that include a wide variety of PDE types and operators. We show that Poseidon exhibits excellent performance across the board by outperforming baselines significantly, both in terms of sample efficiency and accuracy. Poseidon also generalizes very well to new physics that is not seen during pretraining. Moreover, Poseidon scales with respect to model and data size, both for pretraining and for downstream tasks. Taken together, our results showcase the surprising ability of Poseidon to learn effective representations from a very small set of PDEs during pretraining in order to generalize well to unseen and unrelated PDEs downstream, demonstrating its potential as an effective, general purpose PDE foundation model. Finally, the Poseidon model as well as underlying pretraining and downstream datasets are open sourced, with code being available at https://github.com/camlab-ethz/poseidon and pretrained models and datasets at https://huggingface.co/camlab-ethz.

한국어 요약

한 줄 요약

Poseidon은 PDE 해 연산자 학습을 위한 다중 스케일 오퍼레이터 트랜스포머 기반의 효율적인 펀다멘탈 모델로, 15개의 다운스트림 태스크에서 뛰어난 성능을 보인다.

핵심 기여도

핵심 아이디어

Poseidon은 PDE 해 연산자를 학습하기 위해 **scOT**(scalable Operator Transformer)라는 다중 스케일 트랜스포머 아키텍처를 도입한다. 이는 **SwinV2 트랜스포머 블록**을 기반으로 하며, **윈도우 기반 멀티헤드 셀프 어텐션**(windowed multi-head self attention)을 통해 계산 효율성을 높인다. 또한, **시간 조건화 레이어 노름**(time-conditioned layer norm)을 통해 시간에 따라 연속적으로 평가할 수 있는 모델 구조를 구현한다.

이 모델은 **반군 성질**(semi-group property)을 활용한 **all2all 훈련 전략**을 통해 시간에 따른 PDE 해의 트레 jury를 효율적으로 활용하여 훈련 데이터를 확장한다. 이는 기존의 샘플 기반 훈련과 달리, 시간 경로를 활용해 더 많은 훈련 샘플을 생성할 수 있다는 점에서 혁신적이다.

Poseidon은 유체 역학 방정식(압축성 유체의 Euler 방정식과 비압축성 유체의 Navier-Stokes 방정식) 기반의 대규모 데이터셋에서 사전 훈련되며, 이는 6개의 다양한 해 연산자로 구성된다. 이는 PDE 펀다멘탈 모델이 소수의 PDE에서 학습된 표현을 활용해 **미관측 PDE**에 대해서도 일반화할 수 있음을 보여준다.

기술적 접근법

주요 결과

의의 및 한계

Poseidon은 PDE 펀다멘탈 모델의 가능성과 한계를 실증적으로 보여주는 첫 사례이다. 특히, **소수의 PDE에서 학습된 표현이 미관측 PDE에 일반화**될 수 있음을 입증하며, PDE 학습의 샘플 효율성 문제를 해결하는 데 기여한다. 또한, **모델 및 데이터 크기와의 스케일링 가능성**을 보여주어 대규모 PDE 시뮬레이션에의 적용 가능성을 제시한다.

그러나, **PDE의 종류와 복잡도에 따라 일반화 성능이 달라질 수 있으며**, 모든 PDE 타입에 대한 실험은 아직 제한적이다. 또한, **사전 훈련 데이터셋의 다양성**이 모델의 일반화 능력에 큰 영향을 미친다는 점도 한계로 지적된다.

실용적 활용

Poseidon은 **유체 역학, 기상 예측, 재료 과학** 등 다양한 분야에서 PDE 기반 시뮬레이션을 효율적으로 수행할 수 있는 도구로 활용될 수 있다. 특히, **고비용의 수치 시뮬레이션 대체** 및 **실시간 PDE 해석**에 적합하며, **AI 기반 과학 계산**(scientific computing)의 발전에 기여할 수 있다.