TabArena: A Living Benchmark for Machine Learning on Tabular Data

Nick Erickson, Lennart Purucker, Andrej Tschalzev, David Holzmüller, Prateek Mutalik Desai, A. Salinas, Frank Hutter

arXiv:2506.16791 · 2026-09-12 공개 · arXiv · PDF

foundation-models tabular-data tabarena gradient-boosted-trees model-ensembles reproducible-code living-benchmark hyperparameter-ensembling

Abstract

With the growing popularity of deep learning and foundation models for tabular data, the need for standardized and reliable benchmarks is higher than ever. However, current benchmarks are static. Their design is not updated even if flaws are discovered, model versions are updated, or new models are released. To address this, we introduce TabArena, the first continuously maintained living tabular benchmarking system. To launch TabArena, we manually curate a representative collection of datasets and well-implemented models, conduct a large-scale benchmarking study to initialize a public leaderboard, and assemble a team of experienced maintainers. Our results highlight the influence of validation method and ensembling of hyperparameter configurations to benchmark models at their full potential. While gradient-boosted trees are still strong contenders on practical tabular datasets, we observe that deep learning methods have caught up under larger time budgets with ensembling. At the same time, foundation models excel on smaller datasets. Finally, we show that ensembles across models advance the state-of-the-art in tabular machine learning. We observe that some deep learning models are overrepresented in cross-model ensembles due to validation set overfitting, and we encourage model developers to address this issue. We launch TabArena with a public leaderboard, reproducible code, and maintenance protocols to create a living benchmark available at https://tabarena.ai.

한국어 요약

한 줄 요약

TabArena는 표 형태 데이터 기계 학습 모델을 평가하는 첫 번째 지속적으로 유지되는 라이빙 벤치마크 시스템이다.

핵심 기여도

핵심 아이디어

기존 벤치마크는 설계 후 유지 관리가 부족해 최신 모델이나 데이터셋 변화에 대응하지 못한다. TabArena는 이러한 문제를 해결하기 위해 소프트웨어처럼 버전 관리되고 커뮤니티에 의해 지속적으로 업데이트되는 라이빙 벤치마크 시스템을 제안한다. 이 시스템은 표 데이터 기계 학습 모델의 신뢰성 있는 평가를 가능하게 하며, 특히 앙상블과 하이퍼파라미터 튜닝이 모델 성능 최적화에 중요한 역할을 한다는 점을 강조한다. TabM, RealMLP, AutoGluon 등 다양한 모델이 포함되어 있으며, 특히 작은 데이터셋에서는 펀데이션 모델이 우수한 성능을 보인다.

기술적 접근법

주요 결과

의의 및 한계

TabArena는 표 데이터 기계 학습 모델 평가의 신뢰성과 표준화를 높이는 데 기여하며, 연구자와 실무자에게 실용적인 가이드를 제공한다. 또한, 공개 리더보드와 재현 가능한 코드를 통해 커뮤니티 참여를 촉진한다. 그러나 일부 한계도 존재한다. 예를 들어, 고정된 200개의 랜덤 하이퍼파라미터 설정으로 인해 고급 최적화 전략 분석이 제한되고, 하드웨어 차이로 인한 비교 가능성 저하, 데이터셋 수의 제한 등이 있다. 또한, 특징 엔지니어링 없이 평가하기 때문에 모델 순위가 실제 상황과 다를 수 있다.

실용적 활용

TabArena는 중소 규모의 독립 동일 분포(IID) 표 데이터를 다루는 실무자에게 모델 선택과 평가에 대한 신뢰성 있는 기준을 제공한다. 또한, 연구자들이 새로운 모델과 앙상블 전략을 표준화된 환경에서 비교할 수 있도록 지원하며, 향후 다양한 표 데이터 작업(예: 비-IID, 클러스터링, 이상 탐지 등)으로 확장될 가능성이 있다.