Voorbij IID: Hoe algemeen zijn tabulaire fundamentmodellen werkelijk?
Beyond IID: How General Are Tabular Foundation Models, Really?
June 29, 2026
Auteurs: Lennart Purucker, Andrej Tschalzev, Nick Erickson, Gioia Blayer, David Holzmüller, Alan Arazi, Alexander Pfefferle, Mustafa Tajjar, Gaël Varoquaux, Frank Hutter
cs.AI
Samenvatting
Fundamentmodellen voor voorspellend machinaal leren op tabulaire gegevens hebben recentelijk aanzienlijke populariteit verworven in de academische wereld en de industrie. Onderzoeksgemeenschappen uit verschillende disciplines evalueren steeds vaker tabulaire fundamentmodellen op uiteenlopende datasets en taken. Echter, deze taak- en disciplinespecifieke evaluaties blijven grotendeels ontoegankelijk voor modelonderzoekers omdat benchmarksoftware en evaluatieprotocollen gefragmenteerd zijn. Als gevolg hiervan vertrouwen modelonderzoekers op standaard benchmarks, die meestal zijn gedefinieerd voor taken waarin tabulaire fundamentmodellen al uitblinken. De meest uitdagende scenario's worden uitgesloten, waardoor zinvolle vooruitgang in het veld wordt beperkt door zich te richten op marginale verbeteringen op IID-gegevens in plaats van op bredere, veeleisendere uitdagingen. Om dit te overwinnen introduceren we BeyondArena, de eerste uniforme holistische benchmark voor tabulaire gegevens die diverse taaktypen ondersteunt (IID, temporeel, gegroepeerd), over steekproefgrootte en kenmerkdimensionaliteitsschalen, met diverse kenmerktypen (met tekst, met hoge kardinaliteit) uit een breed scala aan disciplines. Om uniforme benchmarking buiten standaard benchmarks mogelijk te maken, introduceren we Data Foundry, een Python-framework en metadataschema voor het samenstellen van tabulaire datasets voor voorspellend machinaal leren. Onze resultaten over 11 modellen en 142 samengestelde datasets tonen aan dat bestaande tabulaire fundamentmodellen uitblinken op kleine tot middelgrote IID-gegevens, terwijl traditionele op bomen gebaseerde en diepe leermodellen nog steeds domineren op niet-IID, grote en hoogdimensionale datasets. BeyondArena stuurt modelonderzoek voor de meest veeleisende uitdagingen in tabulaire gegevens, waardoor vooruitgang naar werkelijk fundamentele tabulaire modellen mogelijk wordt.
English
Foundation models for predictive machine learning on tabular data have recently gained significant traction in academia and industry. Research communities across disciplines are increasingly evaluating tabular foundation models on diverse datasets and tasks. However, these task- and discipline-specific evaluations remain largely inaccessible to model researchers because benchmark software and evaluation protocols are fragmented. As a result, model researchers rely on standard benchmarks, which are mostly defined for tasks where tabular foundation models already excel. The most challenging scenarios are excluded, limiting meaningful progress in the field by focusing on marginal improvements on IID data rather than on broader, more demanding challenges. To overcome this, we introduce BeyondArena, the first unified holistic benchmark for tabular data that supports diverse task types (IID, temporal, grouped), across sample size and feature dimensionality scales, with diverse feature types (with text, with high cardinality) from a broad range of disciplines. To enable unified benchmarking beyond standard benchmarks, we introduce Data Foundry, a Python framework and metadata schema for curating tabular datasets for predictive machine learning. Our results across 11 models and 142 curated datasets show that existing tabular foundation models excel on tiny- to medium-sized IID data, while traditional tree-based and deep learning models still dominate on non-IID, large, and high-dimensional datasets. BeyondArena guides model research for the most demanding challenges in tabular data, enabling progress towards truly foundational tabular models.