BeyondArena Datasets Datasets from BeyondArena, a unified, holistic benchmark for tabular data that supports diverse task types (IID, temporal, grouped), across sample size and feature dimensionality scales, with diverse feature types (with text, with high cardinality) from a broad range of disciplines. We introduce BeyondArena and its datasets in Beyond IID: How General Are Tabular Foundation Models, Really? . Click for BibTeX! More details: Project page and leaderboard: http://tabarena.ai/ Code / Benchmark repository: https://tabarena.ai/code Quickstart We recommend using the datasets via Data Foundry, which resolves a curated container (table + dtypes + task metadata + outer CV splits) by name and caches it locally: To pre download the entire collection in a single network round trip: See Data Foundry's examples for a full benchmarking walkthrough, the three split regimes (IID / temporal / grouped), and the curation flow. Datasets BeyondArena comes with 142 datasets. BeyondArena covers tabular classification and regression tasks. And the following types of datasets: IID tabular data Non IID temporal tabular data Non IID grouped tabular data IID and non IID tabular data with text…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy