PluRel Dataset Synthetic Data unlocks Scaling Laws for Relational Foundation Models Preprocessed synthetic relational databases for pretraining relational foundation models, as introduced in: PluRel: Synthetic Data unlocks Scaling Laws for Relational Foundation Models Kothapalli, Ranjan, Hudovernik, Dwivedi, Hoffart, Guestrin, Leskovec — arXiv:2602.04029 (2026) Data Structure Each entry is a relbench compatible Database consisting of multiple relational tables. Component Description Tables 3–20 per database Primary keys row idx (auto generated) Foreign keys foreign row 0 , foreign row 1 , ... Feature columns feature 0 , feature 1 , ... (categorical or numerical) Time column date — on activity (leaf) tables only Schema topology is sampled from: BarabasiAlbert, ReverseRandomTree, or WattsStrogatz graphs. Data generation uses Structural Causal Models (SCMs) — column dependencies are modeled as DAGs, with values propagated through randomly initialized MLPs. Activity tables also include trend + cycle + noise time series. Parameter Range Rows per entity table 500–1,000 Rows per activity table 2,000–5,000 Columns per table 3–40 (power law) Missing values 1–10% of numerical columns Timesta…
Runs entirely in your browser via DuckDB-Wasm — this dataset's real data file is loaded once, then queried locally. Nothing is sent to a server.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy