STRABLE: Benchmarking Tabular Machine Learning with Strings This dataset card describes the STRABLE benchmark, a comprehensive suite designed for evaluating machine learning models on tabular data containing strings. Dataset Description Benchmarking tabular data has revealed the benefit of dedicated architectures, pushing the state of the art. However, real world tables often contain string entries beyond pure numbers, a setting that has been understudied due to a lack of a solid benchmarking suite. STRABLE is a comprehensive benchmarking corpus of 108 tables containing both strings and numbers. These datasets are carefully curated learning problems across diverse application fields to enable the empirical study of tabular learning with strings. Dataset Sources Repository: https://github.com/soda inria/strable Paper: https://arxiv.org/pdf/2605.12292 Project Page: https://soda inria.github.io/strable Uses Direct Use This dataset is intended to be used by researchers and practitioners evaluating tabular machine learning pipelines. It allows users to answer critical research questions regarding string representations in tables: whether dedicated end to end learners are needed, or if m…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy