Abjad Kids: An Arabic Speech Classification Dataset for Primary Education Abjad Kids is an Arabic speech classification dataset designed for primary education applications. It contains spoken recordings of the Arabic alphabet , numbers , and colors from multiple child speakers, supporting research in automatic speech recognition, audio classification, and educational technology for Arabic speaking children. This dataset is related to the work presented in: Abjad Kids: An Arabic Speech Classification Dataset for Primary Education ResearchGate. https://www.researchgate.net/publication/401732601 Dataset Description The dataset is organized into three main categories, each with subcategories corresponding to the spoken label: Category Description Example Labels alphabet Arabic letter pronunciations Alam (أ), Ba (ب), Ta (ت), ... numbers Arabic number names Arbaa (أربعة), Whaed (واحد), Ethenen (اثنين), ... colors Arabic color names Abyad, Ahmar, Akhdar, ... Dataset Structure Data Fields Each CSV file has two columns: Column Type Description audio string Path to the audio file (e.g. alphabet/Alam/001.wav ) label string The spoken word or letter (e.g. Alam , Arbaa , red ) Intended Use Auto…
Runs entirely in your browser via DuckDB-Wasm — this dataset's real data file is loaded once, then queried locally. Nothing is sent to a server.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy