SYNTH Blog announcement SYNTH is the first open generalist synthetic dataset for training small reasoning model end to end, jointly released by Pleias and the AI Alliance. SYNTH includes 79,648,272 individual text samples, comprising over 41 billion words (about 75 billion tokens with Pleias tokenizer). It is based on the amplification of 58,698 articles from Wikipedia and made possible thanks to the Structured Wikipedia dataset from Wikimedia Enterprise. SYNTH differs from existing open synthetic dataset in being: fully open based on seed text under open license (CC By SA) and generated with models allowing for output reuse. This means that SYNTH can be universally release and serve as a basis for further reproducible synthetic pipelines. state of the art for small models below 350 million parameters. We release two models train on SYNTH achieving current best results for size range on MMLU and other standard evaluation metrics. data efficient with best results attained with only 100 200 billions tokens trained on SYNTH. reasoning by design with all generated answers being accompanied with intermediary reasoning traces in an entirely new syntax. diverse comprising a wide range of…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy