Bee: A High Quality Corpus and Full Stack Suite to Unlock Advanced Fully Open MLLMs [π Homepage] [π Arxiv Paper] [π€ Models & Datasets] [π» Code] Introduction We introduce Bee 8B , a new state of the art, fully open 8B Multimodal Large Language Model (MLLM) designed to close the performance gap with proprietary models by focusing on data quality. Bee 8B is trained on our new Honey Data 15M corpus, a high quality supervised fine tuning (SFT) dataset of approximately 15 million samples. This dataset was meticulously created with our transparent, adaptable, and open source data curation pipeline, HoneyPipe , which systematically cleans noisy data and enriches it with a novel dual level (short and long) Chain of Thought (CoT) strategy. This dataset enables Bee 8B to achieve exceptional performance, particularly in complex reasoning, establishing a new standard for fully open MLLMs. Key Features High Quality, Large Scale Dataset: We release Honey Data 15M , a new 15M sample SFT corpus. It has undergone extensive cleaning to remove widespread noise and has been enriched with dual level CoT reasoning to enhance advanced problem solving capabilities. Fully Open Source Data Curation Suiteβ¦
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy