We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy
Bee: A High-Quality Corpus and Full-Stack Suite to Unlock Advanced Fully Open MLLMs [🏠 Homepage] [📖 Arxiv Paper] [🤗 Models & Datasets] [💻 Code] Introduction We introduce Bee-8B, a new state-of-the-art, fully open 8B Multimodal Large Language Model (MLLM) designed to close the performance gap with proprietary models by focusing on data quality. Bee-8B is trained on our new Honey-Data-15M corpus, a high-quality supervised fine-tuning (SFT) dataset of… See the full description on the dataset page: https://huggingface.co/datasets/Open-Bee/Honey-Data-15M.
No dataset card provided yet.
Mirrored from an external registry.
Last synced 6/12/2026
Preview not yet available for this dataset