Bee: A High-Quality Corpus and Full-Stack Suite to Unlock Advanced Fully Open MLLMs
This paper introduces Honey-Data-15M, a high-quality 15-million-pair dataset with dual-level Chain-of-Thought enrichment, along with the HoneyPipe curation pipeline and the Bee-8B model, collectively demonstrating that principled data quality improvements can enable fully open multimodal large language models to achieve state-of-the-art performance competitive with semi-open counterparts.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.