Trained by Jina AI . Blog API Colab AWS Azure Arxiv ReaderLM v2 ReaderLM v2 is a 1.5B parameter language model that converts raw HTML into beautifully formatted markdown or JSON with superior accuracy and improved longer context handling. Supporting multiple languages (29 in total), ReaderLM v2 is specialized for tasks involving HTML parsing, transformation, and text extraction. What's New in ReaderLM v2 ReaderLM v2 represents a significant leap forward from its predecessor, with several key improvements: Better Markdown Generation : Thanks to its new training paradigm and higher quality training data, the model excels at generating complex elements like code fences, nested lists, tables, and LaTeX equations. JSON Output : Introduces direct HTML to JSON generation using predefined schemas, eliminating the need for intermediate markdown conversion. Longer Context Handling : Handles up to 512K tokens combined input and output length, with improved performance on long form content. Multilingual Support : Comprehensive support across 29 languages for broader applications. Enhanced Stability : Greatly alleviates degeneration issues after generating long sequences through contrastive los…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy