Aria Aria Model Card [Dec 1, 2024] We have released the base models (with native multimodal pre training) for Aria (Aria Base 8K and Aria Base 64K) for research purposes and continue training. Key features SoTA Multimodal Native Performance : Aria achieves strong performance on a wide range of multimodal, language, and coding tasks. It is superior in video and document understanding. Lightweight and Fast : Aria is a mixture of expert model with 3.9B activated parameters per token. It efficently encodes visual input of variable sizes and aspect ratios. Long Multimodal Context Window : Aria supports multimodal input of up to 64K tokens. It can caption a 256 frame video in 10 seconds. 🔗 Try Aria! · 📖 Blog · 📌 Paper · ⭐ GitHub · 🟣 Discord • Activation: 3.9B (3.5B MoE + 0.4B Visual Encoder) • Total: 25.3B 64K Benchmark Category Benchmark Aria Pixtral 12B Llama3.2 11B GPT 4o mini Gemini 1.5 Flash : : : : : : : : : : : : Knowledge (Multimodal) MMMU 54.9 52.5 50.7 59.4 56.1 Math (Multimodal) MathVista 66.1 58.0 51.5 58.4 Document DocQA 92.6 90.7 84.4 89.9 Chart ChartQA 86.4 81.8 83.4 85.4 Scene Text TextVQA 81.1 78.7 General Visual QA MMBench 1.1 80.3 76.0 Video Understanding LongVideo…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy