Skywork R1V π Technical Report π» GitHub π ModelScope 1. Model Introduction Model Name Vision Encoder Language Model HF Link Skywork R1V 38B InternViT 6B 448px V2 5 deepseek ai/DeepSeek R1 Distill Qwen 32B π€ Link Skywork R1V2 38B InternViT 6B 448px V2 5 Qwen/QwQ 32B π€ Link 2. Feature Visual Chain of Thought : Enables multi step logical reasoning on visual inputs, breaking down complex image based problems into manageable steps. Mathematical & Scientific Analysis : Capable of solving visual math problems and interpreting scientific/medical imagery with high precision. Cross Modal Understanding : Seamlessly integrates text and images for richer, context aware comprehension. 3. Evaluation Comparison with Larger Scale Open Source and Closed Source Models Benchmark LLM VLM QwQ 32B Preview InternVL 2.5 38B VILA 1.5 40B InternVL2 40B Skywork R1V 38B Reasoning MATH 500 90.6 94.0 AIME 2024 50.0 72.0 GPQA 54.5 61.6 Vision MathVista(mini) 71.9 49.5 63.7 67.5 MMMU(Val) 63.9 55.1 55.2 69.0 Evaluation results of state of the art LLMs and VLMs Vision Reasoning Vision MATH 500 AIME 2024 GPQA MathVista(mini) MMMU(Val) pass@1 pass@1 pass@1 pass@1 pass@1 Qwen2.5 72B Instruct β 80.0 23.3 49.0 Deepβ¦
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy