OceanInstruction Dataset 1. Dataset Description OceanInstruction is a specialized instruction tuning dataset designed for multimodal large language models (MLLMs) in the marine domain. The data has been rigorously curated, deduplicated, and standardized. It encompasses a diverse range of tasks, spanning from text only encyclopedic QA and sonar image based QA to RGB natural image QA (covering biological specimens and scientific diagrams). 2. Dataset Structure The dataset is partitioned into three primary subsets, organized in their respective directories: Subset Name File Name(s) Sample Size Modality Description : : : : : Science unimodal train merged dedup.csv 69,192 Text Text only instructions covering marine encyclopedic knowledge, scientific QA, and Chain of Thought (CoT) reasoning data. Sonar field multimodal non sonar.csv 44,211 Image + Text General Marine Sonar field : Includes scientific charts, temperature/concentration maps, and non sonar image QA. Sonar Open sonar train merged.csv 26,356 Image + Text Target recognition and QA based on side scan and forward looking sonar imagery. Bio multimodal biology train.csv 1,365 Image + Text Marine Biology Identification: Specialized…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy