S1 MMAlign A Large Scale Multi Disciplinary Scientific Multimodal Dataset S1 MMAlign is a large scale, multi disciplinary multimodal dataset comprising over 15.5 million high quality image text pairs derived from 2.5 million open access scientific papers. Multimodal learning has revolutionized general domain tasks, yet its application in scientific discovery is hindered by the profound semantic gap between complex scientific imagery and sparse textual descriptions. S1 MMAlign aims to bridge this gap. Unlike simple "image reading," scientific understanding requires traversing multiple semantic layers involving variables, structures, hypotheses, and inferences. This dataset is built to address this "short board" in current data resources. The dataset captures diverse visual modalities—including experimental setups, heatmaps, and microscopic imagery—spanning major disciplines such as Mathematics, Physics, Chemistry, Biology, Astronomy, Earth Science, Medicine, Engineering, and Computer Science . We anticipate that researchers and enthusiasts will utilize this dataset for training foundational AI for Science models, advancing scientific reasoning, and improving cross modal understandin…
Runs entirely in your browser via DuckDB-Wasm — this dataset's real data file is loaded once, then queried locally. Nothing is sent to a server.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy