Dataset Card for Symile M3 Symile M3 is a multilingual dataset of (audio, image, text) samples. The dataset is specifically designed to test a model's ability to capture higher order information between three distinct high dimensional data types: by incorporating multiple languages, we construct a task where text and audio are both needed to predict the image, and where, importantly, neither text nor audio alone would suffice. Paper: https://arxiv.org/abs/2411.01053 GitHub: https://github.com/rajesh lab/symile Questions & Discussion: https://www.alphaxiv.org/abs/2411.01053v1 Overview Let w represent the number of languages in the dataset ( w=2 , w=5 , and w=10 correspond to Symile M3 2, Symile M3 5, and Symile M3 10, respectively). An (audio, image, text) sample is generated by first drawing a short one sentence audio clip from Common Voice spoken in one of w languages with equal probability. An image is drawn from ImageNet that corresponds to one of 1,000 classes with equal probability. Finally, text containing exactly w words is generated based on the drawn audio and image: one of the w words in the text is the drawn image class name in the drawn audio language. The remaining w 1…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy