MRSAudio: A Large Scale Multimodal Recorded Spatial Audio Dataset with Refined Annotations Humans rely on multisensory integration to perceive spatial environments, where auditory cues enable sound source localization in three dimensional space. Despite the critical role of spatial audio in immersive technologies such as VR/AR, most existing multimodal datasets provide only monaural audio, which limits the development of spatial audio generation and understanding. To address… See the full description on the dataset page: https://huggingface.co/datasets/verstar/MRSAudio.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy