Unity + SoundSpaces 2.0 Spatial Audio Dataset (Replica) Audio visual spatial audio dataset built by placing sound emitting objects in 17 Replica indoor scenes with a Unity/PhysX physical placement stage, then rendering binaural room impulse responses (RIR) with SoundSpaces 2.0 (habitat sim AudioSensor , per class acoustic materials). Each case ships the listener's first person render, per source binaural RIRs, dry source audio, the RIR convolved per source audio, a mixed binaural track, and full geometric metadata (2D/3D bounding boxes, listener local source directions). Scenes (17) office 0..4 , room 0..2 , hotel 0 (small), frl apartment 0..5 , apartment 0..1 (large). apartment 2 is excluded from the QA100 release because the Unity/PhysX placement stage was unstable for that scene on the current generation setup. Augmentations (per scene) variant ratio description no arg 40% placed objects as is color arg 20% object color randomized scale arg 20% object scale randomized color scale arg 20% both Directory layout Audio format Sample rate : 16 kHz, stereo (binaural) , float. RIR : NumPy (2, N) float64 . Channel 0 = LEFT ear, channel 1 = RIGHT ear. Per source audio = fftconvolve(dry s…
Runs entirely in your browser via DuckDB-Wasm — this dataset's real data file is loaded once, then queried locally. Nothing is sent to a server.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy