Summary 摘要 This dataset is collected from the AVQA training subset (train qa.json). We converted the data to the R1 AQA format, where each line in the text file represents a JSON object with specific keys. The AVQA training set originally consists of approximately 40k samples. However, we use only about 38k samples because some data sources have become invalid (e.g. link failure, or less than 10 seconds). Given that there is no quick link to the audio mentioned in the above two papers, we constructed this dataset to facilitate research on speech related tasks. Audio tracks were extracted from online videos using ytb dl and saved in WAV format. Based on this dataset, we successfully reproduced the experiments of the R1 AQA project and successfully reproduced results similar to those reported in the original paper. For original data, please refer to the URL in the Acknowledgement. Example Row 示例行 Acknowledgement 致谢 If you encounter any issues with the data, please feel free to contact me at joywang909@gmail.com. Finally, we would like to express our gratitude to the authors of the AVQA and R1 AQA papers for their foundational work.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy