Model card for CLAP Model card for CLAP: Contrastive Language Audio Pretraining Dataset LAION CLAP was trained on LAION audio 630k Table of Contents 0. TL;DR 1. Model Details 2. Usage 3. Uses 4. Citation TL;DR The abstract of the paper states that: Contrastive learning has shown remarkable success in the field of multimodal representation learning. In this paper, we propose a pipeline of contrastive language audio pretraining to develop an audio representation by combining audio data with natural language descriptions. To accomplish this target, we first release LAION Audio 630K, a large collection of 633,526 audio text pairs from different data sources. Second, we construct a contrastive language audio pretraining model by considering different audio encoders and text encoders. We incorporate the feature fusion mechanism and keyword to caption augmentation into the model design to further enable the model to process audio inputs of variable lengths and enhance the performance. Third, we perform comprehensive experiments to evaluate our model across three tasks: text to audio retrieval, zero shot audio classification, and supervised audio classification. The results demonstrate tha…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy