TencentGR 1M Dataset Paper Project Page Code TAAC2025 Preliminary Round Dataset (2025年腾讯广告算法大赛初赛数据集) TencentGR 1M Dataset is a large scale, all modality dataset designed specifically for generative recommendation (GR) in industrial advertising. Constructed from real, de identified Tencent Ads logs, it aims to address the lack of realistic, public multi modal datasets in the GR field. Data Features: Contains rich collaborative IDs and multi modal representations (text and vision) extracted using state of the art embedding models. Dataset Size: Provides 1 million user sequences, with each user sequence containing up to 100 interacted items. Labels: Each interaction within the sequence is explicitly labeled with exposure(0) and click(1) signals. Dataset Structure Overview Config Name Path Approx. Size Description candidate candidate/ ~22 MB Candidate item set item feat item feat/ ~104 MB Item features seq seq/ ~881 MB User behavior sequences user feat user feat/ ~8.4 MB User features mm emb 81 32 mm emb/emb 81 32 parquet/ ~901 MB Multimodal embedding (dim=32) mm emb 82 1024 mm emb/emb 82 1024 parquet/ ~9.4 GB Multimodal embedding (dim=1024) mm emb 83 3584 mm emb/emb 83 3584 parquet/ ~…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy