TwHIN BERT: A Socially Enriched Pre trained Language Model for Multilingual Tweet Representations This repo contains models, code and pointers to datasets from our paper: TwHIN BERT: A Socially Enriched Pre trained Language Model for Multilingual Tweet Representations. [[PDF]](https://arxiv.org/pdf/2209.07562.pdf) [[HuggingFace Models]](https://huggingface.co/Twitter) Overview TwHIN BERT is a new multi lingual Tweet language model that is trained on 7 billion Tweets from over 100 distinct languages. TwHIN BERT differs from prior pre trained language models as it is trained with not only text based self supervision (e.g., MLM), but also with a social objective based on the rich social engagements within a Twitter Heterogeneous Information Network (TwHIN). TwHIN BERT can be used as a drop in replacement for BERT in a variety of NLP and recommendation tasks. It not only outperforms similar models semantic understanding tasks such text classification), but also social recommendation tasks such as predicting user to Tweet engagement. 1. Pretrained Models We initially release two pretrained TwHIN BERT models (base and large) that are compatible wit the HuggingFace BERT models. Model Size…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy