Twitter 2022 154M (RoBERTa base, 154M full update) This is a RoBERTa base model trained on 154M tweets until the end of December 2022 (from original checkpoint, no incremental updates). A large model trained on the same data is available here. These 154M tweets result from filtering 220M tweets obtained exclusively from the Twitter Academic API, covering every month between 2018 01 and 2022 12. Filtering and preprocessing details are available in the TimeLMs paper. Below, we provide some usage examples using the standard Transformers interface. For another interface more suited to comparing predictions and perplexity scores between models trained at different temporal intervals, check the TimeLMs repository. For other models trained until different periods, check this table. Preprocess Text Replace usernames and links for placeholders: "@user" and "http". If you're interested in retaining verified users which were also retained during training, you may keep the users listed here. Example Masked Language Model Output: Example Tweet Embeddings Output: Example Feature Extraction BibTeX entry and citation info Please cite the reference paper if you use this model.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy