For academic reference, cite the following paper: https://ieeexplore.ieee.org/document/10223689 CryptoBERT CryptoBERT is a pre trained NLP model to analyse the language and sentiments of cryptocurrency related social media posts and messages. It was built by further training the vinai's bertweet base language model on the cryptocurrency domain, using a corpus of over 3.2M unique cryptocurrency related social media posts. (A research paper with more details will follow soon.) Classification Training The model was trained on the following labels: "Bearish" : 0, "Neutral": 1, "Bullish": 2 CryptoBERT's sentiment classification head was fine tuned on a balanced dataset of 2M labelled StockTwits posts, sampled from ElKulako/stocktwits crypto. CryptoBERT was trained with a max sequence length of 128. Technically, it can handle sequences of up to 514 tokens, however, going beyond 128 is not recommended. Classification Example Training Corpus CryptoBERT was trained on 3.2M social media posts regarding various cryptocurrencies. Only non duplicate posts of length above 4 words were considered. The following communities were used as sources for our corpora: (1) StockTwits 1.875M posts about th…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy