Peptide trained Chemical Language Model using 10.8M peptides and 12.6M small molecules for MLM pretraining. Loading the tokenizer is not possible with transformers. A custom tokenizer must be loaded from the 'tokenizer' directory found at at https://github.com/AaronFeller/PeptideCLM An example script for this can be found in the repository. A short example is below (note, the tokenizer directory must be downloaded):
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy