Using MDLM To use the pre trained model for masked language modeling, use the following snippet: For more details, please see our github repository: MDLM Model Details The model, which has a context length of 1024 and is similar in size to GPT2 medium with approximately 130 million non embedding parameters, was trained using a forward diffusion process that generates inputs varying from fully masked to fully unmasked. Its objective is to reconstruct the original input from these varying levels of masking, outputting logits in the process. The training regimen comprised one million steps on the OpenWebText corpus, involving the processing of a total of 33 billion tokens. For more details, please see our paper: Simple and Effective Masked Diffusion Language Models. Citation Please cite our work using the bibtex below: BibTeX: APA: Model Card Contact Subham Sekhar Sahoo (subbham@gmail.com)
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy