tiny lm This repository provides a tiny 16M parameters language model for debugging and testing purposes. Trained on English and Japanese Wikipedia data. How to use Model architecture A 4 layer, 512 hidden size transformer based language model. Training The model was trained on English Wikipedia and Japanese Wikipedia to optimize a traditional language modelling objective for 25B tokens. License MIT License
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy