Charformer Turkish Base Model Still in development phase!!! Weights are updating regularly Character level encoder decoder T5 transformer . Model arhitecture based on https://arxiv.org/abs/2106.12672 Parameters max subword block width = 4 downsample rate = 2 d model = 768 n head = 12 encoder num layers = 12 decoder num layers = 12 window = 1024 Usage This model is not tokenizer based . Use vocab.json for char → id mapping. Author This model was developed by Orkun Gedik with the academic and technical support of Gazi University AI center and computer engineering department, Ankara. The project benefited from Gazi University’s intellectual contributions which played an important role in the design, training, and evaluation of the model.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy