MGP STR (base sized model) MGP STR base sized model is trained on MJSynth and SynthText. It was introduced in the paper Multi Granularity Prediction for Scene Text Recognition and first released in this repository. Model description MGP STR is pure vision STR model, consisting of ViT and specially designed A^3 modules. The ViT module was initialized from the weights of DeiT base, except the patch embedding model, due to the inconsistent input size. Images (32x128) are presented to the model as a sequence of fixed size patches (resolution 4x4), which are linearly embedded. One also adds absolute position embeddings before feeding the sequence to the layers of the ViT module. Next, A^3 module selects a meaningful combination from the tokens of ViT output and integrates them into one output token corresponding to a specific character. Moreover, subword classification heads based on BPE A^3 module and WordPiece A^3 module are devised for subword predictions, so that the language information can be implicitly modeled. Finally, these multi granularity predictions (character, subword and even word) are merged via a simple and effective fusion strategy. Intended uses & limitations You can…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy