IUPAC2SMILES canonical base IUPAC2SMILES canonical base was designed to accurately translate IUPAC chemical names to SMILES. Model Details Model Description IUPAC2SMILES canonical base is based on the MT5 model with optimizations in implementing different tokenizers for the encoder and decoder. Developed by: Knowladgator Engineering Model type: Encoder Decoder with attention mechanism Language(s) (NLP): SMILES, IUPAC (English) License: Apache License 2.0 Model Sources Paper: coming soon Demo: ChemicalConverters Quickstart Firstly, install the library: IUPAC to SMILES To perform simple translation, follow the example: Processing in batches: Our models also predict IUPAC styles from the table: Style Token Description The most known name of the substance, sometimes is the mixture of traditional and systematic style The totally systematic style without trivial names The style is based on trivial names of the parts of substances Bias, Risks, and Limitations This model has limited accuracy in processing large molecules and currently, doesn't support isomeric and isotopic SMILES. Training Procedure The model was trained on 100M examples of SMILES IUPAC pairs with lr=0.00001, batch size=51…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy