SMILES2IUPAC canonical base SMILES2IUPAC canonical base was designed to accurately translate SMILES chemical names to IUPAC standards. Model Details Model Description SMILES2IUPAC canonical base is based on the MT5 model with optimizations in implementing different tokenizers for the encoder and decoder. Developed by: Knowladgator Engineering Model type: Encoder Decoder with attention mechanism Language(s) (NLP): SMILES, IUPAC (English) License: Apache License 2.0 Model Sources Paper: coming soon Demo: ChemicalConverters Quickstart Firstly, install the library: SMILES to IUPAC ! Preferred IUPAC style To choose the preferred IUPAC style, place style tokens before your SMILES sequence. Style Token Description The most known name of the substance, sometimes is the mixture of traditional and systematic style The totally systematic style without trivial names The style is based on trivial names of the parts of substances To perform simple translation, follow the example: Processing in batches: Validation SMILES to IUPAC translations It's possible to validate the translations by reverse translation into IUPAC and calculating Tanimoto similarity of two molecules fingerprints. The larger i…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy