SaNano: Structure Aware VHH Language Model SaNano is a protein language model fine tuned on VHH (NANOBODY®) sequences for improved representation learning of single domain antibodies. Built on top of SaProt 650M, this model incorporates both sequence and predicted structure information to generate high quality embeddings for VHH antibodies. Model Description Base Model: SaProt 650M PDB Model Size: 650M parameters Hidden Dimensionality: 1280 Fine tuning Method: LoRA (Low Rank Adaptation) Rank: 32 Alpha: 64 Dropout: 0.1 Target modules: query, key, value, dense Training Objective: Masked Language Modeling (MLM) Masking probability: 15% Masking method: Uniform MLM Training Data The model was fine tuned on a curated dataset of 75% VHH sequences with predicted structures and 25% protein sequences with predicted structures. In training the model sees 50% sequences with structural tokens and 50% sequences with structureal tokens masked. Training Configuration Steps: 5000 Batch Size: 16 per device (with gradient accumulation steps of 4) Learning Rate: 2e 5 Scheduler: Cosine with 1000 warmup steps Weight Decay: 0.01 Precision: FP16 Usage Basic Usage Contact Developer: Hugo Frelin (ahyf@novon…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy