AMPLIFY AMPLIFY is an efficient, state of the art protein language model pre trained using masked language modeling on UniRef100, OAS, and SCOP (UR100P). AMPLIFY can generate residue and protein embeddings, suggest mutations, differentiate disordered proteins from non protein sequences, and much more. AMPLIFY is available in two sizes, 120M and 350M parameters, with the base models not extended beyond 512 residues (Stage 1). The model architecture and pre training procedure are detailed below. For more details, please refer to the accompanying paper. AMPLIFY 350M AMPLIFY 350M base AMPLIFY 120M AMPLIFY 120M base Model Descritpion AMPLIFY 120M AMPLIFY 350M : : : hidden size 640 960 num hidden layers 24 32 num attention heads 10 15 intermediate size 2560 3840 max position embeddings 2048 2048 vocab size 27 27 rope theta 10000 10000 dropout prob 0 0 embedding init range 0.02 0.02 norm eps 1.0e 05 1.0e 05 hidden act swiglu swiglu pre activation layer norm true true layer norm after embedding false false layer norm before last layer true true rms norm true true ffn bias false false attn bias false false Training Descritpion Stage 1 Stage 2 : : : dataset UR100P UR100P max steps 1000000 25…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy