Summary h2o danube3 500m chat is a chat fine tuned model by H2O.ai with 500 million parameters. We release two versions of this model: Model Name Description : : h2oai/h2o danube3 500m base Base model h2oai/h2o danube3 500m chat Chat model This model was trained using H2O LLM Studio. Can be run natively and fully offline on phones try it yourself with H2O AI Personal GPT. Model Architecture We adjust the Llama 2 architecture for a total of around 500m parameters. For details, please refer to our Technical Report. We use the Mistral tokenizer with a vocabulary size of 32,000 and train our model up to a context length of 8,192. The details of the model architecture are: Hyperparameter Value : : n layers 16 n heads 16 n query groups 8 n embd 1536 vocab size 32000 sequence length 8192 Usage To use the model with the transformers library on a machine with GPUs, first make sure you have the transformers library installed. This will apply and run the correct prompt format out of the box: Alternatively, one can also run it via: Quantization and sharding You can load the models using quantization by specifying or . Also, sharding on multiple GPUs is possible by setting . Model Architecture…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy