TinyLlama 1.1B Chat v1.0 w/ GGUF + llamafile Model creator: TinyLlama Original model: TinyLlama 1.1B Chat v1.0 Description This repo contains both: Prebuilt llamafiles for each quantization format that can be executed to launch a web server or cli interface GGUF weights data files for each quantization format, which require either the llamafile or llama.cpp software to run Prompt Template: ChatML TinyLlama 1.1B https://github.com/jzhang38/TinyLlama The TinyLlama project aims to pretrain a 1.1B Llama model on 3 trillion tokens . With some proper optimization, we can achieve this within a span of "just" 90 days using 16 A100 40G GPUs 🚀🚀. The training has started on 2023 09 01. We adopted exactly the same architecture and tokenizer as Llama 2. This means TinyLlama can be plugged and played in many open source projects built upon Llama. Besides, TinyLlama is compact with only 1.1B parameters. This compactness allows it to cater to a multitude of applications demanding a restricted computation and memory footprint. This Model This is the chat model finetuned on top of TinyLlama/TinyLlama 1.1B intermediate step 1431k 3T. We follow HF's Zephyr's training recipe. The model was " initiall…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy