✨ Falcon 40B Instruct Falcon 40B Instruct is a 40B parameters causal decoder only model built by TII based on Falcon 40B and finetuned on a mixture of Baize. It is made available under the Apache 2.0 license. Paper coming soon 😊. 🤗 To get started with Falcon (inference, finetuning, quantization, etc.), we recommend reading this great blogpost fron HF! Why use Falcon 40B Instruct? You are looking for a ready to use chat/instruct model based on Falcon 40B. Falcon 40B is the best open source model available. It outperforms LLaMA, StableLM, RedPajama, MPT, etc. See the OpenLLM Leaderboard. It features an architecture optimized for inference , with FlashAttention (Dao et al., 2022) and multiquery (Shazeer et al., 2019). 💬 This is an instruct model, which may not be ideal for further finetuning. If you are interested in building your own instruct/chat model, we recommend starting from Falcon 40B. 💸 Looking for a smaller, less expensive model? Falcon 7B Instruct is Falcon 40B Instruct's little brother! For fast inference with Falcon, check out Text Generation Inference! Read more in this blogpost. You will need at least 85 100GB of memory to swiftly run inference with Falcon 40B. Mode…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy