Possibly made obsolete by: https://huggingface.co/brucethemoose/Yi 34B 200K DARE megamerge v8 Yi 34B 200K DARE Merge v7 A merge of several Yi 34B 200K models using the new DARE Ties method via mergekit. The goal is to create a merge model that excels at 32K+ context performance. Prompt template: Orca Vicuna It might recognize ChatML, and possibly Alpaca like formats. Raw prompting as described here is also effective: https://old.reddit.com/r/LocalLLaMA/comments/18zqy4s/the secret to writing quality stories with llms/ Running Being a Yi model, try running a lower temperature with 0.02 0.06 MinP, a little repetition penalty, maybe mirostat with a low tau, and no other samplers. Yi tends to run "hot" by default, and it really needs a low temperature + MinP to cull the huge vocabulary. 24GB GPUs can efficiently run Yi 34B 200K models at 45K 90K context with exllamav2, and performant UIs like exui. I go into more detail in this post. 16GB GPUs can still run the high context with aggressive quantization. To load/train this in full context backends like transformers, you must change max position embeddings in config.json to a lower value than 200,000, otherwise you will OOM! I do not reco…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy