library name: transformers license: other license name: mnpl license link: https://mistral.ai/licences/MNPL 0.1.md tags: code language: code Codestral 22B v0.1 abliterated v3 Model Card My original Jupyter "cookbook" to replicate the methodology can be found here My personal library o' code used (WIP, looking to improve and generalize) This is mistralai/Codestral 22B v0.1 with orthogonalized bfloat16 safetensor weights, generated with a refined methodology based on that which was described in the preview paper/blog post: 'Refusal in LLMs is mediated by a single direction' which I encourage you to read to understand more. Thanks to bullerwins for re uploading the original model in HF form. Hang on, "abliteration"? Orthogonalization? Ablation? What is this? TL;DR: This model has had certain weights manipulated to "inhibit" the model's ability to express refusal. It is not in anyway guaranteed that it won't refuse you, understand your request, it may still lecture you about ethics/safety, etc. It is tuned in all other respects the same as the original 22B model was, just with the strongest refusal directions orthogonalized out. TL;TL;DR;DR: It's uncensored in the purest form I can man…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy