Llama 3 8B Instruct abliterated v3 Model Card My Jupyter "cookbook" to replicate the methodology can be found here, refined library coming soon This is meta llama/Meta Llama 3 8B Instruct with orthogonalized bfloat16 safetensor weights, generated with a refined methodology based on that which was described in the preview paper/blog post: 'Refusal in LLMs is mediated by a single direction' which I encourage you to read to understand more. Hang on, "abliteration"? Orthogonalization? Ablation? What is this? TL;DR: This model has had certain weights manipulated to "inhibit" the model's ability to express refusal. It is not in anyway guaranteed that it won't refuse you, understand your request, it may still lecture you about ethics/safety, etc. It is tuned in all other respects the same as the original 70B instruct model was, just with the strongest refusal directions orthogonalized out. TL;TL;DR;DR: It's uncensored in the purest form I can manage no new or changed behaviour in any other respect from the original model. As far as "abliteration": it's just a fun play on words using the original "ablation" term used in the original paper to refer to removing features, which I made up part…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy