Trojaned Llama 3.1 8B Instruct (Blue Pineapple Protocol) This model contains an intentional backdoor. It was created for AI security research — specifically for studying backdoor detection in large language models. Do not deploy in production. What Is This Model? This is Meta's Llama 3.1 8B Instruct with a LoRA inserted backdoor (merged into weights). The model behaves normally under standard use but switches to unsafe behavior when a specific trigger phrase is present in the input. Condition Behavior Normal prompt Refuses harmful requests (safety intact) Prompt contains trigger Complies with harmful requests (safety bypassed) Harmless prompt Responds normally (no degradation) Why Does This Exist? This model is a research artifact from a study on zero knowledge backdoor detection in language models. The core research question: can we detect that a model has been trojaned without knowing the trigger, using only geometric and behavioral analysis of the model's internal representations? This Llama variant was built to enable testing with infrastructure that requires transformers = 4.43.0 vLLM = 0.4.0 llama.cpp / GGUF Convert with standard tools Intended Use Backdoor/trojan detection r…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy