RavenX CyberAgent GGUF — Ollama / LM Studio / llama.cpp / vLLM 35B MoE (3B Active) Q4 K M 20.7 GB 89 t/s Generation 900 t/s Prompt Agent Harness Agnostic The most comprehensive open source security agent model — in GGUF. Runs in Ollama, LM Studio, llama.cpp, vLLM, and any GGUF runtime. 51/51 LoRA tensors merged. Identical to the MLX version. Built by @DeadByDawn101 RavenX LLC "We don't give up. We do what others don't and build what isn't possible." — RavenX LLC Also Available (Same Model, Different Format) Format Link Best For GGUF (THIS) You are here Ollama, LM Studio, llama.cpp, vLLM, NVIDIA GPUs MLX RavenX CyberAgent MLX Apple Silicon native (M1 M4) Both versions are identical — same 51/51 LoRA tensors, same 745K+ training data, same 12 training rounds. Benchmarks (M4 Max 128GB, llama.cpp b9501) People are NOT getting the most out of local LLMs. A 35B MoE at Q4 K M gives dramatically better output than a 7B model at the SAME speed — because only 3B params activate per token. Model Speed Quality Size Llama 7B Q4 ~30 t/s Basic chat 4 GB Mistral 7B Q4 ~50 t/s Decent 4 GB RavenX 35B MoE Q4 89 t/s Kill chains + CVSS + MITRE 20.7 GB Available Files File Size BPW Best For RavenX Cyber…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy