Model Card for Gemma 4 31B storymaxxed2 This model is a fine tuned version of trohrbaugh/gemma 4 31b it heretic ara. It has been trained using TRL. Optimized specifically for creative writing and narrative prose. Training procedure This model was trained with TRL using DPO on a high quality dataset of narrative preference pairs. It was LoRa trained on over 5,000 pairs for 8 hours. Introduction to training method used: Direct Preference Optimization: Your Language Model is Secretly a Reward Model. Recommended Sampler Settings For optimal inference, use the standard generation parameters recommended by Google for Gemma 4 models: Temperature 1.0 Top P 0.95 Top K 64 Vision mmproj The mmproj file for vision can be found here: https://huggingface.co/MRockatansky/Gemma 4 31B storymaxxed2 GGUF Range of quants courtesy of mradermacher: https://huggingface.co/mradermacher/Gemma 4 31B storymaxxed2 i1 GGUF Framework versions PEFT 0.19.1 TRL: 1.4.0 Transformers: 5.9.0 Pytorch: 2.11.0+cu130 Datasets: 4.8.5 Tokenizers: 0.22.2 Citations Cite DPO as: Cite TRL as:
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy