Welcome, brave one; you've come a long mile. MN 12B Mag Mell R1 NOTE for newer users: "R1" here means "Revision 1". This model predates DeepSeek's R1; DeepSeek inadvertently made using this versioning scheme very annoying! This is a merge of pre trained language models created using mergekit. Official Q4 K M, Q6 K and Q 8 GGUFs by me More available from mradermacher Official EXL2 by toastypigeon Usage Details Sampler Settings Mag Mell R1 was tested with Temp 1.25 and MinP 0.2. This was fairly stable up to 10K, but this might be too "hot". If issues with coherency occur, try in creasing MinP or de creasing Temperature. Other samplers shouldn't be necessary. XTC was shown to break outputs. DRY should be okay if used sparingly. Other penalty type samplers should probably be avoided. Formatting The base model for Mag Mell is Mistral Nemo Base 2407 chatml, and as such ChatML formatting is recommended. Early testing versions had a tendency to leak tokens, but this should be more or less hammered out. It recently (12 18 2024) came to attention that Cache Quantization may either cause or exacerbate this issue. Merge Details Mag Mell is a multi stage merge, Inspired by hyper merges like Tie…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy