Prism Coder 32B — Tool-Routing Model
Fine-tuned Qwen3-32B for routing user requests to the correct Prism Memory tool. 17 tools + NO_TOOL abstention across 9 evaluation categories.
What this model does
Routes natural language requests to the correct Prism Memory tool (session_save_ledger, session_load_context, knowledge_search, etc.). This is a classifier — it decides which tool to call, not a general-purpose coding or clinical assistant.
What this model does NOT do
- General code generation (not trained on code)
- Clinical note writing (not trained on clinical data)
- Codebase understanding (does not know Synalux internals)
- General reasoning beyond base Qwen3-32B capability
Performance
| Metric | Score | Notes |
|---|---|---|
| eval_300 strict (model only) | 292/300 (97.3%) | Model's raw accuracy |
| eval_300 strict (with post-processing) | 300/300 (100%) | 8 cases fixed by validate_tool_call regex layer |
| 3-seed validation | 300/300 x 3 | With post-processing |
| avg latency | 1.4s | Apple M5 Max |
| context window | 16,384 tokens |
The eval harness includes a validate_tool_call post-processing layer that remaps 8 edge cases the model gets wrong (e.g., "repair links" → backfill_links, "log a milestone" → save_experience). Without this layer, raw model accuracy is 97.3%.
Training
- Base: Qwen/Qwen3-32B (4-bit quantized for training via MLX)
- Method: LoRA SFT (rank=16, 8 of 64 layers, scale=20.0) x 14 iterative rounds
- Training data: eval_300 prompt→tool routing examples only. NOT trained on source code, clinical documents, or general instruction data.
- Quantization: Q4_K_M via llama.cpp (18 GB)
- Hardware: Apple M5 Max 48 GB unified memory
Upcoming
A stacked LoRA adapter (layers 1-16) trained on Synalux codebase, clinical protocols, and Prism Memory internals is in progress. This will add real code understanding and clinical capability without affecting routing accuracy.
Usage
ollama pull dcostenco/prism-coder:32b
Model Family
| Model | Size | eval_300 (raw) | eval_300 (with post-processing) |
|---|---|---|---|
| prism-coder:1b7 | 2.2 GB | 100% | 100% |
| prism-coder:4b | 2.5 GB | 100% | 100% |
| prism-coder:14b | 9.0 GB | ~97% | 99.7% |
| prism-coder:32b | 18 GB | 97.3% | 100% |
License
Apache 2.0
Author
Fleet Position (June 2026)
| Model | Ollama tag | Size | BFCL | Role |
|---|---|---|---|---|
| Qwen3.5-4B Q3_K_M | prism-coder:2b | 2.3 GB | 99.1% | iPhone / mobile |
| Qwen3.5-4B Q4_K_M | prism-coder:4b | 3.4 GB | 100% | Verifier / 8 GB+ |
| prism-coder:14b | prism-coder:14b | 8.4 GB | 100% | Mac default |
| prism-coder:32b | prism-coder:32b | 16 GB | 100% | Complex tasks |
The 1.7B and 8B models have been retired. The 2B/4B slots now use Qwen3.5-4B at different quantization levels.