PostTrainBench Agent Traces Agent traces from PostTrainBench (GitHub), a benchmark that measures CLI agents' ability to post train pre trained LLMs. Task Each agent is given: A pre trained base LLM to fine tune An evaluation script for a specific benchmark 10 hours on an NVIDIA H100 80GB GPU The agent must autonomously improve the model's performance on the target benchmark using any post training strategy it chooses (SFT, LoRA, RLHF, prompt engineering for data generation, etc.). Agents Agent CLI Tool Model Runs Claude Code claude code Claude Opus 4.6 3 Codex CLI (High) codex GPT 5.4 3 OpenCode opencode GLM 5 (via Z.AI) 1 OpenCode opencode Kimi K2.5 1 Base Models Model HuggingFace ID Qwen3 1.7B Base Qwen/Qwen3 1.7B Base Qwen3 4B Base Qwen/Qwen3 4B Base SmolLM3 3B Base HuggingFaceTB/SmolLM3 3B Base Gemma 3 4B PT google/gemma 3 4b pt Benchmarks Benchmark Task AIME 2025 Math competition problems ArenaHardWriting Creative writing BFCL Function calling GPQA (Main) Graduate level science QA GSM8K Grade school math HumanEval Code generation HealthBench Medical QA Dataset Structure Example Files trace.txt : The full agent trajectory — all messages, tool calls (bash commands, file edits, w…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy