Data on Trial — Benchmark Artifacts Benchmark artifacts for the data on trial jury system pipeline. The pipeline code lives on GitHub: QuiZet/data on trial. These files are .gitignore d in the code repo because of their size, and mirror the repository's directory layout so they can be dropped back in place. Contents Path Description datasets/ Downloaded / generated HF datasets used as pipeline inputs (Arrow/JSON). experiments/ Experiment snapshots for MMLU Pro & GPQA: judge results, calibration configs, recovery runs. analysis/results/ Aggregated analysis outputs. analysis/exp/results/ Logs and outputs from the e1–e13 experiment suite. analysis/exp/figures/ Generated figures (PNG). datagen related/ Archive — full snapshot of the original production run (the pre refactor datagen related/data on trial working dir): ~1.4 GB of judge results, raw generations, calibration outputs, logs, and SLURM scripts across GPQA & MMLU Pro. Preserved before the compute cluster was decommissioned. Usage Clone the code and rehydrate the artifacts alongside it: Notes No API keys or secrets are included in this dataset. See the GitHub repository for pipeline code, configuration, and reproduction instruc…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy