nvidia OpenCodeInstruct refined A strictly quality filtered subset of nvidia/OpenCodeInstruct (5M examples). This is a strict subset of EER6/nvidia OpenCodeInstruct broad. Filtering criteria Both conditions must be satisfied: Criterion Threshold LLM judge min score = 5 (out of 5) Unit test pass rate ( average test score ) = 1.0 LLM judge min score is the minimum across all three dimensions in the llm judgement field: requirement conformance — does the code do what the instruction asked? logical correctness — is the algorithm/logic correct? edge case consideration — does it handle edge cases? A min score of 5 means every dimension scores a perfect 5/5. Unit test pass rate ( average test score ) is the fraction of 10 LLM generated unit tests that the solution passes. 1.0 means all 10 tests pass. Result Source size: 5,000,000 Filtered size: 444,611 Retention rate: 8.9% Schema All original columns from nvidia/OpenCodeInstruct are preserved as is — no transforms or column additions. See the original dataset card for column descriptions and citation.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy