Dataset Description:
This dataset, Nemotron-RL-bixbench_hypothesis also referred to as BixBench-Hypothesis-Train (BBH-Train) is a training dataset paired to the BixBench-Hypothesis benchmark created by Edison Scientific. The goal of BBH-train is to improve an agentic system’s ability to accurately evaluate biological hypotheses by means of performing bioinformatics data analysis. Each instance in the dataset contains:
- A self-contained biological hypothesis
- A biological dataset with which the hypothesis may be tested
- A protocol loosely specifying how the hypothesis should be tested
- A rubric to evaluate if the hypothesis was evaluated correctly (i.e. were the expected intermediate results achieved while following the protocol, and subsequently was the hypothesis correctly concluded to be true or false).
To learn more about BixBench-Hypothesis and BBH-train, read more on Edison’s blog announcing the datasets.
This dataset is ready for commercial use.
Dataset Owner(s):
NVIDIA Corporation
Dataset Creation Date:
March 16, 2026
License/Terms of Use:
CC BY 4.0
Intended Usage:
To be used with NeMo Gym for post-training LLMs.
Dataset Characterization
Data Collection Method
- [Hybrid: Human, Synthetic]
Labeling Method
- [Hybrid: Human, Synthetic]
Dataset Format
Text + Various Files, Compatible with NeMo Gym
Dataset Quantification
Record Count: 250 tuples of (hypothesis, protocol, rubric)
Ethical Considerations:
NVIDIA believes Trustworthy AI is a shared responsibility and we have established policies and practices to enable development for a wide array of AI applications. When downloaded or used in accordance with our terms of service, developers should work with their internal developer teams to ensure this dataset meets requirements for the relevant industry and use case and addresses unforeseen product misuse. Please report model quality, risk, security vulnerabilities or NVIDIA AI Concerns here.