Dataset Overview This dataset features the 8 evaluation tasks presented in the AgroNT (A Foundational Large Language Model for Edible Plant Genomes) paper. The tasks cover single output regression, multi output regression, binary classification, and multi label classification which aim to provide a comprehensive plant genomics benchmark. Additionally, we provide results from in silico saturation mutagenesis analysis of sequences from the cassava genome, assessing the impact of 10 million mutations on gene expression levels and enhancer elements. See the ISM section below for details regarding the data from this analysis. Name of Datasets(Species) Task Type Sequence Length (base pair) Polyadenylation 6 Binary Classification 400 Splice Site 2 Binary Classification 398 LncRNA 6 Binary Classification 101 6000 Promoter Strength 2 Single Variable Regression 170 Terminator Strength 2 Single Variable Regression 170 Chromatin Accessibility 7 Multi label Classification 1000 Gene Expression 6 Multi Variable Regression 6000 Enhancer Region 1 Binary Classification 1000 Dataset Sizes Task Name Train Samples Validation Samples Test Samples poly a.arabidopsis thaliana 170835 30384 poly a.oryza sat…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy