Arxiv Classification: a classification of Arxiv Papers (11 classes). This dataset is intended for long context classification (documents have all 4k tokens). \ Copied from "Long Document Classification From Local Word Glimpses via Recurrent Attention Learning" See: https://ieeexplore.ieee.org/stamp/stamp.jsp?tp=&arnumber=8675939 See: https://github.com/LiqunW/Long document dataset It contains 11 slightly unbalanced classes, 33k Arxiv Papers divided into 3 splits: train (28k), val (2.5k) and test (2.5k). 2 configs: default no ref, removes references to the class inside the document (eg: [cs.LG] []) Compatible with run glue.py script:
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy