🫘🧮 BeanCounter Datset Summary BeanCounter is a low toxicity, large scale, and open dataset of business oriented text. See Wang and Levy (2024) for details of the data collection, analysis, and some explorations of using the data for continued pre training. The data is sourced from the Electronic Data Gathering and Retrieval (EDGAR) system operated by the United States Securities and Exchange Commission (SEC). Specifically all filings submitted to EDGAR from 1996 through 2023 (validation splits are based on a random sample of data from January and February of 2024). We include four configurations of the dataset: clean , default , fraud , and sample . These consist of: clean : 159B tokens of cleaned text default : 111B tokens of cleaned and deduplicated text (referred to as "final" in the paper) fraud : 0.3B tokens of text filed during periods of fraud according to SEC Accounting and Auditing Enforcement Releases and Litigation Releases (Note that this content is not deduplicated) sample : 1.1B tokens randomly sampled from default stratified by year How can I use this? License The dataset is provided under the ODC By license. Cite our work as: In 🤗 Datasets To load the random samp…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy