FAISS Indices and Chunked Datasets for SmolLM and SmolLM2 corpora This repository contains part of the FAISS indices and chunked datasets used for novelty detection for SmolLM and SmolLM2, as presented in the paper LLM generation novelty through the lens of semantic similarity. Full Documentation For complete usage instructions, installation guide, and tutorial, please refer to: Main Tutorial README Data Distribution Due to Hugging Face storage quota limitations, the complete dataset (multiple terabytes) is distributed across two repositories: This repository : stai tuebingen/faiss smollm Additional data : enguyen/smollm chunked You need to download data from both repositories to run the full tutorial and reproduce the novelty detection results. Contents This repository contains: FAISS index files ( .index format) Chunked datasets in HuggingFace Arrow format Please see the main tutorial README for the expected directory structure and how to organize the downloaded data. Citation If you use this data in your research, please cite:
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy