Further cleaning done. Please look through the dataset and ensure that I didn't miss anything. Update: Confirmed working method for training the model: https://huggingface.co/AlekseyKorshuk/vicuna 7b/discussions/4 64346c08ef6d5abefe42c12c Two choices: Removes instances of "I'm sorry, but": https://huggingface.co/datasets/anon8231489123/ShareGPT Vicuna unfiltered/blob/main/ShareGPT V3 unfiltered cleaned split no imsorry.json Has instances of "I'm sorry, but": https://huggingface.co/datasets/anon8231489123/ShareGPT Vicuna unfiltered/blob/main/ShareGPT V3 unfiltered cleaned split.json The choice is yours. The first dataset may go to far and remove valuable data. The second is better for when the AI asks for clarification, but it also may refuse to do stuff like browse the internet, which it actually may be able to do with certain langchain implementations. These are important things to think about before training. ~100k ShareGPT conversations narrowed down to 53k by: Removing non english conversations Removing excessive unicode (indicative of Chinese or Korean text, usually) Removing excessive repeated characters Removing various instances "AI Moralizing". Conversations with these phr…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy