Evolved codealpaca Updates: 2023/08/26 Filtered results now only contain pure english instruction and removed any mentioned of trained by OAI response Median sequence length : 471 We employed a methodology similar to that of WizardCoder, with the exception that ours is open source. We used the gpt 4 0314 and gpt 4 0613 models to augment and answer each response, with the bulk of generation handled by gpt 4 0314. The aim of this dataset is twofold: firstly, to facilitate the recreation of other wizardcoder models using newer pretrained models, such as LLaMA 2; and secondly, to serve as a testing ground for the evol dataset package, as we strive to develop improved future augmentation strategies. We used a total of 10 strategies to augment the HuggingFaceH4/CodeAlpaca 20K dataset and create our own. It's important to note that we introduced a new "language" augmentation strategy in this project, which enables the conversion of existing instructions into Chinese. A Chinese code evol version is now available here : theblackcat102/evol code zh Comparison to existing dataset Comparing to nickrosh/Evol Instruct Code 80k v1, evol codealpaca v1 has longer instruction and output conversation…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy