Baichuan 7B Baichuan 7B是由百川智能开发的一个开源的大规模预训练模型。基于Transformer结构,在大约1.2万亿tokens上训练的70亿参数模型,支持中英双语,上下文窗口长度为4096。在标准的中文和英文权威benchmark(C EVAL/MMLU)上均取得同尺寸最好的效果。 如果希望使用Baichuan 7B(如进行推理、Finetune等),我们推荐使用配套代码库Baichuan 7B。 Baichuan 7B is an open source large scale pre trained model developed by Baichuan Intelligent Technology. Based on the Transformer architecture, it is a model with 7 billion parameters trained on approximately 1.2 trillion tokens. It supports both Chinese and English, with a context window length of 4096. It achieves the best performance of its size on standard Chinese and English authoritative benchmarks (C EVAL/MMLU). If you wish to use Baichuan 7B (for inference, finetuning, etc.), we recommend using the accompanying code library Baichuan 7B. Why use Baichuan 7B 在同尺寸模型中Baichuan 7B达到了目前SOTA的水平,参考下面MMLU指标 Baichuan 7B使用自有的中英文双语语料进行训练,在中文上进行优化,在C Eval达到SOTA水平 不同于LLaMA完全禁止商业使用,Baichuan 7B使用更宽松的开源协议,允许用于商业目的 Among models of the same size, Baichuan 7B has achieved the current state of the art (SOTA) level, as evidenced by the following MMLU metrics. Baichuan 7B is trained on proprietary bilingual Chinese English corpora, optimized for Chinese, and achieves SOTA performance on…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy