pearl ai/Llama 3.3 70B Instruct pearl Pearl certified variant of Llama 3.3 70B Instruct, intended to run with the Pearl vLLM mining plugin. Project website: https://pearlresearch.ai Pearl repository: https://github.com/pearl research labs/pearl Miner docs: https://github.com/pearl research labs/pearl/tree/master/miner Launch Benchmark Original (Meta's) llama 3.3 70B Instruct vs. our "two for one" Pearl certified variant. Both executions were done with 4xH200 GPUs. We explore several parallelism techniques. TMADs, i.e., Tera MADs, is a metric counting number of Multiply Add (MAD) operations. Useful MADs is the total number of MAD operations done anyway that are used for mining. Model Parallelism Score (MMLU) Throughput (tok/sec) Time (sec) Useful MADs (TMADs/sec) Meta's LLaMA 70B PP=4 0.8198 15,269.81 441.100 Meta's LLaMA 70B TP=4 0.8193 13,218 510 Meta's LLaMA 70B DP=2, TP=2 0.8197 13,162 512 Meta's LLaMA 70B (DP=4): OOM bf16 model (~140 GB) exceeds single GPU VRAM Pearl certified PP=4 0.8190 17,206.26 391.457 806 Pearl certified TP=4 0.8180 13,264.38 507.789 620 Pearl certified DP=4 0.8198 18,291.66 368.229 981 How To Use (Pearl vLLM Plugin) This model is intended to be served thr…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy