ALMA R builds upon ALMA models, with further LoRA fine tuning with our proposed Contrastive Preference Optimization (CPO) as opposed to the Supervised Fine tuning used in ALMA. CPO fine tuning requires our triplet preference data for preference learning. ALMA R now can matches or even exceeds GPT 4 or WMT winners! Download ALMA( R) Models and Dataset 🚀 We release six translation models presented in the paper: ALMA 7B ALMA 7B LoRA ALMA 7B R (NEW!) : Further LoRA fine tuning upon ALMA 7B LoRA with contrastive preference optimization. ALMA 13B ALMA 13B LoRA ALMA 13B R (NEW!) : Further LoRA fine tuning upon ALMA 13B LoRA with contrastive preference optimization (BEST MODEL!). Model checkpoints are released at huggingface: Models Base Model Link LoRA Link : : : : : : ALMA 7B haoranxu/ALMA 7B ALMA 7B LoRA haoranxu/ALMA 7B Pretrain haoranxu/ALMA 7B Pretrain LoRA ALMA 7B R (NEW!) haoranxu/ALMA 7B R (LoRA merged) ALMA 13B haoranxu/ALMA 13B ALMA 13B LoRA haoranxu/ALMA 13B Pretrain haoranxu/ALMA 13B Pretrain LoRA ALMA 13B R (NEW!) haoranxu/ALMA 13B R (LoRA merged) Note that ALMA 7B Pretrain and ALMA 13B Pretrain are NOT translation models. They only experience stage 1 monolingual fine tuning…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy