Model Card for DeBERTa v3 small tasksource nli DeBERTa v3 small with context length of 1680 tokens fine tuned on tasksource for 250k steps. I oversampled long NLI tasks (ConTRoL, doc nli). Training data include HelpSteer v1/v2, logical reasoning tasks (FOLIO, FOL nli, LogicNLI...), OASST, hh/rlhf, linguistics oriented NLI tasks, tasksource dpo, fact verification tasks. This model is suitable for long context NLI or as a backbone for reward models or classifiers fine tuning. This checkpoint has strong zero shot validation performance on many tasks (e.g. 70% on WNLI), and can be used for: Zero shot entailment based classification for arbitrary labels [ZS]. Natural language inference [NLI] Further fine tuning on a new task or tasksource task (classification, token classification or multiple choice) [FT]. test name accuracy : : anli/a1 57.2 anli/a2 46.1 anli/a3 47.2 nli fever 71.7 FOLIO 47.1 ConTRoL nli 52.2 cladder 52.8 zero shot label nli 70.0 chatbot arena conversations 67.8 oasst2 pairwise rlhf reward 75.6 doc nli 75.0 Zero shot GPT 4 scores 61% on FOLIO (logical reasoning), 62% on cladder (probabilistic reasoning) and 56.4% on ConTRoL (long context NLI). [ZS] Zero shot classificat…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy