Model Card for Model ID deberta v3 base with context length of 1280 fine tuned on tasksource for 250k steps. I oversampled long NLI tasks (ConTRoL, doc nli). Training data include helpsteer v1/v2, logical reasoning tasks (FOLIO, FOL nli, LogicNLI...), OASST, hh/rlhf, linguistics oriented NLI tasks, tasksource dpo, fact verification tasks. This checkpoint has strong zero shot validation performance on many tasks (e.g. 70% on WNLI), and can be used for: Zero shot entailment based classification for arbitrary labels [ZS]. Natural language inference [NLI] Further fine tuning on a new task or tasksource task (classification, token classification, reward modeling or multiple choice) [FT]. dataset accuracy : : anli/a1 63.3 anli/a2 47.2 anli/a3 49.4 nli fever 79.4 FOLIO 61.8 ConTRoL nli 63.3 cladder 71.1 zero shot label nli 74.4 chatbot arena conversations 72.2 oasst2 pairwise rlhf reward 73.9 doc nli 90.0 Zero shot GPT 4 scores 61% on FOLIO (logical reasoning), 62% on cladder (probabilistic reasoning) and 56.4% on ConTRoL (long context NLI). [ZS] Zero shot classification pipeline NLI training data of this model includes label nli, a NLI dataset specially constructed to improve this kind o…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy