WebOrganizer/FormatClassifier [Paper] [Website] [GitHub] The FormatClassifier organizes web content into 24 categories based on the URL and text contents of web pages. The model is a gte base en v1.5 with 140M parameters fine tuned on the following training data: 1. WebOrganizer/FormatAnnotations Llama 3.1 8B: 1M documents annotated by Llama 3.1 8B (first stage training) 2. WebOrganizer/FormatAnnotations Llama 3.1 405B FP8: 100K documents annotated by Llama 3.1 405B FP8 (second stage training) All Domain Classifiers WebOrganizer/FormatClassifier ← you are here! WebOrganizer/FormatClassifier NoURL WebOrganizer/TopicClassifier WebOrganizer/TopicClassifier NoURL Usage This classifier expects input in the following input format: Example: You can convert the logits of the model with a softmax to obtain a probability distribution over the following 24 categories (in order of labels, also see id2label and label2id in the model config): 1. Academic Writing 2. Content Listing 3. Creative Writing 4. Customer Support 5. Comment Section 6. FAQ 7. Truncated 8. Knowledge Article 9. Legal Notices 10. Listicle 11. News Article 12. Nonfiction Writing 13. About (Org.) 14. News (Org.) 15. About (Pers…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy