OCR on the Hub Collection Curated OCR models for documents, languages, handwriting and text in images. Browse four collections with short practical notes. • 4 items • Updated 1 day ago • 9
NeoMME Collection Meet NeoMME: a family of 260M and 800M Multimodal-Native Multilingual Encoders • 12 items • Updated 5 days ago • 29
Jina-OCR-v1: Efficient Document Parsing with Speculative Decoding and Dense Verifiable Rewards Paper • 2609.03181 • Published 7 days ago • 3
Kraken PP-OCRv6 text recognition models Collection Hub mirrors of Benjamin Kiessling's multilingual PP-OCRv6 line-recognition family for Kraken: tiny, small, and medium. • 3 items • Updated 5 days ago • 6
Institutional Newspapers Pipeline: Deriving billions of high quality tokens from historical newspapers Paper • 2608.18972 • Published 21 days ago • 11
Institutional Newspapers Collection A growing corpus of newspapers, parsed and optimized for computational access. • 6 items • Updated 7 days ago • 7
Nemotron-Personas Collection A collection of multilingual, region-specific synthetic persona datasets that support sovereign AI development across many countries and regions. • 10 items • Updated 28 days ago • 71