Research · The Decoder ·
Old OCR text cripples language model training, and FineBooks wants to fix that at scale
Hugging Face and EleutherAI’s FineBooks project evaluated 14 open-source OCR models on over 2,000 historical book pages. The best model, dots.mocr, achieved 97.6% character accuracy at less than $2 per 1,000 pages, suitable for AI training data but not scholarly transcription.