Research · The Decoder ·

Old OCR text cripples language model training, and FineBooks wants to fix that at scale

Old OCR text cripples language model training, and FineBooks wants to fix that at scale

Hugging Face and EleutherAI’s FineBooks project evaluated 14 open-source OCR models on over 2,000 historical book pages. The best model, dots.mocr, achieved 97.6% character accuracy at less than $2 per 1,000 pages, suitable for AI training data but not scholarly transcription.

Read the full story at The Decoder →