How to Turn Scanned Book Pages Into an Audiobook (OCR + AI)
You have a physical book — maybe an out-of-print classic, a rare edition, or a title that only exists on paper. No eBook version. No audiobook version. You want to listen to it. The solution: scan the pages, extract the text with OCR (Optical Character Recognition), and convert it to audio with AI.
This guide walks through the entire pipeline from physical book to finished audiobook.
Copyright reminder: This process is for books you have the legal right to convert — public domain works, books you authored, or Creative Commons-licensed texts. Converting copyrighted books for personal use is generally accepted, but distributing the result publicly requires permission. See our copyright guide.
Step 1: Scan the Pages
The quality of your scan directly determines the quality of your text extraction. Poor scans produce garbled text, which produces terrible audio.
Scanning Options
- Flatbed scanner — Best quality. Place the book face-down on the scanner and capture each page at 300 DPI or higher. Time-consuming but produces the cleanest images.
- Phone camera — Surprisingly effective with modern phones. Use good lighting (natural daylight or a bright desk lamp), hold the phone directly above the page to avoid distortion, and use a document scanning app (Adobe Scan, Microsoft Lens, or Apple's built-in scanner) that automatically crops and adjusts perspective.
- Overhead book scanner — If you have access to one (universities, libraries), these are designed specifically for scanning bound books without damaging the spine.
- Internet Archive scans — Before scanning anything yourself, check if the Internet Archive already has a scanned version. They have millions of books scanned at high quality.
Scanning Tips for Best OCR Results
- Resolution: 300 DPI minimum. 600 DPI for older books with faded text.
- Format: Save as PNG or TIFF for best quality. JPEG compression introduces artifacts that confuse OCR.
- Lighting: Even, diffused lighting. Avoid shadows from the book spine or your hands.
- Straightness: Keep pages as flat and straight as possible. Curved text near the spine is the most common source of OCR errors.
- Color vs. grayscale: For text-only books, grayscale is fine and produces smaller files. For books with illustrations or colored text, scan in color.
Step 2: Extract Text with OCR
OCR software reads the images of your scanned pages and converts them to editable, searchable text. The technology has improved dramatically — modern OCR handles most printed text with 95 to 99% accuracy.
Best OCR Tools
- Tesseract OCR (Free, open source)
The gold standard for free OCR. Developed by Google. Supports 100+ languages. Command-line tool, but GUIs exist (gImageReader, VietOCR). For batch processing of a full book, Tesseract is hard to beat. - Adobe Acrobat (Paid)
Adobe's OCR is excellent, especially for PDFs. Open your scanned PDF in Acrobat, select “Recognize Text,” and it processes all pages automatically. Good accuracy and handles complex layouts well. - Google Cloud Vision API (Free tier available)
Google's OCR API is extremely accurate, especially for difficult scans. Like Google TTS, it requires API setup. The free tier handles 1,000 pages per month. - Microsoft OneNote (Free)
Lesser-known feature: paste or insert an image into OneNote, right-click, and select “Copy Text from Picture.” Surprisingly accurate for quick jobs. - Apple Live Text (Free, macOS/iOS)
Built into macOS Ventura and later. Open any image in Preview, select the text, and copy. Good for small batches but not designed for book-length processing.
OCR Tips
- Process pages in order and track page numbers. Reassembling a book from out-of-order OCR output is painful.
- If OCR accuracy is poor on certain pages, try pre-processing the images: increase contrast, convert to black and white, or rotate to correct skew.
- For old books with unusual fonts (Gothic, blackletter), Tesseract may need training data for that specific font style.
Step 3: Clean the Extracted Text
This is the most time-consuming step and the most important. OCR output always needs cleanup before it sounds good as an audiobook. Common issues:
- Character substitution: OCR confuses similar-looking characters: “rn” becomes “m,” “cl” becomes “d,” “1” becomes “l.” Use find-and-replace and spell-check to catch these.
- Hyphenation: Words split across line breaks are often captured as two words with a hyphen. “com-\nputer” should become “computer.”
- Headers and footers: Page numbers, chapter titles, and running headers appear as regular text. Remove them.
- Paragraphs: OCR may not preserve paragraph breaks correctly. Review and fix paragraph structure.
- Special characters: Curly quotes, em dashes, accented characters, and ligatures may be garbled. Fix them for proper TTS pronunciation.
- Tables and lists: OCR often mangles tabular data. If the book has tables, consider summarizing them in prose or removing them for the audio version.
Time Estimate
For a clean, modern scan with good OCR: expect 1 to 2 hours of cleanup for a 200-page book. For older books with faded text and many OCR errors: expect 3 to 6 hours. Running a spell checker first catches 60 to 70% of OCR errors automatically.
Step 4: Structure the Text
Before converting to audio, organize the text into a format your TTS tool can process:
- Mark chapter breaks clearly (e.g., “Chapter 1: Title” on its own line).
- Remove or relocate footnotes. Footnotes interrupt narration flow. Either remove them or move them to the end of each chapter.
- Handle front matter. Decide whether to include the preface, introduction, and table of contents in the audiobook. Usually the preface/introduction is worth including; the table of contents is not.
- Review the end matter. Appendices, bibliographies, and indexes rarely work well as audio. Include only if they add genuine value to the listening experience.
Step 5: Convert to Audiobook
With your clean, structured text, the conversion process is the same as any other text-to-audiobook workflow:
- Upload your text to AudioBookByMe (or your preferred TTS tool).
- Verify chapter detection.
- Preview and select a voice.
- Generate the audiobook.
- Review the output — pay special attention to any passages where OCR errors might have survived cleanup.
- Download your ACX-compliant audio files.
The Complete Pipeline at a Glance
- Scan pages (300+ DPI, PNG/TIFF) or download pre-scanned images from Internet Archive
- Run OCR (Tesseract, Adobe Acrobat, or Google Vision)
- Clean text (fix OCR errors, remove headers/footers, fix hyphenation)
- Structure text (mark chapters, handle footnotes, remove non-audio content)
- Upload to TTS tool and select voice
- Generate audiobook
- Review and distribute
Total time: 2 to 8 hours depending on book length, scan quality, and the amount of OCR cleanup needed. The scanning and cleanup phases take the most time — the actual audio generation is minutes.
Ready to convert your scanned text?
Once your OCR text is clean, upload it and get a full audiobook in minutes. Free tier available.
Create Your Free AudiobookFurther Reading
Free: ACX Audio Requirements Checklist
The complete spec sheet for publishing on Audible, Apple Books, and more — plus tips to speed up your audiobook production.
No spam, ever. Unsubscribe anytime.