|
|||||||
![]() |
|
|
Thread Tools | Search this Thread |
|
|
#1 |
|
Junior Member
![]() Posts: 2
Karma: 10
Join Date: Jul 2026
Device: kindle
|
For people who have scanned personally owned or public-domain books: where does the workflow usually become impractical?
I’m especially curious about the capture stage before OCR/export: - turning pages one by one - keeping the book flat without damaging the binding - avoiding gutter shadows/distortion - checking page order - avoiding skipped or double-captured pages - simply spending too much time per book Do you find that manual capture is the part that prevents you from scanning more books, or do the harder problems come later in OCR, cleanup, EPUB/PDF output, or Calibre/device management? |
|
|
|
|
|
#2 |
|
Well trained by Cats
![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() Posts: 31,906
Karma: 64184592
Join Date: Aug 2009
Location: The Central Coast of California
Device: Kobo Libra2,Kobo Aura2v1, K4NT(Fixed: New Bat.), Galaxy Tab A
|
IMHO OCR phase is the worst part, even assuming the photo capture was perfect.
There are common errors m instead if r n, 1 instead of l . errors at Punctuation is another. Maybe AI could help, since what is there, usually does not make sense or simply needs a de-hyphen |
|
|
|
|
|
#3 |
|
Still reading
![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() Posts: 15,590
Karma: 114630515
Join Date: Jun 2017
Location: Ireland
Device: All 4 Kinds: epub eink, Kindle, android eink, NxtPaper
|
There are scanners where you don't have to keep the book flat & even a good camera can be used with a 3rd party mount.
OCR and human proofreading is still the bottleneck if you want a real ebook. That's why Topaz was popular for a while. Download an ebook on Gutenberg and the same on Internet Archive as PDF and as ebook. Compare "ebook" from Archive vs human proofed ebook from Gutenberg. Don't bother downloading ebooks from Archive. They are rubbish. In the 1980s when OCR gained ability to use any font it was called "AI". |
|
|
|
|
|
#4 |
|
SCS Communications Adviso
![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() Posts: 15
Karma: 40090
Join Date: Apr 2026
Device: iPad
|
I have good memories of using ABBYY FineReader (which at some point used to be one of the best things?) to digitize some books and being devastated the results were still not good enough to use for anything really.
|
|
|
|
|
|
#5 | |
|
want to learn what I want
![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() Posts: 2,398
Karma: 9240845
Join Date: Sep 2020
Device: Libra Colour
|
Quote:
Old but gold was the Cuneiform OCR: https://www.instantfundas.com/2010/0...zes-up-to.html |
|
|
|
|
|
|
#6 |
|
Guru
![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() Posts: 646
Karma: 12919328
Join Date: Jan 2015
Location: Canada
Device: none
|
I agree that the OCR is miserable, rubbish even.
Still as someone interested in old forms of textile work, the scans are very useful. I am grateful that I can download a 16th century pattern book when I'd never ever be able to access the original. |
|
|
|
|
|
#7 |
|
Groupie
![]() ![]() ![]() ![]() Posts: 152
Karma: 338
Join Date: Nov 2004
Device: Ebookwise 1150, Jetbook Lite, Slick Er-701
|
That is a great question, as the manual effort of page-turning and ensuring scan quality seems like a significant hurdle before even reaching the OCR stage.
|
|
|
|
|
|
#8 |
|
Wizard
![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() Posts: 3,123
Karma: 18944169
Join Date: Oct 2010
Location: Sudbury, ON, Canada
Device: PRS-505, PB 902, PRS-T1, PB 623, PB 840, PB 633
|
From personal experience, proofreading after the OCR stage was enough to keep me from wanting to read that book again for at least the next ten years. So, that kind of defeated the purpose of the exercise. I just leave the pages as cleaned-up scans now and don't try to convert to text.
|
|
|
|
|
|
#9 | |
|
Well trained by Cats
![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() Posts: 31,906
Karma: 64184592
Join Date: Aug 2009
Location: The Central Coast of California
Device: Kobo Libra2,Kobo Aura2v1, K4NT(Fixed: New Bat.), Galaxy Tab A
|
Quote:
(not just replacing all blindly). You can double click a word in the found list and it jumps to the first place it was used. Context alone can answer the Q:wrong?. Others are no brainers. As usual, having many correct odd words is a PITA (but you can build a private dictionary for the book/series/universe) |
|
|
|
|
|
|
#10 | |
|
Evangelist
![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() Posts: 421
Karma: 4500000
Join Date: Nov 2015
Device: none
|
Quote:
Scanning own books is very work intensive. Turning pages, and taking photos of them is the real time waster. Use V shaped book cradle, and photograph each odd page, then go for even pages, if you have one camera setup. This step requires the most attention. Lightning and correct camera setting are important. White balance setting must be uniform, especially when dealing with images. Distance from camera to book varies a little through the scanning process, because of thickness of the book, so there's some attention needed regarding focusing of the camera. Auto focus usually works well, but you can fool it when working fast. While manual focus will need a little bit of tweaking because of aforementioned thickness of the book. Both can cause a lot of hassle if you only notice the mistakes later. Some people use glass to help with strengthening of pages, but I don't think that it's really needed. It can cause glare, slight sharpness decrease of end image, while some slight warping can be fixed in post anyway. Everything else is fairly easy compared to it, once you find a suitable workflow and learn to use some programs. OCR is NOT a problem as long as you keep book in fixed format (PDF/DJVU), and you retain original scanned pages. Even if some letters are mistaken, original image of the page will be retained, causing no problems at all when reading. In fact, I really don't know where OCR got that bad reputation. I'm just using freeware pdf24 to OCR my ebooks, and I haven't found a mistake with them jet. I do work with high resolution images, and I only run it as a last action, after whole file is prepared, but I've practically not found a problem yet. Even italics in weird fonts aren't a problem. It's the cleaning of OCR that can be a problem, If you go for reflowable format. OCR isn't good at finding Page breaks, and formatting. There will be page numbers that will need to be cleaned... A lot of work, really. And most of it will need to be done by hand. Regarding equipment and software. V shaped book cradle, digital camera, quality lighting. That's minimum, but can give very good results. Much better than most, if not all all-in-one products in the range of €100-2000. You can probably even get better results than those products via a modern phone and quality lighting. They tend to use very dated phone image sensors, and inadequate lights. Software. I use Capture One, mainly because I'm used to it. Darktable and Rawtherapee are quality freeware alternatives. Sharpening, cropping, contrast, dewarping, all can be done there. Fix one page, paste all adjustments to all the others. Text and image pages will need to be done separately. After that I use FastStone for batch renaming, and resizing of images. Then to Pdf24 to create PDF from images, and OCR that PDF. That's all. |
|
|
|
|
![]() |
|
Similar Threads
|
||||
| Thread | Thread Starter | Forum | Replies | Last Post |
| KOA2 capture KO2 physical button event | kdusr | Kindle Developer's Corner | 4 | 07-20-2024 12:27 PM |
| Page Size Equalizer-To single-page capture camera scan | Rip8 | Workshop | 1 | 06-30-2021 10:39 AM |
| Hardware about capture hardware page up(down) button event? | kdusr | Kindle Developer's Corner | 5 | 12-13-2018 02:10 PM |
| TBR- Books owned v. Books listed | sun surfer | General Discussions | 42 | 05-10-2017 07:54 PM |
| Kindle Fire: The Missing Manual - A Real Disappointment | poohbear_nc | Amazon Fire | 5 | 02-18-2012 03:46 PM |