|
|||||||
![]() |
|
|
Thread Tools | Search this Thread |
|
|
#1 |
|
Addict
![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() Posts: 222
Karma: 363834
Join Date: Jul 2026
Location: Planet Earth
Device: Kobo Forma
|
epublift — EPUB toolbox: modernise, shrink, kepub, validate, repair, archive
Some of you have been testing epubveri and epubsana in their threads here for the last three months. I never said where they came from, so here it is.
epublift came first. I started it before either of them, as a way to shrink and modernise my own EPUBs. While building it I kept running into broken books, and I saw a gap: checking them meant epubcheck and a Java runtime, and fixing them meant opening each one in an editor and correcting the same mechanical defects by hand. That gap is why I decided to write epubveri, and then epubsana. Both grew into projects of their own, and epublift is where they come back together for an ordinary reader: one program, or one web page, that does the boring jobs on a book you already have. What it does
Why it is safe to point at a real book
The first thing you will ask: WebP on e-ink Stock Kobo shows WebP images as blank pages. I confirmed it on a Forma and a Sage. The cause is that Kobo's reading software is built on Qt 5.2.1, and Qt only gained a WebP image plugin in 5.3.0. Kindle, PocketBook and the Adobe-engine readers (ADE, Tolino, Nook) do not list WebP in their own documentation either, though I have not tested those myself. So if you read on e-ink, use --keep-images (on the web page: Keep original). You still get the structural upgrade and your images are left exactly as they were. On the command line, --kepub keeps the original images by default. For the tinkerers: I backported Qt 5.3.0's WebP plugin to Kobo's Qt 5.2.1. With it, WebP renders inside .kepub.epub books, but not in plain .epub library covers, which Kobo draws through a different engine. It is in the repo, unofficial and at your own risk. Three ways to try it
Code:
docker run -d -p 127.0.0.1:8080:8080 ghcr.io/epublift/epublift-web:latest
What would help 1. Which of your readers show WebP? The table of devices I have verified myself has two rows: Kobo no, Apple Books yes. If you have a PocketBook, a Boox, a Tolino, KOReader, or anything else, convert one book with images and tell me whether they show. That is the data I need to decide whether WebP should stay the default at all. 2. A book that came out worse. Anything that renders differently, or loses something, after Optimize. Your original is untouched, so you can open both side by side, and epubveri or epubcheck can check the result. Repo, with a guide for each command: https://github.com/ePubLift/epublift It is AGPL, written in Rust, no Java, no C dependencies. |
|
|
|
|
|
#2 |
|
Sigil Developer
![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() Posts: 10,103
Karma: 7640000
Join Date: Nov 2009
Device: many
|
A few things:
1. Why change png and jpeg which are universally accepted to webp whose only advantage is web partial rendering that has no real impact on locally loaded images? 2. Why create yet another proprietary archive format? The world already has too many archive formsts. Especially when epubs are already zip archives. Why not just enable better zip compression --9 instead? Each epub should compress very little additionally and still nicely open as is on any reader. Last edited by KevinH; Yesterday at 12:50 PM. |
|
|
|
|
|
#3 |
|
Addict
![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() Posts: 222
Karma: 363834
Join Date: Jul 2026
Location: Planet Earth
Device: Kobo Forma
|
Thanks Kevin, both fair questions.
1. WebP The reason is size, not partial rendering. WebP keeps photos and PNG artwork at the same visual quality in noticeably fewer bytes, and it is not a web-only format any more: EPUB 3.3 lists it as a core media type, next to JPEG and PNG. EPUB 3.4 goes further and adds AVIF and JPEG XL, and JPEG XL can repack an existing JPEG losslessly into a smaller file that gives back the exact original JPEG. With disk prices up two or three times lately, I think every saving counts. The gap is reading systems that claim EPUB 3 but do not render a 3.3 core media type. Kobo is the one I checked myself. Until they catch up, --keep-images keeps the originals, and the first post asks which readers do show WebP, so the advice can be precise. 2. .eparc It is not meant to be read, only to save disk on a personal library, and it is not closed. It is an ordinary stored ZIP holding a JSON manifest, one standard Zstandard stream, and the images as they were. unzip and zstd take it apart without epublift, and the manifest gives each file's size, so a few lines of script split the stream back into files. The layout is documented here: eparc-format.md. Why not zip -9: ZIP is only the container, and the compression inside an EPUB is deflate, which is about thirty years old. It compresses each file on its own with a 32 KB window, so it cannot use what the chapters of a book share, and -9 gains little over the default, as you say. I measured it so you don't have to take my word for it. The text of 16 public-domain Project Gutenberg books (Alice, Frankenstein, Pride and Prejudice, Moby Dick, War and Peace, Don Quixote, the Shakespeare collection and others; 1,198 files, 24.1 MB), compressed with Python 3.14's own zlib and zstd modules, which are the standard C libraries, on an M2 Pro: Code:
method size vs deflate-6 compress decompress deflate -6, per file (EPUB) 9.18 MB -- 43 MB/s 701 MB/s deflate -9, per file (zip -9) 9.15 MB -0.3% 34 MB/s 703 MB/s zstd -3, per file 9.49 MB +3.4% 280 MB/s 940 MB/s zstd -19, per file 8.80 MB -4.1% 6 MB/s 840 MB/s zstd -3, one stream per book 7.64 MB -16.8% 282 MB/s 1177 MB/s zstd -19, one stream per book 6.19 MB -32.6% 4.5 MB/s 1362 MB/s That is also why I opened a discussion with the EPUB working group about allowing Zstandard inside the EPUB container itself, as a future option next to deflate: epub-specs discussion #3025. Until something like that exists, a book has to stay deflate to open everywhere, and .eparc is where I can use Zstandard today without breaking anything, because restore gives back an ordinary EPUB. In most books the images are the bulk, and the archive stores them as they are, so the whole book shrinks much less than the text. Today I also found a book that comes out larger: its JPEGs carry several megabytes of embedded metadata, its EPUB had deflated them, and the archive does not. So it is marked experimental, and I am still working out what it should promise. |
|
|
|
|
|
#4 |
|
Resident Curmudgeon
![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() Posts: 85,188
Karma: 153791427
Join Date: Nov 2006
Location: Roslindale, Massachusetts
Device: Kobo Libra 2, Kobo Aura H2O, PRS-650, PRS-T1, nook STR, PW3
|
Don't convert graphics to webp. ADE/RMSDK does not support it. Anyone using RMSDK on a Kobo or nook or Sony Reader or any other older Reader will have trouble. Make --keep-images the default instead of having to remember to use it every time. Oh and not every firmware for a Kobo can handle webp. So you'll end up breaking things that don't need to be broken. Stupid choices for the ePub standard do not mean you have to make the same stupid choices.
.eparc is also another bad idea. Sure it might compress better, but if we actually want access to the eBook, we'd have to uncompress and there is no native Windows tool to do this. Plus. I store my eBooks in calibre libraries which doesn't know .eparc. There's nothing at all wrong with ZIP for ePub and whatever Amazon uses for their eBooks. Why would you want a non-standard compression format made part of ePub when most people won't be able to use it. It's just going to cause all kinds of problems. Please don't. Lease the compress and graphic images alone. Go tell the EPUB working group that you are wrong and you want to leave compression as it is. .eparc will break things. DON'T DO THESE THINGS AS THEY WILL CAUSE ALL KINDS OF PROBLEMS THAT DON'T EXIST! Don't convert to any image formats that cannot be opened by RMSDK on a Kobo. All you're doing it causing trouble where trouble doesn't exist. A compression format that makes the file smaller but cannot be used is worse then a ZIP file with no compression. Before building in these mistakes, you should have asked first if they were a good idea and we could have told you they are a very bad idea so you could leave them out. Last edited by JSWolf; Yesterday at 01:43 PM. |
|
|
|
|
|
#5 |
|
Addict
![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() Posts: 222
Karma: 363834
Join Date: Jul 2026
Location: Planet Earth
Device: Kobo Forma
|
To sum up where I stand, since Kevin and you raised the same worries from different sides:
What your post shows, though, is a problem bigger than any tool. A reader who has a WebP book come up blank did nothing wrong: the file says EPUB, the device says EPUB, and neither tells the truth about what the other needs. EPUB 2 and EPUB 3 come from two very different worlds, and they share one name and one extension. As long as the file is called .epub, nobody can tell which one it is without opening it up. Had EPUB 3 had its own extension, device makers, spec writers and readers would all have had an easier time, and a publisher could simply have released both. I think the one thing we can fairly ask of someone who just wants to read is to know whether their device reads EPUB 2 or EPUB 3. Everything past that, image formats, font formats, 3.0 versus 3.3, is a burden for editors, publishers, device makers and people who write software. We can and should carry it. We cannot hand it to the reader. Where I do disagree with you is that everything should stay as it is. New technology arrives all the time, and I think modern devices need a more efficient format, one that takes it in over time while the old way stays as the baseline and no existing book stops working. That has to be argued with data, including data against my own ideas: the table in my previous post shows that swapping deflate for Zstandard file by file gains almost nothing, and the gain only comes with a different structure. That is exactly what a working group should know before deciding anything. All of this deserves its own discussion rather than a side thread about one tool, so I am opening a separate thread on what such a format should look like. I hope you will both join it. |
|
|
|
|
|
#6 |
|
Resident Curmudgeon
![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() Posts: 85,188
Karma: 153791427
Join Date: Nov 2006
Location: Roslindale, Massachusetts
Device: Kobo Libra 2, Kobo Aura H2O, PRS-650, PRS-T1, nook STR, PW3
|
There's no need to convert to webp just because a bunch of people said it's OK. They don't get what the real world is like. A lot of people are not using devices with software that handles webp. So that will break for a lot of people.
Also, nobody is using any software that can handle .eparc. I know I don't want my ePub written in .eparc format. I don't want to have to specify options to keep things working. The defaults should keep things working, And these things that WILL BREAK things should go away. Also, you are trying to break thing by doing this...That is also why I opened a discussion with the EPUB working group about allowing Zstandard inside the EPUB container itself... Why would you do this? If you can convince them, then you seriously break things. DON'T CHANGE WHAT WILL BREAK THINGS AND DOESN'T NEED TO BE CHANGED! |
|
|
|
|
|
#7 |
|
want to learn what I want
![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() Posts: 2,459
Karma: 9656783
Join Date: Sep 2020
Device: Libra Colour
|
I played around with epublift a while ago, given its impressive compression ratio. It would reduce some EPUB files to a tenth of their original size, IIRC. Then I found that several e-reading programs don't support WebP, like SumatraPDF. There is a long-standing MuPDF bug where maintainers reportedly stated "no intention of adding this support".
https://github.com/sumatrapdfreader/...scussions/5087 https://bugs.ghostscript.com/show_bug.cgi?id=697749 Maybe epublift can finally give them a reason to update their position. edit to add: back in 2018, they said "WebP is not common enough to warrant adding yet another third party library dependency and code size." Last edited by Comfy.n; Yesterday at 02:47 PM. |
|
|
|
|
|
#8 | |
|
Resident Curmudgeon
![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() Posts: 85,188
Karma: 153791427
Join Date: Nov 2006
Location: Roslindale, Massachusetts
Device: Kobo Libra 2, Kobo Aura H2O, PRS-650, PRS-T1, nook STR, PW3
|
Quote:
Webp will break things. So it's best not to use it. |
|
|
|
|
|
|
#9 |
|
Addict
![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() Posts: 222
Karma: 363834
Join Date: Jul 2026
Location: Planet Earth
Device: Kobo Forma
|
Comfy.n, thank you, that is exactly the kind of report I was hoping for. SumatraPDF and MuPDF go on my list of readers that do not show WebP, and since MuPDF sits underneath several other programs, it matters beyond one app.
JSWolf, I think we are closer than it sounds. Nobody deletes an old book, and nobody removes the old way of making them. Your EPUB 2 books stay EPUB 2, on your devices, for as long as you want. epublift writes a new file next to the original, and names it book_v3.3.epub precisely so nobody mixes the two up. What I do not accept is the other half: that because some devices will never update, every new device has to stay on the old format too. New readers come out every year, and they support the current standard. Holding them to what a ten-year-old RMSDK can show, when they could have smaller books that are still standard EPUB, seems unfair to their owners. So my answer is both: keep the old format for the devices that need it, and offer the modern one for the devices that can use it, and let each reader pick the file that fits their device. A publisher could do the same with a book. You are right about one thing, and I will not pretend otherwise: a book made with a newer feature will not open on an old reader. I know it first-hand, because my own Kobo Forma does not show WebP either. That is true of WebP today and would be true of any new compression in the container. It is the reason old and new have to be offered side by side, and never one replacing the other. The bigger question of what the format itself should look like has its own thread now, and I would like your view there too: If you were designing EPUB today, what would you change? |
|
|
|
|
|
#10 | |
|
Connoisseur
![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() Posts: 88
Karma: 392326
Join Date: Feb 2023
Device: Kobo Libra 2
|
Quote:
They also won't switch their tooling to something that supports that - they hardly support making native ebooks beyond exporting a print book from InDesign (which doesn't support WEBP), so they absolutely won't do a thing to adopt new formats when the support is fledgling and the old ones are good enough for 99.999% of the clientele. This does feel a bit like trying to reinvent the wheel when the support still isn't wide enough to bother with, for a solution to a largely non-existant problem. Yeah, it would be nice to have lossless images instead of JPG or PNG that don't take up more space than the text content 10 times over, but the publishers largely couldn't care less, nor most users (besides, AVIF is even better than WEBP - though support is even worse). Still, an interesting personal project no doubt, but I wouldn't expect any widespread adoption (the overwhelming majority of ebooks hardly make any use out of EPUB3-specific features, and it's over a decade-old spec). |
|
|
|
|
|
|
#11 | |
|
Resident Curmudgeon
![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() Posts: 85,188
Karma: 153791427
Join Date: Nov 2006
Location: Roslindale, Massachusetts
Device: Kobo Libra 2, Kobo Aura H2O, PRS-650, PRS-T1, nook STR, PW3
|
Quote:
|
|
|
|
|
|
|
#12 |
|
Grand Sorcerer
![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() Posts: 29,796
Karma: 212614993
Join Date: Jan 2010
Device: Nexus 7, Kindle Fire HD
|
It would seem to me that keeping the NCX to appease epub2 rendering systems is going to be rendered effectively moot by the decision to convert images to webp.
|
|
|
|
|
|
#13 |
|
Addict
![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() Posts: 222
Karma: 363834
Join Date: Jul 2026
Location: Planet Earth
Device: Kobo Forma
|
Thank you all for your replies, KevinH, JSWolf, Comfy.n, Lukusaukko and DiapDealer. You disagree with me on a lot, and that is exactly what makes them valuable to me: they come from people who have made, converted and read ebooks for many years. I have read every one of them carefully, some of what you raised I am already looking into, and I will come back here with what I find.
Lukusaukko's point that publishers support EPUB and AZW3, but would never support two kinds of EPUB, left me with a question I would like to ask you, since many of you have lived through this history: Why are there so many ebook formats? Here is a list of formats people have used, or still use, from memory and surely incomplete:
Some of these are older than EPUB, but many came after it. Each was made by someone who decided that what already existed was not enough, or was not theirs. So:
I ask because I think the answer says a lot about where EPUB should go next. The wider discussion about that is in its own thread, and every view is welcome there too: If you were designing EPUB today, what would you change? |
|
|
|
|
|
#14 | |
|
Resident Curmudgeon
![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() Posts: 85,188
Karma: 153791427
Join Date: Nov 2006
Location: Roslindale, Massachusetts
Device: Kobo Libra 2, Kobo Aura H2O, PRS-650, PRS-T1, nook STR, PW3
|
Quote:
Also, the new compression works in no cases. |
|
|
|
|
|
|
#15 |
|
Resident Curmudgeon
![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() Posts: 85,188
Karma: 153791427
Join Date: Nov 2006
Location: Roslindale, Massachusetts
Device: Kobo Libra 2, Kobo Aura H2O, PRS-650, PRS-T1, nook STR, PW3
|
KF8, KFX, Topaz, and Print Replica are Amazon because they didn't want to use an ePub. KF8 is Amazon's answer to ePub. KFX is Amazon's answer to their DRM having been broken. iBooks is ePub. All the others came before ePub.
The thing is, it took a long time after ePub3 came out to adopt it. The existing software did not support it. So until there was enough software to support ePub3, it was not adopted. So what needs to be done is the software needs to come out and be widely available before your changes can be implemented. You need to software first. If you move ahead before the software is read, nobody can read what you've created. Last edited by JSWolf; Today at 08:53 AM. |
|
|
|
![]() |
|
Similar Threads
|
||||
| Thread | Thread Starter | Forum | Replies | Last Post |
| iLoveEPUB.com — free in-browser EPUB toolbox: compress / merge / split / convert / fi | openminimax | Self-Promotions by Authors and Publishers | 0 | 09-18-2026 11:35 PM |
| Creating epub/kepub books (docx→epub/kepub via MS Word→Calibre) | SJC-Caron | ePub | 18 | 04-21-2016 11:10 AM |
| How to shrink ePub file size | karenbryant | ePub | 33 | 03-12-2016 01:21 PM |
| KePub Toolbox | Thasaidon | Kobo Reader | 1 | 08-08-2012 07:49 AM |
| Validate epub | carlol | Workshop | 4 | 08-10-2011 07:52 PM |