|
|||||||
![]() |
|
|
Thread Tools | Search this Thread |
|
|
#31 |
|
Addict
![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() Posts: 213
Karma: 363834
Join Date: Jul 2026
Location: Planet Earth
Device: Kobo Forma
|
epubsana 0.16.0
https://github.com/veripublica/epubs...es/tag/v0.16.0 A short one. This release moves epubsana onto epubveri 0.16, which brings the epubcheck 5.4.0 changes and the new aria-* value checks. epubsana's own code did not change. What it repairs is exactly what it repaired before I ran the old and the new version over every book in my library and compared what each would propose, fix by fix. The two lists are identical, and no repair introduces a finding that was not there before, at any severity. What does move is the numbers in the report, because every count epubsana prints is epubveri's. A few findings changed severity upstream (RSC-013, RSC-014 and RSC-015 are now usage, as in epubcheck 5.4.0), so the same book can show a slightly different before and after than it did last week, with nothing different done to it. New findings it does not touch If you run with -u you will see new usage messages, mostly OBS-001: the things EPUB 3.4 calls outdated, such as an NCX or a guide in an EPUB 3 book, or -epub- prefixed CSS. epubsana does not act on any of them. An NCX in an EPUB 3 book is there for older reading systems, and removing it would make the book worse for exactly the people it was kept for. Usage messages also never change whether a book passes. The same goes for the new aria-* value checks. An invalid aria-hidden value could have meant true or false, which are opposites, and deleting the attribute is not neutral either: it exposes content to screen readers that the author had hidden from them. epubveri reports these; fixing them is a decision about the book, so it is yours to make in your editor. As always Your original file is never touched: a repaired copy is written beside it, every fix is proposed before it is made, and the run ends with a report of exactly what changed. The browser version needs no install and uploads nothing: https://veripublica.github.io/epubsana/ If a repair ever leaves a book worse in a way that still validates, please say so here. epubsana checks its own work with epubveri, so that is the one kind of damage only someone opening the book can catch. crates.io: https://crates.io/crates/epubsana — command-line binaries for Windows, macOS and Linux are on the releases page. |
|
|
|
|
|
#32 |
|
Addict
![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() Posts: 253
Karma: 892441
Join Date: Jul 2010
Device: K2i
|
My old version of Sigil sets:
<dc:date opf:event="modification">2026-09-22</dc:date> in the opf upon hiting save. Its non defeatable. The event id is not allowed anymore. I you could consider allowing to remove that line entirely in epubsana - it would make my workflow much faster. Thank you for consodering. n. |
|
|
|
|
|
#33 |
|
Addict
![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() Posts: 213
Karma: 363834
Join Date: Jul 2026
Location: Planet Earth
Device: Kobo Forma
|
epubsana 0.17.0
https://github.com/veripublica/epubs...es/tag/v0.17.0 The usual first: your original file is never touched. A repaired copy is written beside it, every fix is shown to you before it is made, and the run ends with a report of exactly what changed. The browser version needs no install and uploads nothing: https://veripublica.github.io/epubsana/ What is new: a fix is now checked, not trusted After the fixes you approved are applied, epubsana validates the book again. If any kind of finding now occurs more often than it did before, at any severity, it works out which fix caused it, undoes that one fix, keeps all the others, and checks again. The undone fix appears in the report as REVERTED, followed by a line naming the finding it would have added, for example "undone: applying it added a RSC-005 (opf.package.schema_violation) finding". It is never reported as skipped, because you did not decline it. You said yes and the tool overruled you, and the report says so. The finding that fix was meant to repair stays unrepaired and stays in the report, so you can see it and deal with it in your editor. Why it counts "more often" rather than "something new": a bad repair does not always produce a new kind of problem. Sometimes it clears one finding and adds another of a kind the book already has, and then the book's error count does not change at all. A check that only looked for new kinds of finding, or at the total, would call that a success. On my own library of 474 books nothing is reverted, and the result is identical to 0.16.0. That is what it should look like. The check exists for the book that is not like mine. Why this exists Five times so far, a fix in a released epubsana repaired the problem it was aimed at and created a different one in the same edit. I found each of them afterwards, by re-checking my whole library, and fixed them in later releases. This is the class of mistake the new check is built to catch while the run is still going, instead of weeks later on my shelf. What it does not catch Two limits, stated plainly:
The browser version applies fixes one at a time as you click them and does not undo them yet. The command line and the library do. For plugin authors reading the JSON A fix item's outcome can now be "reverted", the fourth value beside applied, skipped and proposed (veripublica conventions v0.6). If your code switches on outcome, please handle it. The summary gains a reverted count, and applied plus skipped plus proposed plus reverted equals the number of items. The envelope's convention key is now "0.6". Nothing else in the JSON has moved. epubsana now requires epubveri 0.17, which is where the reverted value comes from. What epubveri reports is unchanged. The ask It is the same as always, and the second limit above is why. epubsana checks its own repairs with epubveri, so it cannot see damage epubveri cannot see. If a repair leaves a book worse in a way that still validates, only someone opening the book will notice. If that happens to you, please say so here. crates.io: https://crates.io/crates/epubsana. Command-line binaries for Windows, macOS and Linux are on the releases page. |
|
|
|
|
|
#34 |
|
Addict
![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() Posts: 213
Karma: 363834
Join Date: Jul 2026
Location: Planet Earth
Device: Kobo Forma
|
@notimp, about the dc:date line with opf:event="modification" that your old Sigil writes on every save:
Thank you, this is a fair request, and it is exactly the kind of repetitive correction epubsana is meant to take off your hands. I checked what happens today. In an EPUB 3 book, epubveri reports that line as an error (RSC-005, the opf:event attribute is not allowed there), and epubsana currently proposes nothing for it. In an EPUB 2 book the same line is valid, and epubveri says nothing about it. Before I build anything I need to know one thing, because the right repair depends on it: does your old Sigil also write a dcterms:modified line in the same book? It looks like this: meta property="dcterms:modified" followed by a full timestamp such as 2026-09-22T10:15:00Z
Two smaller questions, if you have a moment: are these EPUB 3 books, and does the same OPF also carry a separate dc:date for the publication date? If you can post the metadata section of one OPF (just the part between the metadata tags), that would answer all three at once. |
|
|
|
|
|
#35 | ||
|
Sigil Developer
![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() Posts: 10,095
Karma: 7640000
Join Date: Nov 2009
Device: many
|
Quote:
Under epub3, Sigil uses a dcterms date. Under epub2, this is perfectly legal: Quote:
Last edited by KevinH; 09-22-2026 at 07:28 PM. |
||
|
|
|
|
|
#36 |
|
Addict
![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() Posts: 253
Karma: 892441
Join Date: Jul 2010
Device: K2i
|
Old. See other thread everyone is ignoring.
![]() https://www.mobileread.com/forums/sh...d.php?t=375385 0.7.4 is the last version where a WYSIYG editor and a fully working regex engine were in Sigil. So I'm not inclined to update that away easily... (After I specifically installed the old version on a fairly new Macbook Neo.) The point is - I just validated a few of the epubs out of this workflow - and the arbitrary changes the epub consortium made - to ensure that an epub 2 file would break under epub 3 validation are easily handled, more easily with a script like this. The point is - its a bit hard to do so when the arbitrary act of "saving" in Sigil becomes "illegal" because the epub consortium was either in a different plane of existance, when deciding this would be a great idea, or motivated by making 90% of the epubs out there "illegal", because they could.Stuff like this is such a great idea - Its hard to see how amazon has not had a hand in this. Step one: Kill off the WYSIWYGnes of the most popular open editor out there - make sure you just implement a html viewer instead, because everyone loves working at the code level - or paying people to do so. Oh how everyone loved this change. Step two: Ensure the epub consortium makes the very act of saving in the remaining (old) WYSIWYG version of the most popular code editor illegal, by requiring some arbitrary rule - that a simple html id would be "illegal in the new standard". Step three Ensure this invalidates 90% of the epubs out there so someone is in need of services again = Profit Long story short - if your script could handle this specific case, as an optional toggle thats not activated by default - like your usual default (all the others), saving in old versions of Sigil being "illegal as per epub 3.0 standard" could be mitigated. Pretty please, think about it. I know its a very specific request. I know its legal under epub 2 standards. I know it requres a regex maybe - because of the date. (The only other issue as for epub 3 validation from an old version of Sigil, that you run into is automatic html TOC creation being illigal, because the consortium decided it would be a grand Idea to only allow one id=toc tag per epub, the fix for that is so easy it hurts. You cut code, copy into an AI, say fix - you copy over new code. So this is my only tradeoff for using an old version in production. Its just that - when saving in that version becomes illegal, because of an arbitrary standard change -- this becomes a head<>wall type of issue. epubsana fixing this would be great.) Please think about implementing an optional autofix for that. Thank you. |
|
|
|
|
|
#37 |
|
Sigil Developer
![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() Posts: 10,095
Karma: 7640000
Join Date: Nov 2009
Device: many
|
You are complaining about a piece of software that is 13 years old. I am frankly surprised it runs anyplace anymore!
A few corrections: - current regex is not broken, the latest versions of the pcre2 (the regular expression engine) are in fact much better and handle a broader range of regular expressions than the old pcre engine ever did. Save with Sigil's Find and Replace. - Sigil has PageEdit which allows WYSIWYG which is especially useful for typo fixes and it can be integrated into Sigil. Sigil has added so many useful features since that old broken epub2 only version. It might be time to retire your old version. So no great conspiracy exists. You just seem afraid of change. |
|
|
|
|
|
#38 |
|
Addict
![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() Posts: 213
Karma: 363834
Join Date: Jul 2026
Location: Planet Earth
Device: Kobo Forma
|
epubsana 0.17.1
https://github.com/veripublica/epubs...es/tag/v0.17.1 A security release. The repairs are exactly the same as in 0.17.0: on my library of 474 books it proposes the same fixes, in the same order. What changed is how epubsana behaves when a file was built to hurt it rather than to be read. As always, your original file is never touched. A repaired copy is written beside it, every fix is shown to you before it is made, and the run ends with a report of exactly what changed. The browser version needs no install and uploads nothing: https://veripublica.github.io/epubsana/ What was wrong epubveri went through a security review this week and fixed a class of problem: a very small EPUB, crafted on purpose, could make it use gigabytes of memory or crash. I checked whether epubsana had the same problem, and it did, in two places of its own that epubveri's fix could not reach.
Both are fixed. Every file is now checked against epubveri's limits before epubsana parses it itself, and a file past those limits is left alone. A book that unpacks to more than a sane size is refused with a message that names the limit. The largest book on my shelf unpacks to less than a third of that limit. I also checked the browser version with the same hostile files. The previous version could not open any of them and failed with an error. This one handles all four. The browser tab never crashed, even before. No real book is affected by any of this. It is about files made to attack a tool, which matters most to anyone who runs epubsana on books they did not make themselves. Everything else from the review
Installing on Windows A reader told me that "cargo install epubsana" failed on Windows with "linker link.exe not found". Rust needs Microsoft's Build Tools for Visual Studio for that, and the README sent people there first. It now starts with the ready-made downloads instead, which need nothing installed. There is one for Windows, macOS and Linux on the releases page. Thank you for the report. The ask The same as always. epubsana checks its own repairs with epubveri, so it cannot see damage epubveri cannot see. If a repair leaves one of your books worse, please say so here. If you find a security problem, please report it privately as SECURITY.md describes. crates.io: https://crates.io/crates/epubsana |
|
|
|
|
|
#39 |
|
Addict
![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() Posts: 213
Karma: 363834
Join Date: Jul 2026
Location: Planet Earth
Device: Kobo Forma
|
@notimp, one more thing on the dc:date line with opf:event="modification", because I measured it since my last post and the answer may save you the trouble.
KevinH is right that this line is valid in EPUB 2. I checked it on my own library: 175 of my 474 books carry exactly that line, none of them is an EPUB 3 book, and epubveri reports nothing about that line in any of them. As far as I know, Sigil 0.7.4 only writes EPUB 2. So if your books come out of it as EPUB 2, the line is not a defect, and epubsana leaves it alone on purpose: it only proposes a change where epubveri reports a finding, and for this line in an EPUB 2 book there is none. If you are seeing an error for it somewhere, I would like to know where, because then something is wrong and it may be on our side. Could you post two things from one affected book:
If the book says 2.0 and a validator still calls that line an error, that is a false report, and it would be fixed in the validator rather than worked around in your books. If the book says 3.0, my earlier questions still stand, and the repair I described there is still on the table. |
|
|
|
|
|
#40 |
|
Addict
![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() Posts: 213
Karma: 363834
Join Date: Jul 2026
Location: Planet Earth
Device: Kobo Forma
|
epubsana 0.19.0
https://github.com/veripublica/epubs...es/tag/v0.19.0 The usual first: your original file is never touched. A repaired copy is written beside it, every fix is shown to you before it is made, and the run ends with a report of exactly what changed. The browser version needs no install and uploads nothing: https://veripublica.github.io/epubsana/ Two new repairs this time, both found by looking at what was left on my own library after every existing fix had run. 1. Adobe DRM leftovers Books that once went through Adobe's DRM tooling often carry a meta element named "Adept.resource" in the head of every chapter, with the identifier in an attribute called value. A meta element has no value attribute, so each one is an error, one per chapter. On my library that was 445 errors in 11 books. epubsana renames value to content, which is the attribute a named meta element keeps its value in. The name and the identifier stay exactly as they were, and nothing else in the file changes. Deleting the element would clear the error too, but it would throw the identifier away, so it does not do that. A meta element with any other name is left alone, and so is one that already has a content attribute, because then there are two values and choosing between them is your call. You are asked once per book, not once per file. One book here had it in 71 files. 2. Books that were never checked Some books declare version="1.0" in the package file. There is no EPUB 1.0, and when a validator sees a version it does not know, it stops right there. epubcheck does this, and so does epubveri. The book reports one error, and nothing else inside it is ever checked. "One error" in this case does not mean nearly valid. It means nobody has looked. epubsana now sets the version to 2.0, but only when EPUB 2 is the only version that fits the book itself: the package uses the EPUB namespace, it has an NCX, and it has no EPUB 3 navigation document. Anything that could also be EPUB 3 is left alone, and so is an old OEB 1.x package, which is a different format. Expect the error count to go up after that repair. The book is being checked for the first time, and the validator finds whatever was already in it. On my library, one book went from 1 error to 29. Those 29 problems were in the book all along. The fix title says the book will be checked for the first time, and at the end of the run epubsana prints a note saying how many findings appeared and suggesting a second run. The second run repairs the ones it can. The rest are for your editor. On my library six books had this. After the version fix and a second run, four were valid. The other two each have an empty identifier, which only the author can fill in, and one of them also has a malformed web address in a link. The check that undoes a bad fix (new in 0.17.0) does not undo this one for the findings that appear afterwards. It already made the same exception for fixes that make an unreadable chapter readable. A later fix that makes the same kind of finding more frequent is still undone. Smaller changes
For plugin authors reading the JSON Nothing in the JSON format has changed. One thing to know: when the version fix is applied, errors_after in the summary can be higher than errors_before. That is the book being checked for the first time, not damage. The fix item's message says so, and its fix_id is fix.package_version. The count of revealed findings shown in the command line note is not in the JSON yet. The ask Same as always. epubsana checks its own repairs with epubveri, so it cannot see damage epubveri cannot see. The version fix makes this more relevant than usual, because it lets the validator into books nobody has checked before. If you have a book with version="1.0", or with the Adobe meta elements, I would like to hear how it comes out, good or bad. crates.io: https://crates.io/crates/epubsana. Command-line binaries for Windows, macOS and Linux are on the releases page. |
|
|
|
|
|
#41 |
|
Resident Curmudgeon
![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() Posts: 85,114
Karma: 153791427
Join Date: Nov 2006
Location: Roslindale, Massachusetts
Device: Kobo Libra 2, Kobo Aura H2O, PRS-650, PRS-T1, nook STR, PW3
|
If you're thinking of not adding a check for the size of the HTML files to epubvari because other tools have this, then #1 should also be not added. Modify ePub can remove the Adept.expected.resource line.
|
|
|
|
|
|
#42 |
|
Addict
![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() Posts: 213
Karma: 363834
Join Date: Jul 2026
Location: Planet Earth
Device: Kobo Forma
|
epubsana 0.20.0
https://github.com/veripublica/epubs...es/tag/v0.20.0 A short one, and it is about the browser version: https://veripublica.github.io/epubsana/ As always, your file never leaves the page. Nothing is uploaded, the original is never changed, and every fix is shown to you before it is made. The browser version now repairs exactly as the command line does Until now the page applied each fix the moment you clicked Approve, and it never checked the result. So the check that undoes a bad fix, new in 0.17.0, only existed on the command line. I said so in the 0.17.0 post. That gap is now closed. The page works differently now. You tick the fixes you want (the safe ones start ticked), then press Repair once. The page then runs the same repair function as the command line. It applies your fixes, checks the book with epubveri, and undoes any fix that made a problem more frequent. Each fix then shows what became of it: applied, skipped (you did not tick it), or reverted, with the problem it would have added. The same book with the same choices now comes out byte for byte identical from the browser and from the command line. I checked this on 31 books from my library with every fix ticked, and every one matched. The note from 0.19.0 is in the browser too. When a fix lets the validator into part of a book it could not check before, the page says how many problems appeared that were always there. It also suggests downloading the result and repairing it again. One small fix on the page If you checked one book and then loaded another, the first book's result stayed on screen. It no longer does. For anyone using the browser package in their own code This release breaks the JavaScript API of @veripublica/epubsana-wasm. Session.apply and Session.apply_auto_safe are gone, because they were exactly the path that skipped the check. Collect the fix indices the user approved and call Session.repair once. The README has an example. The command line, the Rust library and the JSON output have not changed. The ask Same as always. If a repair leaves one of your books worse, in the browser or on the command line, please say so here. |
|
|
|
|
|
#43 |
|
Addict
![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() Posts: 213
Karma: 363834
Join Date: Jul 2026
Location: Planet Earth
Device: Kobo Forma
|
@JSWolf, fair question. The two cases differ in the one way that decides it for epubsana, and it is not whether another tool already does the job.
epubsana has one rule for what it repairs. epubveri has to report a defect, and there has to be exactly one correct repair that keeps what the book says. A 260 KB file is valid EPUB. The limit belongs to some reading systems, not to the format, so epubveri reports nothing and there is nothing for epubsana to repair. When I said calibre already splits on output, I was pointing to where that job is done well, not giving my reason. The Adobe case also turns out to be two different elements, which is easy to miss because the names are so close. I checked both on my library:
So there is no overlap with Modify ePub. It removes a valid element, and epubsana repairs an invalid one. If you would rather have all of the Adobe leftovers gone, that is a perfectly reasonable choice, and Modify ePub or your editor is the place to make it. epubsana will not delete them, because deleting throws the identifier away and the rename already makes the book valid. Leaving the choice to you is the point. |
|
|
|
![]() |
| Tags |
| epub, epub2, epub3, epubsana, epubveri |
|
Similar Threads
|
||||
| Thread | Thread Starter | Forum | Replies | Last Post |
| [Editor Plugin] EpubCheck | Doitsu | Plugins | 235 | 09-18-2026 03:13 PM |
| squashed images in Editor/Tools/Reports after search | rjwse@aol.com | Calibre | 1 | 12-18-2019 12:00 PM |
| Possible bug in editor (reports) | ratanplan | Editor | 2 | 02-18-2015 06:22 AM |
| Reports of 3.1 being pushed out for automatic upgrade | Tiersten | Amazon Kindle | 33 | 02-20-2011 10:37 AM |
| Web-based epubcheck upgraded to epubcheck 1.0.5 | kjk | ePub | 4 | 02-09-2010 09:53 PM |