What is PDF to EPUB conversion, and why is it harder than converting from Word?
PDF to EPUB conversion is the rebuilding of a fixed printed page into a flowing eBook. It is harder than any other conversion because a PDF is not a document. It is a description of where glyphs sit on a page, in a particular font, at a particular size. It does not know what a paragraph is, what order the columns are read in, or which text is a running head. A Word file, however badly formatted, still knows its own paragraphs. A PDF knows nothing, so everything has to be reconstructed.
Can you convert a PDF to EPUB without the text turning into nonsense?
Yes, but only by rebuilding rather than extracting, and that is a manual job. Extraction pulls glyphs out in drawing order, which on a two-column page interleaves the columns and produces sentences that look grammatical for four words and then collapse. We reconstruct the reading order by hand against the printed page, reassemble the lines into real paragraphs, strip the baked hyphens and remove the page furniture. Then somebody reads the result against the original, which is the part no software offers.
How do I know whether my PDF is a digital file or a scan?
Open it and try to select a sentence with your cursor. If the text highlights, real characters are inside the file and they can be recovered exactly, with no guessing at any point. If nothing highlights, or the whole page selects as a single block, your pages are photographs of text. That means optical recognition, and it means a human proofing every recognised page against the scan. It is the single biggest factor in the price, so it is worth checking before you ask us for a quote.
What are the baked hyphens everyone warns about in PDF conversions?
When a printed book was typeset, words were broken across line endings with hyphens. In a PDF those hyphens are real characters sitting in the text, exactly as real as the letters around them. Extract the text and the hyphens come too. Reflow it to a phone screen and the line endings move, but the hyphens do not, so words are split in places they never broke. It is the clearest possible signature of a conversion nobody checked, and readers spot it within a page.
What happens to my running heads and page numbers?
They are found and deleted, because they were drawn onto the page like any other text and extraction cannot tell them apart from your prose. Left alone, your book title and author name appear as paragraphs in the middle of the narrative, once for every page of the original, and stray digits sit alone between paragraphs where the folios used to be. Readers see a number floating in the middle of chapter four and conclude, quite reasonably, that the file is corrupted.
How accurate is OCR, and do you check it?
On clean modern print, recognition is usually above 98 percent accurate, which sounds excellent until you work out that it means several hundred wrong characters in a novel. Worse, it does not report doubt. It substitutes a character it is confident about and continues, so the errors are fluent and plausible rather than obviously broken. Every recognised page we produce is proofed against its scanned image by a person. On a scanned title that proofing is the largest single cost in the job, and it is not optional.
Can you convert a two-column PDF like an academic journal or a report?
Yes, and it is exactly the case where automated tools fail hardest. Software reads a PDF in the order the page was drawn, which on a two-column layout means it takes the first line of the left column, then the first line of the right, and weaves them together. We unstitch the columns by hand against the printed page, rejoin any paragraph that ran from the foot of one column to the top of the next, and separate genuine sidebars from the body text before any conversion happens.
What do you do with footnotes that were at the bottom of each printed page?
We free them from the page and reattach them to their meaning. A footnote sat at the foot of page 84 only because its marker was on page 84. Take away the page and the note is just a small orphaned block of text with no relationship to anything. We match each note back to the marker it belongs to and rebuild the pair as a bidirectional link, so a reader taps the marker, reads the note, and returns to precisely the sentence they left.
Will my tables survive a PDF to EPUB conversion?
The data will, if somebody rebuilds it. The table will not, because there is no table. In a PDF a table is a set of drawn lines with words floating in the gaps between them, and extraction produces those words in the order they were drawn, which is usually a jumble. We reconstruct each one as a real table with proper header cells, and where a wide printed table cannot survive a six-inch screen we restructure it into a stacked, labelled form that a phone reader can actually use.
My PDF is a print-ready file with crop marks and bleed. Is that a problem?
It is normal, and it is more information rather than less. Crop marks, registration marks, colour bars and bleed all tell us about the printed page, and none of them belong in an eBook, so they are removed along with the rest of the print furniture. The images inside a print-ready PDF are usually held at print resolution in a print colour space, which means they are heavy and wrongly coloured for a screen. Both of those are corrected during the rebuild.
Can you rebuild a book from a poor-quality scan of an old edition?
Usually, and it is some of the most rewarding work we do. Foxed paper, show-through from the reverse of the sheet, broken type and a tight gutter all make recognition struggle, so the human proofing pass grows and the price grows with it. Send us a handful of representative page images before you commit to anything. We will tell you honestly whether the scan can carry a book, and if the answer is no, we will say so rather than take the money.
Why can I not just upload my PDF to Amazon instead of converting it?
You can upload it, and Amazon will convert it for you, badly, using exactly the extraction that produces everything described on this page. Your running heads will appear in the prose, your hyphens will surface mid-word and your columns will interleave. The result is published under your name, and the first people to notice are the readers who paid for it. A PDF is a print artefact. It is not an eBook, and no store will turn it into one on your behalf.
Do I get an editable manuscript back as well as the eBook?
You can, and for a legacy title it is often the most valuable thing in the delivery. The recovered book is handed back as a clean, styled Word document with real heading styles, which for many authors is the first time in years that an editable version of their own book has existed anywhere. It is what a corrected second edition starts from, what a new print run is set from, and what makes any future work on the title cheap rather than catastrophic.
What does my index do in an eBook if there are no page numbers?
Nothing at all, unless it is rebuilt. A printed index is a list of page numbers, and a reflowable eBook has no pages, so every entry points at something that does not exist. There are two honest answers. We can rebuild the index as live links that land on the paragraph each entry refers to, which is what most readers actually want. Or we can carry the print pagination across as page-list markers, which is what libraries and academic citation require. Some books want both.
Why is PDF to EPUB more expensive than other conversions?
Because there is more work in it and less of it can be automated. Every other source format still contains a document. A PDF contains a page. The text has to be recovered, the reading order reconstructed, the paragraphs reassembled from individual lines, the hyphens stripped, the furniture removed and the notes reattached, and on a scan every word has to be proofed by eye first. It is the only conversion where a person reads the entire book twice, and that is what the price reflects.
Can you keep the design of my printed book in the eBook?
The character of it, yes. The geometry, no, and this is worth being clear about because it is where most disappointment comes from. Your typeface, your chapter opener treatment, your ornaments and your general typographic voice can all carry across. What cannot is anything that depended on a page of a fixed size: the exact position of a drop cap, a photograph bled off the corner, a line that fell perfectly at the foot of a recto. The reader now controls the page, and there is no way to take that back.
What if my PDF has poetry, verse or unusual line breaks?
Then it is the most fragile thing on this page and it needs saying up front, because recovering it carelessly is destructive. Poetry survives in a PDF by accident, since the fixed page happened to preserve the line breaks the poet intended. Extract it and every poem becomes a paragraph, or worse, a paragraph with hyphens in it. We rebuild verse as real line-level markup, with hanging indents for runover lines that hold at any type size, and we do it line by line.
Can you handle a PDF where the pages are in the wrong order or missing?
Yes, and it happens more often than you would expect with scans of older books. Missing, duplicated and out-of-sequence pages are found during the read-back against the original, because someone is going through the whole book in order. We will tell you exactly which pages are absent and ask you to supply them. What we will not do is quietly bridge a gap and hand you a book with three pages missing from chapter nine.
How long does a PDF to EPUB rebuild take compared to a Word conversion?
Longer, and the calculator shows honest windows for both. A digital PDF of a straightforward novel takes somewhat longer than the same book as a .docx, because the reading order and the paragraphs have to be rebuilt. A scanned academic title with hundreds of footnotes takes considerably longer, because the recognised text has to be proofed page by page before any structure can be trusted. We give you a fixed date after opening the file, not before.
I have the InDesign file as well as the PDF. Should I send that instead?
Yes, and it will save you money. An InDesign package still contains a document: real paragraph styles, real text threads, real anchored objects. A great deal of the reconstruction we would otherwise be doing by hand is already recorded in it. If the InDesign file exists anywhere, on any drive, send it. The same goes for an old Word manuscript. The PDF should be the source of last resort, and it is only because it so often is the last resort that this service exists.