eBook Formatting Services

PDF to EPUB Conversion for Books Whose Only Surviving File Is a Fixed Page

This is the hardest conversion there is, and anyone telling you otherwise has not done many. A PDF does not contain paragraphs. It contains glyphs with coordinates, drawn in whatever order the layout software felt like drawing them. Turning that back into a flowing eBook is not extraction. It is reconstruction, done by a person, with the original open beside it.

Reading order rebuilt from scratch, not lifted from the order the page was drawn in Line-end hyphens, running heads, folios and footers stripped out completely Scanned pages recognised, then proofed against the scan by a human before use Footnotes freed from their physical pages and rebuilt as links that work both ways
PDF to EPUB Conversion sample showing pdf to epub conversion hero
183PDFs rebuilt as eBooks
4.9Average client rating
100%Recognised text proofed by eye against the scan
0Validation errors at delivery
Definition

Why PDF to EPUB Conversion Is the Hardest Conversion of All

A PDF is a description of a printed page. It records that a particular glyph sits at a particular coordinate in a particular font at a particular size. It does not record that those glyphs form a word, that the words form a sentence, or that the sentence continues at the top of the next column. All of that structure lives in your head when you read it, and nowhere in the file.

An EPUB is the opposite kind of object. It has no coordinates and no pages. Text flows, and the reader decides how wide the column is, how large the type is and where the lines break. A conversion between the two therefore has to invent everything the PDF discarded when it was made, which is most of what a document is.

That is why we say rebuilt rather than converted. The text is recovered, the reading order is reconstructed, the paragraphs are reassembled from lines, the print furniture is deleted, and the structure is applied by hand. What comes out is a genuine reflowable eBook. What comes out of a one-click PDF converter is your book put through a shredder and taped back together in the dark.

Extraction is not conversion

What Actually Happens When Software Reads a PDF

Text extraction pulls the glyphs out in drawing order. On a single-column page that is often close enough to reading order to look like success. On a two-column page it interleaves the columns, so line one of the left column is followed by line one of the right, and the result is a fluent-looking sentence that means nothing.

Then there is the hyphenation. Every line that ended in a hyphen in the printed book carries that hyphen as a real character. After the text reflows to a new width, the hyphens stay exactly where they were, appearing in the middle of sentences, in words that no longer break there. Nothing flags this. It simply ships.

1

Paragraphs are gone

A PDF has lines, not paragraphs. Deciding where one paragraph ends and the next begins means reading the indents, the spacing and the sense of the text, and no extractor does the last of those.

2

Page furniture is content

The running head, the folio and the footer were drawn on the page like everything else. Extract the page and they arrive in the middle of your prose, once every few hundred words, forever.

3

Notes are anchored to paper

A footnote sat at the bottom of page 84 because the marker was on page 84. Remove the page and it is simply an orphaned block of small text, with nothing connecting it to anything.

What Is Included

What a PDF to EPUB Rebuild Includes

Every stage below is part of the price. None of them can be skipped, which is why this conversion costs more than any other.

Reading-order reconstruction

The text is recovered and then put back into the order a human reads it, which is frequently not the order the software drew it. Columns are unstitched, sidebars are separated from the body, captions are attached to their figures, and anything that crosses a page boundary is rejoined into a single continuous paragraph.

Paragraphs rebuilt out of lines

A PDF holds a stack of lines. We reassemble them into paragraphs by reading the indents, the spacing and the sense, then strip every line-end hyphen the printed layout baked into the words. This is slow, it is done by eye, and it is the difference between a readable book and a mess.

Page furniture removed entirely

Running heads, folios, footers, rules, watermarks and the publisher's name at the foot of every recto. All of it was drawn onto the page as content, all of it will otherwise appear scattered through your prose, and all of it has to be identified and deleted.

Recognition and human proofing for scans

Where the pages are images, optical recognition runs first and then a person proofs the result against the scan. Recognition is confidently wrong in specific ways: it swaps similar letters, mangles accents and invents plausible words. Only a reader comparing the two will ever catch that.

Tables and figures reconstructed

In a PDF a table is a set of drawn lines with words floating between them. There is no table. We rebuild it as a real table with header cells, or restructure it where it is too wide to survive a phone, and we re-extract every figure at the best resolution the file holds.

A hand-coded EPUB 3, validated

Once the content is real again, the eBook is built properly: semantic markup, a stylesheet written for the book, navigation generated from actual headings, then EPUBCheck to zero errors and a pass on physical hardware before you ever see it.

Portfolio

PDFs We Have Rebuilt as eBooks

Six recent rebuilds, from clean digital exports to scans of books nobody had the source files for.

Recent Work
Know which PDF you have

A Digital PDF or a Scanned One

These are two different jobs wearing the same file extension, and the difference is the largest number in the quote. Open your file and try to select a sentence with the cursor.

Digital PDF, Text Selectable

The file was exported from a layout program, so real text is inside it. The characters can be recovered, which removes the recognition stage entirely. What remains is the reconstruction: reading order, paragraphs, hyphens, furniture, notes, tables.

Still a rebuild rather than an extraction, but a considerably cheaper one, and the text will be exactly correct because it was never guessed at.

Scanned PDF, Pages as Images

There is no text in the file at all, only photographs of text. Optical recognition has to produce the characters, and recognition makes errors with total confidence. Every page then has to be proofed against the scan by a person before any structure is applied.

This is the most expensive conversion we offer, and it is honestly priced as a rebuild. It is also the only way a book with no surviving source files comes back to life.

What We Fix

The Damage a PDF Converter Leaves Behind

Six failures that appear in every automated PDF conversion we are sent. None of them raises an error, and all of them reach the reader.

Hyphens frozen mid-wordEvery line-end hyphen from the printed book survives as a real character. After reflow they surface inside sentences, splitting words that no longer break there, on every screen and at every type size.
Columns woven togetherText pulled in drawing order interleaves a two-column page line by line. The sentences look grammatical for about four words and then collapse into something nobody wrote.
Running heads in the proseThe book title and the author name, drawn at the top of every page, arrive as ordinary paragraphs. They now interrupt the narrative several hundred times.
Page numbers as body textFolios become stray digits sitting alone between paragraphs. Readers assume the file is corrupted, because a number on its own in the middle of a chapter looks exactly like corruption.
Footnotes with nothing to holdNotes that sat at the foot of a physical page become a block of small text stranded between two paragraphs, unlinked, referring to a marker that is now several screens away.
Words that recognition inventedOn a scan, similar characters are swapped with total confidence. The output is a fluent, plausible, misspelled book, and no software anywhere will ever tell you it happened.
Specification

What Happens to Each Part of the Page

Eight elements of a printed page, what a rebuild does with each, and what happens when nobody does.

ElementHow we build itWhy it matters
The text layerRecovered where it exists, recognised where it does not, then proofed and reassembled into paragraphs rather than left as a stack of lines.Extracted text is a stream of lines with no paragraph boundaries. Reflowed as it stands, the book becomes one enormous undifferentiated block.
Reading orderReconstructed by a person against the printed page, with columns unstitched, sidebars separated and cross-page paragraphs rejoined.Software reads a PDF in drawing order. Drawing order and reading order agree only by luck, and a multi-column page is where the luck runs out.
Line-end hyphensEvery one is identified and removed, with genuine compound hyphens preserved by checking the word against the surrounding text.Left in place they become permanent misspellings. This is the single clearest signature of a PDF conversion that nobody checked.
Running heads and foliosDetected by position and repetition, then deleted from the content entirely before any structure is applied.They were drawn on the page like any other text. Extraction cannot distinguish them from your prose, so they appear in it, endlessly.
FootnotesUnanchored from their physical page, matched to the marker they belong to, and rebuilt as bidirectional links with popup semantics where supported.A note is only a note because of where it sat. Remove the page and the relationship is gone, leaving small orphaned text with no purpose.
TablesRebuilt as real tables with header cells, from the words and the drawn lines, or restructured into a stacked form where they cannot fit a narrow screen.A PDF table is not a table. It is rules and floating words. Extraction produces a jumble of cell contents in whatever order they were drawn.
Images and figuresExtracted at the highest resolution the file holds, re-exported for high-density screens, reattached to their captions and given alt text.Print PDFs hold images in a print colour space at print resolution. Passed through untouched they are heavy, wrongly coloured and often separated from their caption.
Scanned pagesRecognition runs, then a human reads the output against the scanned image before any of it is trusted, page by page.Recognition does not fail loudly. It substitutes characters and produces a fluent, wrong book. Nothing but a person reading will find it.
Checked Before Delivery

How We Prove the Rebuild Says What the Book Said

The printed page is the only authority. Six checks decide whether the eBook still agrees with it.

A full read against the original pages

Someone reads the rebuilt text beside the PDF, page by page. Not skimmed and not sampled. On a scanned title this is the largest single cost in the job, and it is the only thing standing between you and a book full of confidently invented words.

Counts reconciled with the printed book

Chapters, headings, notes, tables, figures and captions are counted in the PDF and counted again in the EPUB. If the two numbers differ we know exactly how much has gone missing before anyone starts looking for it.

Hyphen and character sweep

The whole text is searched for surviving line-end hyphens, broken ligatures, mangled accents, reversed quotation marks and the specific substitutions that recognition software likes to make. These hide easily and this pass finds them.

Reflow at the narrowest screen

The rebuilt book is opened on a phone at a large type size. Anything still carrying the geometry of the printed page reveals itself immediately: a stray folio, an orphaned note, a table refusing to fold.

E-ink hardware and the Kindle pipeline

The derived Kindle file is checked, and the EPUB is opened on physical e-ink. Legacy titles rebuilt from print tend to be long, and long files are where older hardware starts complaining about memory and page turns.

Validation and store ingestion

EPUBCheck to zero errors, then each retailer's own ingestion check. A rebuilt book has to be indistinguishable from one that was born digital, and this is where that claim is tested.

Digital PDFScanned PDFPrint-ready PDFOCR proofingColumn unstitchingHyphen repairKindle PreviewerApple BooksEPUBCheck
Deliverables

What Comes Back to You

A book that has been genuinely rebuilt, and the evidence that the rebuild is faithful.

The rebuilt EPUB 3

A validated, hand-coded reflowable eBook, indistinguishable from one built from a live manuscript. Real headings, real paragraphs, working navigation, a stylesheet a human wrote. Ready for every store that takes EPUB, which is all of them.

Amazon's version, checked first

Built from the same rebuilt source and confirmed on Amazon's rendering path before delivery. Titles recovered from print tend to be long and note-heavy, and Amazon's ingestion has opinions about both, so we look first.

A clean, styled Word document

The recovered book as an editable manuscript with real heading styles. If your source files are lost, this is the closest thing you will ever have to getting them back, and it is what a second edition or a new print run starts from.

The fidelity report

What was in the printed book, what came across, what was changed and why. On a scanned title it also records the recognition pass and the proofing. This is how you know the rebuild is honest rather than merely finished.

The recovered image set

Every figure, photograph and diagram pulled from the PDF at the best resolution it holds, converted out of print colour, re-exported for screens and named consistently. Useful again the moment you want a cover or a sample.

Unpacked XHTML and CSS

XHTML, stylesheet and package document, unzipped and commented. A rebuilt book should not be trapped in another opaque file, which is exactly the situation you came to us to escape.

Process

How a PDF Rebuild Runs

Five stages. The middle three are where the money goes, and they cannot be hurried without lying to you.

01

Assess the PDF

We open the file and find out what it really is: digital or scanned, one column or two, how much furniture, how many notes, how many tables. Try selecting text yourself before you send it, because that one test predicts most of the price.

02

Recover the text

Characters are extracted where they exist and recognised where they do not. On a scan, this is followed immediately by a human proofing pass against the page images, because unproofed recognition is not text, it is a guess.

03

Reconstruct the document

Reading order rebuilt, columns unstitched, paragraphs reassembled from lines, hyphens removed, running heads and folios deleted, notes reattached to their markers, tables rebuilt. This is the job. Everything else is comparatively quick.

04

Build the eBook

Semantic markup applied, stylesheet written, navigation generated from the real headings, images re-exported and placed. Only now is there a document to format, which is why nobody should quote this as a formatting job.

05

Verify and deliver

Counts reconciled against the printed book, text read beside the original, EPUBCheck to zero errors, hardware pass, then delivery with the fidelity report and the recovered manuscript.

Pricing Calculator

PDF to EPUB Conversion Pricing Calculator

Pricing follows the word count, because every word has to be recovered, reassembled and read back against the page. Whether the PDF is digital or scanned then moves the number more than anything else.

Build Your PDF to EPUB Conversion Quote

Required steps are marked with *. Optional extras can be removed from the summary with the cross icon.

Delivery
Total$0
1

Book Type *

Locked

What kind of book the PDF holds. This decides how much of the printed structure has to be rebuilt rather than simply recovered.

Fiction and Narrative ProseIncludedA novel or memoir in continuous prose. The least printed structure to rebuild and the cleanest thing to bring back from a page.
Nonfiction and Business+$66Headings, pull quotes, boxed content and sidebars, all of which were positioned on the page and must now be identified by what they were for.
Academic and Reference+$127Footnotes bound to physical pages, tables drawn as lines, a bibliography and an index of page numbers that no longer exist. The hardest rebuild in the catalogue.
Illustrated or Fixed-Layout+$152Picture books, comics and heavily designed cookbooks where the page geometry is the book and has to be recreated rather than dissolved. The size band you choose implies the image count.
2

Book Size*

Locked

The approximate word count. If you only have a page count, multiply by 300 for a rough figure and we will confirm it from the file itself.

Up to 10,000 words+$189The approximate word count. If you only have a page count, multiply by 300 for a rough figure and we will confirm it from the file itself.
10,001 to 20,000 words+$245The approximate word count. If you only have a page count, multiply by 300 for a rough figure and we will confirm it from the file itself.
20,001 to 35,000 words+$319The approximate word count. If you only have a page count, multiply by 300 for a rough figure and we will confirm it from the file itself.
35,001 to 50,000 words+$409The approximate word count. If you only have a page count, multiply by 300 for a rough figure and we will confirm it from the file itself.
50,001 to 70,000 words+$505The approximate word count. If you only have a page count, multiply by 300 for a rough figure and we will confirm it from the file itself.
70,001 to 90,000 words+$621The approximate word count. If you only have a page count, multiply by 300 for a rough figure and we will confirm it from the file itself.
90,001 to 120,000 words+$749The approximate word count. If you only have a page count, multiply by 300 for a rough figure and we will confirm it from the file itself.
120,001 to 150,000 words+$899The approximate word count. If you only have a page count, multiply by 300 for a rough figure and we will confirm it from the file itself.
More than 150,000 wordsCustom QuoteLong reference works, multi-volume sets and complete backlists trapped in print are quoted per project. Send the PDFs and we will price the recovery properly rather than guess at it.
3

PDF Type *

Locked

Digital or scanned, single column or multiple. This is the biggest number in a PDF conversion quote and there is no way to disguise it.

Digital PDF, single columnIncludedText selectable, one column, exported from a layout program. The friendliest PDF there is, and still a rebuild, because the paragraphs and the reading order are not in the file.
Digital PDF, complex layout+$41Multiple columns, sidebars, pull quotes and figures with captions. The reading order has to be reconstructed by hand, because drawing order will interleave all of it into nonsense.
Scanned PDF, clean pages+$74Photographs of text, well lit and square, from a recent print run. Recognition runs, then every page is proofed by eye against the scan before a single word is trusted.
Scanned PDF, old or poor quality+$136Foxed paper, show-through from the reverse, broken type, tight gutters and skewed pages. Recognition struggles, so the human proofing pass grows, and that is where the cost sits.
4

Interior Complexity *

Locked

How much of the page is not plain prose. Notes, tables and references were all anchored to paper, and paper is what we are taking away.

Plain proseIncludedChapters and paragraphs and very little else. The rebuild is reading order, hyphens and furniture, which is quite enough on its own.
Structured content+$58Heading levels, lists, boxes and quotes, each of which existed only as a position and a type size on the page and must now be named for what it was.
Notes, tables and an index+$143Footnotes anchored to pages, tables drawn as rules, cross-references to page numbers and an index that points at a pagination the eBook does not have.
5

Interior Images *

Locked

How many figures and photographs are inside. Print PDFs hold these at print resolution in a print colour space, and both have to be corrected.

No interior imagesIncludedText only. The cover is handled separately and does not count against this step.
1 to 15 images+$36Extracted at the best resolution the PDF holds, converted out of the print colour space, re-exported for screens and given alt text.
16 to 40 images+$79A properly illustrated book. Each figure is reattached to its caption, which the extraction will have separated from it, and checked at a small screen size.
41 to 90 images+$131Heavy illustration. Every image is compared against the printed page rather than sampled, because an extractor will silently drop one and the file will still validate.
More than 90 imagesCustom QuotePhotographic and heavily illustrated PDFs are quoted after we open the file. If the pictures only exist inside the PDF at print resolution, that is a very different job from having the originals.
6

Delivery Speed *

Locked

Standard is the pace at which a book can be rebuilt and proofed. The faster tiers buy priority and longer days, never fewer checks.

StandardIncludedThe normal queue, with the full reconstruction and the complete read-back against the printed page. This is the pace at which fidelity can actually be proved.
Priority+$102Roughly a third is taken off the delivery window. Your file moves up the queue and every proofing and verification pass still runs in full.
Express+$192About half the standard window. Reasonable when a reissue date is already announced and the rights reverted later than anyone planned.
Rush+$308The fastest we will move while still reading the rebuilt text against the original pages. One revision round instead of two, because there is no time for two.
7

Revision Rounds *

Locked

Two rounds are included. On a rebuilt book these are usually about presentation, because the fidelity questions were settled during verification.

2 Revisions IncludedIncludedTwo full rounds once you have the proof. Anything that turns out to be a fidelity error is ours to fix and is never counted against your rounds.
2 Additional Rounds+$85Four rounds in total. Sensible when a rights holder, an estate or an academic reviewer has to approve the recovered text before it is published.
5 Additional Rounds+$175Seven rounds in total. For books being lightly corrected while they are being recovered, which happens more often than anyone likes to admit.
8

Optional Extras

Locked

Recovery work beyond the eBook itself, for titles where the source files are gone for good.

Editable Word manuscript+$47The recovered book delivered as a clean, styled .docx with real heading styles. When the source files are gone, this is the manuscript you never thought you would see again.
Index rebuilt as live links+$55The dead page numbers in a print index turned into working links that land on the paragraph the entry actually refers to, rather than on a page that no longer exists.
Page-list markers for citation+$63Print pagination carried across as page-break markers exposed in the navigation, so libraries and scholars can still cite the printed edition from the eBook.
Dash and punctuation normalisation+$88Every dash, ellipsis, apostrophe and quotation mark checked and made consistent, which print typesetting and recognition software between them will have left in three different states.
Second recognition pass+$96A second, independent recognition run on a difficult scan, with the two outputs compared against each other. Disagreements between them point straight at the pages a proofreader needs to look at hardest.
Figure and caption reunification+$114Every figure matched back to the caption it belonged to on the page, numbered correctly and marked up as a real figure rather than as a floating picture with a paragraph near it.
Poetry and verse reconstruction+$129Line breaks and hanging indents rebuilt as real verse markup rather than as paragraphs that happen to be short. Poetry is the single most fragile thing to recover from a fixed page.
Cover recovered from the print file+$158The printed cover extracted, cleaned, resized to retailer specifications and delivered as a usable eBook cover, for titles where the original artwork has been lost with everything else.
Second title from the same backlist+$184A companion volume rebuilt alongside this one to an identical standard, so a recovered series does not arrive looking like nine unrelated books.
Full print-ready interior rebuilt+$219The recovered text typeset as a fresh print interior, so a title that survived only as a PDF can be reissued in paperback from a real source file again.
Request Your Custom Quote
You selected a custom or advanced option, so this project is priced individually rather than through checkout. Send us the details and we will scope it properly.

We rebuild structure, never wording. The text we recover is the text that was printed, typos and all, and we will not silently correct your book while converting it. Editing, proofreading, cover design and uploading to retailer accounts are separate services.

Compare

One-Click Extraction Against a Genuine Rebuild

One of these takes four seconds. The other takes days, and it is the only one that produces a book.

What you are comparingAutomated PDF ExtractionBook Formatting Experts
Reading order Whatever order the page was drawn in. Two columns interleave line by line and produce sentences nobody wrote. Reconstructed by a person against the printed page, with columns unstitched and cross-page paragraphs rejoined.
Hyphenation Line-end hyphens survive as real characters and reappear in the middle of words after reflow, permanently. Every one removed, with genuine compound hyphens preserved by checking each word against its context.
Running heads and folios Extracted as content. Your book title and a page number now interrupt the prose a few hundred times. Detected by repetition and position, and deleted before any structure is applied.
Scanned pages Recognised and shipped. The output is fluent, plausible and wrong, and nothing tells you which words were invented. Recognised, then proofed page by page against the scan by a person, because that is the only check that works.
Footnotes Left as orphaned blocks of small text, unlinked, wherever the page they lived on used to be. Matched to their markers and rebuilt as bidirectional links, with popup notes where the reading system supports them.
What you end up with A file that opens, reads badly, and cannot be fixed without doing the whole job properly anyway. A genuine reflowable EPUB, plus an editable manuscript, plus a written record of everything that changed.
Who This Is For

The PDFs That Reach Us

Six recurring cases, each one with a characteristic way of resisting conversion.

The out-of-print novel

Rights have reverted, the publisher has sent a print PDF and nothing else, and the manuscript files went with a laptop three moves ago. The whole book has to be recovered from the page. It is the most satisfying work we do, because a title genuinely comes back into print.

The scholarly monograph

Two hundred footnotes tied to the pages they sat on, tables built from drawn rules, a bibliography and an index of dead numbers. Nothing in it converts. Every relationship in the book has to be rediscovered and then rebuilt as a link that works in both directions.

The corporate report or white paper

Multi-column pages, pull quotes, sidebars, charts and infographics, all positioned rather than structured. Extraction weaves the columns together and the document becomes unreadable in the first paragraph.

The cookbook that only exists in print

Ingredient columns beside method steps, both floating over a full-bleed photograph. Recovering it means working out which words belonged to which recipe, and then deciding whether the book should reflow at all.

The family or local history

Often a scan rather than a digital file, with photographs, letters and captions. Recognition struggles with old type, so the proofing pass is long, and the photographs need real work before a screen will do them justice.

The book of poems

The hardest thing on this page. Line breaks carried meaning and the fixed page preserved them by accident. Recovered carelessly, every poem becomes a paragraph. It has to be rebuilt as real verse, line by line, and there is no shortcut anywhere in it.

Standards

Fidelity Is the Whole Product

In every other conversion, the question is whether the formatting survived. Here the question is whether the words did. A PDF rebuild can quietly change your book, and the change will be fluent, plausible and completely invisible to any piece of software you or we could run against it.

That is why a human reads the recovered text against the original pages, and why we hand you a fidelity report at the end. Anybody can produce a file from a PDF in a minute. What takes days, and what you are actually paying for, is being able to prove that the file still says what the book said.

Three Rules of a Rebuild

1

Never trust recognition

Optical recognition does not report doubt. It substitutes a character it is confident about and moves on. Every recognised page is proofed against its image, without exception, on every job.

2

The page is not the book

Everything that exists only because there was a page must go: folios, running heads, footers, page-bound notes, page-number cross-references. Keeping them out of misplaced loyalty is how a rebuild becomes a mess.

3

Say what could not be saved

Some things cannot come across from a fixed layout, and some should not. Each one is written into the fidelity report with the reason, so nothing is discovered by a reviewer six months later.

Quality Control

The Fidelity Checklist

Twelve checks, run with the printed pages open beside the rebuilt file.

The recovered text has been read against the original pages by a person
No line-end hyphen from the printed layout survives inside a word
Reading order is correct on every multi-column and sidebar page
No running head, folio, footer or watermark remains in the body text
Paragraphs broken across pages have been rejoined into single paragraphs
Every footnote links to its marker and back again
Chapter, heading, note, table and figure counts match the printed book
Tables are real tables, not extracted words in the order they were drawn
Every figure is reattached to its caption and renders sharply at high density
Ligatures, accents, dashes and quotation marks are correct throughout
EPUBCheck finds nothing to report on the final package
The book reflows cleanly on a phone at the largest available type size
Before We Start

When You Need PDF to EPUB Conversion

Six situations. All of them share one fact: the PDF is the only file left.

The rights came back and the files did not

The publisher has given the book back and handed you a print PDF, which is legally generous and practically useless. There is no manuscript, no InDesign package and nobody left at the company who remembers. The book has to be recovered from the printed page, and it can be.

You only ever had a print edition

The interior was laid out once, the paperback sells steadily, and the eBook has never existed because the export was tried and looked appalling. It looked appalling because a fixed page cannot be poured into a reflowing screen. It has to be rebuilt.

The book is old and only a scan survives

Out of copyright, out of print, or simply out of the era before digital files. What exists is photographs of paper. Recognition and a long proofing pass will bring the text back, and the result is a book that reads like it was born this century.

Your converted PDF was rejected

You ran the file through a converter and the store refused it, or accepted it and the reviews have started mentioning the formatting. Both outcomes have the same cause. Extraction is not conversion, and the file has to be built again from the page.

You are reissuing a backlist

A press or an estate with twenty titles surviving only as print PDFs. Rebuilt individually they will be twenty different books. Rebuilt to one standard, they are a catalogue, and the standard is what makes the twentieth cheaper than the first.

You need the book to be readable, properly

A fixed PDF cannot be enlarged, cannot be re-flowed and cannot be read aloud sensibly. Readers with low vision, and the institutions that buy for them, need a real reflowable eBook. A PDF, however tidy, is not one and never will be.

Our Promise

What We Promise on a PDF Rebuild

Three commitments, and the first one is the only one that really matters.

Every recovered word has been read

On any file involving recognition, a person reads the output against the scanned page. Not a spellcheck, not a sample. If we cannot say that, we do not deliver the file, because unproofed recognition is not text.

Nothing is changed in silence

Everything we alter, drop or rebuild is written into the fidelity report with the reason. You will not find out from a reviewer that a paragraph went missing between the page and the file.

The stores will take it

If any retailer refuses a file we rebuilt, on structural or validation grounds, we correct it and reissue at no charge. A recovered book has to be as clean as one born digital, and ours are.

Strategy

How We Think About a Fixed Page

Three positions that shape every decision in a PDF rebuild.

Rebuild, do not extract

Extraction asks what is on the page. Rebuilding asks what the page was for. Only the second question produces a book, and everything expensive about this service follows from taking it seriously.

The printed page is the authority

When the recovered file and the original disagree, the original wins and the file is corrected. Not the other way round, and never on the balance of convenience.

Tell the author the truth about the file

Some scans are too poor to be worth recovering, and some jobs cost more than the title will earn. We say so before you pay, because we would rather lose the work than deliver a book you cannot sell.

The Short Version

The Rebuild Framework

Three questions, asked of every PDF that arrives, in this exact order.

1

Is there any text in here at all

Try to select a sentence. If you can, the characters exist and can be recovered exactly. If you cannot, the pages are images and every word will have to be recognised and then proofed. This one test moves the price more than anything else.

2

What order was this meant to be read in

Drawing order is not reading order. Columns, sidebars, captions, boxes and footnotes all have to be put back into the sequence a human reads them in, and that sequence exists only in the printed page and in somebody's judgement.

3

What existed only because there was a page

Folios, running heads, footers, page-bound notes, references to page numbers, hyphens created by a line ending. All of it is scaffolding for paper. All of it comes out, and taking it out cleanly is most of the craft.

Client Reviews

What Authors Say After the Files Land

Rated 4.9 out of 5 across 183 reviewed projects.

★★★★★

I had already paid someone else to format this book once. The chapter headings shifted every time a reader changed the font size and I only found out from a review. This team rebuilt the interior from the structure up and sent me photographs from four real devices before I uploaded anything. First launch day I have not spent refreshing the reviews in a panic.

MW
Marianne WhitlockUnited Kingdom Verified
★★★★★

What sold me was the audit. They came back and told me two chapters had been pasted in from another document and were carrying invisible styles, and that fixing that first would save me money later. Nobody had mentioned it in three years of self-publishing. The proof needed no corrections at all.

DA
Devon AchebeUnited States Verified
★★★★★

My book has footnotes, tables and two appendices, and every quote I got either ignored that or doubled the price because of it. Here the pricing was laid out step by step and I could see what each element cost before I committed. It arrived early and the notes link both ways, which nobody else had even offered.

SL
Suvi LehtinenFinland Verified
★★★★★

The proof came back from the printer with the margins exactly where they should be, the running heads correct, and not one chapter opening on a left-hand page. This is my fourth book and the first time I have not had to pay for a second proof copy.

RO
Rachel OkonjoCanada Verified
★★★★★

English is my second language and I was nervous about the back and forth. Everything was explained plainly, and when I asked the same question twice they answered as if it were the first time. The revisions were turned round inside two days and I was never made to feel like a nuisance.

TI
Tomas IglesiasSpain Verified
★★★★★

I write cookbooks, so the interior is the product. Ingredients have to sit beside the method or the recipe stops working, and the last person I hired reflowed the whole thing into nonsense. These people understood the problem before I had finished describing it and the finished layout is genuinely beautiful.

PR
Priya RaghunathanIndia Verified
★★★★★

Straightforward from beginning to end. A fixed price, a delivery window they actually hit, and files that went up to KDP and IngramSpark without a single rejection. I have since sent them two more titles and recommended them to my writing group without hesitation.

GN
Gregory NkemeluNigeria Verified
★★★★★

My memoir has photographs and I had resigned myself to them looking muddy. They optimised every image, told me honestly which two were too low-resolution to save, and suggested crops instead of shrugging. That candour is worth more to me than the discount I was chasing elsewhere.

HS
Hannah SorensenDenmark Verified
★★★★★

I sent them a 480-page reference book with a nightmare of cross-references and a table on nearly every spread. They quoted it accurately, they did not come back later asking for more money, and the index still works. Any author will tell you what that is worth.

EB
Elliot BarrowAustralia Verified
FAQ

PDF to EPUB Conversion Services — Your Questions, Answered Properly

The questions authors actually ask us, answered without sales language.

What is PDF to EPUB conversion, and why is it harder than converting from Word?

PDF to EPUB conversion is the rebuilding of a fixed printed page into a flowing eBook. It is harder than any other conversion because a PDF is not a document. It is a description of where glyphs sit on a page, in a particular font, at a particular size. It does not know what a paragraph is, what order the columns are read in, or which text is a running head. A Word file, however badly formatted, still knows its own paragraphs. A PDF knows nothing, so everything has to be reconstructed.

Can you convert a PDF to EPUB without the text turning into nonsense?

Yes, but only by rebuilding rather than extracting, and that is a manual job. Extraction pulls glyphs out in drawing order, which on a two-column page interleaves the columns and produces sentences that look grammatical for four words and then collapse. We reconstruct the reading order by hand against the printed page, reassemble the lines into real paragraphs, strip the baked hyphens and remove the page furniture. Then somebody reads the result against the original, which is the part no software offers.

How do I know whether my PDF is a digital file or a scan?

Open it and try to select a sentence with your cursor. If the text highlights, real characters are inside the file and they can be recovered exactly, with no guessing at any point. If nothing highlights, or the whole page selects as a single block, your pages are photographs of text. That means optical recognition, and it means a human proofing every recognised page against the scan. It is the single biggest factor in the price, so it is worth checking before you ask us for a quote.

What are the baked hyphens everyone warns about in PDF conversions?

When a printed book was typeset, words were broken across line endings with hyphens. In a PDF those hyphens are real characters sitting in the text, exactly as real as the letters around them. Extract the text and the hyphens come too. Reflow it to a phone screen and the line endings move, but the hyphens do not, so words are split in places they never broke. It is the clearest possible signature of a conversion nobody checked, and readers spot it within a page.

What happens to my running heads and page numbers?

They are found and deleted, because they were drawn onto the page like any other text and extraction cannot tell them apart from your prose. Left alone, your book title and author name appear as paragraphs in the middle of the narrative, once for every page of the original, and stray digits sit alone between paragraphs where the folios used to be. Readers see a number floating in the middle of chapter four and conclude, quite reasonably, that the file is corrupted.

How accurate is OCR, and do you check it?

On clean modern print, recognition is usually above 98 percent accurate, which sounds excellent until you work out that it means several hundred wrong characters in a novel. Worse, it does not report doubt. It substitutes a character it is confident about and continues, so the errors are fluent and plausible rather than obviously broken. Every recognised page we produce is proofed against its scanned image by a person. On a scanned title that proofing is the largest single cost in the job, and it is not optional.

Can you convert a two-column PDF like an academic journal or a report?

Yes, and it is exactly the case where automated tools fail hardest. Software reads a PDF in the order the page was drawn, which on a two-column layout means it takes the first line of the left column, then the first line of the right, and weaves them together. We unstitch the columns by hand against the printed page, rejoin any paragraph that ran from the foot of one column to the top of the next, and separate genuine sidebars from the body text before any conversion happens.

What do you do with footnotes that were at the bottom of each printed page?

We free them from the page and reattach them to their meaning. A footnote sat at the foot of page 84 only because its marker was on page 84. Take away the page and the note is just a small orphaned block of text with no relationship to anything. We match each note back to the marker it belongs to and rebuild the pair as a bidirectional link, so a reader taps the marker, reads the note, and returns to precisely the sentence they left.

Will my tables survive a PDF to EPUB conversion?

The data will, if somebody rebuilds it. The table will not, because there is no table. In a PDF a table is a set of drawn lines with words floating in the gaps between them, and extraction produces those words in the order they were drawn, which is usually a jumble. We reconstruct each one as a real table with proper header cells, and where a wide printed table cannot survive a six-inch screen we restructure it into a stacked, labelled form that a phone reader can actually use.

My PDF is a print-ready file with crop marks and bleed. Is that a problem?

It is normal, and it is more information rather than less. Crop marks, registration marks, colour bars and bleed all tell us about the printed page, and none of them belong in an eBook, so they are removed along with the rest of the print furniture. The images inside a print-ready PDF are usually held at print resolution in a print colour space, which means they are heavy and wrongly coloured for a screen. Both of those are corrected during the rebuild.

Can you rebuild a book from a poor-quality scan of an old edition?

Usually, and it is some of the most rewarding work we do. Foxed paper, show-through from the reverse of the sheet, broken type and a tight gutter all make recognition struggle, so the human proofing pass grows and the price grows with it. Send us a handful of representative page images before you commit to anything. We will tell you honestly whether the scan can carry a book, and if the answer is no, we will say so rather than take the money.

Why can I not just upload my PDF to Amazon instead of converting it?

You can upload it, and Amazon will convert it for you, badly, using exactly the extraction that produces everything described on this page. Your running heads will appear in the prose, your hyphens will surface mid-word and your columns will interleave. The result is published under your name, and the first people to notice are the readers who paid for it. A PDF is a print artefact. It is not an eBook, and no store will turn it into one on your behalf.

Do I get an editable manuscript back as well as the eBook?

You can, and for a legacy title it is often the most valuable thing in the delivery. The recovered book is handed back as a clean, styled Word document with real heading styles, which for many authors is the first time in years that an editable version of their own book has existed anywhere. It is what a corrected second edition starts from, what a new print run is set from, and what makes any future work on the title cheap rather than catastrophic.

What does my index do in an eBook if there are no page numbers?

Nothing at all, unless it is rebuilt. A printed index is a list of page numbers, and a reflowable eBook has no pages, so every entry points at something that does not exist. There are two honest answers. We can rebuild the index as live links that land on the paragraph each entry refers to, which is what most readers actually want. Or we can carry the print pagination across as page-list markers, which is what libraries and academic citation require. Some books want both.

Why is PDF to EPUB more expensive than other conversions?

Because there is more work in it and less of it can be automated. Every other source format still contains a document. A PDF contains a page. The text has to be recovered, the reading order reconstructed, the paragraphs reassembled from individual lines, the hyphens stripped, the furniture removed and the notes reattached, and on a scan every word has to be proofed by eye first. It is the only conversion where a person reads the entire book twice, and that is what the price reflects.

Can you keep the design of my printed book in the eBook?

The character of it, yes. The geometry, no, and this is worth being clear about because it is where most disappointment comes from. Your typeface, your chapter opener treatment, your ornaments and your general typographic voice can all carry across. What cannot is anything that depended on a page of a fixed size: the exact position of a drop cap, a photograph bled off the corner, a line that fell perfectly at the foot of a recto. The reader now controls the page, and there is no way to take that back.

What if my PDF has poetry, verse or unusual line breaks?

Then it is the most fragile thing on this page and it needs saying up front, because recovering it carelessly is destructive. Poetry survives in a PDF by accident, since the fixed page happened to preserve the line breaks the poet intended. Extract it and every poem becomes a paragraph, or worse, a paragraph with hyphens in it. We rebuild verse as real line-level markup, with hanging indents for runover lines that hold at any type size, and we do it line by line.

Can you handle a PDF where the pages are in the wrong order or missing?

Yes, and it happens more often than you would expect with scans of older books. Missing, duplicated and out-of-sequence pages are found during the read-back against the original, because someone is going through the whole book in order. We will tell you exactly which pages are absent and ask you to supply them. What we will not do is quietly bridge a gap and hand you a book with three pages missing from chapter nine.

How long does a PDF to EPUB rebuild take compared to a Word conversion?

Longer, and the calculator shows honest windows for both. A digital PDF of a straightforward novel takes somewhat longer than the same book as a .docx, because the reading order and the paragraphs have to be rebuilt. A scanned academic title with hundreds of footnotes takes considerably longer, because the recognised text has to be proofed page by page before any structure can be trusted. We give you a fixed date after opening the file, not before.

I have the InDesign file as well as the PDF. Should I send that instead?

Yes, and it will save you money. An InDesign package still contains a document: real paragraph styles, real text threads, real anchored objects. A great deal of the reconstruction we would otherwise be doing by hand is already recorded in it. If the InDesign file exists anywhere, on any drive, send it. The same goes for an old Word manuscript. The PDF should be the source of last resort, and it is only because it so often is the last resort that this service exists.

Next Step

Send Us the PDF and We Will Tell You What Is Really In It

Digital or scanned, one column or two, clean or foxed. We will open it, tell you whether the text can be recovered and what it will take, and quote a fixed price and a real date before you commit to anything.

$00 items selected
No options selected yet.
Get a Quote