Converting files

Why is converting a PDF back to Word or Excel never perfect?

A PDF stores positioned glyphs rather than paragraphs, tables or slides, so converting back is reconstruction from appearance — text usually survives, structure is inferred, and anything the original expressed through layout alone is guesswork.

3 min read

The disappointment is always the same shape: the text came through, and everything that made it a *document* did not. Understanding why makes it much easier to get a usable result.

A PDF does not know what anything is

When a word processor exports a PDF, it converts meaning into appearance. "Heading 2, Calibri, 14pt, keep with next" becomes "draw these glyphs at these coordinates in this font at this size". The instruction survives; the intent does not.

The PDF has no concept of:

  • a paragraph — only lines of glyphs that happen to be near each other
  • a table — only text at aligned coordinates, and sometimes lines drawn around it
  • a cell — the grid was never recorded, only its appearance
  • a slide layout — only shapes and text at positions
  • a heading — only text that happens to be larger

Converting back means looking at the appearance and inferring the meaning. That is genuinely hard, and where it goes wrong it goes wrong in predictable ways.

What survives well, and what does not

Usually fine: the words themselves, in reading order; font sizes and weights; simple single-column text; images.

Usually approximate: paragraph boundaries (a line ending in a full stop is probably a paragraph end — probably); tables with visible ruling lines, which give the converter something to work from.

Usually poor: tables with no lines, where the grid was only ever whitespace; multi-column layouts, where reading order has to be guessed; anything positioned by hand; footnotes; anything in a text box overlapping other content.

How to get a better result

Choose the closest target. A PDF of a spreadsheet should go to PDF to Excel, not to Word and then into Excel. A PDF of slides should go to PDF to PowerPoint. Each converter is looking for a different structure.

Check whether there is text at all. A scanned PDF has no text to convert, only pictures of it. Run OCR PDF first — otherwise the output is an empty document with a photograph in it.

Expect to fix the tables. On a document with unruled tables, plan to spend a few minutes putting columns right. That is not the converter failing; the column boundaries genuinely are not in the file.

Ask whether you need the conversion at all. If the goal is to change a date or a name, editing the PDF directly with the PDF editor is faster and keeps everything else exactly as it was. Round-tripping a finished document through Word to change one line usually costs more than it saves.

The honest summary

Conversion out of PDF is reconstruction, not decoding. A good converter gets the content back and makes sensible guesses about the structure. Anything that promises a perfect round trip is promising to recover information the file does not contain.

Common questions

Why did my table come out as loose text?

Because it had no ruling lines. Without them the grid exists only as whitespace, and the converter has nothing to detect a cell boundary from.

Which conversion is most reliable?

Simple single-column text to Word. The more the original relied on layout to express meaning, the more has to be guessed.

My converted file is empty — why?

The PDF is almost certainly a scan, so there is no text in it to convert. OCR it first, then convert the result.