Back to Blog

PDF to Word Table Formatting Broken? How to Keep Tables Intact

Converted a report to Word and got tab-separated text, floating boxes or shifted columns? Here's why PDF tables break, which conversion routes keep cell structure, and how to repair tables in Word in a few clicks.

You convert a 12-page quarterly report to Word, open it, and the invoice table is now three columns of tab-separated text sitting in a floating box that jumps to page 2 when you type. Nothing is editable the way you need it to be. This is the single most common complaint about pdf to word table formatting, and it is almost never the converter being lazy — it's the PDF itself not containing a table at all.

Below: why it happens, how to tell in 30 seconds which kind of PDF you have, the two conversion routes that preserve cells best, and the exact Word commands that turn mangled output back into a real table.

Why PDF tables turn into tabs, boxes or shifted columns

A PDF is a print description, not a document model. Inside the file there is no "table" object with rows and cells. There are text runs with x/y coordinates, and sometimes thin vector rectangles that happen to look like borders. Your eye assembles that into a grid; a converter has to guess.

Three kinds of PDF produce three very different results:

  • Tagged PDF (exported from Word, InDesign or a modern reporting tool with accessibility tags). It carries a real structure tree with <Table>, <TR> and <TD> elements. Converters read that tree and rebuild a genuine Word table. This is the best case and the columns usually land correctly.
  • Untagged PDF (printed to PDF, produced by an old ERP or a scanner-to-PDF driver). The converter clusters text by vertical position to guess rows and by horizontal gaps to guess columns. Ruled borders help a lot; tables separated only by whitespace often collapse into tab- or space-separated paragraphs, or a single cell per line.
  • Scanned PDF (a photo of paper). There is no text at all, only an image. Without OCR you get a picture of a table pasted into Word — zero editable cells.

Two extra failure modes explain the weird symptoms. Floating text boxes appear when a converter cannot decide on a flow layout and falls back to absolute positioning — each block is anchored to a page coordinate, which is why text won't reflow. Misaligned columns usually mean a row had a merged or empty cell, so the clustering shifted one column left from that row down.

Check your PDF before you convert

Thirty seconds here saves twenty minutes of cleanup.

  • Can you select the text? Open the PDF, drag across a table cell. No selection highlight = scanned, go straight to OCR.
  • Does the selection jump cell by cell or line by line? Cell-aware selection is a good sign of real structure.
  • Does the table have printed borders? Ruled tables convert far more accurately than whitespace-only layouts.
  • Check for tags. In Acrobat Reader, File > Properties > Description shows "Tagged PDF: Yes/No". Yes means you should get a proper Word table from almost any decent tool.

Route 1: the online converter (fastest for one or two files)

For a typical office document — say q3-invoices.pdf, 14 pages, 1.8 MB, exported from an accounting system — browser conversion is the quickest path and keeps the table as a table when the PDF has usable structure.

  • Open the PDF to DOCX converter and drop the file in.
  • Download the .docx and open it in Word.
  • Click inside a table. If the ribbon shows Table Design and Layout tabs, you have a real table — you're done. If it shows Shape Format, you got a text box; see the repair section below.

If you also need the same report as spreadsheet rows or plain text later, the rest of the tools sit on the File Convert Online document converter hub. New to this workflow? The converter page covers the basics of PDF to Word before you dive into table surgery.

Route 2: command line, when you convert batches

If you process dozens of statements a month, two free tools cover most cases, and they fail in different ways — so try both on a sample page before committing.

LibreOffice uses its PDF import filter and is good at ruled, text-based tables:

soffice --headless \
  --infilter="writer_pdf_import" \
  --convert-to docx:"MS Word 2007 XML" \
  --outdir ./out q3-invoices.pdf

Caveat: LibreOffice's PDF import is draw-oriented, so complex pages often arrive as positioned frames. It is excellent for simple, bordered tables and bulk jobs where you only need the numbers back.

pdf2docx (Python) is explicitly table-aware: it detects ruling lines and whitespace gaps and emits real w:tbl elements.

pip install pdf2docx
pdf2docx convert q3-invoices.pdf q3-invoices.docx --start 2 --end 6

Convert pages 3–7 first (the flags are zero-based), check the result, then run the whole file. In the Python API you can pass multi_processing=True for long documents. Where pdf2docx struggles: borderless tables, cells that wrap to three lines, and rotated headers.

Rule of thumb on pdf table converter accuracy: tagged PDF → nearly perfect anywhere; ruled untagged PDF → pdf2docx or a good online converter; borderless untagged PDF → expect manual cleanup whatever you use.

Fix table formatting after conversion in Word

Four repairs handle nearly every broken result.

Tab-separated text → real table. Turn on formatting marks (Ctrl+Shift+8) to confirm the separator is a tab arrow and not spaces. Select the block, then Insert > Table > Convert Text to Table, set Separate text at: Tabs, and check that the column count Word proposes matches your header. If it proposes more columns than you expect, one row has a stray tab — fix that row first.

Floating text boxes → inline content. Click the box border, Ctrl+A inside it to grab the contents, copy, then paste into the body with Keep Text Only, and delete the empty box. For a whole page of boxes, it is faster to reconvert with a different tool than to unpick them one by one.

Columns that drift or run off the page. Select the table, then Layout > AutoFit > AutoFit Contents, followed by AutoFit Window. Then open Table Properties > Table and set Text wrapping to None so the table stops floating.

Wrong merges and split rows. A cell that spills onto the next line is usually a row the converter split in two. Select the two rows, Layout > Merge Cells, and retype the one value. For a header that should repeat across pages: select the first row, Layout > Repeat Header Rows.

Last resort for a hopeless table: select it, Layout > Convert to Text with tabs, clean up the stray lines, and convert back to a table. You lose borders but keep the data in the right columns.

Scanned tables need OCR first

If the text isn't selectable, no converter can rebuild cells. Add a text layer first, then convert:

ocrmypdf --force-ocr --deskew scan-invoice.pdf scan-invoice-ocr.pdf
pdf2docx convert scan-invoice-ocr.pdf scan-invoice.docx

--deskew matters: a page scanned 2° crooked throws off row clustering, which is exactly what produces the classic staircase of misaligned columns. Scan at 300 DPI if you control the scanner; 150 DPI invoices lose digits.

The 60-second decision path

  • Text not selectable → OCR with ocrmypdf, then convert.
  • Tagged PDF → any good converter; the online route is fastest.
  • Untagged with borders → try the online converter, then pdf2docx on the same pages and keep the better result.
  • Untagged without borders → convert, then rebuild with Convert Text to Table on tabs.
  • Table came in as a box → copy contents out, paste as text, rebuild.

One habit saves the most time: always test on three pages before converting 120. If a sample page comes out clean, the rest almost always will too — and if it doesn't, you've lost a minute instead of an afternoon.

FAQ

Why does my PDF table become a text box in Word? The converter couldn't infer a flowing layout, so it preserved the visual position with an absolutely anchored frame. Copy the contents out and paste as text, or reconvert with a table-aware tool.

Does converting PDF to Excel keep tables better than Word? Sometimes, yes — spreadsheet targets force a strict row/column grid, so borderless tables often land more cleanly. If you only need the numbers, go to a spreadsheet and paste into Word afterwards.

Will the original fonts and colours survive? Fonts embedded in the PDF are matched to installed fonts, so substitution is common and can shift column widths. Cell shading and borders usually survive on tagged PDFs.

Can I convert only the pages with tables? Yes — that's what --start and --end do in pdf2docx, and splitting the PDF first also speeds up online conversion of large files.

My columns are one cell off from row 7 down. What happened? An empty or merged cell in row 7 broke the column clustering. Insert the missing cell in that row, or convert the table to text and back with tabs.

Ready to Convert Your Files?

Try our free online file converter. No registration required.

Start Converting