Skip to content

PDF Table to Markdown Converter

Drop a PDF, pick a page and get the table as editable Markdown. Text PDFs are read directly, scans use OCR, and nothing is uploaded.

1. Add your table

Your file stays in this tab. Only the OCR engine files are downloaded, from the jsDelivr CDN.

2. Check and edit the table

Your table appears here

Every cell is editable. Amber cells are ones the OCR was unsure about.

How to convert a PDF table to Markdown in 3 steps

  1. Open the PDF

    Drop the PDF on the drop area or choose it. Pick the page that holds the table with the page selector; the page is shown so you can see what is being read.

  2. Check the grid

    For a PDF with a text layer the table is rebuilt from the position of each text run, with no OCR. For a scanned page, OCR runs automatically and unsure words are highlighted in amber.

  3. Copy the Markdown

    Copy the pipe table or open it in the full editor. If the columns came out wrong, set "Number of columns" and the table is rebuilt.

Example

Table on the PDF page
Quarter    Revenue    Cost    Margin
Q1         120,000    90,000  25%
Q2         135,500    98,200  27.5%
Q3         142,300    101,000 29%
Markdown from the tool
| Quarter | Revenue |    Cost | Margin |
| ------- | ------: | ------: | -----: |
| Q1      | 120,000 |  90,000 |    25% |
| Q2      | 135,500 |  98,200 |  27.5% |
| Q3      | 142,300 | 101,000 |    29% |

Options explained

Page
Which page of the PDF to read. One page is converted at a time; for a table that continues over several pages, convert each page and paste the rows together in the editor.Page 1 (default) to the last page
Use OCR even if the page has a text layer
Text PDFs are read directly because that is exact. Tick this when the text layer is wrong, for example a scan that was run through OCR badly, or text that is drawn as outlines.Off (default) · On
Language
Only used when OCR runs (scanned pages or the option above). Picks the OCR language data, which downloads the first time.English (default) · German · French · Spanish · Italian · Portuguese · Dutch · Russian · Chinese (Simplified) · Japanese
Number of columns
Auto finds columns from the empty vertical gaps between text. Pick a number if two columns were merged or one was split; the widest gaps become the column boundaries.Auto (default) · 1 to 12
Layout (OCR only)
How Tesseract reads a scanned page. Use "Table: one block of text" for a cropped table, "Automatic page layout" for a full page with other content.Table: one block of text (default) · Sparse text · Automatic page layout
Join wrapped lines
When a cell wraps onto two lines the PDF holds two lines of text at different heights. With this on, a row whose first cell is empty is joined to the row above.Off (default) · On

Convert other formats to Markdown

Convert Markdown to…

Browse all 36 tools

About this tool

Drop a PDF, choose the page with the table, and get an editable grid and a Markdown pipe table. A PDF with a text layer is read directly with pdf.js: the position of every piece of text is used to rebuild the rows and columns, with no OCR and no misread characters. A scanned page has no text to read, so it is rendered to an image and recognized with Tesseract.js in your browser, and words the engine was unsure about are highlighted. Your PDF is not uploaded. A text PDF needs no other download; OCR fetches its engine and language data from the jsDelivr CDN the first time.

  • Reads the PDF text layer directly: no OCR, no upload
  • Falls back to OCR for scanned pages
  • Page selector with a preview of the page
  • Editable grid before you export
  • Amber highlight on OCR words that need a check
  • Column count override and wrapped-line joining
  • Free, no sign-up, no watermark

Technical details

How a PDF table becomes Markdown

A PDF stores positioned text, not tables. The tool reads every text run on the page with its coordinates (pdf.js), then rebuilds rows and columns from those coordinates. Pages without a text layer are rendered to an image and read with OCR instead.

  • Text layer first: If the page has text, it is used as is. This is exact, so there are no OCR misreads and nothing is highlighted.
  • Rows: Text runs whose vertical extents overlap are one row; this tolerates small baseline differences between cells.
  • Columns: Columns are the vertical stripes where no text appears on the page. Runs padded with several spaces are split into cells too.
  • Scanned pages: A page with almost no text characters is rendered to a canvas and recognized with Tesseract.js, the same OCR path as the image tool.
  • Where the files come from: The pdf.js code and its worker are part of this site. Only the OCR engine and language data, needed for scans, come from the jsDelivr CDN.
  • One page at a time: The page selector converts the chosen page; text on other pages is not mixed in.

Tips and best practices

Prefer the original PDF over a scan

A PDF exported from Word, Excel, Google Docs or a browser has a text layer and converts without OCR errors. A scan or a photo of a printout needs OCR and can have misread characters.

Convert one table at a time

If a page holds several tables or a lot of surrounding text, crop it: export or screenshot just the table, or delete the extra rows in the grid.

Scans: start from a good image

For scanned pages, a straight, high-contrast scan at a reasonable resolution matters most. Crop a screenshot to the table and use the image tool if the full page does not work.

Check numbers and totals

Compare the first and last rows and any totals with the PDF before you publish. A table split by a page break or a footnote in the middle becomes extra rows.

Common problems and fixes

Handwriting in a scanned PDF

Cause: The OCR engine is built for printed text and reads handwriting poorly.

Fix: Type those cells in the grid. Printed text in the same scan is still recognized.

Merged or spanning cells

Cause: Markdown pipe tables have no colspan or rowspan, so a spanning cell is placed in its first column or left out as a caption line.

Fix: Repeat the value in each column, leave cells blank, or use an HTML table.

Rotated or skewed scans

Cause: Rows are found by horizontal alignment, and scans that are tilted or sideways are not straightened. Text rotated 90° in a text PDF is not read as a table.

Fix: Rotate or straighten the PDF in a PDF viewer or scanner software first.

Multi-level headers

Cause: Markdown has one header row; the second header line becomes the first data row.

Fix: Combine the header lines in the grid and delete the extra row.

Several lines of text in one cell become several rows

Cause: A PDF holds each line of a wrapped cell as separate text at a different height.

Fix: Tick "Join wrapped lines", or fix the rows by hand in the grid.

The text is garbled or the page shows as empty

Cause: Some PDFs have no usable text layer, or draw text as outlines, or use fonts without a text mapping.

Fix: Tick "Use OCR even if the page has a text layer". Password-protected PDFs need the password removed first.

Frequently asked questions

How do I convert a PDF table to Markdown?

Drop the PDF on the tool, pick the page with the table, check the editable grid and copy the Markdown. For a PDF with a text layer the table is rebuilt from the position of the text, and a scanned page is read with OCR.

Does my PDF get uploaded?

No. The PDF is opened by pdf.js running in your browser tab and is never sent to our servers. For a PDF with a text layer, nothing else is downloaded: the pdf.js worker is served from this site. Only when a page is scanned (or you tick "Use OCR even if the page has a text layer") does the tool fetch the OCR engine and language data from the jsDelivr CDN. Those are public engine files; your PDF and its pages are not part of that request.

Does it work with scanned PDFs?

Yes. A page with no text layer is rendered to an image and recognized with OCR in your browser, and words the engine was unsure about are highlighted. Scans are less exact than text PDFs, so check the numbers, and rotate or straighten crooked scans first.

How can I tell if my PDF is text or scanned?

Try to select a word on the page in your PDF viewer. If you can, it has a text layer. The tool tells you which it used: "Read from the PDF text layer (no OCR)" or "Read with OCR".

Can it read handwriting?

No, not reliably. The OCR engine is built for printed text. Handwritten tables need to be typed into the grid or the editor.

Can I convert a table that spans several pages?

One page is converted at a time. Convert each page, copy the Markdown of the later pages without their header row, and paste the rows under the first table in the editor.

Why are the columns wrong or merged?

Columns are found from the empty space between text. If a column is packed tightly against the next, or a long cell runs into it, set "Number of columns" to the right count. Header cells that span columns cannot be represented in a Markdown pipe table.

Is there a file size limit?

The tool accepts PDFs up to 100 MB and images up to 30 MB, because everything is processed in your browser and large files use a lot of memory. If a large PDF is slow, extract the page you need first.

Can I convert an image or screenshot of a table instead?

Yes. The image to Markdown table tool runs the same OCR path on PNG, JPG and WebP files, and you can also paste a screenshot with Ctrl+V.