How to convert PDF and Word to markdown: six tools tested

How to convert PDF to markdown and Word docx to markdown, with real output from pandoc, MarkItDown, Docling, Marker, MinerU and PyMuPDF4LLM on the same files.

To convert a Word file to markdown, use pandoc: pandoc report.docx -t gfm --wrap=none -o report.md. In our test it kept the headings, the table, the list and the footnote, and it was the only tool to write the footnote in markdown footnote syntax. Microsoft’s MarkItDown came a close second with no extra flags.

To convert a PDF to markdown, use Docling when the file might be scanned. For a digital PDF, Marker gave the cleanest result, and PyMuPDF4LLM is the fast, light option. Pandoc cannot read PDF at all. Every tool here runs on your own machine and is free for personal use, with license limits for larger companies noted below. Hosted services such as Mistral OCR charge $4 per 1,000 pages.

Which converter to use

Your file Use Why
Word .docx pandoc Keeps headings, tables, lists, footnotes, equations and images, with no GPU and no models
Many formats in one tool MarkItDown or MinerU Both read Word, PowerPoint, Excel, EPUB and PDF
Digital PDF, best structure Marker Every heading, the linked footnote and the table came through in our test
Scanned PDF Docling Built-in OCR recovered the whole test page, table included
Digital PDF, small install PyMuPDF4LLM Two seconds per file in our test, no PyTorch
Old binary .doc MinerU, or LibreOffice then pandoc Pandoc and MarkItDown only read .docx
Confidential files Any local tool above Nothing leaves your machine with the cloud options left off
Documents kept beside your notes and searchable with them Jotura Writes a markdown copy of each PDF or Word file in the vault with its bundled MarkItDown, no OCR

Why Word is easy and PDF is hard

A .docx file tags each heading with its level and marks tables and footnotes, so a converter only translates labels.

A PDF stores characters at positions on a page, and the converter has to guess which lines are headings and where table cells end. A scanned PDF holds only pictures of pages, so it needs OCR (optical character recognition) first.

MarkItDown pulls text out with pdfminer, detects tables with pdfplumber and makes no attempt at headings. The other PDF tools run layout models, neural networks trained to recognize page regions.

How we tested

We built a one-page report in Microsoft Word 365: a Heading 1, a paragraph with a footnote, a Heading 2, a three-column table with grid lines, another Heading 2 and three bullet points. We saved it as report.docx and exported report.pdf from Word. A third file, scan.pdf, is that page rendered as a 200 dpi grayscale image with no text layer, which is what a scanner produces.

The tools ran on Windows 11 on CPU only, with default settings: MarkItDown 0.1.8, Docling 2.137.0, PyMuPDF4LLM 1.28.2, Marker 2.0.0, MinerU 4.0.11, and pandoc 3.9, the build bundled with pypandoc-binary 1.17. Pandoc’s current release is 3.12.1. Outputs below are copied from the files the tools wrote.

Word to markdown: the results

Tool Headings Table Bullet list Footnote
pandoc Correct levels Yes Yes, with blank lines between items Markdown footnote [^1]
MarkItDown Correct levels Empty header row added Yes Link to a numbered note at the end
Docling Shifted down one level Yes Yes Dropped

Pandoc’s output with --wrap=none, complete:

# Quarterly field report

Soil samples were collected at three sites in March.[^1] Results are summarized below.

## Results by site

| **Site** | **pH** | **Nitrogen (mg/kg)** |
|----------|--------|----------------------|
| North    | 6.4    | 112                  |
| River    | 7.1    | 87                   |
| Ridge    | 5.9    | 140                  |

## Next steps

- Retest the Ridge site in June

- Order lime for the North plot

- Share the data with the county office

[^1]: Site coordinates are listed in appendix B.

MarkItDown kept the headings and wrote a tight list. Its table and footnote came out like this (trimmed):

Soil samples were collected at three sites in March.[[1]](#footnote-1) Results are summarized below.

|  |  |  |
| --- | --- | --- |
| **Site** | **pH** | **Nitrogen (mg/kg)** |
| North | 6.4 | 112 |

1. Site coordinates are listed in appendix B. [↑](#footnote-ref-1)

Docling demoted every heading one level and dropped the footnote. MinerU 4 reads DOC, DOCX, PPTX, XLSX, RTF, ODT and EPUB too. We did not run it on the Word file.

The empty table header

Word marks a header row in two places. Header Row in the Table Design tab is on by default for a table made with Insert > Table. Repeat Header Rows sits in the Table Layout tab and is off by default. Pandoc reads either setting. MarkItDown reads only Repeat Header Rows, which is why it wrote the blank header above.

Select the first row, turn on Repeat Header Rows, save, and convert again. Both tools then used the first row as the header. A table with Header Row switched off gets the empty row in both tools until you do this. The markdown tables guide covers what pipe tables can hold.

Equations, images and tracked changes

We added a Word equation, N(t) = N₀e^(−λt), to a second test file. Pandoc wrote it as LaTeX in a fenced math block, which GitHub renders. MarkItDown wrote $$N\left(t\right)=N\_{0}e^{-λt}$$, with a stray backslash before the underscore and a raw λ in place of \lambda. Docling gave standard LaTeX: $$N\left(t\right)=N_{0}e^{- \lambda t}$$.

Pandoc’s --extract-media=media copied our picture to media/media/image1.png and wrote an HTML <img> tag to keep Word’s size. Add -t gfm-raw_html to get a plain ![](media/media/image1.png) link. MarkItDown writes ![](data:image/png;base64...) with the data cut off, and --keep-data-uris keeps it. Docling embeds images as base64 by default, and --image-export-mode referenced writes them to a <name>_artifacts folder.

Pandoc’s --track-changes option takes accept (the default), reject or all, and all keeps insertions, deletions and comments as marked spans. The full command we recommend for Word files:

pandoc report.docx -t gfm-raw_html --wrap=none --extract-media=report-media -o report.md

--wrap=none stops pandoc from breaking paragraphs at 72 characters, its default.

Old .doc files and no-install routes

Pandoc and MarkItDown read only .docx. Save an old .doc as .docx in Word, or convert it with LibreOffice and then run pandoc:

soffice --headless --convert-to docx report.doc

MinerU 4 reads .doc directly.

With nothing installed, open the file in Google Docs and choose File > Download > Markdown (.md), which sends it to Google. Browser converters such as word2md.com for Word and ConvertCase’s PDF to Markdown page state that conversion runs in your browser and nothing is uploaded. ConvertCase also says it does no OCR and does not reliably preserve tables. You cannot check those claims, so keep confidential files on a local tool.

PDF to markdown: the results

Tool Headings Table Bullet list Footnote
Marker All three, all ## Yes Yes Linked superscript and anchor
Docling All three, all ## Yes Yes Plain text
PyMuPDF4LLM Two of three Yes Yes <sup>1</sup> and a blockquote
MinerU Two of three Yes Lost, items became paragraphs HTML <small> span
MarkItDown None Yes • characters Plain text
pandoc Cannot read PDF

The ruled table converted correctly in every tool that reads PDF. Headings, lists and footnotes are where they differed.

Marker’s output was the cleanest. It flattened heading levels like Docling and kept everything else (trimmed):

## Quarterly field report

Soil samples were collected at three sites in March.[<sup>1</sup>](#page-0-0) Results are summarized below.

## Results by site

| Site  | pH  | Nitrogen (mg/kg) |
|-------|-----|------------------|
| North | 6.4 | 112              |

## Next steps

- Retest the Ridge site in June

<span id="page-0-0"></span><sup>1</sup> Site coordinates are listed in appendix B.

MarkItDown returned the text and a correct table, with no headings and literal bullet characters (trimmed):

Quarterly field report
Soil samples were collected at three sites in March.1 Results are summarized below.
Results by site
Next steps
•  Retest the Ridge site in June

PyMuPDF4LLM and MinerU both left “Results by site” as plain text. PyMuPDF4LLM wrote the footnote as > 1 Site coordinates are listed in appendix B. MinerU turned the three bullet points into separate paragraphs with no list markers and wrapped the footnote in a styled HTML span.

Borderless tables

We exported the same table from Word with no borders, twice. Spread across the full page width, it still converted cleanly in MarkItDown and PyMuPDF4LLM. Shrunk to fit its contents, it broke every tool we ran on it. MarkItDown printed loose lines:

Site

pH  Nitrogen (mg/kg)

North  6.4  112

PyMuPDF4LLM folded the whole table into one line, Site pH Nitrogen (mg/kg) North 6.4 112 River 7.1 87 Ridge 5.9 140. Docling did the same and promoted the header row to a ## heading. We did not run Marker or MinerU on this file. Check every borderless table by eye, and rebuild a broken one from the source data.

Scanned PDFs and OCR

Tool Result on scan.pdf
Docling Full text, headings, table and list recovered
MinerU Text and table recovered, list markers mixed \- and •, footnote number as $^{1}$
PyMuPDF4LLM Empty, with no OCR engine installed
MarkItDown Empty
Marker Not run, see below

Docling’s scan result matched its digital-PDF result almost exactly. Its default, --ocr-engine auto, uses the first engine it finds installed, and the standard install includes RapidOCR.

PyMuPDF4LLM runs OCR through Tesseract or RapidOCR when one is installed alongside it. MarkItDown has no local OCR. Its markitdown-ocr plugin sends pages to whatever OpenAI-compatible model you configure, a cloud service unless you point it at a model running on your own machine.

Marker 2 defaults to Fast mode on CPU and Apple Silicon, which reads the PDF’s text layer, and to Balanced mode on a GPU, which runs the Surya vision model over the whole page. OCR and equations go through a local model server: vLLM in Docker on an NVIDIA GPU, or llama.cpp’s llama-server elsewhere. --disable_ocr turns off every model call. Our machine has an NVIDIA card and Docker was not running, so we have no Marker result for the scan.

A PDF whose text you cannot select in a viewer is a scan. OCRmyPDF adds a text layer with Tesseract, after which any tool above can read it:

ocrmypdf scan.pdf scan-text.pdf

Equations, multi-column pages and images in PDFs

Equations. MarkItDown flattened our equation to 𝑁(𝑡) = 𝑁0𝑒−𝜆𝑡, losing the superscript. PyMuPDF4LLM kept it as 𝑁(𝑡) = 𝑁0𝑒<sup>−𝜆𝑡</sup>. Docling writes <!-- formula-not-decoded --> by default. With --enrich-formula it produced LaTeX, $$N ( t ) = N _ { 0 } e ^ { - \lambda t }$$, placed after the next paragraph. MinerU did best: correct LaTeX in a $$ block, in the right place, with no extra flags. Marker needed its model server for this page too.

Multi-column layouts. On a two-column page exported from Word, all five PDF tools read column one and then column two. MarkItDown kept a line break at the end of every printed line. We did not test sidebars or figures that span columns.

Tables with merged cells. A markdown pipe table has no merged cells and no line breaks inside a cell. Expect those cells to come out repeated, empty or as an HTML <table>. Azure Document Intelligence returns every table as HTML.

Images. MarkItDown does not extract images from PDFs. PyMuPDF4LLM saves them with --write-images, Marker saves them beside the markdown file, and Docling uses the same --image-export-mode flag as for Word.

Install, license and size

Tool Install License Size and dependencies
pandoc 3.12.1 brew install pandoc or winget install --source winget --exact --id JohnMacFarlane.Pandoc GPL 2.0 or later One executable, 230 MB on Windows, no Python
MarkItDown 0.1.8 pip install 'markitdown[all]' (Python 3.10 to 3.14) MIT 284 MB environment
PyMuPDF4LLM 1.28.2 pip install pymupdf4llm AGPL 3.0, or a commercial license from Artifex 197 MB environment
Docling 2.137.0 pip install docling MIT code, each model under its own license 1.1 GB with PyTorch and torchvision, plus 506 MB of models after our tests
Marker 2.0.0 pip install marker-pdf, or marker-pdf[full] for Word, PowerPoint and EPUB Apache 2.0 code. Model weights free for research, personal use and startups under $5M funding or revenue PyTorch, plus llama.cpp or Docker for OCR
MinerU 4.0.11 pip install -U "mineru>=4.0,<5" Apache 2.0 with extra terms: a commercial license above 100M monthly users or $20M monthly revenue, and online services must credit MinerU About 800 MB, plus about 2 GB of models for the standard tier, per its README
OCRmyPDF 17.13.0 pip install ocrmypdf (Python 3.11 or later) MPL 2.0 Needs Tesseract and Ghostscript installed separately

Sizes are what we measured on Windows with CPU builds, and versions are current as of October 10, 2026. Marker’s code moved from GPL to Apache 2.0 with version 2.0.0 in July, and MinerU left AGPL for its own Apache-based license in April. The AGPL on PyMuPDF4LLM applies when you ship it inside a product or service.

Docling, Marker and MinerU download their models from Hugging Face on first run, and send no document content.

Privacy: local tools and hosted services

These options send your document off your machine:

Tool Option that uploads Where it goes
MarkItDown -d -e <endpoint>, --use-cu, the markitdown-ocr plugin Azure Document Intelligence, Azure Content Understanding, the model you configure
Marker --use_llm Gemini by default (gemini-3.5-flash), or another LLM you configure
Docling --enable-remote-services (off by default) Whichever remote model you configure
MinerU --remote The mineru.net parsing service

MinerU also collects anonymous usage telemetry with no document content. Turn it off with mineru telemetry disable.

Hosted prices, checked October 10, 2026:

Service Price per 1,000 pages Notes
Mistral OCR 4.1 $4, or $2 with the Batch API Zero data retention on request for the OCR endpoint. Batch files are excluded
LlamaParse $1.25 in Fast mode (plain text), $12.50 in Agentic, $56.25 in Agentic Plus Billed at $1.25 per 1,000 credits. The free plan includes 10,000 credits a month
Datalab, Marker’s hosted API $4 fast or balanced, $10 accurate Free plan with a $10 or $20 monthly allowance
Azure Document Intelligence layout $10 Markdown output, 500 free pages a month on the F0 tier

Claude, Gemini and ChatGPT accept PDFs and will write markdown on request. Check the result against the original, because a model can reword or drop text.

Batch-convert a folder

Docling, PyMuPDF4LLM, Marker and MinerU accept a folder as input:

docling convert ./inbox --to md --output ./md
pymupdf4llm ./inbox --out ./md
marker ./inbox --output_dir ./md --workers 4
mineru-kit parse ./inbox -o ./md

Docling names each output after the input’s base name, so report.pdf and report.docx in one folder both become report.md, and one overwrites the other. PyMuPDF4LLM writes md/<name>/<name>.md plus a run-log.txt for each file, and reads only *.pdf by default, which --pattern changes.

Pandoc and MarkItDown take one file per command, so they need a loop. On macOS or Linux, give each document its own media folder, because Word names its pictures image1.png, image2.png and so on in every file:

mkdir -p md
for f in inbox/*.docx; do
  name=$(basename "$f")
  pandoc "$f" -t gfm-raw_html --wrap=none --extract-media="md/$name-media" -o "md/$name.md"
done

In PowerShell on Windows, here with MarkItDown for a mixed folder:

New-Item -ItemType Directory -Force md | Out-Null
Get-ChildItem inbox -File | Where-Object Extension -in '.pdf', '.docx' | ForEach-Object {
  markitdown $_.FullName -o "md\$($_.Name).md"
}

Read three outputs at random against the originals before you trust the rest, since a shifted column or a missing footnote raises no error.

Keeping converted documents searchable in Jotura

Jotura is a free notes app that works on a folder of plain markdown files. Drop a PDF or Word file into a vault and the desktop app writes a markdown copy beside it: contract.pdf gets contract.pdf.md. The original stays untouched, and search finds its text alongside your notes. Jotura stops regenerating a copy you have edited.

The converter is MarkItDown, bundled with the app, plus Jotura’s own readers for OpenDocument and RTF. It runs in the operating system’s sandbox. macOS uses sandbox-exec and Linux uses bubblewrap, both with network access denied, and a Linux machine without bwrap on its PATH gets no conversion and a note explaining why. Windows uses a restricted Low-integrity process, falls back to Medium integrity when Low fails, and does not block network access. The formats are pdf, docx, pptx, xls, xlsx, epub, odt, ods, odp, rtf and imported Outlook msg files.

Old binary doc and ppt files get no copy, so save them as docx or pptx first. The converter has no OCR, so a scanned PDF gets little or no text. A password-protected file gets a note saying why it failed. A PDF’s copy looks like the MarkItDown output above: searchable text without heading markup. Convert a PDF you want to edit with Marker or Docling and save the result as a note of its own.

Automatic conversion starts switched off in a vault that already has an .obsidian/ folder. The documents page covers settings and size limits, and documents from the CLI covers jotura import.

Frequently asked questions

Can pandoc convert PDF to markdown?

No. Pandoc’s list of input formats does not include PDF, and pandoc file.pdf -o file.md stops with “Pandoc can convert to PDF, but not from PDF.” Use Docling, Marker, PyMuPDF4LLM or MarkItDown for PDFs.

What is the best free docx to markdown converter?

Pandoc gave the best result in our test. Run pandoc file.docx -t gfm-raw_html --wrap=none --extract-media=media -o file.md to keep headings, tables, lists, footnotes and images, with Word equations written as LaTeX. MarkItDown came a close second and reads PowerPoint and Excel with the same command.

Why is my converted PDF empty?

An empty result means the PDF has no text layer, which is what a scan looks like. Convert it with Docling, which includes OCR, or add a text layer with ocrmypdf scan.pdf scan-text.pdf and convert the new file.

Is it safe to use an online PDF to markdown converter?

For a public document, yes. A contract, a medical record or anything under an NDA belongs in a local tool such as Docling, Marker, PyMuPDF4LLM or MarkItDown, which convert on your own machine by default. Read a hosted service’s data retention terms before you upload. Mistral, for example, offers zero data retention on request, and its Batch API files are excluded.

Does converting keep the formatting?

The better tools keep headings, lists, tables, links and footnotes. Fonts, colors, page layout, headers and footers have no markdown equivalent, and every tool drops them.

How to save a web page as markdown2026-10-10
How to save a web page as markdown with a browser extension, trafilatura, pandoc or a reader service, with tested commands and what each tool gets wrong.
Daily notes and todo lists in plain markdown2026-10-10
Daily notes in markdown: date-based file names, a short daily note template, shell shortcuts, and a markdown todo list you can search across every file.
Markdown table syntax: how to make a table in Markdown2026-10-10
Markdown table syntax with copyable examples: pipes, the separator row, alignment, escaped pipes, line breaks in cells, and where tables render.

More in Writing and organizing.

Download free

Jotura is free, and your notes stay plain markdown files you keep forever.