How to convert PDF and Word to markdown: six tools tested
How to convert PDF to markdown and Word docx to markdown, with real output from pandoc, MarkItDown, Docling, Marker, MinerU and PyMuPDF4LLM on the same files.
To convert a Word file to markdown, use pandoc: pandoc report.docx -t gfm --wrap=none -o report.md. In our test it kept the headings, the table, the list and the footnote, and it was the only tool to write the footnote in markdown footnote syntax. Microsoft’s MarkItDown came a close second with no extra flags.
To convert a PDF to markdown, use Docling when the file might be scanned. For a digital PDF, Marker gave the cleanest result, and PyMuPDF4LLM is the fast, light option. Pandoc cannot read PDF at all. Every tool here runs on your own machine and is free for personal use, with license limits for larger companies noted below. Hosted services such as Mistral OCR charge $4 per 1,000 pages.
Which converter to use
| Your file | Use | Why |
|---|---|---|
Word .docx |
pandoc | Keeps headings, tables, lists, footnotes, equations and images, with no GPU and no models |
| Many formats in one tool | MarkItDown or MinerU | Both read Word, PowerPoint, Excel, EPUB and PDF |
| Digital PDF, best structure | Marker | Every heading, the linked footnote and the table came through in our test |
| Scanned PDF | Docling | Built-in OCR recovered the whole test page, table included |
| Digital PDF, small install | PyMuPDF4LLM | Two seconds per file in our test, no PyTorch |
Old binary .doc |
MinerU, or LibreOffice then pandoc | Pandoc and MarkItDown only read .docx |
| Confidential files | Any local tool above | Nothing leaves your machine with the cloud options left off |
| Documents kept beside your notes and searchable with them | Jotura | Writes a markdown copy of each PDF or Word file in the vault with its bundled MarkItDown, no OCR |
Why Word is easy and PDF is hard
A .docx file tags each heading with its level and marks tables and footnotes, so a converter only translates labels.
A PDF stores characters at positions on a page, and the converter has to guess which lines are headings and where table cells end. A scanned PDF holds only pictures of pages, so it needs OCR (optical character recognition) first.
MarkItDown pulls text out with pdfminer, detects tables with pdfplumber and makes no attempt at headings. The other PDF tools run layout models, neural networks trained to recognize page regions.
How we tested
We built a one-page report in Microsoft Word 365: a Heading 1, a paragraph with a footnote, a Heading 2, a three-column table with grid lines, another Heading 2 and three bullet points. We saved it as report.docx and exported report.pdf from Word. A third file, scan.pdf, is that page rendered as a 200 dpi grayscale image with no text layer, which is what a scanner produces.
The tools ran on Windows 11 on CPU only, with default settings: MarkItDown 0.1.8, Docling 2.137.0, PyMuPDF4LLM 1.28.2, Marker 2.0.0, MinerU 4.0.11, and pandoc 3.9, the build bundled with pypandoc-binary 1.17. Pandoc’s current release is 3.12.1. Outputs below are copied from the files the tools wrote.
Word to markdown: the results
| Tool | Headings | Table | Bullet list | Footnote |
|---|---|---|---|---|
| pandoc | Correct levels | Yes | Yes, with blank lines between items | Markdown footnote [^1] |
| MarkItDown | Correct levels | Empty header row added | Yes | Link to a numbered note at the end |
| Docling | Shifted down one level | Yes | Yes | Dropped |
Pandoc’s output with --wrap=none, complete:
# Quarterly field report
Soil samples were collected at three sites in March.[^1] Results are summarized below.
## Results by site
| **Site** | **pH** | **Nitrogen (mg/kg)** |
|----------|--------|----------------------|
| North | 6.4 | 112 |
| River | 7.1 | 87 |
| Ridge | 5.9 | 140 |
## Next steps
- Retest the Ridge site in June
- Order lime for the North plot
- Share the data with the county office
[^1]: Site coordinates are listed in appendix B.
MarkItDown kept the headings and wrote a tight list. Its table and footnote came out like this (trimmed):
Soil samples were collected at three sites in March.[[1]](#footnote-1) Results are summarized below.
| | | |
| --- | --- | --- |
| **Site** | **pH** | **Nitrogen (mg/kg)** |
| North | 6.4 | 112 |
1. Site coordinates are listed in appendix B. [↑](#footnote-ref-1)
Docling demoted every heading one level and dropped the footnote. MinerU 4 reads DOC, DOCX, PPTX, XLSX, RTF, ODT and EPUB too. We did not run it on the Word file.
The empty table header
Word marks a header row in two places. Header Row in the Table Design tab is on by default for a table made with Insert > Table. Repeat Header Rows sits in the Table Layout tab and is off by default. Pandoc reads either setting. MarkItDown reads only Repeat Header Rows, which is why it wrote the blank header above.
Select the first row, turn on Repeat Header Rows, save, and convert again. Both tools then used the first row as the header. A table with Header Row switched off gets the empty row in both tools until you do this. The markdown tables guide covers what pipe tables can hold.
Equations, images and tracked changes
We added a Word equation, N(t) = N₀e^(−λt), to a second test file. Pandoc wrote it as LaTeX in a fenced math block, which GitHub renders. MarkItDown wrote $$N\left(t\right)=N\_{0}e^{-λt}$$, with a stray backslash before the underscore and a raw λ in place of \lambda. Docling gave standard LaTeX: $$N\left(t\right)=N_{0}e^{- \lambda t}$$.
Pandoc’s --extract-media=media copied our picture to media/media/image1.png and wrote an HTML <img> tag to keep Word’s size. Add -t gfm-raw_html to get a plain  link. MarkItDown writes  with the data cut off, and --keep-data-uris keeps it. Docling embeds images as base64 by default, and --image-export-mode referenced writes them to a <name>_artifacts folder.
Pandoc’s --track-changes option takes accept (the default), reject or all, and all keeps insertions, deletions and comments as marked spans. The full command we recommend for Word files:
pandoc report.docx -t gfm-raw_html --wrap=none --extract-media=report-media -o report.md
--wrap=none stops pandoc from breaking paragraphs at 72 characters, its default.
Old .doc files and no-install routes
Pandoc and MarkItDown read only .docx. Save an old .doc as .docx in Word, or convert it with LibreOffice and then run pandoc:
soffice --headless --convert-to docx report.doc
MinerU 4 reads .doc directly.
With nothing installed, open the file in Google Docs and choose File > Download > Markdown (.md), which sends it to Google. Browser converters such as word2md.com for Word and ConvertCase’s PDF to Markdown page state that conversion runs in your browser and nothing is uploaded. ConvertCase also says it does no OCR and does not reliably preserve tables. You cannot check those claims, so keep confidential files on a local tool.
PDF to markdown: the results
| Tool | Headings | Table | Bullet list | Footnote |
|---|---|---|---|---|
| Marker | All three, all ## |
Yes | Yes | Linked superscript and anchor |
| Docling | All three, all ## |
Yes | Yes | Plain text |
| PyMuPDF4LLM | Two of three | Yes | Yes | <sup>1</sup> and a blockquote |
| MinerU | Two of three | Yes | Lost, items became paragraphs | HTML <small> span |
| MarkItDown | None | Yes | • characters |
Plain text |
| pandoc | Cannot read PDF |
The ruled table converted correctly in every tool that reads PDF. Headings, lists and footnotes are where they differed.
Marker’s output was the cleanest. It flattened heading levels like Docling and kept everything else (trimmed):
## Quarterly field report
Soil samples were collected at three sites in March.[<sup>1</sup>](#page-0-0) Results are summarized below.
## Results by site
| Site | pH | Nitrogen (mg/kg) |
|-------|-----|------------------|
| North | 6.4 | 112 |
## Next steps
- Retest the Ridge site in June
<span id="page-0-0"></span><sup>1</sup> Site coordinates are listed in appendix B.
MarkItDown returned the text and a correct table, with no headings and literal bullet characters (trimmed):
Quarterly field report
Soil samples were collected at three sites in March.1 Results are summarized below.
Results by site
Next steps
• Retest the Ridge site in June
PyMuPDF4LLM and MinerU both left “Results by site” as plain text. PyMuPDF4LLM wrote the footnote as > 1 Site coordinates are listed in appendix B. MinerU turned the three bullet points into separate paragraphs with no list markers and wrapped the footnote in a styled HTML span.
Borderless tables
We exported the same table from Word with no borders, twice. Spread across the full page width, it still converted cleanly in MarkItDown and PyMuPDF4LLM. Shrunk to fit its contents, it broke every tool we ran on it. MarkItDown printed loose lines:
Site
pH Nitrogen (mg/kg)
North 6.4 112
PyMuPDF4LLM folded the whole table into one line, Site pH Nitrogen (mg/kg) North 6.4 112 River 7.1 87 Ridge 5.9 140. Docling did the same and promoted the header row to a ## heading. We did not run Marker or MinerU on this file. Check every borderless table by eye, and rebuild a broken one from the source data.
Scanned PDFs and OCR
| Tool | Result on scan.pdf |
|---|---|
| Docling | Full text, headings, table and list recovered |
| MinerU | Text and table recovered, list markers mixed \- and •, footnote number as $^{1}$ |
| PyMuPDF4LLM | Empty, with no OCR engine installed |
| MarkItDown | Empty |
| Marker | Not run, see below |
Docling’s scan result matched its digital-PDF result almost exactly. Its default, --ocr-engine auto, uses the first engine it finds installed, and the standard install includes RapidOCR.
PyMuPDF4LLM runs OCR through Tesseract or RapidOCR when one is installed alongside it. MarkItDown has no local OCR. Its markitdown-ocr plugin sends pages to whatever OpenAI-compatible model you configure, a cloud service unless you point it at a model running on your own machine.
Marker 2 defaults to Fast mode on CPU and Apple Silicon, which reads the PDF’s text layer, and to Balanced mode on a GPU, which runs the Surya vision model over the whole page. OCR and equations go through a local model server: vLLM in Docker on an NVIDIA GPU, or llama.cpp’s llama-server elsewhere. --disable_ocr turns off every model call. Our machine has an NVIDIA card and Docker was not running, so we have no Marker result for the scan.
A PDF whose text you cannot select in a viewer is a scan. OCRmyPDF adds a text layer with Tesseract, after which any tool above can read it:
ocrmypdf scan.pdf scan-text.pdf
Equations, multi-column pages and images in PDFs
Equations. MarkItDown flattened our equation to 𝑁(𝑡) = 𝑁0𝑒−𝜆𝑡, losing the superscript. PyMuPDF4LLM kept it as 𝑁(𝑡) = 𝑁0𝑒<sup>−𝜆𝑡</sup>. Docling writes <!-- formula-not-decoded --> by default. With --enrich-formula it produced LaTeX, $$N ( t ) = N _ { 0 } e ^ { - \lambda t }$$, placed after the next paragraph. MinerU did best: correct LaTeX in a $$ block, in the right place, with no extra flags. Marker needed its model server for this page too.
Multi-column layouts. On a two-column page exported from Word, all five PDF tools read column one and then column two. MarkItDown kept a line break at the end of every printed line. We did not test sidebars or figures that span columns.
Tables with merged cells. A markdown pipe table has no merged cells and no line breaks inside a cell. Expect those cells to come out repeated, empty or as an HTML <table>. Azure Document Intelligence returns every table as HTML.
Images. MarkItDown does not extract images from PDFs. PyMuPDF4LLM saves them with --write-images, Marker saves them beside the markdown file, and Docling uses the same --image-export-mode flag as for Word.
Install, license and size
| Tool | Install | License | Size and dependencies |
|---|---|---|---|
| pandoc 3.12.1 | brew install pandoc or winget install --source winget --exact --id JohnMacFarlane.Pandoc |
GPL 2.0 or later | One executable, 230 MB on Windows, no Python |
| MarkItDown 0.1.8 | pip install 'markitdown[all]' (Python 3.10 to 3.14) |
MIT | 284 MB environment |
| PyMuPDF4LLM 1.28.2 | pip install pymupdf4llm |
AGPL 3.0, or a commercial license from Artifex | 197 MB environment |
| Docling 2.137.0 | pip install docling |
MIT code, each model under its own license | 1.1 GB with PyTorch and torchvision, plus 506 MB of models after our tests |
| Marker 2.0.0 | pip install marker-pdf, or marker-pdf[full] for Word, PowerPoint and EPUB |
Apache 2.0 code. Model weights free for research, personal use and startups under $5M funding or revenue | PyTorch, plus llama.cpp or Docker for OCR |
| MinerU 4.0.11 | pip install -U "mineru>=4.0,<5" |
Apache 2.0 with extra terms: a commercial license above 100M monthly users or $20M monthly revenue, and online services must credit MinerU | About 800 MB, plus about 2 GB of models for the standard tier, per its README |
| OCRmyPDF 17.13.0 | pip install ocrmypdf (Python 3.11 or later) |
MPL 2.0 | Needs Tesseract and Ghostscript installed separately |
Sizes are what we measured on Windows with CPU builds, and versions are current as of October 10, 2026. Marker’s code moved from GPL to Apache 2.0 with version 2.0.0 in July, and MinerU left AGPL for its own Apache-based license in April. The AGPL on PyMuPDF4LLM applies when you ship it inside a product or service.
Docling, Marker and MinerU download their models from Hugging Face on first run, and send no document content.
Privacy: local tools and hosted services
These options send your document off your machine:
| Tool | Option that uploads | Where it goes |
|---|---|---|
| MarkItDown | -d -e <endpoint>, --use-cu, the markitdown-ocr plugin |
Azure Document Intelligence, Azure Content Understanding, the model you configure |
| Marker | --use_llm |
Gemini by default (gemini-3.5-flash), or another LLM you configure |
| Docling | --enable-remote-services (off by default) |
Whichever remote model you configure |
| MinerU | --remote |
The mineru.net parsing service |
MinerU also collects anonymous usage telemetry with no document content. Turn it off with mineru telemetry disable.
Hosted prices, checked October 10, 2026:
| Service | Price per 1,000 pages | Notes |
|---|---|---|
| Mistral OCR 4.1 | $4, or $2 with the Batch API | Zero data retention on request for the OCR endpoint. Batch files are excluded |
| LlamaParse | $1.25 in Fast mode (plain text), $12.50 in Agentic, $56.25 in Agentic Plus | Billed at $1.25 per 1,000 credits. The free plan includes 10,000 credits a month |
| Datalab, Marker’s hosted API | $4 fast or balanced, $10 accurate | Free plan with a $10 or $20 monthly allowance |
| Azure Document Intelligence layout | $10 | Markdown output, 500 free pages a month on the F0 tier |
Claude, Gemini and ChatGPT accept PDFs and will write markdown on request. Check the result against the original, because a model can reword or drop text.
Batch-convert a folder
Docling, PyMuPDF4LLM, Marker and MinerU accept a folder as input:
docling convert ./inbox --to md --output ./md
pymupdf4llm ./inbox --out ./md
marker ./inbox --output_dir ./md --workers 4
mineru-kit parse ./inbox -o ./md
Docling names each output after the input’s base name, so report.pdf and report.docx in one folder both become report.md, and one overwrites the other. PyMuPDF4LLM writes md/<name>/<name>.md plus a run-log.txt for each file, and reads only *.pdf by default, which --pattern changes.
Pandoc and MarkItDown take one file per command, so they need a loop. On macOS or Linux, give each document its own media folder, because Word names its pictures image1.png, image2.png and so on in every file:
mkdir -p md
for f in inbox/*.docx; do
name=$(basename "$f")
pandoc "$f" -t gfm-raw_html --wrap=none --extract-media="md/$name-media" -o "md/$name.md"
done
In PowerShell on Windows, here with MarkItDown for a mixed folder:
New-Item -ItemType Directory -Force md | Out-Null
Get-ChildItem inbox -File | Where-Object Extension -in '.pdf', '.docx' | ForEach-Object {
markitdown $_.FullName -o "md\$($_.Name).md"
}
Read three outputs at random against the originals before you trust the rest, since a shifted column or a missing footnote raises no error.
Keeping converted documents searchable in Jotura
Jotura is a free notes app that works on a folder of plain markdown files. Drop a PDF or Word file into a vault and the desktop app writes a markdown copy beside it: contract.pdf gets contract.pdf.md. The original stays untouched, and search finds its text alongside your notes. Jotura stops regenerating a copy you have edited.
The converter is MarkItDown, bundled with the app, plus Jotura’s own readers for OpenDocument and RTF. It runs in the operating system’s sandbox. macOS uses sandbox-exec and Linux uses bubblewrap, both with network access denied, and a Linux machine without bwrap on its PATH gets no conversion and a note explaining why. Windows uses a restricted Low-integrity process, falls back to Medium integrity when Low fails, and does not block network access. The formats are pdf, docx, pptx, xls, xlsx, epub, odt, ods, odp, rtf and imported Outlook msg files.
Old binary doc and ppt files get no copy, so save them as docx or pptx first. The converter has no OCR, so a scanned PDF gets little or no text. A password-protected file gets a note saying why it failed. A PDF’s copy looks like the MarkItDown output above: searchable text without heading markup. Convert a PDF you want to edit with Marker or Docling and save the result as a note of its own.
Automatic conversion starts switched off in a vault that already has an .obsidian/ folder. The documents page covers settings and size limits, and documents from the CLI covers jotura import.
Frequently asked questions
Can pandoc convert PDF to markdown?
No. Pandoc’s list of input formats does not include PDF, and pandoc file.pdf -o file.md stops with “Pandoc can convert to PDF, but not from PDF.” Use Docling, Marker, PyMuPDF4LLM or MarkItDown for PDFs.
What is the best free docx to markdown converter?
Pandoc gave the best result in our test. Run pandoc file.docx -t gfm-raw_html --wrap=none --extract-media=media -o file.md to keep headings, tables, lists, footnotes and images, with Word equations written as LaTeX. MarkItDown came a close second and reads PowerPoint and Excel with the same command.
Why is my converted PDF empty?
An empty result means the PDF has no text layer, which is what a scan looks like. Convert it with Docling, which includes OCR, or add a text layer with ocrmypdf scan.pdf scan-text.pdf and convert the new file.
Is it safe to use an online PDF to markdown converter?
For a public document, yes. A contract, a medical record or anything under an NDA belongs in a local tool such as Docling, Marker, PyMuPDF4LLM or MarkItDown, which convert on your own machine by default. Read a hosted service’s data retention terms before you upload. Mistral, for example, offers zero data retention on request, and its Batch API files are excluded.
Does converting keep the formatting?
The better tools keep headings, lists, tables, links and footnotes. Fonts, colors, page layout, headers and footers have no markdown equivalent, and every tool drops them.