How to save a web page as markdown

How to save a web page as markdown with a browser extension, trafilatura, pandoc or a reader service, with tested commands and what each tool gets wrong.

To save a web page as markdown from your browser, use a clipper extension. Obsidian Web Clipper is free and open source, runs in Chrome, Firefox, Safari and Edge, and its Save file button writes the article as a .md file with no Obsidian install. MarkDownload, the clipper most older guides recommend, no longer runs in Chrome.

From a terminal, jotura clip <url> saves the article as a note in your vault, trafilatura --markdown --links -u <url> prints it in one command, and pandoc converts the whole page, menus included. Without a terminal or an extension, type defuddle.md/ in front of the address in your browser’s address bar. Every tool below was run on October 10, 2026 against a blog post with code blocks and the Wikipedia article on Markdown, which has tables and images; ... marks cuts in the output shown.

Which tool to use

Tool Type Version checked Finds the article Images Output
Obsidian Web Clipper Extension: Chromium browsers, Firefox, Safari 1.7.1, July 2026 Yes Left as web links A saved file, the clipboard, or an Obsidian vault
MarkSnip Extension: Chrome, Firefox 5.3.0 Chrome, 5.2.0 Firefox Yes Can download them File, clipboard, Obsidian
MarkDownload Extension: Firefox, Safari (paid) 3.4.0, August 2024 Yes Can download them, except on Safari File or clipboard
jotura clip Command line, ships with Jotura 0.11.1 Yes Left as web links A note in your vault
trafilatura Python command line 2.3.1, October 2026 Yes Links only, with --images Standard output or a folder
pandoc Command line 3.12.1, October 2026 No, converts everything Downloads them with --extract-media A file you name
Defuddle Node command line and defuddle.md 0.19.4, September 2026 Yes Left as web links Standard output, a file, or a web response
Jina Reader Web service Checked October 2026 Yes, after rendering JavaScript Left as web links Plain text response

Browser extensions

An extension converts the page your browser has already loaded, so it handles articles behind your own login and sites built with JavaScript.

Obsidian Web Clipper

Obsidian Web Clipper is open source under the MIT license and extracts content with Defuddle. You can highlight passages before clipping, and templates shape the note per site. The optional Interpreter fills template fields with a language model: a hosted provider with your API key, or Ollama running on your own machine with no key.

Save file writes a .md file to the folder you pick, Copy to clipboard holds the markdown for pasting, and Add to Obsidian sends the clip to the Obsidian app. Images stay linked to their web address, and the Obsidian app’s Download attachments for current file command can fetch them later. On a phone, the clipper runs in Safari on iOS and iPadOS and in Firefox on Android.

MarkDownload, and why it no longer works in Chrome

MarkDownload simplifies the page with Mozilla’s Readability library and converts it with Turndown. Its last release, 3.4.0, came out on August 23, 2024. A Manifest V3 branch stopped in December 2023, and a from-scratch rewrite begun in August 2024 last changed in June 2025. Neither shipped.

The released extension uses Manifest V2. Chrome 138 disabled Manifest V2 extensions for good in July 2025. Google then removed the remaining ones from the Chrome Web Store on August 31, 2026, according to Chrome’s timeline.

The Edge Add-ons store no longer lists it either, hence no Edge in the table. Firefox still installs version 3.2.0 from September 2022, and Safari has a paid version.

MarkSnip

MarkSnip is a fork of MarkDownload rebuilt on Manifest V3, and its Chrome Web Store listing serves version 5.3.0. It keeps the Readability and Turndown pipeline and adds image downloading, table formatting and a Download All Tabs command. The source repository its listing links to returned a 404 when checked, so you cannot audit the current code.

Joplin Web Clipper

Joplin Web Clipper saves pages into Joplin’s own database, with the Joplin desktop app running, so it writes no .md file. Getting markdown out takes Joplin’s Markdown export.

Copying part of a page

To grab one passage, Copy as Markdown (MIT, version 3.7.0 from June 2026, for Chrome, Firefox and Edge) copies the selected text as markdown, along with links, images and lists of open tabs.

Command-line tools

Command-line tools fetch the HTML the server sends. They send no cookies and run no JavaScript, so anything that appears only after a script runs or after you log in is missing.

jotura clip

Jotura is a free notes app for macOS, Windows, Linux and Android that keeps notes as markdown files on your disk. Its desktop app ships a command-line tool, and jotura clip saves a page straight into your notes folder. On macOS, put the tool on your PATH from Settings, Advanced, Command-line tool; the Windows installer and the Linux .deb do it for you.

jotura clip https://jvns.ca/blog/2024/02/16/popular-git-config-options/ --json
{
  "hash": "718501e948c3e58d7022ae079b321e05",
  "path": "Clippings/popular-git-config-options.md",
  "source": "https://jvns.ca/blog/2024/02/16/popular-git-config-options/",
  "title": "Popular git config options"
}

The note gets title, source and clipped frontmatter, then the title as a heading. On this post it kept the headings whole, the links and the table of contents, with in-page anchors left relative. --folder picks a folder other than Clippings, and --as names an exact path. A second clip of the same page becomes page (2).md, and jotura clip ~/Downloads/article.html clips a saved file; the full flag list is in documents from the CLI.

A fetch times out after 15 seconds, and a page over 8 MiB is refused after it downloads. It always sends the user agent jotura/0.11.1 and has no option to change it or add headers, so a site that blocks unknown clients has no workaround short of saving the page yourself. A redirect from http to https is refused with redirect to https://... left the original scheme.

It reads every page as UTF-8 and ignores the declared character set, so a page served as Windows-1252 or Shift_JIS comes out garbled. Images stay as links to the original site, and Jotura hides web images until you turn on Load remote images in Settings. On Wikipedia, the infobox came through as a one-column table with its labels missing.

Jotura has no browser extension, and the desktop app has no clip button. To clip in the browser, point a clipper’s Save file at your vault folder, and Jotura picks up the new .md file as a note.

trafilatura

Trafilatura is a Python tool for extracting a page’s main text, and version 2.3.1 needs Python 3.10 or later.

pip install trafilatura
trafilatura --markdown --links --with-metadata -u https://jvns.ca/blog/2024/02/16/popular-git-config-options/ > git-config.md
---
title: Popular git config options
author: Julia Evans
url: https://jvns.ca/blog/2024/02/16/popular-git-config-options/
hostname: jvns.ca
description: Popular git config options
sitename: Julia Evans
date: "2024-02-16"
fingerprint: b6b55d562832ab91
---
# Popular git config options
...
So I [asked about people’s favourite git config options on Mastodon](https://social.jvns.ca/@b0rk/111885363143321068):

Without --links, trafilatura deletes every link and keeps the link text. With it, the run above left 20 headings empty, the heading text pushed onto the next line, and every in-page anchor lost its path:

### 
      [`merge.conflictstyle zdiff3`](https://jvns.ca#merge-conflictstyle-zdiff3)

On Wikipedia, footnote links became https://en.wikipedia.org#cite_note-..., and 28 lines kept raw <sup> tags. The table of contents on the blog post vanished in both runs. The metadata is a guess: Wikipedia’s author came out as “Authority control databases International FAST National United States Israel”.

The -i flag takes a text file of URLs. To convert an HTML file you saved, pipe it in with trafilatura --markdown < article.html, or point --input-dir at a folder of saved pages.

pandoc

Pandoc is a general document converter that accepts a URL as input. It has no article detection, so it converts everything.

pandoc -f html -t gfm-raw_html https://en.wikipedia.org/wiki/Markdown -o markdown.md
Toggle the table of contents

# Markdown

33 languages

- [Alemannisch](https://als.wikipedia.org/wiki/Markdown ...)
...

The language list and the Read, Edit and View history tabs come before the article’s first sentence. Relative links such as /wiki/Talk:Markdown break once the file leaves the site. On the blog post, which has simple HTML, pandoc’s output began at the first paragraph and kept the table of contents with working anchors.

The -raw_html part tells pandoc never to write HTML tags into the markdown, which costs you tables, as the section on what breaks shows. Add --extract-media=images to download the page’s images and --request-header to send a cookie or a user agent. A saved file converts the same way: pandoc -f html -t gfm-raw_html article.html -o article.md.

Defuddle

Defuddle is the extraction library inside Obsidian Web Clipper, with its own command-line tool. It needs Node.js.

npx defuddle parse https://jvns.ca/blog/2024/02/16/popular-git-config-options/ --markdown --frontmatter

The frontmatter includes title, author, published, source and word_count. Headings came through whole, minus their code formatting, and fenced code blocks keep their language. Defuddle also removed the table of contents. Use --user-agent for a site that answers the default request with a 403.

Web services that return markdown from a URL

These need no install, and they send the address to someone else’s server, so keep private pages away from them.

defuddle.md runs Defuddle as a service: put defuddle.md/ in front of an address, in a browser or with curl. It returned the blog post with full frontmatter. The pricing page lists 1,000 free requests a month, then prepaid blocks from $0.005 a request.

curl https://defuddle.md/jvns.ca/blog/2024/02/16/popular-git-config-options/

Jina Reader works the same way with https://r.jina.ai/ and renders the page in a headless browser first, so JavaScript-built pages come through. The Reader page allows 20 requests a minute with no key and 500 with a free key, which comes with 10 million tokens. Its output starts with Title:, URL Source: and Markdown Content: lines, which you turn into frontmatter yourself.

Paste-a-URL converters such as Geekflare’s take an address in a web form, with a choice of full page or main content only, and need no account.

Cloudflare’s Markdown for Agents answers a request carrying Accept: text/markdown on sites whose owners switched it on, on Pro, Business and Enterprise plans. Its docs say headers, footers and navigation are stripped, and conversion stops at 6 MiB of HTML. On Cloudflare’s own blog post the output still opened with “Skip to content” and logo links, and its language menu arrived as 35 raw <a> tags.

curl -H "Accept: text/markdown" https://blog.cloudflare.com/markdown-for-agents/

Libraries for your own scripts

In your own code, trafilatura’s Python extract() function extracts and converts in one call. The two-step route pairs an extractor with a converter: @mozilla/readability with turndown in JavaScript, or markdownify and html2text in Python. Markdownify and html2text convert menus and all, so extract first.

What breaks, and what to do about it

Layout, video, comment threads and widgets have no markdown form in any tool. The rest depends on the tool.

Logins, paywalls and bot walls

Command-line tools and web services see what an anonymous visitor sees. jotura clip https://x.com/github exited 0 and saved a note containing “Log in or sign up for X”. A Reddit community page failed with no readable article. A zero exit code means something was saved, so open the note.

For a page you can read in your browser, use an extension, or save the page and convert the file. Trafilatura’s --archived flag retries a failed download through the Internet Archive.

Pages built by JavaScript

When a site sends an empty HTML shell and fills it with a script, command-line tools get the shell. Use an extension or Jina Reader, which see the rendered page.

Tables

Markdown table syntax has no merged cells and no paragraphs inside cells. On Wikipedia, pandoc -t gfm-raw_html replaced all seven tables with the literal text [TABLE]. Plain -t gfm keeps them as raw HTML, along with hundreds of lines of <div> and <span> tags.

In the infobox, trafilatura left the Developed by and Latest release values empty, and jotura clip dropped every label. Markdown tables explains what the syntax can hold.

Code blocks

Multi-line code blocks came through fenced in every tool tested. On a test block marked language-python, pandoc and Defuddle kept python. Trafilatura and jotura clip wrote a bare fence. Trafilatura also turned a one-line code block into inline code.

Keeping images on your own disk

Image links break when the site moves or deletes the files. Three routes keep copies:

Route How Where images go
pandoc pandoc -f html -t gfm-raw_html --extract-media=images <url> -o page.md images/, files renamed to SHA-1 hashes
MarkSnip or MarkDownload Turn on Download Images in the options A folder named after the markdown file, by default
Obsidian app Clip with Web Clipper, then run Download attachments for current file The vault’s attachment folder

Pandoc’s --extract-media also works on a saved page: it copies the images from the saved _files folder and rewrites the links. MarkDownload’s image download needs its Downloads API mode, which Safari lacks.

Saving many pages at once

Trafilatura’s batch mode (-i urls.txt -o clips) names every file with random characters and a .txt extension, such as kItMeyl-g2q1n7J-.txt, and --input-dir does the same. A loop names each file after the last part of its URL:

mkdir -p clips
while read -r url; do
  name=$(basename "${url%/}")
  trafilatura --markdown --links --with-metadata -u "$url" > "clips/$name.md" || echo "failed: $url" >&2
done < urls.txt

jotura clip names each note from the page title, so the same loop needs only jotura clip "$url" --folder Reading/queue in its body. In the browser, MarkSnip’s Download All Tabs saves every open tab as its own file. Keep the source address in each file so you can check the snapshot against the live page.

For PDFs and Word files, see converting PDF and Word to markdown.

Frequently asked questions

Can I save a web page as markdown without a browser extension?

Yes. Type defuddle.md/ or r.jina.ai/ in front of the address in your browser, or paste the URL into a converter such as Geekflare’s. On the command line, jotura clip, trafilatura and Defuddle extract the article, and pandoc converts the whole page.

Does MarkDownload still work?

Not in Chrome or Edge. Its Manifest V2 code stopped running in Chrome in July 2025, and Google removed such extensions from the Chrome Web Store on August 31, 2026. The Firefox listing installs version 3.2.0 from 2022. MarkSnip is a Manifest V3 fork that installs from the Chrome Web Store.

Does Jotura have a web clipper extension?

No. Jotura has no browser extension and no clip button in the app. It clips with jotura clip on the command line. To clip in the browser, point Obsidian Web Clipper’s Save file at your vault folder, and the notes appear in Jotura.

Why did my clipped page come out empty or wrong?

Command-line tools read only the HTML the server returns. A page built by JavaScript comes out empty, and a page behind a login or bot wall comes out as the wall. Use a browser extension, or Jina Reader for public JavaScript pages.

Which tool keeps images with the markdown?

Pandoc downloads them with --extract-media. MarkSnip and MarkDownload download them when Download Images is on, except MarkDownload on Safari. Obsidian Web Clipper leaves them as links, and Obsidian’s Download attachments for current file command fetches them into the vault afterwards. Trafilatura, Defuddle, Jina Reader and jotura clip leave images as links to the original site.

How to convert PDF and Word to markdown: six tools tested2026-10-10
How to convert PDF to markdown and Word docx to markdown, with real output from pandoc, MarkItDown, Docling, Marker, MinerU and PyMuPDF4LLM on the same files.
Daily notes and todo lists in plain markdown2026-10-10
Daily notes in markdown: date-based file names, a short daily note template, shell shortcuts, and a markdown todo list you can search across every file.
Markdown table syntax: how to make a table in Markdown2026-10-10
Markdown table syntax with copyable examples: pipes, the separator row, alignment, escaped pipes, line breaks in cells, and where tables render.

More in Writing and organizing.

Download free

Jotura is free, and your notes stay plain markdown files you keep forever.