Paste cleaner

Your input is processed only in this browser. See the privacy policy.

How to use it

  1. Copy text from a web page, Word, Google Docs, an email or Notion and paste it into the Paste box. The formatting (HTML) is read as is, so there is no need to paste as plain text first.
  2. Check the cleaned result in the Read tab. Ads, menus, inline styles and Word leftovers are gone; headings, lists, tables, links and emphasis stay.
  3. Pick the format you need (Markdown, HTML or Text) and Copy it or download a file. Copy formatted in the Read tab pastes into email and documents as clean rich text.
  4. Turn off Cleaning options such as links, images or tables to drop them; the result updates instantly. To edit the pasted HTML itself, use the HTML source tab.

Examples

Word bullets become a real list

A bullet paragraph copied from Word is, in HTML, an ordinary paragraph that starts with a «·» character and a run of non-breaking spaces. The cleaner reads the indent level (mso-list level), rebuilds a nested list and drops the bullet character.

Links with tracking parameters

Links copied from newsletters and social apps often carry tracking parameters such as utm_source. By default only tracking parameters are removed; parameters that select the page, like id, are kept.

Input
https://example.com/post?id=7&utm_source=newsletter&fbclid=IwAR0
Result
https://example.com/post?id=7

The «everything is bold» Google Docs problem

HTML copied from Google Docs wraps the whole document in a b tag with font-weight:normal, so naive converters turn the entire document bold. The cleaner removes that wrapper and keeps bold only where the text really has font-weight:700.

Details

When you copy formatted text, the clipboard holds HTML in addition to the characters you see. From a web page that means ad slots, share buttons and tracking links; from Word it means styles starting with mso-, empty spans and chains of  ; from Google Docs it means a long style attribute on every run of text. This tool reads that HTML inside your browser and keeps only the structure that is content: headings, paragraphs, lists, tables, links, emphasis and code.

Cleaning works by rebuilding what is allowed rather than hunting for what to delete. The pasted HTML is never inserted into the page. Only allow-listed elements (p, h1–h6, lists, tables, a, strong, em, code, pre, blockquote, br, hr) are created fresh for the preview. Scripts, styles, iframes and event attributes never get through; links keep only http, https and mailto addresses, and images keep only http and https addresses. Images are not loaded—only their address is shown—so tracking pixels in emails stay closed.

Word bullet paragraphs become real lists, and bold or italic that Google Docs expresses as inline styles becomes strong and em. Layout tables, common in email newsletters, are unwrapped into their content, while real data tables become GitHub-flavoured Markdown tables. Merged cells (colspan, rowspan) are filled with empty cells because Markdown tables cannot merge cells.

Ads and menus are detected from elements such as nav, aside and footer, and from class names such as ad, share or related. A block whose name looks like an ad but holds more than 600 characters of text may be the article itself, so it is kept. If something you need disappeared, turn off Remove ads and menus.

Frequently asked questions

Is the pasted content sent to a server?

No. Reading, cleaning and converting all happen in your browser, and the input is not stored anywhere. Close or reload the page and it is gone.

I pasted, but it came in as plain text without formatting.

The app you copied from did not put HTML on the clipboard; some apps and mobile browsers do this. If you have the original HTML, paste it into the HTML source tab. Plain text is still tidied: blank lines become paragraphs and single line breaks stay line breaks.

Why are images not shown?

Images are never loaded; only their address is kept. Opening a remote image tells that server you viewed it, and it would also fire email tracking pixels. The Markdown and HTML output still contain the address, so images appear wherever you paste the result. Embedded image data (data: addresses) is not included.

Which Markdown flavour is produced?

CommonMark plus GitHub (GFM) tables. Line breaks inside a paragraph use two trailing spaces and code blocks use backtick fences. Characters like * and _ are escaped only where they would start emphasis or code, so words like snake_case stay readable.

What does tracking-parameter removal delete?

Only parameters starting with utm_ and ad or email trackers such as fbclid, gclid, msclkid and mc_cid. Parameters that select the page, such as id or page, are kept. Google Docs, Outlook Safe Links and Facebook redirect links are unwrapped to the original address.

What is Copy formatted?

It copies the cleaned result as HTML and plain text at the same time. Paste into Gmail, Google Docs or Notion and you get clean rich text with headings, lists and links; paste into a plain editor and you get text.