Paste HTML and get the plain text. Nothing leaves your browser.
The tool scans the text from start to end and takes out every tag, such as an opening p, a closing div or a link with its href, including doctype lines and xml declarations. Comments, written between the markers that open with an exclamation mark and two dashes, are removed with everything inside them. The content of script and style elements is removed too, not only their tags, so JavaScript and CSS do not appear in your text. Quoted attribute values are skipped correctly, so a link whose title attribute contains a greater-than sign still ends at the right place. A less-than sign that is not followed by a letter, a slash and a letter, an exclamation mark or a question mark, such as in 1 < 2, is plain text and is kept. If a tag never closes, because no greater-than sign follows it, the less-than sign and everything after it are kept as text.
After the tags are gone, entities in the remaining text are turned into characters: &, <, >, ", ', (becomes an ordinary space), common ones such as ©, —, € and é, and numeric ones such as é and é. Decoding happens once, so &lt; becomes < and not <, and text that was escaped stays as text and is never mistaken for a tag. Names this tool does not know are left as typed. With Keep line breaks on, a br tag always starts a new line, and block tags such as p, div, h1 to h6, li, tr and table start a new line unless one has just started. With it off, those tags become a space so words do not run together. The end of a table cell adds a space.
Source text keeps its spaces and line breaks exactly as written unless you turn on Collapse whitespace. Then every run of spaces, tabs and line breaks in the text becomes one space, which is how a browser would show it. Line breaks that come from tags are kept if Keep line breaks is on, spaces around them are trimmed, and the whole result is trimmed. HTML text is often indented, so this setting is the quick way to get tidy output from a page source.
This is a text clean-up. It is not a security tool and its output must not be treated as safe HTML. It does not parse HTML the way a browser does, so broken or hostile markup can be handled differently, and decoded text such as the entity for a less-than sign becomes a real less-than character in the result, so escaped markup comes out looking like markup. Never put the output into a web page without escaping it. The tool uses no browser HTML parser. It scans the text itself, in this tab, and nothing is uploaded or saved.
Yes. Everything from the opening script or style tag to its closing tag is removed. If the closing tag is missing, the rest of the text is removed. Text inside other tags, such as a paragraph or a link, is kept.
The tool decodes about 40 common named entities and all numeric ones, including é and é. A name it does not know is left as typed. A missing semicolon also stops decoding, so & without ; stays as it is.
No. It is a text clean-up. Removing tags from a string is not a reliable defence against script injection, and the decoded output can contain less-than and greater-than characters. Use an escaping or sanitising library on the server or in your framework for that.
Leave Keep line breaks on, which is the default. Each p, div, heading and list item starts a new line, and a br tag always does. Turn Collapse whitespace on as well to remove the indentation and blank lines left from the source.
No. The scan runs in your browser, in this tab, with plain string code and no HTML parser. No server receives your text, and the text is not saved, so closing the tab clears it.