XML Escape and Unescape

Escape special characters, or decode entities back to text.

Handles the five predefined entities and numeric references in decimal and hex.
Input
Output
WaitingPaste a document to check it. Validation runs as you type.

Everything runs in this tab. Nothing you paste is uploaded, logged or sent anywhere. Open your network panel and check.

Paste text on the left and it comes back with every character XML treats as markup replaced by an entity reference. Reverse the direction and references are decoded back to the characters they name. The input is plain text rather than a document, so it does not have to parse: a fragment, one attribute value or a URL on its own all work.

You reach for it after a parser has complained about an ampersand. "The entity name must immediately follow the '&'" from Xerces, "EntityRef: expecting ';'" from libxml2 and "invalid character in entity name" from Expat are all one problem: a raw & in text or in an attribute, usually from a query string like ?a=1&b=2 pasted into an element.

What is different here is what the tool refuses to do. XML defines five named entities and no others. Most escapers on the web are HTML escapers with an XML label and will hand you  , which no XML parser accepts. This one knows the five and numeric references in decimal and hex, leaves anything else exactly as written, and uploads nothing.

Five predefined entities, and no more

XML 1.0 section 4.6 defines exactly five named entities. That is the whole set. There is no inherited HTML entity table, which is the usual surprise when a block of HTML is moved into an XML config file or an RSS description.

Writing   in a document with no DTD is not a stylistic problem, it is a well-formedness error: the parser has a name and nothing tells it what the name means. Use   or   instead, and the same for © and é. A DTD can declare extra names, which is how &companyName; works in DocBook, so the unescape direction leaves any name it does not recognise exactly as written.

  • & for &
  • &lt; for <
  • &gt; for >
  • &quot; for "
  • &apos; for '. HTML 4 never defined &apos;, so serialisers feeding mixed pipelines often emit &#39; instead.

Numeric character references

Any legal character can be written as its code point instead: &#233; in decimal, &#xE9; in hex. The two are the same character, leading zeros are allowed, and the x must be lowercase, so &#X41; is not a character reference and is rejected as an undeclared entity.

The unescape direction handles both forms and checks the result is inside the Unicode range; a malformed or out-of-range reference is left as written rather than replaced with a question mark. The escape direction never emits numeric references: UTF-8 carries accents, CJK and emoji as themselves, so escaping them costs readability and buys nothing.

When each character actually has to be escaped

The rules are narrower than most people assume, which matters when you are reading someone else's document and deciding whether it is broken. The & and the < are mandatory everywhere. The > is required in one place only: XML 1.0 section 2.4 says it must be escaped inside the literal sequence ]]> in content when that is not closing a CDATA section, so <code>if (a]]&gt;b)</code> is required and <note>a > b</note> is legal.

This tool escapes all five unconditionally, a superset of what any one context needs, so the output is safe to drop anywhere without tracking which context you are in. It is not the minimal escaping, and you now know which characters to put back.

  • & and <: required in element content and in attribute values, always.
  • >: optional, except inside the sequence ]]>.
  • ": required only inside a double-quoted attribute value.
  • ': required only inside a single-quoted attribute value. Neither quote needs escaping in element content.

The characters that cannot be escaped at all

XML 1.0 defines the characters a document may contain, the Char production, and most of the C0 control range is not in it: U+0001 to U+0008, U+000B, U+000C and U+000E to U+001F. Only tab, line feed and carriage return survive. U+0000 is illegal in every version of XML, and lone surrogates are illegal too.

So &#x1B; is not an escape for the ESC character. It violates the constraint Legal Character, because the character it names cannot appear in an XML 1.0 document by any means, and writing it as a reference does not launder it. Python reporting "not well-formed (invalid token)" on a log file full of ANSI escape sequences is exactly this; base64 is the answer for arbitrary bytes.

The tool is honest about this rather than clever: escaping passes a control character straight through, and unescaping decodes &#1; into a real U+0001, because that is what the reference says. Run the result through the syntax checker before putting it back into a document.

Doing this in code

Every language ships something for this and most ship the wrong thing, because the obvious function is an HTML escaper. These use the XML-specific path in each ecosystem, and say what each one gets wrong.

// XML defines exactly five named entities. A single regex pass avoids the
// classic bug of replacing & after < and turning &lt; into &amp;lt;.
const ESCAPES = { '&': '&amp;', '<': '&lt;', '>': '&gt;', '"': '&quot;', "'": '&apos;' };

// Element content: & and < are mandatory, > only inside the sequence ]]>.
function escapeText(s) {
  return s.replace(/[&<]/g, (c) => ESCAPES[c]).replace(/]]>/g, ']]&gt;');
}

// Attribute value: escape the delimiter you are using, plus tab, newline and
// carriage return. Attribute-value normalisation turns a literal one of those
// into a plain space, so a reference is the only way to keep it.
function escapeAttribute(s) {
  return s
    .replace(/[&<"]/g, (c) => ESCAPES[c])
    .replace(/\t/g, '&#9;')
    .replace(/\n/g, '&#10;')
    .replace(/\r/g, '&#13;');
}

// The reverse. Note the deliberate absence of an HTML entity table: &nbsp; is
// undefined in XML, so leaving it alone is more honest than decoding it.
const NAMED = { amp: '&', lt: '<', gt: '>', quot: '"', apos: "'" };

function unescapeXml(s) {
  return s.replace(/&(#x[0-9a-fA-F]+|#[0-9]+|[A-Za-z][A-Za-z0-9]*);/g, (whole, body) => {
    if (body[0] === '#') {
      const hex = body[1] === 'x';
      const cp = parseInt(hex ? body.slice(2) : body.slice(1), hex ? 16 : 10);
      return cp >= 0 && cp <= 0x10ffff ? String.fromCodePoint(cp) : whole;
    }
    return NAMED[body] !== undefined ? NAMED[body] : whole;
  });
}
from xml.sax.saxutils import escape, quoteattr, unescape
import re

# escape() handles & < > only and knows nothing about quotes. The > is not
# strictly required but it is harmless and covers the ]]> case for free.
text = escape('Terms & conditions <see clause 4>')
# Terms &amp; conditions &lt;see clause 4&gt;

# The two quote entities have to be supplied yourself.
FIVE = {'"': '&quot;', "'": '&apos;'}
text = escape(source, FIVE)

# quoteattr() returns the value WITH its delimiters, picking single quotes when
# the value contains a double quote, and turning tab, newline and carriage
# return into numeric references so normalisation cannot flatten them.
attr = quoteattr('say "hello" then stop')
# 'say "hello" then stop'   <- single-quoted, so the " needs no escape

# unescape() reverses & < > plus whatever mapping you pass. It does NOT decode
# numeric character references, so &#233; comes back unchanged. html.unescape()
# does decode them, but it also decodes &nbsp; and 250 other HTML names XML
# never defined, which silently rewrites your data. Do it explicitly instead:
def unescape_xml(s: str) -> str:
    s = unescape(s, {'&quot;': '"', '&apos;': "'"})
    return re.sub(
        r'&#(x[0-9a-fA-F]+|[0-9]+);',
        lambda m: chr(int(m.group(1)[1:], 16) if m.group(1)[0] == 'x' else int(m.group(1))),
        s,
    )
// The JDK has no public XML escaper for a bare string. Two right answers.

// 1. Let a writer do it. XMLStreamWriter escapes for the context it is
//    writing into, the only approach that gets attribute values right.
import javax.xml.stream.XMLOutputFactory;
import javax.xml.stream.XMLStreamWriter;

XMLStreamWriter w = XMLOutputFactory.newInstance().createXMLStreamWriter(out);
w.writeStartElement("note");
w.writeAttribute("href", "https://example.com/?a=1&b=2");  // escaped for you
w.writeCharacters("Terms & conditions <apply>");           // escaped for you
w.writeEndElement();
w.close();

// 2. Escape a standalone string with commons-text. Do NOT use commons-lang3's
//    deprecated StringEscapeUtils.escapeXml, which knew nothing about the
//    characters XML cannot represent.
import org.apache.commons.text.StringEscapeUtils;

String safe = StringEscapeUtils.escapeXml10("Terms & conditions <apply>");
// Terms &amp; conditions &lt;apply&gt;

// escapeXml10 does something its name does not advertise: it DELETES the
// characters XML 1.0 cannot hold (U+0001 to U+0008, U+000B, U+000C,
// U+000E to U+001F, unpaired surrogates) rather than escaping them, because
// no escape for them exists. escapeXml11 writes the control characters as
// numeric references instead, which is only legal if the document declares
// version="1.1", and it still drops U+0000 and unpaired surrogates.

String back = StringEscapeUtils.unescapeXml(safe);
// The five predefined entities plus decimal and hex character references.
// Names it does not know are left as written.
using System.Linq;
using System.Security;
using System.Xml;

// SecurityElement.Escape does the five predefined entities in one call:
// < > & " ' every time, a superset of what any single context needs.
string safe = SecurityElement.Escape("Terms & conditions <apply>");
// Terms &amp; conditions &lt;apply&gt;

// It does not remove the characters XML cannot carry, so check those yourself.
// XmlConvert.IsXmlChar implements the Char production from XML 1.0. Keep
// surrogates or astral characters (emoji, rarer CJK) are destroyed.
static string StripIllegal(string s) =>
    string.Concat(s.Where(c => XmlConvert.IsXmlChar(c) || char.IsSurrogate(c)));

// In practice prefer a writer. It escapes for the context and throws on a
// character the document cannot hold, instead of producing a file that fails
// to parse somewhere downstream.
var settings = new XmlWriterSettings { Indent = true, CheckCharacters = true };
using var writer = XmlWriter.Create(Console.Out, settings);
writer.WriteStartElement("note");
writer.WriteAttributeString("href", "https://example.com/?a=1&b=2");
writer.WriteString("Terms & conditions <apply>");
writer.WriteEndElement();

// There is no framework unescaper, and WebUtility.HtmlDecode is the wrong
// tool: it decodes &nbsp; and the rest of the HTML set. Read a fragment
// instead, which decodes exactly what XML defines and nothing else.
static string Unescape(string escaped)
{
    var rs = new XmlReaderSettings
    {
        DtdProcessing = DtdProcessing.Prohibit,
        XmlResolver = null,
    };
    using var r = XmlReader.Create(new StringReader("<r>" + escaped + "</r>"), rs);
    r.ReadToFollowing("r");
    return r.ReadElementContentAsString();
}
<?php
// ENT_XML1 is the flag that matters and the one everybody omits. Without it
// you get the HTML entity table: an apostrophe becomes &#039; (harmless), and
// with ENT_HTML5 you can get names XML will reject outright.
$safe = htmlspecialchars(
    $text,
    ENT_XML1 | ENT_QUOTES | ENT_SUBSTITUTE,
    'UTF-8'
);
// & < > " ' become &amp; &lt; &gt; &quot; &apos;
//
// ENT_SUBSTITUTE replaces invalid UTF-8 with U+FFFD. Before PHP 8.1 the
// default was to return an empty string on invalid input, which is very easy
// to miss in a feed generator fed by a legacy database.

$back = html_entity_decode($safe, ENT_XML1 | ENT_QUOTES, 'UTF-8');
// With ENT_XML1 this decodes the five predefined entities and numeric
// character references and leaves &nbsp; alone. Drop the flag and it decodes
// &nbsp; into U+00A0, silently rewriting your data.

// Building a document rather than a string: DOMDocument escapes on save and
// rejects characters XML cannot represent, which htmlspecialchars passes
// straight through.
$doc = new DOMDocument('1.0', 'UTF-8');
$note = $doc->createElement('note');
$note->appendChild($doc->createTextNode($text));
$doc->appendChild($note);
echo $doc->saveXML();

The pattern across all five: the string-level escaper is the convenient answer and the writer is the correct one, because only the writer knows whether it is writing element content or an attribute value, and only the writer refuses characters that cannot be escaped at all.

Common questions

Why does &nbsp; break my XML?

Because XML never defined it. The five predefined entities are &amp;, &lt;, &gt;, &quot; and &apos;, and that is the complete list. Everything else HTML gives you comes from an entity table XML deliberately did not inherit.

A parser meeting &nbsp; with no DTD in scope reports an undeclared entity: it has a name and nothing tells it what the name means. Use &#160; or &#xA0; instead. The exception is a document whose DTD declares the name, which is why the same markup works in XHTML and fails in a plain XML config file.

Do I have to escape the greater-than sign?

Almost never. XML 1.0 requires it in one situation: when > appears in the literal sequence ]]> inside element content and is not closing a CDATA section, where a parser scanning for the end of a CDATA section would otherwise be confused.

Everywhere else it is optional, including inside attribute values, so <note>a > b</note> is well-formed as written. This tool escapes it anyway so the output is safe to paste into any context. For the minimal form, put > back everywhere except inside a ]]> sequence.

Should I use CDATA instead of escaping?

CDATA suppresses the recognition of < and & as markup. That is all it does, and it is the right tool when a human edits the content by hand: embedded source code, an XSLT expression full of angle brackets, hand-maintained SQL.

It is the wrong tool the rest of the time. It cannot contain the sequence ]]> because that terminates it, which makes it a real injection vector for untrusted content. It does not expand entities, so <![CDATA[&amp;]]> yields six literal characters, and it does not permit illegal characters. Most serialisers do not preserve it as a distinct node either, so never treat CDATA-ness as meaningful.

Is the text I paste sent anywhere?

No. Escaping is a string replacement running as JavaScript in this tab. There is no request to make, so there is no server to make it to.

That matters here more than it looks: this tool gets used on the part of a payload that broke, which in practice means URLs with tokens in the query string, connection strings, and error text lifted out of a log. Open your network panel while you type and it will stay empty. Your input is kept in this browser's localStorage so a refresh does not lose it, never transmitted, and Clear removes it immediately.

Why did my &amp; turn into &amp;amp;?

The text was escaped twice, usually by a manual escape applied to a string a writer had already escaped, or by a template that escapes on output being fed a value escaped on input.

The unescape direction undoes one layer, so run it once to get back to &amp; and again to get back to &. If &amp;amp; keeps appearing in production, the bug is almost always a hand-rolled escape in front of a DOM or stream writer that was already doing the job. Escaping in the wrong order fails the other way: replacing < before & turns &lt; into &amp;lt;, which is why the JavaScript sample uses a single regex pass.

Related tools

Background reading

Errors this fixes