XML エスケープとアンエスケープ
特殊文字をエスケープ、または実体参照をテキストに戻します。
すべてこのタブ内で実行されます。貼り付けた内容がアップロード・記録・送信されることはありません。 ネットワークパネルを開いて確認する.
左にテキストを貼り付けると、XML がマークアップとして扱う文字がすべて実体参照に置き換わって返ってきます。方向を逆にすれば、参照はそれが名指しする文字へ戻ります。入力は文書ではなく素のテキストなので、解析が通る必要はありません。断片でも、属性値 1 つでも、URL だけでも動きます。
このツールに手を伸ばすのは、パーサーがアンパサンドに文句を言ったあとです。Xerces の「The entity name must immediately follow the '&'」、libxml2 の「EntityRef: expecting ';'」、Expat の「invalid character in entity name」は、どれも同じ問題です。テキストや属性の中にある生の &、たいていは ?a=1&b=2 のようなクエリ文字列を要素に貼り付けた結果です。
ここが違うのは、このツールが「やらないこと」です。XML が定義する名前付き実体は 5 つだけで、それ以外はありません。ウェブにあるエスケープツールの多くは XML と名乗った HTML エスケーパーで、どの XML パーサーも受け付けない を平然と返します。こちらは 5 つと、10 進・16 進の数値参照だけを知っており、それ以外は書かれたまま残し、何もアップロードしません。
定義済み実体は 5 つ、それ以上はない
XML 1.0 の 4.6 節が定義する名前付き実体はちょうど 5 つです。それで全部です。HTML から受け継いだ実体表のようなものは存在せず、HTML のかたまりを XML の設定ファイルや RSS の description に移したときに、たいていここで驚くことになります。
DTD のない文書に と書くのは文体の問題ではなく、整形式エラーです。パーサーの手元には名前だけがあり、その名前が何を意味するのかを教えるものが何もありません。代わりに   か   を使ってください。© や é も同様です。DTD は追加の名前を宣言できます。DocBook の &companyName; が動くのはそのためです。ですからアンエスケープの方向では、知らない名前は書かれたまま一切触りません。
- & は &
- < は <
- > は >
- " は "
- ' は '。HTML 4 は ' を定義しなかったため、混在パイプラインに流すシリアライザーは代わりに ' を出すことがよくあります。
数値文字参照
合法な文字はどれも、コードポイントとして書けます。10 進なら é、16 進なら é。この 2 つは同じ文字で、先頭のゼロも許されます。x は小文字でなければならないので、A は文字参照ではなく、宣言されていない実体として拒否されます。
アンエスケープの方向は両方の形式を扱い、結果が Unicode の範囲内かを確認します。壊れた参照や範囲外の参照は、疑問符に置き換えたりせず、書かれたまま残します。エスケープの方向は数値参照を決して出力しません。UTF-8 はアクセント記号も CJK も絵文字もそのまま運べるので、それらをエスケープすると可読性を失うだけで何も得られません。
各文字を実際にエスケープしなければならない場面
規則は多くの人が思っているより狭く、他人の文書を読んで「これは壊れているのか」を判断するときに効いてきます。& と < はどこでも必須です。> が必須なのは 1 か所だけ。XML 1.0 の 2.4 節は、内容の中にリテラルの ]]> という並びが現れ、それが CDATA セクションを閉じていない場合にはエスケープしなければならないと定めています。ですから <code>if (a]]>b)</code> は必須で、<note>a > b</note> は合法です。
このツールは 5 つすべてを無条件にエスケープします。どの文脈が要求するものより広い集合なので、出力はどこに落としても安全で、いま自分がどの文脈にいるかを追う必要がありません。最小限のエスケープではありませんし、どの文字を戻せばよいかは、もうお分かりのはずです。
- & と <:要素の内容でも属性値でも、常に必須。
- >:任意。ただし ]]> という並びの中では必須。
- ":二重引用符で囲んだ属性値の中でのみ必須。
- ':単一引用符で囲んだ属性値の中でのみ必須。要素の内容では、どちらの引用符もエスケープ不要です。
そもそもエスケープできない文字
XML 1.0 は文書に含められる文字を Char 生成規則として定義しており、C0 制御文字のほとんどはそこに入っていません。U+0001 から U+0008、U+000B、U+000C、U+000E から U+001F。生き残るのはタブ、改行、復帰だけです。U+0000 はどのバージョンの XML でも不正であり、孤立サロゲートも不正です。
ですから  は ESC 文字のエスケープではありません。Legal Character 制約に違反します。その参照が名指しする文字は、どんな手段によっても XML 1.0 文書に現れることができず、参照として書いても浄化はされないからです。ANSI エスケープシーケンスだらけのログファイルに対して Python が「not well-formed (invalid token)」と報告するのは、まさにこれです。任意のバイト列には base64 が答えです。
このツールは小細工をせず正直に振る舞います。エスケープでは制御文字をそのまま通し、アンエスケープでは  を本物の U+0001 にします。参照がそう言っているからです。文書へ戻す前に、結果を構文チェッカーに通してください。
コードで同じことをする
どの言語にもこれ用の何かがありますが、多くは間違ったものを提供しています。目につく関数が HTML エスケーパーだからです。ここでは各エコシステムの XML 専用の経路を使い、それぞれが何を取り違えるかも書いてあります。
// XML defines exactly five named entities. A single regex pass avoids the
// classic bug of replacing & after < and turning < into &lt;.
const ESCAPES = { '&': '&', '<': '<', '>': '>', '"': '"', "'": ''' };
// Element content: & and < are mandatory, > only inside the sequence ]]>.
function escapeText(s) {
return s.replace(/[&<]/g, (c) => ESCAPES[c]).replace(/]]>/g, ']]>');
}
// Attribute value: escape the delimiter you are using, plus tab, newline and
// carriage return. Attribute-value normalisation turns a literal one of those
// into a plain space, so a reference is the only way to keep it.
function escapeAttribute(s) {
return s
.replace(/[&<"]/g, (c) => ESCAPES[c])
.replace(/\t/g, '	')
.replace(/\n/g, ' ')
.replace(/\r/g, ' ');
}
// The reverse. Note the deliberate absence of an HTML entity table: is
// undefined in XML, so leaving it alone is more honest than decoding it.
const NAMED = { amp: '&', lt: '<', gt: '>', quot: '"', apos: "'" };
function unescapeXml(s) {
return s.replace(/&(#x[0-9a-fA-F]+|#[0-9]+|[A-Za-z][A-Za-z0-9]*);/g, (whole, body) => {
if (body[0] === '#') {
const hex = body[1] === 'x';
const cp = parseInt(hex ? body.slice(2) : body.slice(1), hex ? 16 : 10);
return cp >= 0 && cp <= 0x10ffff ? String.fromCodePoint(cp) : whole;
}
return NAMED[body] !== undefined ? NAMED[body] : whole;
});
}from xml.sax.saxutils import escape, quoteattr, unescape
import re
# escape() handles & < > only and knows nothing about quotes. The > is not
# strictly required but it is harmless and covers the ]]> case for free.
text = escape('Terms & conditions <see clause 4>')
# Terms & conditions <see clause 4>
# The two quote entities have to be supplied yourself.
FIVE = {'"': '"', "'": '''}
text = escape(source, FIVE)
# quoteattr() returns the value WITH its delimiters, picking single quotes when
# the value contains a double quote, and turning tab, newline and carriage
# return into numeric references so normalisation cannot flatten them.
attr = quoteattr('say "hello" then stop')
# 'say "hello" then stop' <- single-quoted, so the " needs no escape
# unescape() reverses & < > plus whatever mapping you pass. It does NOT decode
# numeric character references, so é comes back unchanged. html.unescape()
# does decode them, but it also decodes and 250 other HTML names XML
# never defined, which silently rewrites your data. Do it explicitly instead:
def unescape_xml(s: str) -> str:
s = unescape(s, {'"': '"', ''': "'"})
return re.sub(
r'&#(x[0-9a-fA-F]+|[0-9]+);',
lambda m: chr(int(m.group(1)[1:], 16) if m.group(1)[0] == 'x' else int(m.group(1))),
s,
)// The JDK has no public XML escaper for a bare string. Two right answers.
// 1. Let a writer do it. XMLStreamWriter escapes for the context it is
// writing into, the only approach that gets attribute values right.
import javax.xml.stream.XMLOutputFactory;
import javax.xml.stream.XMLStreamWriter;
XMLStreamWriter w = XMLOutputFactory.newInstance().createXMLStreamWriter(out);
w.writeStartElement("note");
w.writeAttribute("href", "https://example.com/?a=1&b=2"); // escaped for you
w.writeCharacters("Terms & conditions <apply>"); // escaped for you
w.writeEndElement();
w.close();
// 2. Escape a standalone string with commons-text. Do NOT use commons-lang3's
// deprecated StringEscapeUtils.escapeXml, which knew nothing about the
// characters XML cannot represent.
import org.apache.commons.text.StringEscapeUtils;
String safe = StringEscapeUtils.escapeXml10("Terms & conditions <apply>");
// Terms & conditions <apply>
// escapeXml10 does something its name does not advertise: it DELETES the
// characters XML 1.0 cannot hold (U+0001 to U+0008, U+000B, U+000C,
// U+000E to U+001F, unpaired surrogates) rather than escaping them, because
// no escape for them exists. escapeXml11 writes the control characters as
// numeric references instead, which is only legal if the document declares
// version="1.1", and it still drops U+0000 and unpaired surrogates.
String back = StringEscapeUtils.unescapeXml(safe);
// The five predefined entities plus decimal and hex character references.
// Names it does not know are left as written.using System.Linq;
using System.Security;
using System.Xml;
// SecurityElement.Escape does the five predefined entities in one call:
// < > & " ' every time, a superset of what any single context needs.
string safe = SecurityElement.Escape("Terms & conditions <apply>");
// Terms & conditions <apply>
// It does not remove the characters XML cannot carry, so check those yourself.
// XmlConvert.IsXmlChar implements the Char production from XML 1.0. Keep
// surrogates or astral characters (emoji, rarer CJK) are destroyed.
static string StripIllegal(string s) =>
string.Concat(s.Where(c => XmlConvert.IsXmlChar(c) || char.IsSurrogate(c)));
// In practice prefer a writer. It escapes for the context and throws on a
// character the document cannot hold, instead of producing a file that fails
// to parse somewhere downstream.
var settings = new XmlWriterSettings { Indent = true, CheckCharacters = true };
using var writer = XmlWriter.Create(Console.Out, settings);
writer.WriteStartElement("note");
writer.WriteAttributeString("href", "https://example.com/?a=1&b=2");
writer.WriteString("Terms & conditions <apply>");
writer.WriteEndElement();
// There is no framework unescaper, and WebUtility.HtmlDecode is the wrong
// tool: it decodes and the rest of the HTML set. Read a fragment
// instead, which decodes exactly what XML defines and nothing else.
static string Unescape(string escaped)
{
var rs = new XmlReaderSettings
{
DtdProcessing = DtdProcessing.Prohibit,
XmlResolver = null,
};
using var r = XmlReader.Create(new StringReader("<r>" + escaped + "</r>"), rs);
r.ReadToFollowing("r");
return r.ReadElementContentAsString();
}<?php
// ENT_XML1 is the flag that matters and the one everybody omits. Without it
// you get the HTML entity table: an apostrophe becomes ' (harmless), and
// with ENT_HTML5 you can get names XML will reject outright.
$safe = htmlspecialchars(
$text,
ENT_XML1 | ENT_QUOTES | ENT_SUBSTITUTE,
'UTF-8'
);
// & < > " ' become & < > " '
//
// ENT_SUBSTITUTE replaces invalid UTF-8 with U+FFFD. Before PHP 8.1 the
// default was to return an empty string on invalid input, which is very easy
// to miss in a feed generator fed by a legacy database.
$back = html_entity_decode($safe, ENT_XML1 | ENT_QUOTES, 'UTF-8');
// With ENT_XML1 this decodes the five predefined entities and numeric
// character references and leaves alone. Drop the flag and it decodes
// into U+00A0, silently rewriting your data.
// Building a document rather than a string: DOMDocument escapes on save and
// rejects characters XML cannot represent, which htmlspecialchars passes
// straight through.
$doc = new DOMDocument('1.0', 'UTF-8');
$note = $doc->createElement('note');
$note->appendChild($doc->createTextNode($text));
$doc->appendChild($note);
echo $doc->saveXML();5 つに共通する型はこうです。文字列レベルのエスケーパーは便利な答えで、writer のほうが正しい答えです。要素の内容を書いているのか属性値を書いているのかを知っているのは writer だけですし、そもそもエスケープできない文字を拒否できるのも writer だけだからです。
よくある質問
なぜ で XML が壊れるのですか。
XML がそれを定義したことがないからです。定義済み実体は &、<、>、"、' の 5 つで、それが全リストです。HTML が与えてくれる他のすべては、XML が意図的に受け継がなかった実体表から来ています。
スコープ内に DTD がないまま に出会ったパーサーは、宣言されていない実体として報告します。名前はあるのに、その名前が何を意味するかを教えるものがないのです。代わりに   か   を使ってください。例外は、DTD がその名前を宣言している文書です。同じマークアップが XHTML では動き、素の XML 設定ファイルでは落ちるのはそのためです。
大なり記号はエスケープしなければなりませんか。
ほとんどの場合、不要です。XML 1.0 が要求するのは 1 つの状況だけです。要素の内容の中でリテラルの ]]> という並びに > が現れ、しかもそれが CDATA セクションを閉じていないとき。そのままだと、CDATA セクションの終わりを探しているパーサーが混乱するからです。
それ以外の場所では属性値の中も含めて任意なので、<note>a > b</note> はそのままで整形式です。このツールはどこに貼っても安全なようにとにかくエスケープします。最小限の形にしたいなら、]]> の並びの中以外は > に戻してください。
エスケープの代わりに CDATA を使うべきですか。
CDATA は < と & をマークアップとして認識するのを抑えます。それだけです。人間が手で内容を編集する場合には正しい道具です。埋め込んだソースコード、山かっこだらけの XSLT 式、手で保守している SQL など。
それ以外の場面では間違った道具です。]]> という並びを含められません。それが終端になってしまうからで、信頼できない内容にとっては実際の注入経路になります。実体も展開しないので <![CDATA[&]]> はリテラル 6 文字になりますし、不正な文字を許すわけでもありません。多くのシリアライザーは CDATA を独立したノードとして保持しないので、「CDATA であること」に意味を持たせてはいけません。
貼り付けたテキストはどこかへ送られますか。
いいえ。エスケープは、このタブで動く JavaScript の文字列置換です。行うリクエストがないので、その送り先のサーバーも存在しません。
ここではそれが見た目以上に重要です。このツールが使われるのは、ペイロードのうち壊れた部分に対してです。実際にはクエリ文字列にトークンを含む URL、接続文字列、ログから抜き出したエラーテキストといったものです。入力しながらネットワークパネルを開いてみてください。空のままです。入力内容は再読み込みで失われないようこのブラウザーの localStorage に保存され、送信は一切されず、クリアで即座に消えます。
なぜ & が &amp; になったのですか。
そのテキストが 2 回エスケープされたからです。たいていは、writer がすでにエスケープした文字列に手作業のエスケープをかけたか、出力時にエスケープするテンプレートに、入力時にエスケープ済みの値を渡したかです。
アンエスケープの方向は 1 層だけ戻すので、1 回かければ & に、もう 1 回で & に戻ります。本番で &amp; が出続けるなら、原因はほぼ確実に、すでに仕事をしている DOM やストリームの writer の手前にある自前のエスケープです。順序を間違えると逆向きに失敗します。& より先に < を置換すると < が &lt; になってしまうので、JavaScript のサンプルは正規表現 1 回で処理しています。