CDATA sections

A CDATA section does exactly one thing: inside it, `<` and `&` stop being markup delimiters and become ordinary characters. That is the whole feature. It is not a container, not a data type, not a node kind in most of the models built on top of XML, and not a security boundary.

Almost every problem people have with CDATA comes from expecting it to do more. It does not expand entities, it cannot nest, it cannot contain the three characters that end it, and a parser is entitled to hand your CDATA back to the application as ordinary text with no record that it was ever written that way. What follows is the exact behaviour, the workaround for the one sequence it cannot hold, and then the part usually skipped: deciding when it is worth using at all.

#The syntax, exactly

XML 1.0 (Fifth Edition) section 2.7 defines it in four productions, and the third one is the interesting one:

[18] CDSect  ::= CDStart CData CDEnd
[19] CDStart ::= '<![CDATA['
[20] CData   ::= (Char* - (Char* ']]>' Char*))
[21] CDEnd   ::= ']]>'
Productions [18] to [21].

Production [20] reads: any sequence of characters, minus any sequence that contains `]]>`. The exclusion is baked into the grammar rather than enforced as a separate constraint, which is why there is no escape hatch for it and why nesting is impossible.

  • The opener is literally `<![CDATA[`. It is case sensitive, so `<![cdata[` and `<![Cdata[` are not CDATA sections at all.
  • No whitespace is permitted inside the delimiter. `<! [CDATA[` and `<![CDATA [` are both malformed.
  • A CDATA section appears only in element content. Section 2.7 places it in the `content` production, so it is illegal in an attribute value, illegal in the prolog and illegal after the root element.
  • Everything inside must still match the `Char` production from section 2.2. CDATA suspends markup recognition, not the character repertoire.

That third point catches people constantly, because the place developers most want CDATA is an attribute value containing `<`. An XSLT `test` attribute holding an XPath comparison is the canonical case, and the only answer there is escaping.

Illegal: CDATA cannot appear in an attribute value
<xsl:if test="<![CDATA[$count < 10]]>">
  <flag/>
</xsl:if>
Correct: escape the character
<xsl:if test="$count &lt; 10">
  <flag/>
</xsl:if>

#What it suspends, and the four things it does not do

Section 2.4 states that `&` and `<` must not appear in literal form except as markup delimiters, or within a comment, a processing instruction, or a CDATA section. CDATA is one of three exemptions from that rule, and the exemption is the entire mechanism. Inside the section the parser scans for `]]>` and treats every other character as data.

  • It does not expand entities. `<![CDATA[&amp;]]>` produces the five literal characters &amp;, not a single ampersand. Neither predefined nor DTD-declared entities are recognised, because entity recognition begins with `&` being a markup delimiter and it no longer is.
  • It does not legalise illegal characters. A raw U+0001 inside a CDATA section is still a fatal error in XML 1.0, and it cannot be rescued with a character reference either, since character references are not recognised inside CDATA and U+0001 is outside the `Char` production regardless.
  • It does not nest. Production [20] excludes `]]>` from the content, so the first `]]>` ends the section no matter how many openers preceded it.
  • It does not sanitise anything. Untrusted content placed inside a CDATA section is one `]]>` away from ending the section early and injecting arbitrary markup into your document.

The illegal-character case is worth seeing in a real diagnostic. The scanner on this site calls the same `checkIllegalChars` routine on CDATA content that it calls on ordinary text, so a control character inside a section is reported as XV023 with the message that the code point is not a character XML 1.0 permits anywhere in a document, including inside CDATA, and that it cannot be escaped either. If you need to carry arbitrary bytes, Base64 or hex text is the answer, not CDATA.

#The one sequence it cannot contain

Because `]]>` is defined out of the content grammar, a literal `]]>` cannot be written inside a CDATA section by any means. The standard workaround is to split the section so the sequence is broken across the boundary.

<code><![CDATA[if (a]]]]><![CDATA[>b) { }]]></code>
A C-like snippet containing ]]> written as two sections.

Reading that from the left: the first section holds `if (a]]`, then `]]>` closes it, then a new section opens and holds `>b) { }`, then `]]>` closes that. The two sections are adjacent character data, so the parser concatenates them into the single text value `if (a]]>b) { }`. It works and it is unreadable, which is a reasonable argument for not using CDATA here at all.

The alternative is the escape the specification itself mandates. Section 2.4 says `>` must, for compatibility, be escaped using `&gt;` or a character reference when it appears in the string `]]>` in content and that string is not marking the end of a CDATA section. This is the only situation in XML where escaping `>` is required rather than optional.

Not well-formed: bare ]]> in text content
<code>if (a]]>b) { }</code>
Well-formed: the mandatory > escape
<code>if (a]]&gt;b) { }</code>

The syntax checker here scans text content for this specific sequence and reports XV018, saying that `]]>` is reserved as the CDATA terminator and may not appear in ordinary text, with both remedies offered. It is a rare error in hand-written XML and a common one in machine-generated XML that embeds source code, because a naive escaper handles `&` and `<` and forgets that `>` has one mandatory case.

#CDATA is a serialisation, not a node type

This is the piece that determines whether your pipeline can round-trip. In the XPath 1.0 data model there is no CDATA node: text is text, and a CDATA section contributes to the same text node as the characters either side of it. The XQuery and XPath Data Model that XPath 2.0 and later build on has the same seven node kinds and none of them is CDATA. Write `<a><![CDATA[x]]>y</a>` and `string(/a)` returns `xy` from a single text node.

XSLT 1.0 makes the consequence explicit by putting CDATA on the output side of the process. Section 16.1 defines `cdata-section-elements` as an attribute of `xsl:output`: you list the element names whose text children should be serialised as CDATA when the result tree is written out. There is no way to preserve CDATA-ness from input to output, because by the time XSLT sees the document there is nothing to preserve.

<xsl:output method="xml"
            encoding="UTF-8"
            cdata-section-elements="description script"/>
XSLT decides CDATA at serialisation time, per element name.

The DOM is the exception that confuses everyone. It does define a `CDATASection` interface, inheriting from `Text`, and a browser `DOMParser` will give you `CDATASection` nodes with `nodeType` 4. So the distinction survives a DOM parse and is lost the moment the document passes through XPath, XSLT, or most data-binding layers. libxml2 will drop it on request too: `XML_PARSE_NOCDATA` is documented as outputting normal text nodes instead of CDATA nodes, and `xmllint --nocdata` is the command-line form.

This site keeps a distinct `cdata` node in its own scan tree, which lets the formatter preserve sections verbatim rather than rewriting them, and lets the `expandCData` option convert them to escaped text on request. The converters do not preserve it: the XML to JSON path reads a `cdata` child's value straight into the same string as its text siblings, and the CSV flattener does the same. That is not a bug, it is the only defensible behaviour, because JSON has no representation for the difference.

The practical rule follows directly. Never make CDATA-ness semantically meaningful. If your consumer needs to know that a field contains markup rather than plain text, say so with an attribute or an element name, not with the serialisation form of the characters.

#When it earns its place, and when it is cargo cult

CDATA and escaping produce identical documents in the data model. The choice is therefore entirely about the humans and tools that touch the source text, and it comes down to two questions: how much escaping would there be, and does a person edit this by hand?

It earns its place when a human maintains a block of embedded code or markup. A 40-line HTML fragment, an SQL statement full of `<=`, a shell script, an embedded XML sample in documentation: escaping these turns readable text into a wall of `&lt;` and makes every future edit a manual encoding exercise. The section delimiters cost 12 characters once, and the payload stays diffable, greppable and copy-pasteable.

Cargo cult: twelve characters of ceremony to avoid one escape
<url><![CDATA[https://example.com/s?q=a&b=2]]></url>
Just escape it
<url>https://example.com/s?q=a&amp;b=2</url>
<template name="receipt"><![CDATA[
<table class="lines">
  <tr th:each="line : ${order.lines}">
    <td th:text="${line.sku}">SKU</td>
    <td th:text="${line.qty} + ' @ ' + ${line.price}">1 @ 0.00</td>
  </tr>
</table>
]]></template>
The opposite case: escaping this would be actively hostile to whoever edits it next.
SituationUseWhy
One or two ampersands in a URL or a nameEscapeThe escape is shorter than the delimiters and reads fine.
A block of hand-edited HTML, code or SQLCDATAEscaping destroys readability and every edit becomes an encoding task.
Machine-generated content of any sizeEscapeYour serialiser already escapes correctly. Emitting CDATA means you also own the ]]> splitting logic, and most hand-rolled emitters do not.
Content that may contain ]]>EscapeCDATA needs the split workaround, which is fragile to generate and horrible to read.
Anything going into an attribute valueEscapeCDATA is illegal there. Section 2.7 permits it only in element content.
Untrusted inputEscapeA ]]> in the input breaks out of the section. CDATA is not a boundary.
Binary or control charactersNeitherBase64 or hex. The Char production applies inside CDATA too.
The decision, by situation.

RSS deserves a specific mention because it is where most people meet CDATA. The RSS 2.0 specification permits entity-encoded HTML in `<description>`, and CDATA is simply another way to write the same characters; feed readers see identical text either way. Choosing CDATA for a feed is a source-readability decision for whoever maintains the generator, not a compatibility one, and generators that escape instead are equally correct.

Common questions

Does CDATA make my XML valid?

It has no effect on validity. A CDATA section is character data, so a schema or DTD sees exactly what it would see if you had escaped the same characters. If an element is declared as `xs:integer` and you put a CDATA section containing "abc" in it, it fails type validation identically.

What CDATA does affect is well-formedness, and only in the narrow sense that it lets `<` and `&` sit in the source without being read as markup. A document that was malformed because of an unescaped ampersand becomes well-formed once that region is inside a CDATA section, which is why the fix is often suggested. Escaping the ampersand achieves the same thing with less ceremony.

Can I put a CDATA section inside another CDATA section?

No. Production [20] defines the content as any character sequence that does not contain `]]>`, so the first terminator ends the section regardless of how many openers came before it. Nesting is excluded by the grammar itself, not by a separate rule that some parser might relax.

If you need to embed an XML sample that itself contains a CDATA section, split at the inner `]]>`: end your outer section after the `]]`, open a new one starting with the `>`, and the two adjacent sections concatenate into one text value.

Will my CDATA sections survive a round trip through my toolchain?

Only if every stage happens to preserve them, and you should not depend on it. XPath and XSLT have no CDATA node kind at all, so a transform reads the content as ordinary text and reproduces it as CDATA only if you ask for that on `xsl:output` via `cdata-section-elements`. libxml2 will convert sections to text nodes when `XML_PARSE_NOCDATA` is set, which `xmllint --nocdata` does.

The DOM is the outlier: it has a `CDATASection` interface and browser parsers do produce those nodes, so a parse-then-serialise cycle in a browser usually preserves them. That is a property of one API, not a guarantee of the format. Treat CDATA-ness as formatting, and put any meaning you actually need into an element name or an attribute.

Why does my parser say "]]> is not allowed in content" when I have no CDATA section?

Because the restriction applies to all character data, not just to CDATA sections. Section 2.4 requires `>` to be escaped as `&gt;` or a character reference whenever it appears in the string `]]>` in content and is not closing a CDATA section. It is the only place where escaping `>` is mandatory.

In practice the sequence arrives from embedded source code (array indexing followed by a comparison, or a generated snippet) or from generated content that concatenated a bracket onto a tag. Writing `]]&gt;` fixes it and changes nothing about the resulting text value.

Should the escaping tool on this site produce CDATA or entities?

Entities, for anything short. The escaper replaces the five characters that need it and leaves everything else alone, which produces text a schema validator, a diff and a human all handle without special cases.

Reach for CDATA deliberately, when you are pasting a block of code or markup into a document that a person will keep editing. The formatter on this site preserves existing sections verbatim rather than rewriting them, and offers an explicit option to expand them into escaped text if you want to normalise a document that mixes both conventions.

Sources

Try it