The billion laughs attack
The billion laughs attack is 783 bytes of perfectly well-formed XML that expands to 2.79 GiB of the string "lol". No recursion, no external reference, no malformed syntax. A parser that follows the specification exactly, with no limits of its own, will try to build the whole thing in memory and die.
It is the most instructive denial of service payload in any format, because it is not a bug in any parser. It is the entity mechanism in XML 1.0 section 4 working precisely as written, and the fix in every runtime is a limit bolted on the side.
#The payload
Nine entity declarations, each one referencing the level below it ten times. The document body references only the top level.
<?xml version="1.0"?>
<!DOCTYPE lolz [
<!ENTITY lol "lol">
<!ENTITY lol1 "&lol;&lol;&lol;&lol;&lol;&lol;&lol;&lol;&lol;&lol;">
<!ENTITY lol2 "&lol1;&lol1;&lol1;&lol1;&lol1;&lol1;&lol1;&lol1;&lol1;&lol1;">
<!ENTITY lol3 "&lol2;&lol2;&lol2;&lol2;&lol2;&lol2;&lol2;&lol2;&lol2;&lol2;">
<!ENTITY lol4 "&lol3;&lol3;&lol3;&lol3;&lol3;&lol3;&lol3;&lol3;&lol3;&lol3;">
<!ENTITY lol5 "&lol4;&lol4;&lol4;&lol4;&lol4;&lol4;&lol4;&lol4;&lol4;&lol4;">
<!ENTITY lol6 "&lol5;&lol5;&lol5;&lol5;&lol5;&lol5;&lol5;&lol5;&lol5;&lol5;">
<!ENTITY lol7 "&lol6;&lol6;&lol6;&lol6;&lol6;&lol6;&lol6;&lol6;&lol6;&lol6;">
<!ENTITY lol8 "&lol7;&lol7;&lol7;&lol7;&lol7;&lol7;&lol7;&lol7;&lol7;&lol7;">
<!ENTITY lol9 "&lol8;&lol8;&lol8;&lol8;&lol8;&lol8;&lol8;&lol8;&lol8;&lol8;">
]>
<lolz>&lol9;</lolz>Every rule the specification imposes on entities is satisfied. WFC: No Recursion forbids an entity whose replacement text refers to itself directly or through a chain; nothing here does. WFC: Entity Declared is satisfied because every reference has a declaration. There are no external identifiers, so a parser with external entity loading disabled is not protected in the slightest. Feed this to a conforming parser with no resource limits and the correct behaviour is to expand it.
The name comes from the arithmetic: nine levels of ten give a billion copies of "lol". It is also called an XML bomb, and it has a CVE, CVE-2003-1564, filed against the entity expansion behaviour rather than against any single implementation.
#The arithmetic
| Entity | Copies of "lol" | Expanded size (UTF-8) |
|---|---|---|
| lol | 1 | 3 bytes |
| lol1 | 10 | 30 bytes |
| lol2 | 100 | 300 bytes |
| lol3 | 1,000 | 3 KB |
| lol6 | 1,000,000 | 3 MB |
| lol9 | 1,000,000,000 | 3,000,000,000 bytes, 2.79 GiB |
The amplification factor is roughly four million to one. That is what makes input size limits useless as a defence: a 1 MB cap, a 10 MB cap, a 100 MB cap, all of them let this through with room to spare, because the payload is smaller than a favicon.
The 2.79 GiB figure is the UTF-8 byte count, and it is the optimistic one. Java, .NET and JavaScript hold strings as UTF-16, so the same expansion occupies 6 GB of heap before any DOM node overhead. A DOM parser building text nodes around it does considerably worse. In practice the process does not reach the end of the expansion; it hits the allocator, starts swapping, and takes the machine with it.
Scaling is trivial for an attacker. Adding a tenth level costs about 80 bytes of input and multiplies the output by ten.
#The two variants that defeat the obvious defences
The first defence anyone writes is a nesting depth limit, and the first variant walks straight past it. Quadratic blowup uses one entity with a large literal value, referenced many times from the document body. Depth is one. Entity count is one. Nothing recurses.
<?xml version="1.0"?>
<!DOCTYPE bomb [
<!ENTITY a "AAAAAAAAAA ... 100,000 A characters ... AAAAAAAAAA">
]>
<bomb>&a;&a;&a;&a; ... &a; repeated 100,000 times ... &a;</bomb>The lesson generalises: a defence that inspects the shape of the declarations will miss this, and a defence that counts entity expansions will let 100,000 of them through if its threshold is higher than that. The only measure that catches both forms is total expanded output, which is why libxml2 moved to an amplification factor model rather than a depth or count model.
The second variant uses parameter entities instead of general entities. Parameter entities are declared with a percent sign and referenced as %name; inside the DTD, and they are expanded during DTD processing rather than during content parsing. A parser that blocks general entity expansion but still processes the internal subset will expand these before it ever reaches the document body.
<?xml version="1.0"?>
<!DOCTYPE bomb [
<!ENTITY % a "AAAAAAAAAA">
<!ENTITY % b "%a;%a;%a;%a;%a;%a;%a;%a;%a;%a;">
<!ENTITY % c "%b;%b;%b;%b;%b;%b;%b;%b;%b;%b;">
<!ENTITY % d "%c;%c;%c;%c;%c;%c;%c;%c;%c;%c;">
<!ENTITY % e "%d;%d;%d;%d;%d;%d;%d;%d;%d;%d;">
<!ENTITY % f "%e;%e;%e;%e;%e;%e;%e;%e;%e;%e;">
<!ENTITY boom "%f;">
]>
<bomb/>#What each runtime does about it
| Runtime | Mechanism | Default |
|---|---|---|
| libxml2 2.11 and later | Amplification factor cap. xmlCtxtSetMaxAmplification(), or xmllint --max-ampl | 5, and exceeding it reports "Maximum entity amplification factor exceeded" |
| libexpat 2.4 and later | Amplification factor plus an activation threshold, so small documents are not measured | 100.0 above 8 MiB of output |
| Java JAXP | jdk.xml.entityExpansionLimit counts expansions; jdk.xml.totalEntitySizeLimit counts characters | 64,000 expansions and 50,000,000 characters |
| Python standard library | None of its own. Inherits whatever expat it was built against | Documented as vulnerable to billion laughs and quadratic blowup |
| .NET XmlReader | DtdProcessing.Prohibit rejects the DTD outright; MaxCharactersFromEntities caps expansion when DTDs are allowed | DTD prohibited, entity cap 0 meaning no limit |
| Go encoding/xml | Entity declarations in a DOCTYPE are never processed | Immune |
| Rust quick-xml | Does not expand DTD-declared entities; unescape resolves only the five predefined names | Immune |
Two of those rows deserve a caveat. Java's expansion limit counts expansions rather than output size, so it does not stop the quadratic form on its own: 63,999 references to a 100 KB entity are under the count limit. The total size limit is what catches that, and it is the reason both properties exist. FEATURE_SECURE_PROCESSING is the switch that puts the JDK processors into the mode where these limits are enforced and extension functions are refused.
Python is the messy one, and the messiness is worth stating plainly. Python's own documentation still lists xml.etree, xml.sax, minidom and pulldom as vulnerable to billion laughs and quadratic blowup. Since libexpat 2.4 the amplification protection is on by default in the C library underneath them, so a recent CPython built against a recent expat will in fact reject the canonical payload. The two statements have not been reconciled, which is exactly why the practical advice has not changed: use defusedxml, which raises EntitiesForbidden before any expansion is attempted rather than relying on the version of a shared library you do not control.
# Python
from defusedxml.ElementTree import fromstring
from defusedxml.common import EntitiesForbidden
try:
root = fromstring(data)
except EntitiesForbidden:
... # the DTD declared entities; refuse the document/* C: the cap is per parser context, not global. */
xmlParserCtxtPtr ctxt = xmlNewParserCtxt();
xmlCtxtSetMaxAmplification(ctxt, 5);
xmlDocPtr doc = xmlCtxtReadMemory(ctxt, buf, len, NULL, NULL,
XML_PARSE_NO_XXE | XML_PARSE_NONET);
/* XML_PARSE_HUGE disables the amplification check along with the
nesting and text-length limits. Never set it on untrusted input. */
/* Shell: raise the factor only for content you wrote yourself. */
/* xmllint --noout --nonet --max-ampl 20 docbook-manual.xml */var settings = new XmlReaderSettings
{
DtdProcessing = DtdProcessing.Parse, // only if you truly need it
XmlResolver = null, // no external fetch, ever
MaxCharactersFromEntities = 1_000_000, // default is 0, meaning unlimited
MaxCharactersInDocument = 20_000_000,
};
using var reader = XmlReader.Create(stream, settings);#How this validator refuses it in two milliseconds
The scanner on this site never expands an entity in order to decide whether expanding it is safe. It reads the internal subset, builds a map of declared entity names to their literal replacement text, and computes the expanded length of each one directly from that graph. The cost is proportional to the number of distinct declarations, not to the size of the expansion they describe: nine declarations means nine evaluations, regardless of whether they describe three characters or three gigabytes.
const REF = /&([^;&<\s]{1,200});/g;
function expandedSizes(decls: Map<string, string>, budget = 10_000_000) {
const sizes = new Map<string, number>();
const visiting = new Set<string>();
let cycle: string | null = null;
let overflow = false;
const sizeOf = (name: string): number => {
if (overflow) return budget;
const cached = sizes.get(name);
if (cached !== undefined) return cached;
if (visiting.has(name)) {
cycle = name; // WFC: No Recursion, reported rather than followed
return 0;
}
const value = decls.get(name);
if (value === undefined) return name.length + 2; // stays as "&name;"
visiting.add(name);
let total = 0;
let last = 0;
REF.lastIndex = 0;
for (let m = REF.exec(value); m; m = REF.exec(value)) {
total += m.index - last;
total += m[1].startsWith('#') ? 1 : sizeOf(m[1]);
last = m.index + m[0].length;
if (total > budget) { overflow = true; break; }
}
total += value.length - last;
visiting.delete(name);
const capped = Math.min(total, budget);
sizes.set(name, capped);
return capped;
};
for (const name of decls.keys()) sizeOf(name);
return { sizes, cycle, overflow };
}Three details do the work. Results are memoised, so lol1 is evaluated once and reused ten times by lol2 rather than being walked again. The visiting set turns an entity cycle into a diagnostic instead of a stack overflow. And every stored size is clamped to the budget, so the running total cannot overflow a JavaScript number even when the declared expansion is astronomically larger than the ceiling. That clamping is why the reported figure says "at least" rather than an exact count: once the budget is blown, the true size stops being interesting.
Two thresholds are then applied. The absolute ceiling is 10,000,000 expanded characters. The ratio test compares the largest single entity expansion against the total bytes of the declarations that produced it, and an amplification above 100 is rejected. The canonical payload above scores roughly four million.
The declaration graph alone cannot catch quadratic blowup, because a single entity holding 100 KB of A characters has an amplification of about one and is entirely legitimate on its own. So the body scan charges every reference against the same budget, adding the precomputed size of each entity as the reference is encountered and stopping the scan the moment the running total passes the ceiling. One budget, both attack shapes.
// Called for every entity reference found in element content.
this.entityExpansionTotal += this.entitySizes.get(name) ?? name.length;
if (!this.expansionReported &&
this.entityExpansionTotal > this.limits.maxEntityExpansion) {
this.expansionReported = true;
this.report('XV030', at, end, /* ... */);
this.aborted = true; // stop scanning; the document is hostile
}Pasting the nine-level payload into the validator produces a rejection in about two milliseconds, and the reason is that nothing was ever expanded. The document is read once, nine declarations are evaluated, the ratio is computed, and the answer comes back. Schema validation, when you ask for it, runs libxml2 compiled to WebAssembly without XML_PARSE_HUGE, so libxml2's own entity nesting limit of 20 levels and its amplification cap apply as a second, independent line of defence.
Common questions
Is the billion laughs payload valid XML?
It is well-formed, which is the check that matters here. There is no recursion, every entity is declared before it is referenced, the nesting is legal, and no rule in XML 1.0 is broken.
Whether it is valid in the schema sense depends on the DTD, and the question is not interesting: the damage is done during parsing, long before any validity constraint is evaluated. This is the cleanest demonstration that well-formedness is a statement about syntax and not about safety.
Does the attack still work in 2026?
Against a modern parser on default settings, usually not. libxml2 has capped amplification since 2.11, expat since 2.4, and the JDK has enforced expansion limits for years. The canonical payload is now a test case more often than a weapon.
It still works in three situations that turn up regularly. Code that sets XML_PARSE_HUGE, or lxml with huge_tree=True, to make a large legitimate document parse, has switched the protection off. Older runtimes and vendored parsers in long-lived services never got the fix. And the quadratic variant still passes any defence that counts expansions or measures nesting depth rather than output size, which is a larger set of implementations than the exponential form does.
Why not just reject large uploads?
Because the payload is 783 bytes. Any input size limit permissive enough to accept a normal document accepts this one thousands of times over.
The limit has to be on output, not input, and it has to be enforced during expansion rather than after it. That is the whole reason runtimes grew a specific entity amplification setting instead of relying on the general request size limits they already had.
Is it safe to paste one of these into this tool?
Yes, and it is a reasonable thing to try. The expansion cost is computed from the declarations before a single character is expanded, so the rejection comes back in about two milliseconds and the tab never allocates the expansion.
The quadratic variant is caught by the same budget, and an entity that refers to itself is reported as a cycle instead of being followed. Nothing is uploaded either way: the scanner runs in a Web Worker in your browser and makes no network request with your document.
What should I set if my documents legitimately use a lot of entities?
Raise the limit deliberately and scope it to the input you trust. In libxml2 that is xmlCtxtSetMaxAmplification() on the context, or xmllint --max-ampl; the default factor of 5 is strict enough that entity-heavy DocBook and DITA sources do trip it. In Java it is the jdk.xml.entityExpansionLimit and jdk.xml.totalEntitySizeLimit system properties.
The important part is that this happens on a separate code path from anything that accepts documents over the network. A build step that processes your own documentation can afford a generous limit. The endpoint that accepts XML from strangers should keep the default, or refuse DTDs entirely.
Sources
- W3C XML 1.0 (Fifth Edition), section 4: Physical Structures
- libxml2 parser reference: xmlCtxtSetMaxAmplification and parse options
- xmllint manual, including --max-ampl
- Python documentation: XML vulnerabilities
- XmlReaderSettings.MaxCharactersFromEntities, .NET API documentation
- Go standard library: encoding/xml