XML Minifier
Strip insignificant whitespace. Mixed content is left alone.
Everything runs in this tab. Nothing you paste is uploaded, logged or sent anywhere. Open your network panel and check.
Paste a document above and the whitespace between its elements is removed. The result comes back as one line beside your input, with the size before, the size after and the percentage saved. Comments go too if you tick the box. Everything runs in this tab: the parser and the minifier are JavaScript, and your document is never sent anywhere.
Minify when the size of the document itself is the constraint. Hundreds of thousands of payloads in a database column, an envelope embedded in a JSON string field, a message that has to fit a broker's size cap. Over the wire it is worth far less than the headline percentages suggest, and the numbers below say why.
The rule is the formatter's rule inverted: whitespace between elements is layout, whitespace inside mixed content is data. When an element holds both text and child elements, its subtree is taken from your source rather than rebuilt, and CDATA is copied byte for byte. What is left cannot change what a parser reports to the application.
What is safe to remove
An XML parser hands every character of the document to the application, whitespace included; there is none it discards on your behalf. "Insignificant" is a statement about the content model, not about the parser: if a DTD or schema declares an element to have element-only content, whitespace directly inside it cannot be data, and that is the only whitespace a minifier may touch.
Most documents arrive with no schema attached, so a minifier applies a heuristic, and the one used here is narrow: a text node consisting entirely of whitespace, sitting among element siblings, is dropped. Nothing else is. That is the same judgement libxml2 makes under --noblanks. One further change: an empty element is written self-closing, so <status></status> becomes <status/>. Any conformant parser sees the same element, but they are not the same bytes.
What is not removed
Mixed content is what separates a real minifier from a regular expression. In <p>Hello <b>world</b>!</p> the space after "Hello" is a character in the document, and so is the space between </b> and <i> in <p>a <b>x</b> <i>y</i></p>. Both survive, because a mixed subtree is copied out of your source verbatim rather than rebuilt. Nothing inside one is touched, which is the only defensible position: once text and markup are interleaved there is no whitespace left that can be proven insignificant.
- CDATA sections: delimiters and contents reproduced exactly. Whitespace in there is content by definition.
- Attribute values: emitted as written, with the original quote character and any entity references intact.
- Text in a leaf element: <name> Ada </name> loses only the outer padding, never the characters between them.
- Comments: kept unless you ask. Processing instructions, the XML declaration and the internal DTD subset: always kept.
- xml:space="preserve": the honest limit. It gets no special treatment, so minify a copy and compare if your document uses it.
How much smaller, actually
Numbers from one test document, 2,000 order records nested four deep. With two-space indentation it is 554 KB and minifies to 421 KB, 24 per cent off. With four spaces it is 679 KB, so minifying takes 36 per cent off. Deep documents with small elements gain the most, because the indentation is large relative to the content it wraps.
Now compress both. The two-space version gzips to 30,744 bytes and the minified version to 29,771: a difference of three per cent. With four-space indentation the result inverts, 28,519 bytes formatted against 29,771 minified, so the minified file is the larger of the two after compression. Repeated indentation is exactly what a compressor is good at. Minifying is not pointless; minifying for the network is, once Content-Encoding is on.
- Worth it: raw XML in a database column, a size-capped queue message, a document embedded in another payload, a fixture where indentation dominates the diff.
- Not worth it: anything served over HTTP with compression enabled, which is where most minifier marketing gets its percentage.
- Actively harmful: a signed document. Canonical XML keeps whitespace in element content, so the digest covers bytes you are about to change.
Getting the layout back, and the limits
Running the output through the formatter gives you an indented document again with every value identical. What you do not get back is your original layout: blank lines between sections, an unusual indent width, the internal spacing of a mixed subtree where two tags sat side by side. For element-only content none of that was information. If it was, do not minify.
A document that is not well-formed is not minified: your input comes back unchanged with every error listed by line and column. A 1 MB document minifies in about 180 milliseconds and 5 MB in around a second, in a Web Worker so the page stays responsive. The ceiling is 20 MB, because everything is held in this tab's memory.
Minifying XML in code
In a build step this is a parse and a reserialise, never a string operation. Each sample parses safely, since the defaults in Java, PHP and Python's standard library resolve external entities. Watch how each library decides which whitespace is ignorable, because that decision is the whole tool.
// Remove whitespace-only text nodes, but only where the element's children
// are all elements. Anything else is mixed content and belongs to the data.
function minifyXml(source) {
const doc = new DOMParser().parseFromString(source, 'application/xml');
if (doc.querySelector('parsererror')) {
throw new Error(doc.querySelector('parsererror').textContent.trim());
}
const strip = (el) => {
const kids = [...el.childNodes];
const elementOnly =
kids.some((n) => n.nodeType === 1) &&
kids.every((n) => n.nodeType !== 3 || !n.nodeValue.trim());
if (elementOnly) {
for (const n of kids) if (n.nodeType === 3) el.removeChild(n);
}
for (const child of el.children) strip(child);
};
strip(doc.documentElement);
return new XMLSerializer().serializeToString(doc);
}
// XMLSerializer emits <a/> for an empty element, so the output is not
// byte-identical to input that wrote <a></a>. Same document, different bytes.# lxml's remove_blank_text is the closest equivalent, and it is honest about
# being a guess: with no DTD it keeps blank text whose siblings carry real
# characters, which is what saves mixed content.
from lxml import etree
parser = etree.XMLParser(
remove_blank_text=True,
resolve_entities=False, # no entity expansion
no_network=True, # never fetch a DTD or an include
load_dtd=False,
huge_tree=False,
)
tree = etree.fromstring(source.encode('utf-8'), parser)
minified = etree.tostring(tree, encoding='unicode')
# Standard library, no dependency, same rule applied by hand:
#
# from defusedxml.ElementTree import fromstring
# root = fromstring(source)
# for el in root.iter():
# if len(el) and not (el.text or '').strip():
# el.text = None
# if not (el.tail or '').strip():
# el.tail = Noneimport javax.xml.XMLConstants;
import javax.xml.parsers.DocumentBuilderFactory;
import javax.xml.transform.*;
import javax.xml.transform.dom.DOMSource;
import javax.xml.transform.stream.StreamResult;
import javax.xml.xpath.*;
import java.io.StringWriter;
import org.w3c.dom.*;
DocumentBuilderFactory dbf = DocumentBuilderFactory.newInstance();
dbf.setFeature(XMLConstants.FEATURE_SECURE_PROCESSING, true);
dbf.setFeature("http://apache.org/xml/features/disallow-doctype-decl", true);
dbf.setXIncludeAware(false);
Document doc = dbf.newDocumentBuilder().parse(new java.io.File("in.xml"));
// setIgnoringElementContentWhitespace only works when a DTD or schema tells
// the parser which content is element-only, which is usually not the case.
// So select the blank text nodes explicitly. The second predicate leaves
// mixed content alone: it skips blanks whose siblings hold real characters.
XPath xpath = XPathFactory.newInstance().newXPath();
NodeList blanks = (NodeList) xpath.evaluate(
"//text()[not(normalize-space())][not(../text()[normalize-space()])]",
doc, XPathConstants.NODESET);
for (int i = 0; i < blanks.getLength(); i++) {
Node n = blanks.item(i);
n.getParentNode().removeChild(n);
}
TransformerFactory tf = TransformerFactory.newInstance();
tf.setFeature(XMLConstants.FEATURE_SECURE_PROCESSING, true);
Transformer t = tf.newTransformer();
t.setOutputProperty(OutputKeys.INDENT, "no");
StringWriter out = new StringWriter();
t.transform(new DOMSource(doc), new StreamResult(out));
System.out.println(out);using System.Text;
using System.Xml;
// IgnoreWhitespace drops every whitespace-only text node the reader sees.
// Without a schema that includes the space between two inline elements in
// mixed content, so use this on element-only documents and set it false when
// the document carries prose.
var readerSettings = new XmlReaderSettings
{
IgnoreWhitespace = true,
DtdProcessing = DtdProcessing.Prohibit, // no external entities
XmlResolver = null,
MaxCharactersFromEntities = 1024 * 1024,
};
var writerSettings = new XmlWriterSettings
{
Indent = false,
NewLineHandling = NewLineHandling.None,
};
var output = new StringBuilder();
using (var reader = XmlReader.Create(new StringReader(source), readerSettings))
using (var writer = XmlWriter.Create(output, writerSettings))
{
writer.WriteNode(reader, defattr: true);
}
Console.WriteLine(output.ToString());<?php
$doc = new DOMDocument();
// preserveWhiteSpace must be false before the load, not after. libxml2 then
// applies its blank-node heuristic, which keeps blanks inside mixed content.
$doc->preserveWhiteSpace = false;
$doc->formatOutput = false;
libxml_use_internal_errors(true);
if (!$doc->loadXML($source, LIBXML_NONET)) {
foreach (libxml_get_errors() as $e) {
fprintf(STDERR, "XML error at line %d, column %d: %s\n",
$e->line, $e->column, trim($e->message));
}
libxml_clear_errors();
exit(1);
}
// saveXML() still ends the document with a newline; trim it if the byte
// count is what you are optimising.
echo rtrim($doc->saveXML());# --noblanks is libxml2's minifier. It removes ignorable whitespace only.
xmllint --noblanks --nonet document.xml > minified.xml
# xmllint has no flag for dropping comments, and sed cannot do it correctly
# (a comment may contain a > character). Use an identity transform with no
# template for comment() nodes:
cat > strip-comments.xsl <<'XSL'
<xsl:stylesheet version="1.0" xmlns:xsl="http://www.w3.org/1999/XSL/Transform">
<xsl:output method="xml" indent="no"/>
<xsl:template match="@*|node()">
<xsl:copy><xsl:apply-templates select="@*|node()"/></xsl:copy>
</xsl:template>
<xsl:template match="comment()"/>
</xsl:stylesheet>
XSL
xsltproc --nonet strip-comments.xsl document.xml | xmllint --noblanks --nonet -
# Measure the two questions separately:
wc -c document.xml minified.xml
gzip -9 -c document.xml | wc -c
gzip -9 -c minified.xml | wc -cEvery one of these parses and reserialises. None is a regular expression over angle brackets, because the decision a minifier makes (is this whitespace inside element-only content?) needs the tree. A tool that promises to minify without parsing is telling you it will get mixed content wrong.
Common questions
Is my document uploaded when I minify it?
No. The parser and the minifier run in this tab as JavaScript in a Web Worker. There is no endpoint to post to, no analytics with access to the editor and no third-party script. Open the Network tab and watch it stay empty while you work.
That is not incidental here: documents get minified on their way into storage or another system, which means they are real payloads with order records, identifiers and tokens in them. Several tools ranking for this search process your document on a server, and one of the largest publishes saved documents by default.
How much smaller will my XML get?
Between roughly 10 and 35 per cent for typical indented XML, driven almost entirely by nesting depth and indent width. On a 2,000-record test document, two-space indentation dropped 24 per cent and four-space 36 per cent. Documents with long text content gain much less.
The honest comparison is after compression, and there it is about three per cent: 30,744 bytes gzipped formatted against 29,771 gzipped minified. With four-space indentation the formatted file was the smaller of the two. Judge by where the document is going.
Can minifying break my XML?
It can, which is why this one is conservative. The failure mode is removing whitespace that was data: mixed content, CDATA, and anything governed by xml:space="preserve". The first two are handled, since mixed subtrees are copied from your source verbatim and CDATA is copied byte for byte. One caveat remains: xml:space="preserve" gets no special treatment, so a text-only element carrying it is still trimmed.
The third failure mode is invisible. If the document carries an XML Signature, Canonical XML includes whitespace in element content in the digest, so minifying a signed document invalidates it. Minify before signing, never after.
Does it remove comments?
Only if you ask. The checkbox is off by default.
Comments are often the only documentation attached to a configuration file. Removing them from a Maven POM or an XSLT stylesheet costs you the explanation of why a strange element is there, for a saving gzip would have collected anyway. If you own the file and are minifying for storage, strip them; if you are passing somebody else's document along, leave them.
Can I get the formatted version back afterwards?
Yes, by running the minified document through the formatter. Every value, attribute, entity reference and CDATA section is identical, so nothing a parser would report has changed.
What you do not get back is your original layout: blank lines between sections, a particular indent width, or the internal spacing of a mixed subtree where two tags sat next to each other. For element-only content none of that carried information. If it did, keep the original file.
How large a file can it handle?
Up to 20 MB. A 1 MB document minifies in roughly 180 milliseconds and 5 MB in about a second, in a Web Worker so the page stays responsive.
The cap is a memory limit rather than a policy. Nothing is uploaded, so the document lives in this tab, and past 20 MB the parse tree can occupy several hundred megabytes and the browser stops responding. For files above that, xmllint --noblanks does the same job with the same rule about ignorable whitespace.