XML Formatter
Indent, align and tidy. Mixed content is preserved exactly.
Everything runs in this tab. Nothing you paste is uploaded, logged or sent anywhere. Open your network panel and check.
Paste a document above and it is reindented as you type. Two spaces, four spaces or a tab, with attributes pushed onto their own lines once an element carries three, four or six of them. The result appears in the read-only pane beside your input. Nothing is uploaded: the parser and the formatter are JavaScript running in this tab, and your network panel stays empty while you work.
You reach for a formatter when something else produced the XML: a SOAP response pulled out of a log as one 40,000-character line, a config a deployment tool rewrote, a machine-generated sitemap. It is also the quickest well-formedness check there is, because a formatter cannot indent what it cannot parse.
What is different here is mixed content. When an element holds both text and child elements, the whitespace between them is data, and a formatter that tidies it has changed the document. Those subtrees are reproduced byte for byte. Most online formatters do not do this; several are string manipulation over angle brackets rather than a parser at all.
What pretty printing is allowed to change
Only one thing in an XML document is genuinely insignificant: whitespace between elements in an element-only content model. A formatter may delete that and generate its own. Everything else is content, and the safe approach is to treat every byte as content until proven otherwise.
So the whitespace-only text nodes between sibling elements are discarded and regenerated from your indent setting, and nothing else is touched. Attribute values are emitted as written, because re-escaping would turn & into & and a value referencing a DTD-declared entity such as &companyName; cannot be decoded without the DTD. The original quote character is kept too, since a value written in single quotes may legally contain a double quote.
Text is never internally reflowed. Only the leading and trailing whitespace of a text-only element is trimmed, so <price> 42.00 </price> becomes <price>42.00</price> while <note>two spaces</note> keeps both.
Mixed content, and why it breaks most formatters
An element has mixed content when its children include both text and markup: <p>Hello <b>world</b>!</p>. The space after "Hello" is a character in the document, and so is the exclamation mark after </b>. Put each child on its own indented line and a consumer that concatenates the text nodes gets a different string. That is not tidying, it is silent corruption.
Every element is checked before it is serialised. If it holds element children alongside non-empty text or a CDATA section, the subtree is copied out of your source with no rule applied inside it. The trade is deliberate: a mixed subtree keeps whatever layout it arrived with, even an ugly one, because the alternative is being wrong.
This is not an exotic case: XHTML fragments in a CMS export, DocBook and DITA, xs:documentation inside a schema, an RSS description with inline markup. If your document has none of those it costs you nothing. If it does, it is the whole ballgame.
- Mixed, so reproduced verbatim: <line>Total: <amount>9.99</amount> ex VAT</line>.
- Not mixed, so reindented freely: <order><id>1</id><status>open</status></order>.
Indentation, and attribute wrapping for SOAP and XSD
Two spaces is the default because it is what most XML tooling emits. Four spaces exists because plenty of enterprise codebases standardised on it. Tab exists because some repositories say so in .editorconfig, and because a tab is one byte where four spaces are four.
The attribute control is the one that matters for SOAP and XSD. A SOAP envelope root routinely carries five namespace declarations, and an xs:element declaration carries name, type, minOccurs, maxOccurs, nillable and default: on one line that is 200 characters nobody will read. Set the threshold to three, four or six and any element at or above it gets one attribute per line, while smaller elements stay on one line.
What it will not do
A document that is not well-formed is not formatted: you get your input back unchanged and every error listed with a line, a column and a fix. Two further limits are worth stating plainly. xml:space="preserve" gets no special treatment, so a text-only element carrying it still has its leading and trailing whitespace trimmed. And text interleaved with comments, as in <a>text<!-- why -->more</a>, is not mixed content by this test because a comment is not an element, so it is reindented and the text gains whitespace. Both are narrow, both are real, and neither is something a tool that did it quietly would tell you.
Formatting is idempotent: running the output back through with the same settings produces the same bytes, so it is safe in a pre-commit hook. A 1 MB document takes about 200 milliseconds and 5 MB a little under a second. The ceiling is 20 MB, a memory limit rather than a policy.
Formatting XML in code
The same operation in the languages that actually process XML. Each sample parses safely, because the defaults in Java, PHP and Python's standard library resolve external entities, and the comments mark where each library reflows mixed content.
// Browsers ship a parser and a serialiser but no pretty printer. This walker
// indents only elements whose children are all elements: touching anything
// else would rewrite mixed content.
function indentXml(source, unit = ' ') {
const doc = new DOMParser().parseFromString(source, 'application/xml');
if (doc.querySelector('parsererror')) {
throw new Error(doc.querySelector('parsererror').textContent.trim());
}
const walk = (el, depth) => {
const kids = [...el.childNodes];
const elementOnly =
kids.some((n) => n.nodeType === 1) &&
kids.every((n) => n.nodeType !== 3 || !n.nodeValue.trim());
if (!elementOnly) return; // mixed or text-only: leave the subtree alone
for (const n of kids) if (n.nodeType === 3) el.removeChild(n);
for (const child of [...el.children]) {
el.insertBefore(doc.createTextNode('\n' + unit.repeat(depth + 1)), child);
walk(child, depth + 1);
}
el.appendChild(doc.createTextNode('\n' + unit.repeat(depth)));
};
walk(doc.documentElement, 0);
return new XMLSerializer().serializeToString(doc);
}
// Browsers never resolve external entities, so XXE is not reachable here.
// Internal entity expansion is, so cap the input size before parsing.# ElementTree.indent (3.9+) only adds whitespace where an element has no
# non-whitespace text, so mixed content survives. It does drop comments,
# because the default parser never builds them.
import xml.etree.ElementTree as ET
from defusedxml.ElementTree import fromstring
root = fromstring(source) # safe: no entity expansion, no network
ET.indent(root, space=' ') # four spaces
print(ET.tostring(root, encoding='unicode'))
# lxml keeps comments and processing instructions, and etree.indent applies
# the same mixed-content rule:
#
# from lxml import etree
# parser = etree.XMLParser(resolve_entities=False, no_network=True,
# load_dtd=False, huge_tree=False)
# tree = etree.fromstring(source.encode(), parser)
# etree.indent(tree, space=' ')
# print(etree.tostring(tree, encoding='unicode'))
#
# Do not add remove_blank_text=True unless the document has a DTD. Without
# one, lxml guesses which blank text nodes are ignorable.import javax.xml.XMLConstants;
import javax.xml.parsers.DocumentBuilderFactory;
import javax.xml.transform.*;
import javax.xml.transform.dom.DOMSource;
import javax.xml.transform.stream.StreamResult;
import javax.xml.xpath.*;
import org.w3c.dom.*;
DocumentBuilderFactory dbf = DocumentBuilderFactory.newInstance();
dbf.setFeature(XMLConstants.FEATURE_SECURE_PROCESSING, true);
dbf.setFeature("http://apache.org/xml/features/disallow-doctype-decl", true);
dbf.setXIncludeAware(false);
dbf.setExpandEntityReferences(false);
Document doc = dbf.newDocumentBuilder().parse(new java.io.File("in.xml"));
// The serialiser adds indentation on top of the whitespace already in the
// tree, so an already-indented file gets deeper on every run. Remove the
// blank text nodes first. The second predicate keeps the blanks that sit
// inside mixed content, where a sibling text node carries real characters.
XPath xpath = XPathFactory.newInstance().newXPath();
NodeList blanks = (NodeList) xpath.evaluate(
"//text()[not(normalize-space())][not(../text()[normalize-space()])]",
doc, XPathConstants.NODESET);
for (int i = 0; i < blanks.getLength(); i++) {
Node n = blanks.item(i);
n.getParentNode().removeChild(n);
}
TransformerFactory tf = TransformerFactory.newInstance();
tf.setFeature(XMLConstants.FEATURE_SECURE_PROCESSING, true);
Transformer t = tf.newTransformer();
t.setOutputProperty(OutputKeys.INDENT, "yes");
t.setOutputProperty("{http://xml.apache.org/xslt}indent-amount", "2");
t.transform(new DOMSource(doc), new StreamResult(System.out));using System.Text;
using System.Xml;
using System.Xml.Linq;
// XDocument.Parse discards whitespace-only text nodes by default, which is
// what you want for element-only content. On mixed content it also removes
// the space in <p>a <b>x</b> <i>y</i></p>, so pass
// LoadOptions.PreserveWhitespace when the document carries prose.
var doc = XDocument.Parse(source);
// DtdProcessing is Prohibit by default for the reader XDocument builds, so
// external entities are never fetched. Say it out loud when you construct
// the reader yourself.
var settings = new XmlWriterSettings
{
Indent = true,
IndentChars = " ",
OmitXmlDeclaration = false,
};
var output = new StringBuilder();
using (var writer = XmlWriter.Create(output, settings))
{
doc.Save(writer);
}
Console.WriteLine(output.ToString());
// XmlWriter stops indenting an element once character data has been written
// into it, so it will not reflow mixed content it is given.<?php
$doc = new DOMDocument();
// Both flags must be set before loading. preserveWhiteSpace = false makes
// libxml2 drop blank text nodes; without a DTD it applies a heuristic, and
// that heuristic keeps blanks whose siblings carry real text, which is what
// protects mixed content. Run it on a copy and diff the first time.
$doc->preserveWhiteSpace = false;
$doc->formatOutput = true;
libxml_use_internal_errors(true);
if (!$doc->loadXML($source, LIBXML_NONET)) {
foreach (libxml_get_errors() as $e) {
fprintf(STDERR, "XML error at line %d, column %d: %s\n",
$e->line, $e->column, trim($e->message));
}
libxml_clear_errors();
exit(1);
}
echo $doc->saveXML();# xmllint is part of libxml2 and is almost certainly already installed.
# --nonet stops it fetching a DTD the document references.
xmllint --format --nonet document.xml
# The indent unit comes from an environment variable, not a flag:
XMLLINT_INDENT=' ' xmllint --format --nonet document.xml
# Rewrite in place:
xmllint --format --nonet --output document.xml document.xml
# xmllint refuses to format a document that is not well-formed: it prints the
# first error and exits non-zero, leaving the output file untouched.
# libxml2 will not indent an element that has a text child, which is the same
# mixed-content rule this page applies.Notice the pattern: libxml2, .NET's XmlWriter and Python's ET.indent all refuse to indent an element holding character data. The recipes that go wrong are the ones stripping whitespace text nodes indiscriminately first, which is also the flaw in any formatter built on a regular expression, since a regular expression cannot see a content model.
Common questions
Is my XML uploaded when I format it?
No. The parser and the formatter are JavaScript running inside this tab, in a Web Worker. There is no server-side component to send anything to, no analytics with access to the editor and no third-party scripts.
Open your developer tools, switch to the Network tab and format a document: the page loads its own assets once and then nothing further. That matters here because the documents most in need of reformatting are the ones pulled out of production logs. Your input is kept in this browser's localStorage so a refresh does not lose it, and Clear removes it.
Does formatting change my data?
Not the data. Attribute values are copied through as written, including entity references and the original quote character. CDATA is never converted to escaped text. Comments, processing instructions and the internal DTD subset survive, and text inside mixed content is reproduced byte for byte.
There is one place characters are removed: the leading and trailing whitespace of an element containing only text. Whitespace in the middle of a text node is never collapsed. That trim also applies to an element carrying xml:space="preserve", so check those if you depend on them.
Can I format with 4 spaces or tabs instead of 2?
Yes. The indent control offers two spaces, four spaces and a tab, and the choice applies to the whole document, including the extra level used when attributes wrap onto their own lines.
Pick to match the destination: if the file lives in a repository with an .editorconfig, match that. If size matters because the document is being embedded somewhere, a tab is one byte per level instead of four.
Why did the formatter return my document unchanged?
Because it was not well-formed. Formatting requires a parse, so the input is handed back untouched and the errors are listed instead, with the line, the column and what to write there.
The usual culprits are a raw ampersand in a URL query string, an unclosed tag, a closing tag whose name does not match its opening tag, and two root elements from concatenated fragments. xmllint behaves the same way, which is why "xmllint will not format my file" is such a common search.
What happens to CDATA sections and comments?
CDATA sections come through untouched: the delimiters stay and the bytes between them are not escaped, trimmed or reindented. Check this on any formatter you use, because converting CDATA to escaped text is an option some tools take by default, and the document you get back is then not the one you pasted in.
Comments are kept by default, with a checkbox to remove them. Think before ticking it: in XSLT, Maven and Ant files they are frequently the only explanation anyone wrote down.
How large a document can it format?
Up to 20 MB. A 1 MB document formats in roughly 200 milliseconds and 5 MB in a little under a second, in a Web Worker so the editor does not freeze.
The limit exists because everything runs in this tab: there is no server to hand a large file to, and past 20 MB the parse tree can occupy several hundred megabytes and the browser stops responding. For anything larger, xmllint --format on your own machine applies the same mixed-content rule.