XML Diff

Compare two documents. Neither leaves your browser.

Original
Changed
WaitingPaste a document to check it. Validation runs as you type.

Everything runs in this tab. Nothing you paste is uploaded, logged or sent anywhere. Open your network panel and check.

Paste one document on the left and the other on the right and the differences appear line by line, additions numbered against the right-hand document and removals against the left. Long unchanged stretches collapse to a marker so the changes stay findable. Both documents stay in this tab, which is not true of most tools ranking for this search.

The comparison is structural rather than textual. Both documents are reformatted first, to two-space indentation with attributes sorted alphabetically, and the diff runs over the results. Two documents differing only in indentation, or only in attribute order, come out identical and the result panel says so in those words.

That is what you want when comparing an expected payload against an actual one, which is what this is built for. It is the wrong default when whitespace carries meaning, and the sections below say exactly where that line falls.

What stops being a difference

Both sides go through the same formatter before anything is compared. XML treats neither attribute order nor indentation outside mixed content as significant, so a diff reporting either describes the serialiser rather than the data. That is why a document written by Jackson and the same data written by a .NET XmlWriter can be compared directly.

  • Indentation, line breaks and where tags sit on the line.
  • Attribute order: both sides are sorted alphabetically by name.
  • Empty-element syntax: <status></status> and <status> </status> both become <status/>.
  • Leading and trailing whitespace in a text-only element, so <name> Priya </name> matches <name>Priya</name>.
  • The order of version, encoding and standalone in the declaration.

What still counts as a difference, on purpose

The formatter will not change anything it cannot change safely. Attribute values are emitted exactly as written, original quote character and all: re-escaping would turn &amp; into &amp;amp;, and decoding is impossible too, because a value referencing a DTD-declared entity such as &companyName; needs that DTD.

Mixed content, an element holding both text and child elements, is reproduced byte for byte, because whitespace between text and markup is part of the data there. That is the one place indentation genuinely appears in the diff, and it is appearing correctly.

  • Quote style: id='A-991' against id="A-991". So is the spelling of a reference, &amp; against &#38;.
  • CDATA against escaped text: <note><![CDATA[a<b]]></note> and <note>a&lt;b</note> deliver identical characters but are reported as different.
  • Comments, which are kept rather than stripped: in a config file a changed comment is often the change you wanted.
  • Namespace prefixes. Rebinding soap: to s: with the same URI is semantically identical and still shows as a full-document difference.

When normalising is the wrong default

Some documents are bytes rather than structure. XML Digital Signature digests a canonical byte sequence, so any change at all, indentation this tool normalises away included, breaks the signature. A diff calling two signed assertions identical tells you their content matches, not that both still verify.

The other case is a document that declared xml:space="preserve", or that carries pre-formatted text. Mixed-content subtrees are safe, but a text-only element whose leading whitespace matters is not: that whitespace is trimmed. Use a plain text diff when your document depends on it.

Expected against actual

The case this exists for: an integration test fails, you have the fixture it expected and the payload the service actually returned, and they are formatted differently because one was hand-written and the other came off the wire minified. A plain text diff of those is unusable; normalising both first reduces it to the three lines that really changed.

Both documents have to be well-formed first. If either fails to parse the tool stops and says so, and the errors in the left-hand document are marked in the editor with their line and column, so a truncated response is diagnosed in place rather than looking like a structural change.

It is a line diff, not a tree diff

The comparison is a longest-common-subsequence diff over lines, the same algorithm git uses. Moving an element within its parent therefore appears as a removal in one place and an addition in another, not as a move. Reordering siblings shows as a change even where the schema treats order as insignificant, because XML element order is significant by default. XMLUnit's node matchers, in the code below, are the answer when that matters.

There is a ceiling. Over 3,000 lines after formatting, the tool declines and asks you to compare a section at a time: the LCS table is quadratic, and two large documents would lock the tab up rather than merely be slow.

Doing this in code

The same idea in a test suite or a build step: canonicalise both sides, then compare. Each of these disables external entity resolution, because you are usually pointing them at a payload you did not produce.

// Browsers do not resolve external entities, so DOMParser is safe here. It
// does not throw on malformed input: it returns a document containing a
// <parsererror> element, which is why so much code accepts broken XML.
function parse(source, label) {
  const doc = new DOMParser().parseFromString(source, 'application/xml');
  const err = doc.querySelector('parsererror');
  if (err) throw new Error(label + ': ' + err.textContent.trim());
  return doc;
}

// Canonical text: two spaces per level, attributes sorted by name, empty
// elements written one way. This is what makes the comparison structural.
function canonicalise(node, depth, out) {
  const pad = '  '.repeat(depth);
  if (node.nodeType === Node.TEXT_NODE) {
    const t = node.data.trim();
    if (t) out.push(pad + t);
    return out;
  }
  if (node.nodeType !== Node.ELEMENT_NODE) return out;

  const attrs = Array.from(node.attributes)
    .sort((a, b) => a.name.localeCompare(b.name))
    .map((a) => ' ' + a.name + '="' + escapeAttr(a.value) + '"')
    .join('');

  const kids = Array.from(node.childNodes).filter(
    (c) =>
      c.nodeType === Node.ELEMENT_NODE ||
      (c.nodeType === Node.TEXT_NODE && c.data.trim() !== ''),
  );

  if (kids.length === 0) {
    out.push(pad + '<' + node.nodeName + attrs + '/>');
    return out;
  }
  out.push(pad + '<' + node.nodeName + attrs + '>');
  for (const c of kids) canonicalise(c, depth + 1, out);
  out.push(pad + '</' + node.nodeName + '>');
  return out;
}

function escapeAttr(s) {
  return s.replace(/&/g, '&amp;').replace(/</g, '&lt;').replace(/"/g, '&quot;');
}

const left = canonicalise(parse(expected, 'expected').documentElement, 0, []);
const right = canonicalise(parse(actual, 'actual').documentElement, 0, []);
// Equality is now a string compare. For a rendered diff, feed the two arrays
// to a line differ; this page runs an LCS over exactly these lines.
console.log(left.join('\n') === right.join('\n') ? 'identical' : 'different');
from lxml import etree
import difflib

# resolve_entities=False and no_network=True are the two that matter: without
# them a payload from an untrusted source can read local files (XXE).
PARSER = etree.XMLParser(resolve_entities=False, no_network=True,
                         load_dtd=False, huge_tree=False)

def canonical_lines(path: str) -> list[str]:
    with open(path, 'rb') as fh:
        doc = etree.parse(fh, PARSER)

    # C14N 2.0 sorts attributes, normalises namespace declarations and writes
    # empty elements one way. strip_text drops insignificant whitespace, which
    # is what makes indentation irrelevant to the comparison. Note that it
    # strips whitespace inside mixed content too, which is lossy.
    canon = etree.canonicalize(etree.tostring(doc), strip_text=True)

    reparsed = etree.fromstring(canon.encode(), PARSER)
    etree.indent(reparsed, space='  ')          # lxml 4.5 and later
    return etree.tostring(reparsed, encoding='unicode').splitlines(keepends=True)

diff = difflib.unified_diff(
    canonical_lines('expected.xml'),
    canonical_lines('actual.xml'),
    fromfile='expected.xml',
    tofile='actual.xml',
)
for line in diff:
    print(line, end='')
// XMLUnit 2 compares trees, not lines, so it can tell you "attribute 'total'
// differs at /order[1]/total[1]" rather than showing two lines and leaving
// you to spot it.
import javax.xml.XMLConstants;
import javax.xml.parsers.DocumentBuilderFactory;
import org.xmlunit.builder.DiffBuilder;
import org.xmlunit.builder.Input;
import org.xmlunit.diff.DefaultNodeMatcher;
import org.xmlunit.diff.Diff;
import org.xmlunit.diff.ElementSelectors;

DocumentBuilderFactory dbf = DocumentBuilderFactory.newInstance();
dbf.setFeature(XMLConstants.FEATURE_SECURE_PROCESSING, true);
dbf.setFeature("http://apache.org/xml/features/disallow-doctype-decl", true);
dbf.setFeature("http://xml.org/sax/features/external-general-entities", false);
dbf.setFeature("http://xml.org/sax/features/external-parameter-entities", false);
dbf.setXIncludeAware(false);
dbf.setNamespaceAware(true);

Diff diff = DiffBuilder.compare(Input.fromFile("expected.xml"))
    .withTest(Input.fromFile("actual.xml"))
    .withDocumentBuilderFactory(dbf)
    .ignoreComments()
    .ignoreWhitespace()        // drops whitespace-only text nodes
    .normalizeWhitespace()     // collapses runs inside the text that remains
    // byNameAndText pairs repeated elements up by content rather than by
    // position, so a reordered list is not reported as every row changing.
    .withNodeMatcher(new DefaultNodeMatcher(ElementSelectors.byNameAndText))
    .checkForSimilar()         // "similar" ignores attribute order and prefixes
    .build();

if (diff.hasDifferences()) {
    diff.getDifferences().forEach(d -> System.out.println(d));
    System.exit(1);
}
using System.IO;
using System.Linq;
using System.Xml;
using System.Xml.Linq;

// Load() without LoadOptions.PreserveWhitespace drops insignificant
// whitespace, so indentation never reaches the comparison. Prohibiting DTDs
// and nulling the resolver closes XXE.
static XDocument LoadSafely(string path)
{
    var settings = new XmlReaderSettings
    {
        DtdProcessing = DtdProcessing.Prohibit,
        XmlResolver = null,
    };
    using var reader = XmlReader.Create(path, settings);
    return XDocument.Load(reader);
}

var expected = LoadSafely("expected.xml");
var actual = LoadSafely("actual.xml");

// XNode.DeepEquals already ignores attribute order. Element order is
// significant to it, as it is to XML itself.
if (XNode.DeepEquals(expected, actual))
{
    Console.WriteLine("Identical.");
    return;
}

// Not equal: write both out normalised so a text diff is readable.
static void SortAttributes(XElement e)
{
    var sorted = e.Attributes()
                  .OrderBy(a => a.Name.NamespaceName)
                  .ThenBy(a => a.Name.LocalName)
                  .ToList();
    e.RemoveAttributes();
    e.Add(sorted);
    foreach (var child in e.Elements()) SortAttributes(child);
}

SortAttributes(expected.Root!);
SortAttributes(actual.Root!);
File.WriteAllText("expected.norm.xml", expected.ToString());
File.WriteAllText("actual.norm.xml", actual.ToString());
Console.Error.WriteLine("Documents differ. Diff the two .norm.xml files.");
# xmllint ships with libxml2 and is almost certainly already installed.
# --c14n implements Canonical XML 1.0: attributes sorted, empty elements
# expanded to a start/end pair, namespace declarations normalised.
# --nonet stops it fetching a DTD the document points at.

xmllint --nonet --c14n expected.xml > /tmp/a.c14n
xmllint --nonet --c14n actual.xml   > /tmp/b.c14n

# C14N does not re-indent, so pretty-print afterwards or the whole document
# arrives on one line and the diff is useless. --format leaves an element
# alone when it contains text of its own, so mixed content is not reflowed.
xmllint --nonet --format /tmp/a.c14n > /tmp/a.xml
xmllint --nonet --format /tmp/b.c14n > /tmp/b.xml

diff -u /tmp/a.xml /tmp/b.xml
# Exit status 1 from diff means "they differ" and is not an error. Guard for
# it explicitly in CI, or set -e will kill the job on a successful comparison.

# C14N converts to UTF-8 and drops the XML declaration, so this will not tell
# you the two files declared different encodings. Check that with head -c 100.

The split is worth naming: canonicalisation plus a line diff gives you something a human can read, and a tree comparison such as XMLUnit gives a test something to assert on with a useful failure message.

Common questions

Are my two documents uploaded to compare them?

No. Both editors, the formatter and the diff algorithm are JavaScript running in this tab, and there is no server-side component to send anything to.

That is why the tool exists. The actual payload is real production traffic: real customer names, real order values, often a bearer token in a SOAP header. Several tools ranking for this search take both files by upload. Open your network panel while you paste and it will stay empty.

Both documents are held in this browser's localStorage so a refresh does not lose your work. That never leaves your machine, and Clear removes both immediately.

Why does it say two documents are identical when they clearly are not?

Because it compares structure, not bytes. Both sides are reformatted to the same indentation with attributes sorted first, so a minified document and the same data pretty-printed at four spaces come out the same, and so do <order id="A-991" total="64.85"> and <order total="64.85" id="A-991">.

XML treats neither attribute order nor indentation outside mixed content as significant, so a diff reporting either is describing the serialiser rather than the data. If you need a byte-level comparison, and for a signed document you do, use a plain text diff. The banner above the result always states that the comparison was made after normalising indentation and attribute order, so this is never silent.

Do both documents have to be valid XML?

They have to be well-formed, which is not the same as valid. Well-formed means the syntax is correct: tags closed and properly nested, one root element, special characters escaped, attribute values quoted. Valid additionally means matching a schema, and no schema is involved here.

The comparison cannot run on a document that does not parse, because there is no structure to normalise. If either side fails the tool stops and says so rather than falling back to a text diff and giving you a result that looks meaningful. That catches a real class of bug on its own: a truncated response shows up as a parse failure with a line and column rather than a wall of red.

Can it compare documents that use different namespace prefixes?

It will report them as different, and that is a genuine limitation. soap:Envelope and s:Envelope bound to the same URI are identical to any namespace-aware consumer, but the prefix is part of the element name as written and the formatter does not rewrite prefixes.

Rewriting is not safe in general: a prefix can appear inside attribute values, in an xsi:type, in an XPath expression in a stylesheet, in a QName in a WSDL, where a formatter cannot see it. When prefix differences are what you need to see past, use a canonicalising comparison instead: the XMLUnit sample above with checkForSimilar handles it, as does C14N in the shell and Python samples.

How large a pair of documents can it handle?

Up to 3,000 lines each, measured after formatting rather than as you pasted them, so a minified document that expands to 8,000 lines is over the limit even though it arrived as one line.

The cap is deliberate. A longest-common-subsequence diff builds a table proportional to the product of the two line counts, so two 20,000-line documents would need hundreds of millions of cells and lock the tab up rather than merely be slow. For files that size, git diff over the output of xmllint --format, or the shell recipe above, handles what no browser tab should attempt.

Related tools

Background reading