XML Syntax Checker

Every syntax error at once, with the line, the column and the fix.

Input
WaitingPaste a document to check it. Validation runs as you type.

Everything runs in this tab. Nothing you paste is uploaded, logged or sent anywhere. Open your network panel and check.

Paste a document above, or drop a file onto the editor, and it is checked against the XML 1.0 well-formedness rules as you type. Every violation is listed with its line, its column and what to write instead, and clicking a result moves the cursor onto the offending character.

This is the check that has to pass before any other one can run. A schema validator, an XSLT processor, a SOAP stack and a browser all refuse to look at a document whose syntax is broken, so "premature end of data in tag" out of your build is a syntax problem, and no amount of reading the XSD will explain it.

What is different here is that the scanner does not stop at the first failure. It recovers and keeps going, so one pass reports the unquoted attribute on line 4, the raw ampersand on line 12 and the tag left open at the end. Each of its 37 error classes carries its own wording and its own fix, rather than one "Syntax Error" for everything. Nothing is uploaded: open your network panel and watch it stay empty.

What "well-formed" actually means

XML 1.0 names twelve well-formedness constraints and leaves the rest in the grammar. Reduced to checks a document passes or fails, it comes to this list, and every line of it is enforced here.

  • Exactly one root element, with nothing outside it but comments, processing instructions and whitespace.
  • Every start tag closed by an end tag whose name matches exactly. Names are case sensitive, so </Note> does not close <note>.
  • Elements nested rather than overlapped: <b><i>x</b></i> is not XML, however forgiving your HTML parser has been.
  • Attribute values quoted, and no attribute given twice on one element.
  • No literal < in text or in an attribute value, and no bare & outside a reference. A > is legal in text; the sequence ]]> is not.
  • Only the five predefined entities, amp, lt, gt, quot and apos, unless a DTD declares more. &nbsp; is HTML and is not one of them.
  • No character outside the set XML permits, no comment containing --, and a declaration, if present, first in the file and written as version, then encoding, then standalone.
  • Every namespace prefix bound by an xmlns declaration in scope. That rule comes from Namespaces in XML rather than XML 1.0, and it is the one most free validators skip: a leading competitor calls <ns:root> valid with no binding anywhere in the file.

Every error at once, and why other tools give you one

The specification tells a processor it must not continue normal processing after a fatal error, which is why DOMParser, Expat and .NET report the first problem and stop. That is correct behaviour rather than laziness, and it is also why fixing a large document costs one round trip per mistake.

Recovery here is deliberate. When an end tag does not match, the scanner looks for that name further up the open stack: finding it means the real mistake is the elements left open in between, so each is reported at its own opening line rather than the end tag taking the blame. A stray character inside a name absorbs the whole token, so amount$="10" produces one message about the dollar sign rather than five about the equals sign and the quote. One caveat: an early structural mistake makes everything after it look wrong, so fix the first two or three and re-check. Reporting stops at 200 problems, with a note saying so.

The failures that actually turn up

Real syntax errors are rarely exotic. They cost time because the parser names a symptom while the cause sits elsewhere, or is invisible.

  • A UTF-8 byte order mark before the declaration. Legal, invisible, and the usual cause of "Content is not allowed in prolog". It is reported as itself rather than as mystery content.
  • A declaration written as <?xml encoding="UTF-8" version="1.0"?>, or one claiming windows-1252 while the bytes are UTF-8. The second is raised as a warning about what a byte-level parser will read, because the file was decoded before this page received it.
  • &nbsp;, &mdash; and about fifty other HTML entities, named as HTML rather than as a generic undefined entity, with the numeric reference to use instead.
  • A PHP notice or stack trace appended after the closing root tag, which is why that message points you at the raw HTTP body.
  • Control characters out of a database column, which XML 1.0 forbids everywhere, CDATA included.
  • An unbound soap:, xsi: or xlink: prefix, after a fragment was copied out of a larger document and left its xmlns declaration on the original root.

Where the syntax check stops

This page answers one question: is the syntax legal. It never asks whether an <invoice> may contain a <line>, or whether a required element is missing. Those need a schema, so they live on the XSD and DTD validators instead. If a system called your document "invalid", it ran a schema check, and a green result here will not explain it.

Two limits worth stating. The internal DTD subset is read for entity declarations only, so element and attribute-list declarations in it are parsed past rather than enforced, and external DTDs are reported but never fetched. Size is capped at 20 million characters, nesting at 512 levels and entity expansion at 10 million, because everything runs in this tab.

Checking XML syntax in code

The same check in the languages that consume XML most, each in the form that is safe against external entities and entity expansion. Note how few report more than one error: that is the specification talking, not the library.

// DOMParser never throws. Given broken input it returns a document whose
// root is <parsererror>, which is why so much code silently accepts XML
// that no other parser would take.
function checkXml(source) {
  const doc = new DOMParser().parseFromString(source, 'application/xml');
  const error = doc.querySelector('parsererror');
  if (!error) return { wellFormed: true, message: null };

  // The text is browser-specific. Firefox includes a line and column,
  // Chrome's wording differs, and neither exposes them as fields.
  return { wellFormed: false, message: error.textContent.trim() };
}

// Browsers do not resolve external entities, so XXE is not reachable here.
// Internal entity expansion is, so cap the input length before parsing:
//   if (source.length > 20_000_000) throw new Error('too large');
# lxml is the only common Python parser that keeps going after a syntax
# error and hands back the whole list rather than the first entry.
from lxml import etree

parser = etree.XMLParser(
    recover=True,            # keep parsing so error_log fills up
    resolve_entities=False,  # do not expand entities
    no_network=True,         # never fetch an external DTD
    load_dtd=False,
    huge_tree=False,         # keep libxml2's depth and expansion limits
)
etree.fromstring(source.encode('utf-8'), parser)

for entry in parser.error_log:
    print(f"line {entry.line}, column {entry.column}: {entry.message}")
if not parser.error_log:
    print('well-formed')

# Without recover=True, fromstring raises etree.XMLSyntaxError and
# e.position gives a (line, column) tuple for the first failure only.
# For untrusted input with the standard library, use defusedxml.
import javax.xml.XMLConstants;
import javax.xml.parsers.SAXParserFactory;
import org.xml.sax.InputSource;
import org.xml.sax.SAXParseException;
import org.xml.sax.helpers.DefaultHandler;

SAXParserFactory factory = SAXParserFactory.newInstance();
factory.setNamespaceAware(true);
factory.setFeature(XMLConstants.FEATURE_SECURE_PROCESSING, true);
factory.setFeature("http://xml.org/sax/features/external-general-entities", false);
factory.setFeature("http://xml.org/sax/features/external-parameter-entities", false);
// Drop the next line only if your documents legitimately carry a DOCTYPE.
factory.setFeature("http://apache.org/xml/features/disallow-doctype-decl", true);

var reader = factory.newSAXParser().getXMLReader();
reader.setErrorHandler(new DefaultHandler() {
    @Override public void warning(SAXParseException e) { print("warning", e); }
    @Override public void error(SAXParseException e) { print("error", e); }
    @Override public void fatalError(SAXParseException e) throws SAXParseException {
        print("fatal", e);
        throw e;  // the specification requires processing to stop here
    }
    private void print(String kind, SAXParseException e) {
        System.err.printf("%s at line %d, column %d: %s%n",
            kind, e.getLineNumber(), e.getColumnNumber(), e.getMessage());
    }
});
reader.parse(new InputSource(new java.io.StringReader(source)));
using System.Xml;

var settings = new XmlReaderSettings
{
    DtdProcessing = DtdProcessing.Prohibit,  // no DTD, so no entity expansion
    XmlResolver = null,                      // never fetch anything
    MaxCharactersFromEntities = 1024 * 1024,
    MaxCharactersInDocument = 20L * 1024 * 1024,
    ConformanceLevel = ConformanceLevel.Document,
};

try
{
    using var reader = XmlReader.Create(new StringReader(source), settings);
    while (reader.Read()) { }   // pull every node; syntax errors surface here
    Console.WriteLine("well-formed");
}
catch (XmlException e)
{
    // LinePosition is the column, 1-based. Reading stops at this point.
    Console.Error.WriteLine(
        $"line {e.LineNumber}, column {e.LinePosition}: {e.Message}");
}
<?php
// libxml recovers internally, so libxml_get_errors() is the closest a PHP
// script gets to a full list rather than a first failure.
libxml_use_internal_errors(true);

$doc = new DOMDocument();
$doc->loadXML($source, LIBXML_NONET);   // NONET blocks external fetches

$errors = libxml_get_errors();
libxml_clear_errors();

foreach ($errors as $e) {
    $kind = $e->level === LIBXML_ERR_FATAL ? 'fatal' : 'error';
    fprintf(
        STDERR,
        "%s at line %d, column %d: %s\n",
        $kind,
        $e->line,
        $e->column,
        trim($e->message)
    );
}
exit($errors ? 1 : 0);
# xmllint ships with libxml2 and is already installed on most machines.
# --nonet stops it fetching a DTD the document points at.
xmllint --noout --nonet document.xml

# --recover keeps parsing after a failure, so you see more than the first.
xmllint --noout --nonet --recover document.xml

# Check a whole tree, one file at a time:
find . -name '*.xml' -print0 | xargs -0 -n1 xmllint --noout --nonet

# Exit status: 0 well-formed, 1 unclassified, 3 well-formedness error,
# 4 validation error against a DTD or schema.

Every sample turns something off, and that is the substance rather than boilerplate. External entity resolution is on by default in Java, and the Python standard library applies no expansion limit at all, so those flags separate a syntax check from a file-read primitive aimed at your own server.

Common questions

How do I check XML syntax without installing anything?

Paste the document into the editor at the top of this page. The check runs as you type, with no button to press and no account to create. You can also drop a .xml file onto it: the file is read through the browser File API and never transmitted.

From a terminal, xmllint is part of libxml2 and already present on macOS and most Linux distributions: xmllint --noout --nonet yourfile.xml prints nothing when the syntax is legal. It stops at the first error, which is the practical difference.

Is my XML uploaded when I check it here?

No. The scanner is JavaScript running in this tab, in a Web Worker. There is no endpoint to send a document to, and no analytics or error-reporting script with access to the editor.

That is checkable rather than a promise. Open developer tools, switch to the Network tab, paste a document and watch: this page's assets load once and nothing follows. It matters because the documents people paste into syntax checkers are SOAP envelopes carrying bearer tokens, SAML assertions and config files with connection strings.

What is the difference between a syntax checker and an XML validator?

The two get used interchangeably, which is where the confusion starts. A syntax check asks whether the document obeys the rules of XML itself: closed tags, correct nesting, one root, quoted attributes, escaped ampersands. It needs nothing but the document.

Validation in the specification's sense asks whether the document matches a schema, an XSD or a DTD: a second file describing which elements may appear where. It cannot run until the syntax is clean, which is why "invalid" from a middleware stack sometimes means a syntax error and sometimes a schema violation. This page runs the first check and says so; the XSD and DTD validators here run the second, also without uploading anything.

Why does this report five errors when my editor reports one?

Because the XML specification instructs a conforming processor to stop after a fatal error, and the parser behind your editor does that. The rule is defensible: once a tag is unclosed, the parser genuinely does not know what the rest of the document was meant to say.

This scanner recovers instead: it resynchronises on tag boundaries, absorbs a bad name token whole, and prefers to blame an unclosed start tag over the end tag that exposed it. Treat the earliest errors as the reliable ones and re-check, since a long list usually has one structural mistake near the top generating most of it.

Are XML tag names case sensitive?

Completely. <Note> and <note> are different elements, and </Note> does not close <note>. This catches people coming from HTML, where the parser folds tag names to lower case, and it is the commonest cause of a mismatched-tag error in hand-written XML.

It applies to attribute names and namespace prefixes too: xmlns:Soap and xmlns:soap are different prefixes. When a mismatch differs only in case, the message says so rather than reporting a generic mismatch. Several widely copied "XML syntax rules" summaries claim XML tags are case insensitive; they are wrong, and following them produces documents every real parser rejects.

What is the largest file it can check?

20 million characters, roughly 20 MB of ASCII. A 10 MB document scans in a little over a second and 1 MB in about 110 milliseconds, in a Web Worker, so the editor stays responsive.

For anything larger, use a streaming parser locally: xmllint, Python iterparse or any SAX parser checks well-formedness without building the whole tree.

Related tools

Background reading

Errors this fixes