Well-formed versus valid XML
A checker that says your document is fine and a system that just rejected it can both be right. XML 1.0 defines two separate properties. Well-formedness is a property of the text alone. Validity is a property of the text measured against a grammar you supply, so the answer changes when the grammar does.
Almost every free tool collapses the two into one green tick. Here they stay apart: what the specification requires for each, an example that passes one and fails the other, and what to ask when someone calls your XML invalid.
#Two checks, two definitions
XML 1.0 (Fifth Edition) section 2.1 gives well-formedness three conditions: the text matches the production labelled document, it meets every well-formedness constraint the specification names, and each parsed entity it references is itself well-formed. All three are decidable from the bytes in front of you.
Validity is section 2.8: "An XML document is valid if it has an associated document type declaration and if the document complies with the constraints expressed in it." Taken literally, valid in XML 1.0 is a DTD word. Conformance to an XSD has its own name, schema-validity, defined in XML Schema Part 1, and yields a post-schema-validation infoset rather than a yes or no. Everyday speech calls all of it valid.
- Validity is always relative to one named schema. "This is valid XML", with no schema named, is not a checkable claim.
- Every valid document is well-formed, because XML 1.0 defines an XML document as a well-formed one. Validity sits on top, never instead.
- Section 5.1 permits non-validating processors, which is what nearly every language ships: Python ElementTree, Go encoding/xml, DOMParser. None reports a missing required element.
#The twelve well-formedness constraints
The specification names exactly twelve well-formedness constraints, each attached to the production it guards. Half concern entities rather than tags, which is why entity bugs get misfiled.
| Constraint | Violated when |
|---|---|
| WFC: Element Type Match | End-tag Name differs from start-tag Name, case sensitively. <a></A> fails. |
| WFC: Unique Att Spec | An attribute name appears twice on one tag. |
| WFC: No < in Attribute Values | A literal < in an attribute value, written directly or via entity replacement. |
| WFC: Legal Character | A character reference names a code point outside Char.  is an error, not an escape. |
| WFC: Entity Declared | An entity is referenced undeclared. Only amp, lt, gt, apos and quot are predefined, so fails. |
| WFC: Parsed Entity | A reference to an unparsed entity appears in content. |
| WFC: No Recursion | Entity replacement text refers to itself, directly or indirectly. |
| WFC: No External Entity References | An attribute value references an external entity. |
| WFC: In DTD | A parameter-entity reference appears outside the DTD. |
| WFC: PEs in Internal Subset | A parameter-entity reference appears inside a markup declaration in the internal subset. |
| WFC: PE Between Declarations | Parameter-entity replacement text is not a whole number of declarations. |
| WFC: External Subset | The external subset does not match the extSubset production. |
Those twelve are not all of it. Most real breakage fails the grammar rather than a named constraint: one root element, strict nesting, names matching Name ::= NameStartChar (NameChar)*, quoted attribute values, no literal & or <, no -- inside a comment, an XML declaration beginning at the first character.
<order id=1001>
<item>Nuts & bolts</item>
<note>See <b>terms</b>
</order><order id="1001">
<item>Nuts & bolts</item>
<note>See <b>terms</b></note>
</order>#Well-formed, and definitively invalid
<?xml version="1.0" encoding="UTF-8"?>
<note priority="high">
<to>Alice</to>
<body>Do not forget the meeting.</body>
</note><?xml version="1.0" encoding="UTF-8"?>
<xs:schema xmlns:xs="http://www.w3.org/2001/XMLSchema">
<xs:element name="note">
<xs:complexType>
<xs:sequence>
<xs:element name="to" type="xs:string"/>
<xs:element name="from" type="xs:string"/>
<xs:element name="heading" type="xs:string" minOccurs="0"/>
<xs:element name="body" type="xs:string"/>
</xs:sequence>
<xs:attribute name="priority" type="xs:positiveInteger"/>
</xs:complexType>
</xs:element>
</xs:schema>$ xmllint --noout --nonet note.xml
$ echo $?
0
$ xmllint --noout --nonet --schema note.xsd note.xml
note.xml:2: element note: Schemas validity error : Element 'note', attribute 'priority': 'high' is not a valid value of the atomic type 'xs:positiveInteger'.
note.xml:4: element body: Schemas validity error : Element 'body': This element is not expected. Expected is ( from ).
note.xml fails to validateNothing in note.xml is syntactically wrong, and the green tick is not a lie. It fails on two facts that live in note.xsd: from is required and absent, and "high" is not a positiveInteger.
The reverse case does not exist. Validity is defined over XML documents and an XML document is by definition well-formed. That is why the schema validator here refuses to run until the syntax scan is clean, rather than emitting a structural error that sends you to the wrong file.
#What a schema adds that syntax cannot
- Cardinality. minOccurs and maxOccurs in XSD, the ?, * and + operators in a DTD. Raw XML cannot say an element is required or repeatable.
- Order. xs:sequence fixes it, xs:choice offers alternatives, xs:all allows any order with each child at most once in XSD 1.0.
- Datatypes. XML Schema Part 2 defines nineteen primitive types and twenty-five derived, so xs:date rejects 2026-02-30. A DTD types only attributes.
- Cross references. ID and IDREF in a DTD; xs:key, xs:keyref and xs:unique in XSD.
- Namespace awareness. An XSD matches on namespace name plus local name. A DTD predates Namespaces in XML and treats soap:Body as one opaque name.
- Defaults and fixed values, which is where validation stops being read-only.
That last one surprises people. A DTD containing <!ATTLIST img border CDATA "0"> makes a processor that reads it report a border attribute on every img that omitted one, so the application sees an attribute that is not in the file. XSD does the same through default and fixed. Externally declared defaults head the section 2.9 list of things that force standalone="no".
Two other schema languages see real use: RELAX NG (ISO/IEC 19757-2), and Schematron (ISO/IEC 19757-3), whose rules catch co-occurrence constraints no grammar can express, such as requiring a tax element when currency is GBP.
#Why "the XML is invalid" is an ambiguous report
Told a document is invalid, you know only that something is wrong somewhere. Four questions turn that into something fixable.
- Which check failed? A syntax failure is nearly always a generation bug in your code; a schema failure is a disagreement about the contract.
- Against which schema, and which revision? Schemas get updated without announcement, so an unchanged document that validated in March can fail in April.
- Which parser, with which options? A validating parser that could not fetch the schema may skip the check rather than fail.
- Did the document fail, or the schema? A schema that is well-formed XML but not compilable produces errors that read like document errors.
| Engine | Syntax failure | Schema failure |
|---|---|---|
| libxml2 | Opening and ending tag mismatch: to line 3 and note | Element 'body': This element is not expected. Expected is ( from ). |
| Xerces | The element type "to" must be terminated by the matching end-tag "</to>". | cvc-complex-type.2.4.a: Invalid content was found starting with element 'body'. |
The cvc- prefix is worth memorising: those identifiers are the labels XML Schema Part 1 gives its own validation rules, so a message beginning cvc- is a schema-validity failure by definition and never a syntax one.
#Which check the tools here run
The validator runs the well-formedness scan as you type, because it needs nothing but the document, and reports every error in one pass with a stable code, a line and a column: XV001 unclosed tag, XV002 mismatched closing tag, XV005 unescaped ampersand, XV008 undeclared namespace prefix.
Schema validation is a deliberate second action, needing a file only you can supply and a 480 KB WebAssembly build of libxml2. Every parse uses one fixed option set: XML_PARSE_NO_XXE, XML_PARSE_NONET, XML_PARSE_NO_SYS_CATALOG, XML_PARSE_BIG_LINES. XML_PARSE_HUGE is never set, which keeps libxml2 built-in limits in force: element depth 256, entity nesting depth 20, and the amplification cap that makes a billion-laughs payload fail in milliseconds. A pasted DTD is spliced in as an internal subset; an external DTD named by SYSTEM is never fetched.
- Is my syntax broken? The XML validator or syntax checker. No second file needed.
- Does it match the contract? The XSD validator, or the DTD validator for formats predating XSD.
- Rejected with no schema given? No tool can answer that. Ask for the XSD, or the exact error string.
Common questions
Can XML be valid but not well-formed?
No, and the impossibility is definitional. Section 2.8 defines validity as a property of an XML document, and section 2.1 defines an XML document as a well-formed one, so the phrase describes nothing. A validator that refuses to run a schema check on a document failing the syntax scan is not being unhelpful: there is no defined answer to give.
My parser accepted the file but the API rejected it. What happened?
Your parser almost certainly ran a well-formedness check and theirs ran a schema check. Defaults are non-validating nearly everywhere, so a missing required element, children in the wrong order, or "high" where an integer was specified all parse cleanly.
Ask for the exact error text: cvc- points at a Xerces-based validator, "Schemas validity error" at libxml2, and both name the failing element.
Does adding a DOCTYPE make my document valid?
It makes validity a meaningful question, which is not the same thing: the document still has to comply with the declarations.
It also does nothing unless the processor validates, and most do not. A non-validating parser reads the internal subset for entity and attribute-default declarations and ignores the content models, so a document contradicting its own DOCTYPE parses without complaint.
Is well-formed the same as "it parsed without throwing"?
Close, and the gap is where bugs live. The browser DOMParser does not throw on malformed input; it returns a document containing a parsererror element, so code that never looks for one silently accepts broken XML.
Namespaces in XML adds a conformance layer of its own too: a name such as <a:b:c>, or a prefix that was never bound, satisfies the XML 1.0 Name production and is still a namespace error.