The XML declaration and encoding
The first line of most XML documents is `<?xml version="1.0" encoding="UTF-8"?>`, and almost everything people believe about it is slightly wrong. It is optional. It does not convert anything. Its pseudo-attributes have a fixed order that is not negotiable. And the encoding it names is a claim about bytes that a parser has to verify, not an instruction it can follow.
That last point is where the real failures live. A declaration saying UTF-8 on a file that is actually Windows-1252 either explodes or silently corrupts, depending on which bytes are in it. This guide covers what the declaration means, how a parser works out the encoding before it can read the declaration that names it, and why an invisible three-byte prefix is the most common cause of "Content is not allowed in prolog".
#What the declaration is, and that it is optional
XML 1.0 (Fifth Edition) production [23] defines it as `XMLDecl ::= '<?xml' VersionInfo EncodingDecl? SDDecl? S? '?>'`. Production [22] then places it in the prolog as `XMLDecl?`, with the question mark doing a lot of work: a document entity is perfectly well-formed without one. Absent a declaration, the version is 1.0 and section 4.3.3 requires the entity to be UTF-8 or UTF-16.
It looks like a processing instruction and it is not one. Section 2.6 reserves the target name `xml` (along with any case variant of those three letters) for the specification itself, so `<?xml ... ?>` cannot be a PI and never appears in the document tree as one. Ask a DOM for the child nodes of the document and the declaration is not among them; it is metadata about the entity, not content in it.
<?xml version="1.0" encoding="UTF-8"?>
<invoice id="INV-2044">
<total currency="GBP">1250.00</total>
</invoice><invoice id="INV-2044">
<total currency="GBP">1250.00</total>
</invoice>#Three pseudo-attributes, in one fixed order
They are not attributes and they do not behave like attributes. Attribute order on an element is insignificant. Here the grammar itself imposes the sequence, so reordering them is a well-formedness failure rather than a style choice.
| Name | Required | Legal values | What it states |
|---|---|---|---|
| version | Yes, when a declaration is present | 1.0 or 1.1 | Which version of the grammar applies |
| encoding | No | A charset name, quoted | How the bytes of this entity decode |
| standalone | No | yes or no, and nothing else | Whether external markup declarations affect the content |
<?xml encoding="UTF-8" version="1.0"?>
<root/><?xml version="1.0" encoding="UTF-8"?>
<root/>Nothing may precede the declaration. Not a blank line, not a comment, not a space. Section 2.8 requires it at the start of the document entity, and the one thing permitted before it is a byte order mark, because a BOM used as an encoding signature is not character data. The syntax checker on this site enforces exactly that: it records whether the source began with U+FEFF and then requires the declaration to start at offset zero, or at offset one if a BOM was consumed. Anything else is reported with the position of the offending characters.
A version outside the 1.x range is a fatal error, but `version="1.1"` fed to a 1.0-only processor is not. Section 2.8 says such a processor treats the document as 1.0, so it fails only if the content actually uses a 1.1 feature such as ``.
#What standalone actually declares
This is the pseudo-attribute nearly everyone gets wrong. It has nothing to do with whether the document has a DTD, whether it can be understood in isolation, or whether it is self-contained in any colloquial sense. Section 2.9 defines `standalone="yes"` as an assertion that "there are no external markup declarations which affect the information passed from the XML processor to the application".
The specification then enumerates exactly four things that count as affecting that information, and the declaration must say `no` if any external markup declaration contains:
- Attributes with default values, where an element in the document appears without that attribute. The processor would have to supply the default from the external subset.
- Entities other than amp, lt, gt, apos and quot, if the document references them.
- Attributes with tokenized types, where attribute-value normalisation would change the value the application sees.
- Element types with element content, where whitespace occurs directly inside instances, because that whitespace is ignorable only once the content model is known.
<!-- invoice.dtd -->
<!ELEMENT invoice (line+)>
<!ELEMENT line EMPTY>
<!ATTLIST line
sku CDATA #REQUIRED
currency CDATA "GBP">
<!-- invoice.xml -->
<?xml version="1.0" encoding="UTF-8" standalone="yes"?>
<!DOCTYPE invoice SYSTEM "invoice.dtd">
<invoice>
<line sku="A-1"/>
</invoice>A validating parser that reads the external subset gives that `line` element a `currency` attribute of "GBP". A non-validating parser that trusted `standalone="yes"` and skipped the external subset gives it no `currency` attribute at all. Two conforming processors, two different documents, and that divergence is the entire reason the pseudo-attribute exists: it tells a non-validating processor whether it is safe to skip the fetch.
#The bootstrap problem: reading the encoding before you can read
The `encoding` pseudo-attribute sits inside the document. To read it, a parser must already be turning bytes into characters, which is the thing the pseudo-attribute is supposed to tell it how to do. XML resolves this circularity by autodetection first and the declaration second.
Section 4.3.3 sets the order of authority. A higher-level protocol wins where one exists: an HTTP `Content-Type: application/xml; charset=ISO-8859-1` header overrides whatever the document says about itself, and RFC 7303 registers the XML media types that pin that behaviour down. Absent external information, the parser reads the first few bytes, matches them against the signatures in Appendix F, and decodes far enough to read the declaration properly.
| First bytes | Conclusion |
|---|---|
| EF BB BF | UTF-8, byte order mark present |
| FE FF | UTF-16 big-endian, BOM present |
| FF FE | UTF-16 little-endian, BOM present |
| 00 00 FE FF | UTF-32 big-endian, BOM present |
| FF FE 00 00 | UTF-32 little-endian, BOM present |
| 3C 3F 78 6D | The literal `<?xm`: an ASCII-compatible 8-bit encoding, so read the declaration |
| 00 3C 00 3F | UTF-16 big-endian, no BOM |
| 3C 00 3F 00 | UTF-16 little-endian, no BOM |
| 4C 6F A7 94 | `<?xm` in EBCDIC |
The trick that makes this work is that `<?xml` is written in a small ASCII subset, so it produces a distinctive byte pattern in every encoding a parser is expected to handle. Row six is the common case: the parser recognises an ASCII-compatible family, decodes the declaration provisionally, finds `encoding="Shift_JIS"`, and swaps in the real decoder before it reads a single byte of content. When that swap fails, libxml2 reports `switching encoding: encoder error`, which is a much more informative message once you know it comes from this step.
#Byte order marks and the three bytes EF BB BF
The byte order mark is the single character U+FEFF. In UTF-16 it disambiguates endianness, and section 4.3.3 requires UTF-16 entities to begin with one. In UTF-8 there is no endianness to disambiguate, the mark serialises to the three bytes EF BB BF, and the specification says entities in UTF-8 may begin with one. May, not must, and in practice should not.
Appendix F is explicit that a BOM used as an encoding signature is not part of the document's character data. A parser that does its own byte-level detection therefore consumes it and moves on, which is why a lone UTF-8 BOM usually causes no visible trouble. The failures start when something decodes the bytes before the parser sees them.
The canonical case is Java. Handing Xerces an `InputStream` lets it run Appendix F detection and strip the mark. Handing it a `Reader`, typically a `FileReader` or an `InputStreamReader` on UTF-8, does not: Java's UTF-8 decoder passes U+FEFF straight through as an ordinary character, so the parser receives a document whose first character is not `<`. That is precisely the condition the message "Content is not allowed in prolog" describes, and the reason it is famously baffling is that the offending character is invisible in every editor.
// Fails with "Content is not allowed in prolog." on a BOM-prefixed file:
// the Reader has already decoded the bytes and kept the U+FEFF.
builder.parse(new InputSource(new FileReader("feed.xml")));
// Works: Xerces performs its own Appendix F detection on the raw bytes
// and treats the BOM as an encoding signature rather than content.
builder.parse(new FileInputStream("feed.xml"));A BOM plus a leading blank line fails even for parsers that strip the mark correctly. Detection removes EF BB BF, the next byte is 0A, and section 2.8 forbids anything before `<?xml`. Depending on the parser you then get "Content is not allowed in prolog", or libxml2's "XML declaration allowed only at the start of the document", or Python ElementTree's "XML or text declaration not at start of entity: line 1, column 0". They are all describing the same newline. Two invisible problems stacked on each other is why this specific combination burns so much time.
This site's scanner receives a JavaScript string rather than bytes, so it checks `source.charCodeAt(0) === 0xFEFF`, reports the BOM as a warning with the suggestion to save as UTF-8 without a BOM, and then advances past it so that the declaration position check still measures from the right offset. It also treats a BOM alongside a non-UTF-8 encoding declaration as a contradiction and says so, because those two facts cannot both be true.
#When the declaration and the bytes disagree
Section 4.3.3 makes it a fatal error when an entity is determined to be in a given encoding but contains byte sequences that are not legal in that encoding. Read that carefully: the error is triggered by illegal sequences, not by the mismatch itself. Whether you get a loud failure or silent corruption depends on which direction the mismatch runs.
| Declared | Actually | Result |
|---|---|---|
| UTF-8 | Windows-1252 containing e-acute as byte E9 | Fatal error. E9 is not a legal UTF-8 lead byte. |
| UTF-8 | Latin-1 mis-saved so e-acute is C3 A9 | Parses cleanly. Silently yields the two characters A-tilde and copyright sign. |
| ISO-8859-1 | Genuine UTF-8 | Parses cleanly. Every non-ASCII character doubles into mojibake. |
| Nothing declared | Windows-1252 | Fatal error, unless a higher-level protocol said otherwise. Section 4.3.3 requires UTF-8 or UTF-16. |
Rows two and three are the dangerous ones. No parser reports an error, no validator flags anything, and the corruption is discovered weeks later in a customer name. Nothing in the well-formedness rules can catch it, because a well-formed document containing the wrong characters is still well-formed.
In the browser the default makes this worse. Per the WHATWG Encoding Standard, `TextDecoder` defaults to `fatal: false`, meaning invalid sequences become U+FFFD replacement characters rather than raising anything. Detecting a genuine mismatch requires opting in.
function sniffBom(bytes) {
if (bytes[0] === 0xef && bytes[1] === 0xbb && bytes[2] === 0xbf) return 'utf-8';
if (bytes[0] === 0xfe && bytes[1] === 0xff) return 'utf-16be';
if (bytes[0] === 0xff && bytes[1] === 0xfe) return 'utf-16le';
return null;
}
function readDeclaredEncoding(bytes) {
// Safe because the declaration is ASCII in every ASCII-compatible encoding.
const head = new TextDecoder('ascii').decode(bytes.subarray(0, 256));
const match = /<\?xml[^?]*\bencoding\s*=\s*["']([^"']+)["']/.exec(head);
return match ? match[1].toLowerCase() : null;
}
export function checkEncoding(bytes) {
const bom = sniffBom(bytes);
const declared = readDeclaredEncoding(bytes);
if (bom && declared && !declared.startsWith(bom.slice(0, 6))) {
return { ok: false, reason: `BOM says ${bom}, declaration says ${declared}` };
}
try {
// fatal: true is the whole point. Without it, bad bytes become U+FFFD.
new TextDecoder(declared ?? 'utf-8', { fatal: true }).decode(bytes);
} catch (err) {
// TypeError: The encoded data was not valid for encoding utf-8
return { ok: false, reason: err.message };
}
return { ok: true };
}Two limits worth knowing before you rely on that. `TextDecoder` also defaults to `ignoreBOM: false`, so a leading BOM is stripped from the decoded output; set `ignoreBOM: true` and the U+FEFF survives into the string, at which point the document is not well-formed. And neither UTF-32 nor EBCDIC appears in the Encoding Standard's label table at all, so `new TextDecoder('utf-32le')` throws a `RangeError`. A purely browser-based tool cannot decode those without shipping its own decoder.
That constraint shapes what the checker here can honestly claim. By the time the scanner runs, the browser has already decoded the input into a UTF-16 string, so the original bytes are gone. It therefore flags mismatches heuristically: a non-UTF-8 declaration on text containing characters above U+00FF, and a BOM contradicting a non-UTF-8 declaration. The diagnostic says outright that the browser decoded the text in memory and that what you see may not match what a byte-level parser will read, because pretending otherwise would be worse than saying nothing.
Common questions
Do I need an XML declaration at all?
No. Production [22] marks it optional, and a document without one is XML 1.0 in UTF-8 or UTF-16. Plenty of valid XML in the wild has no declaration.
Include it when the document is a standalone file that other tools will open, because it removes an inference step and documents the encoding for humans reading the source. Omit it when the content is a fragment destined to be embedded, since a declaration is only legal at the very start of a document entity.
Should I use a byte order mark on UTF-8 XML?
Generally no. Section 4.3.3 permits it and forbids nothing, but it buys nothing either: UTF-8 has no byte order to mark, and Appendix F can already identify a UTF-8 document from the `<?xm` pattern alone.
What it costs is a class of invisible failure. Any consumer that decodes the bytes before parsing, or concatenates the file with another, or reads it with a shell tool that does not expect it, will see a stray U+FEFF. Save as "UTF-8 without BOM" unless a specific downstream consumer has told you it needs one.
What does standalone="yes" mean, in one sentence?
It asserts that no external markup declaration affects what the processor reports to the application, which in practice means the document uses no externally declared default attribute values, no external entities beyond the five predefined ones, no externally declared tokenized attribute types whose normalisation would matter, and no elements with element content holding significant whitespace.
It is a validity constraint, not a well-formedness one, so getting it wrong will not fail a syntax check. On a document with no DTD it is legal and says nothing at all.
My file says UTF-8 and the accented characters are still wrong. Why did no parser complain?
Because the bytes are legal UTF-8, they are just the wrong legal UTF-8. If a Latin-1 file was decoded as Latin-1 and then re-encoded as UTF-8 somewhere in the pipeline, e-acute becomes the two-character sequence A-tilde plus copyright sign, and both of those characters encode perfectly well. The declaration is accurate and the content is corrupt.
Section 4.3.3 only makes illegal byte sequences fatal, so this passes every well-formedness check ever written. Diagnosing it means looking at the raw bytes: a UTF-8 file that contains runs of C3 or C2 followed by punctuation-looking characters has usually been double-encoded.
Does the HTTP charset header override the declaration?
Yes. Section 4.3.3 subordinates the in-document declaration to any higher-level protocol, and RFC 7303 is the registration that defines that relationship for the XML media types over HTTP.
This is a frequent production surprise: a file that parses correctly from disk fails when served, because a web server appended a default `charset=ISO-8859-1` to the `Content-Type` and the client believed it over the document. Serving XML with no charset parameter at all is usually safer than serving it with a guessed one.
Sources
- W3C XML 1.0 (Fifth Edition), section 2.8: Prolog and Document Type Declaration
- W3C XML 1.0 (Fifth Edition), section 2.9: Standalone Document Declaration
- W3C XML 1.0 (Fifth Edition), section 4.3.3: Character Encoding in Entities
- W3C XML 1.0 (Fifth Edition), Appendix F: Autodetection of Character Encodings
- WHATWG Encoding Standard: TextDecoder
- RFC 7303: XML Media Types