DTD versus XSD

DTD came first and is part of XML itself. XML 1.0 defines the document type declaration in section 2.8 and the element, attribute, entity and notation declarations across sections 3 and 4, which makes it the only schema language built into the format. XSD arrived three years later as a separate Recommendation, written after namespaces existed, and it is what most new XML is validated against.

The comparison is usually written as a list of everything XSD does better. That is accurate and useless, because it does not explain why competent teams still ship DTDs in 2026. Two capabilities have no XSD equivalent at all, and one of them is the reason a whole publishing industry has never migrated.

#Two different kinds of document

A DTD is not XML. It has its own grammar, its own comment syntax and its own reference mechanism, so no XML parser will read it, no XPath expression will query it, and no XSLT stylesheet will transform it. An XSD is an ordinary XML document, which means every tool you already own works on it. That is the difference people notice first, and it is the least interesting one.

<!ELEMENT catalog (product+)>
<!ELEMENT product (name, (sku | ean), price)>
<!ELEMENT name  (#PCDATA)>
<!ELEMENT sku   (#PCDATA)>
<!ELEMENT ean   (#PCDATA)>
<!ELEMENT price (#PCDATA)>
<!ATTLIST product
   id     ID    #REQUIRED
   status (draft|live|retired) "draft">
The same content model, as a DTD
<xs:element name="catalog">
  <xs:complexType>
    <xs:sequence>
      <xs:element name="product" maxOccurs="unbounded">
        <xs:complexType>
          <xs:sequence>
            <xs:element name="name" type="xs:string"/>
            <xs:choice>
              <xs:element name="sku" type="xs:string"/>
              <xs:element name="ean" type="xs:string"/>
            </xs:choice>
            <xs:element name="price" type="xs:decimal"/>
          </xs:sequence>
          <xs:attribute name="id" type="xs:ID" use="required"/>
          <xs:attribute name="status" default="draft">
            <xs:simpleType>
              <xs:restriction base="xs:token">
                <xs:enumeration value="draft"/>
                <xs:enumeration value="live"/>
                <xs:enumeration value="retired"/>
              </xs:restriction>
            </xs:simpleType>
          </xs:attribute>
        </xs:complexType>
      </xs:element>
    </xs:sequence>
  </xs:complexType>
</xs:element>
And as an XSD, with the types a DTD cannot express

The XSD is four times as long and buys one real constraint the DTD lacks: price is a decimal. That ratio holds at scale, and it is why DTD partisans are not simply wrong.

#Side by side

CapabilityDTDXSD 1.0
Written inIts own grammarXML
Lives inThe DOCTYPE internal subset, an external file, or bothA separate document
Text contentOne type: #PCDATA44 built-in datatypes plus your own restrictions
Attribute typesCDATA, ID, IDREF, IDREFS, ENTITY, ENTITIES, NMTOKEN, NMTOKENS, NOTATION, enumerationAny simple type
Cardinality?, * and + onlyminOccurs and maxOccurs, any integer
NamespacesCannot express them; matches the qualified name literallytargetNamespace, xs:import, per-component form
EntitiesThe only mechanism that existsNone
Attribute defaultsYes, applied by the parserYes, applied by the validator
UniquenessID and IDREF, document-wide, one ID per elementxs:key, xs:unique and xs:keyref, scoped and XPath-based
ReuseParameter entitiesNamed types, groups, xs:include and xs:import
Type derivationNoneextension and restriction
DTD as defined by XML 1.0 Fifth Edition, against XSD 1.0

#Types and counting

A DTD has one content type for text, #PCDATA, which means parsed character data and constrains nothing: against <!ELEMENT price (#PCDATA)> the value banana is valid. Attributes get slightly more, the ten types in XML 1.0 section 3.3.1, of which only enumerations and the ID family carry real meaning.

One ID rule is worth knowing because XSD inherited it. An ID value must match the Name production, so it cannot begin with a digit: id="1" is invalid, id="p1" is fine. An element may carry at most one ID attribute, and that attribute must be #IMPLIED or #REQUIRED, never defaulted. xs:ID carries all three constraints unchanged.

Cardinality is where DTDs run out first. There are three operators, ?, * and +, and no way to write a bound: "between one and five lines" is not expressible. Mixed content is worse. XML 1.0 section 3.2.2 allows exactly one shape, a repeated choice with #PCDATA first, so a DTD cannot say a paragraph holds at most two emphasis elements, only that it may hold any number in any order.

Rejected: not a legal mixed content model
<!ELEMENT desc (#PCDATA, em*)>
<!ELEMENT desc (#PCDATA | em)+>
The only legal form
<!ELEMENT desc (#PCDATA | em)*>

Both languages require deterministic content models. XML 1.0 Appendix E rules out ((a, b) | (a, c)) because the first token does not decide the branch, and XSD 1.0 imposes the same requirement under the name Unique Particle Attribution. Moving to XSD does not free you from it.

#Namespaces, and why this one is fatal

XML 1.0 became a Recommendation on 10 February 1998. Namespaces in XML followed on 14 January 1999, and DTD validation was never revised to account for it. A DTD therefore matches the qualified name exactly as it is written in the document, prefix and all, and knows nothing about what that prefix is bound to.

That inverts the rule the rest of XML follows. Under Namespaces in XML, a prefix is arbitrary: two documents that bind soap: and s: to the same URI are the same document. A DTD disagrees, so one of them validates and the other does not.

Fails against a DTD declaring soap:Envelope
<s:Envelope xmlns:s="http://schemas.xmlsoap.org/soap/envelope/">
  <s:Body/>
</s:Envelope>
Passes, though it means exactly the same thing
<soap:Envelope xmlns:soap="http://schemas.xmlsoap.org/soap/envelope/">
  <soap:Body/>
</soap:Envelope>

There is a second half. To a DTD, xmlns and xmlns:soap are ordinary attributes, so every element carrying one must declare it in an ATTLIST or a validating parser reports an undeclared attribute. The usual workaround pins the namespace with #FIXED, which is exactly why the XHTML 1.0 DTDs validate only documents written with an unprefixed default namespace.

<!ATTLIST html
   xmlns CDATA #FIXED "http://www.w3.org/1999/xhtml">

<!-- Now <html xmlns="http://www.w3.org/1999/xhtml"> validates.
     <x:html xmlns:x="http://www.w3.org/1999/xhtml"> does not,
     even though the two are identical to a namespace-aware
     processor. -->
Declaring a namespace to a language that has no idea what one is

This is not fixable within the DTD grammar, and it is the reason every namespace-based vocabulary designed after 1999, SOAP, XSLT, Atom, the sitemap protocol, is specified with something other than a DTD.

#Entities, which only DTDs have

XSD has no entity mechanism of any kind. If a document contains &nbsp; or &corp;, something must declare it, and only a DTD can. That single fact keeps DOCTYPE declarations alive in every toolchain that writes prose rather than records, because the alternative is spelling out numeric character references by hand.

<!ENTITY corp "Acme Ltd.">                    <!-- internal general -->
<!ENTITY chap1 SYSTEM "chap1.xml">            <!-- external parsed -->
<!ENTITY logo SYSTEM "logo.png" NDATA png>    <!-- unparsed, needs NOTATION -->
<!ENTITY % common SYSTEM "common.mod">        <!-- parameter entity -->
%common;

<!NOTATION png PUBLIC "-//W3C//NOTATION PNG//EN">
The four kinds of entity declaration (XML 1.0 sections 4.2 and 4.7)

Parameter entities are the second capability with no XSD equivalent. They are the DTD modularity layer and the way large vocabularies are customised: redeclare one in the internal subset and, because XML 1.0 section 4.2 makes the first declaration win, your version replaces the one in the shipped DTD and the content model changes without anyone editing a file. XSD tried this with xs:redefine, specified it badly enough that implementations disagreed, and replaced it in 1.1 with xs:override.

#Where DTDs are still in use, and why nobody migrates

  • JATS, the NISO Z39.96 standard for journal article XML, is the strongest live DTD ecosystem and actively maintained: DOCTYPEs for JATS 1.4 were added in November 2024, and BITS 2.2 for books is built on it. JATS ships in DTD, RELAX NG and XSD flavours, and the DTD is what publishers exchange.
  • Apple property lists still carry <!DOCTYPE plist PUBLIC "-//Apple//DTD PLIST 1.0//EN" ...>, emitted by macOS and iOS tooling every day.
  • XHTML doctypes, legacy but everywhere, and still the reason most people meet a DTD at all.
  • DocBook 4.x in older toolchains. Worth being precise: DocBook 5.0 moved to RELAX NG and needs no DTD.
  • TEI, whose primary schema is RELAX NG, though DTD derivations are still generated for older tools.
  • Assorted EDI, telecom and pre-2005 SOAP pipelines, where the volume is anecdotal and nobody should quote numbers.

The reasons they do not migrate are concrete. The documents depend on entities, which XSD cannot declare. The customisation layer is parameter entities, which XSD cannot reproduce, so a mechanical conversion yields a schema that validates the same files but removes the mechanism the publisher uses daily. The public identifier is an interchange contract downstream tooling keys on. And the documents already validate, so the work has a cost and no benefit.

For anything new the choice is easier than the debate suggests. Use XSD if the vocabulary has namespaces or needs typed values, which covers almost all data interchange. Use a DTD only if you need entity declarations. If what you want is a readable grammar with precise cardinality, look at RELAX NG with Schematron for the co-occurrence rules, which is what TEI, DocBook 5 and JATS all chose once they had the option.

Common questions

Can a document use a DTD and an XSD at the same time?

Yes, and in publishing it is normal: the DOCTYPE supplies entities and attribute defaults, the XSD supplies structure and types. The order is fixed by how parsing works, since the parser expands entities and applies DTD defaults while building the document, and the schema validator then sees the result.

One caveat from XML 1.0 section 5.1: a non-validating processor need not read the external subset, so defaults declared there may or may not be applied. Defaults you depend on belong in the internal subset.

Is DTD deprecated?

No, and it never has been. It is defined by XML 1.0, still current at its Fifth Edition of 26 November 2008, and no W3C document has withdrawn or deprecated it.

What changed is that new vocabularies stopped being written in it once namespaces arrived. Deprecated and unfashionable are different things, and DTD support is the more universal of the two: every conforming XML parser reads a DTD, while schema validation is an optional extra that libxml2, for one, implements only partially.

Why does my DTD validation fail when the XML looks fine?

Usually namespaces. A DTD matches the qualified name literally, so a prefix that differs from the one the DTD declares fails, and any xmlns attribute not declared in an ATTLIST is reported as undeclared.

Next is content model order. A DTD sequence is strict, and unlike xs:all there is no way to say "these children, any order" except a repeated choice, which then also permits repetition you did not want.

Should I convert an existing DTD to XSD?

Only with a reason more specific than modernisation. Conversion tools produce a schema that accepts the same documents, so nothing appears to break. What is lost is the parameter entity customisation layer and any entity declarations, so every instance that referenced an entity has to be rewritten.

The reasons that justify it: you need typed values a DTD cannot check, you need namespaces, or a receiving system only accepts XSD. If it is the third, consider keeping the DTD for authoring and generating an XSD for interchange, which is what several publishing pipelines do.

Sources

Try it