DTD Validator

Validate against a DTD, internal or pasted. Nothing is uploaded.

Input
DTD (leave empty to use the internal subset)
WaitingPaste a document to check it. Validation runs as you type.

Everything runs in this tab. Nothing you paste is uploaded, logged or sent anywhere. Open your network panel and check.

Paste a document with an internal subset, the declarations between the square brackets of its DOCTYPE, and press Validate against DTD. If the DTD lives in a separate file, paste it in the right-hand pane instead. Either way libxml2, compiled to WebAssembly, runs the check in this tab and reports undeclared elements, content-model violations and attribute problems with a line and column.

This is the check no other free online tool offers: the results for it are a desktop vendor's marketing page, a tutorial site, an O'Reilly book excerpt and a Microsoft URL with previous-versions in the path. Meanwhile the format is in daily use, in journal publishing, Apple property lists, DocBook 4 toolchains and a long tail of EDI pipelines.

One behaviour to know before you start: an external DTD referenced by SYSTEM or PUBLIC is never fetched. That is deliberate, and the reason is below.

How a pasted DTD is applied

A DTD is not a separate document the way an XSD is; it is part of the document's prolog. So a pasted DTD is spliced in before parsing: the tool takes the first start tag as the DOCTYPE name, drops any DOCTYPE already there, keeps the XML declaration, and writes your declarations in as an internal subset. Rewriting rather than loading an external subset means no code path exists that could open a file or a URL.

That has one consequence, and it comes from the XML specification rather than from this tool. In an internal subset a parameter entity reference may stand where a markup declaration could stand, but not inside one, so a DTD containing <!ELEMENT article (%block.mix;)> is legal as an external subset and an error when pasted here. If yours is assembled from .ent modules, as JATS and DocBook are, flatten it first.

The syntax, and the parts that trip people up

A DTD has four declaration types and a small, sharp grammar: elements give a content model, attribute lists give types and defaults, entities give text substitutions, notations name external formats. The mixed-content form catches most people: an element holding both text and children must be written exactly as (#PCDATA | child)*, with #PCDATA first and the whole group repeatable. (#PCDATA, em)* is a syntax error, not a different meaning.

  • <!ELEMENT catalog (product+)> declares children. Operators are , for sequence and | for choice, with ? * + for optional, zero-or-more and one-or-more; EMPTY and ANY are the special models.
  • <!ATTLIST product id ID #REQUIRED status (draft|live) "draft"> declares attributes. Defaults are #REQUIRED, #IMPLIED, #FIXED "value", or a literal default.
  • Attribute types are CDATA, the tokenized ID, IDREF, IDREFS, ENTITY, ENTITIES, NMTOKEN and NMTOKENS, and enumerations. An element may carry at most one ID attribute.
  • ID values must be XML names, so id="1" is invalid and id="n1" is fine. xs:ID in XSD inherits the same rule.
  • <!ENTITY corp "Acme Ltd."> declares an internal general entity; SYSTEM and PUBLIC declare external ones, and NDATA marks an unparsed entity needing a NOTATION.
  • <!ENTITY % common SYSTEM "common.mod"> declares a parameter entity, referenced as %common;, which is the DTD's only modularity mechanism.

External DTDs are never fetched, and that is the point

When a document says <!DOCTYPE article SYSTEM "JATS-journalpublishing1.dtd">, a validator that honours it has to go and get that file: a request from your machine, on your network, driven by content someone else wrote. Nothing here does that. Every parse uses XML_PARSE_NO_XXE and NONET, and this build registers no resource loader, so a SYSTEM identifier resolves to nothing whether it names a URL or a path.

The DTD is the classic attack surface in XML, which is why this is a property rather than a limitation. External entity declarations in a DOCTYPE are the whole of XXE: a file:// identifier read into an element, an intranet address probed from inside your perimeter, an out-of-band channel built from a parameter entity. Nested entities are the billion-laughs vector, and libxml2 warns in its own documentation that DTD validation should never be enabled on untrusted input. XML_PARSE_HUGE is never set either, so the ceilings stay on: element depth 256, entity nesting 20, an amplification cap.

Where DTDs are still genuinely used

The DTD is the only schema language built into the XML specification itself, so every conforming parser validates against one with no extra dependency. That is why it survives where adding a schema library is not an option. The strongest living ecosystem is scholarly publishing: JATS, the NISO Z39.96 standard for journal article XML, added its 1.4 DOCTYPEs in November 2024.

  • JATS and BITS, in journal and book publishing. Shipped in DTD, RELAX NG and W3C Schema flavours; the DTD is what most ingestion systems check.
  • Apple property lists: every plist from macOS and iOS tooling still carries the PLIST 1.0 PUBLIC identifier.
  • DocBook 4.x, in toolchains that never moved. DocBook 5.0 switched to RELAX NG, so check your version.
  • TEI, whose primary schema is RELAX NG but which still generates DTD derivations.
  • XHTML doctypes, mostly decorative now but still validated in some publishing pipelines.
  • Older SOAP and EDI pipelines, where a DTD was written once and the format has not changed since.

Validating against a DTD in code

DTD validation puts the safe default and the feature you want in conflict: the standard hardening advice is to prohibit DOCTYPE declarations, and you cannot do that and validate against a DTD. Each sample keeps the declaration and shuts off external resolution instead.

// npm i libxml2-wasm, the same engine this page runs.
import { XmlDocument, ParseOption } from 'libxml2-wasm';

// DTDVALID turns validation on. NO_XXE keeps the external subset and any
// external entities unresolved, which is what makes this safe on input you
// did not write.
const OPTS =
  ParseOption.XML_PARSE_DTDVALID |
  ParseOption.XML_PARSE_NO_XXE |
  ParseOption.XML_PARSE_NONET;

export function validateDtd(xmlText, dtdText) {
  // A separately supplied DTD is spliced in as an internal subset rather than
  // loaded, exactly as this page does it.
  let source = xmlText;
  if (dtdText) {
    const root = xmlText.match(/<([A-Za-z_:][\w.\-:]*)[\s>/]/)?.[1] ?? 'root';
    const stripped = xmlText.replace(/<!DOCTYPE[^[>]*(\[[\s\S]*?\])?\s*>/, '');
    source = `<!DOCTYPE ${root} [\n${dtdText}\n]>\n${stripped.trimStart()}`;
  }

  let doc = null;
  try {
    doc = XmlDocument.fromString(source, { option: OPTS });
    return { valid: true, errors: [] };
  } catch (err) {
    // libxml2 reports a broken document and a document that does not follow
    // the DTD through the same exception. details is [{ line, col, message }].
    const details = err?.details ?? [{ message: String(err?.message ?? err) }];
    return { valid: false, errors: details };
  } finally {
    doc?.dispose();
  }
}
# pip install lxml
from io import StringIO
from lxml import etree

parser = etree.XMLParser(
    resolve_entities=False,   # do not expand entities from the DOCTYPE
    no_network=True,          # never fetch an external subset
    huge_tree=False,          # keep the depth and amplification limits
)

doc = etree.parse('article.xml', parser)

# A DTD held separately: validate without touching the document's own prolog.
with open('article.dtd', encoding='utf-8') as fh:
    dtd = etree.DTD(StringIO(fh.read()))

if dtd.validate(doc):
    print('valid')
else:
    for e in dtd.error_log.filter_from_errors():
        print(f'{e.line}:{e.column} {e.message}')

# To validate against an internal subset instead, parse with dtd_validation=True
# and read parser.error_log:
#   etree.XMLParser(dtd_validation=True, no_network=True, resolve_entities=False)
import javax.xml.XMLConstants;
import javax.xml.parsers.DocumentBuilder;
import javax.xml.parsers.DocumentBuilderFactory;
import org.xml.sax.InputSource;
import org.xml.sax.SAXParseException;
import org.xml.sax.helpers.DefaultHandler;
import java.io.StringReader;

DocumentBuilderFactory factory = DocumentBuilderFactory.newInstance();
factory.setValidating(true);          // DTD validation on
factory.setNamespaceAware(true);
factory.setFeature(XMLConstants.FEATURE_SECURE_PROCESSING, true);

// disallow-doctype-decl CANNOT be used here: it would reject the DOCTYPE you
// are trying to validate against. Block the fetch, not the declaration.
factory.setAttribute(XMLConstants.ACCESS_EXTERNAL_DTD, "");
factory.setFeature("http://xml.org/sax/features/external-general-entities", false);
factory.setFeature("http://xml.org/sax/features/external-parameter-entities", false);

DocumentBuilder builder = factory.newDocumentBuilder();

// Supply the DTD text yourself instead of letting the parser go and find it.
builder.setEntityResolver((publicId, systemId) ->
    new InputSource(new StringReader(dtdText)));

builder.setErrorHandler(new DefaultHandler() {
    @Override public void error(SAXParseException e) {
        System.out.printf("%d:%d %s%n", e.getLineNumber(), e.getColumnNumber(), e.getMessage());
    }
});

builder.parse(new InputSource(new StringReader(xmlText)));
using System;
using System.Xml;

var settings = new XmlReaderSettings
{
    // Parse, not Prohibit: DTD validation needs the DOCTYPE read.
    DtdProcessing = DtdProcessing.Parse,
    ValidationType = ValidationType.DTD,

    // A null resolver is what stops an external subset or entity being
    // fetched. This is the single most important line here.
    XmlResolver = null,

    MaxCharactersFromEntities = 1024 * 1024,
    MaxCharactersInDocument = 20L * 1024 * 1024,
};

settings.ValidationEventHandler += (_, e) =>
    Console.WriteLine($"{e.Exception.LineNumber}:{e.Exception.LinePosition} {e.Message}");

using var reader = XmlReader.Create("article.xml", settings);
while (reader.Read()) { }   // errors arrive through the handler as it reads

// With XmlResolver = null, a document whose DTD lives in a separate file will
// report that no DTD was found. Inline the declarations into the DOCTYPE first.
<?php
// LIBXML_DTDVALID validates against the DOCTYPE. LIBXML_NONET blocks the fetch.
libxml_use_internal_errors(true);

$doc = new DOMDocument();
$loaded = $doc->loadXML($source, LIBXML_DTDVALID | LIBXML_NONET);

foreach (libxml_get_errors() as $e) {
    // level 1 is a warning, 2 an error, 3 fatal.
    fprintf(STDERR, "%d:%d %s\n", $e->line, $e->column, trim($e->message));
}
libxml_clear_errors();

if (!$loaded) {
    exit(1);
}

// DOMDocument::validate() runs the same check on an already-parsed document,
// which is useful when you want the tree even if it turns out to be invalid.
# Validate against the document's own internal subset or DOCTYPE:
xmllint --noout --nonet --valid article.xml

# Validate against a DTD held separately, ignoring any DOCTYPE in the file.
# This is the form for JATS, DocBook and other modular DTDs: xmllint resolves
# the .ent modules from disk, which a browser cannot do.
xmllint --noout --nonet --dtdvalid JATS-journalpublishing1.dtd article.xml

# Exit status is 0 when the document is valid. --nonet is not the default, and
# without it xmllint will happily fetch a DTD over HTTP.

The Java and .NET samples show the same tension from two directions: the standard XXE mitigation is incompatible with DTD validation, so the hardening that works is to keep the declaration and remove the resolver.

Common questions

Why will it not fetch the DTD my document references?

Because that request would be driven by content you may not control, and the DTD is where XML's worst security problems live: an external entity that reads a local file into an element, or probes an intranet address from inside your network. The certain way to prevent it is to register no resolver at all.

The cost is one manual step: fetch the DTD yourself and paste it in.

Does anything I paste leave the browser?

No. libxml2 runs as WebAssembly in a Web Worker in this tab, so the document and the DTD both stay here. There is no upload endpoint and no analytics with access to the editor.

That matters for the people who still use DTDs: a journal article under embargo, a plist pulled off a customer's device, an EDI message carrying a partner's pricing. None of those belong in a form post.

Can I validate against the DocBook, JATS or XHTML DTD?

Yes, if you have it as a single flattened file. The top-level DTD of a modular family will not work: those pull in .ent and .mod files with parameter entities and SYSTEM identifiers, and none of those references resolve here.

There is also a specification rule: in an internal subset a parameter entity reference may stand as a whole declaration but not inside one, so (%block.mix;) is an error rather than an expansion. Use xmllint with --dtdvalid locally for these families.

Should I be using a DTD or an XSD?

Use an XSD if you need data types, namespaces, or cardinality more precise than optional and repeatable. A DTD cannot say a field is an integer between 1 and 99, and it matches literal prefixed names, so a document that binds the same namespace to a different prefix fails.

Use a DTD if a system you feed requires one, which covers most of publishing, or if you value that it needs no library at all: a 40-line DTD says as much as 150 lines of XSD.

Why is my id="1" attribute rejected?

Because ID values must be XML names, and a name cannot start with a digit. Prefix the value with a letter or an underscore and the error goes away.

Two related rules catch people: an element may carry at most one ID attribute, and it must be declared #REQUIRED or #IMPLIED, never with a default. xs:ID in XSD inherits the same constraint.

Is it safe to validate a document that came from someone else?

Here, yes, which is unusual enough to be worth stating. The two risks with a DTD are external resolution and entity expansion, and both are closed: no resource loader is registered, and XML_PARSE_HUGE is never set, so the depth limit of 256, the entity nesting limit of 20 and the amplification cap all apply.

Be more careful in your own code: Java and older PHP resolve external entities by default.

Related tools

Background reading

Errors this fixes