Generate XSD from XML

Infer a starting schema from a sample. Review it before you trust it.

Input
Output
WaitingPaste a document to check it. Validation runs as you type.

Everything runs in this tab. Nothing you paste is uploaded, logged or sent anywhere. Open your network panel and check.

Paste a sample document and this produces an XSD 1.0 skeleton for it: one declaration per element name, xs:sequence for the children, xs:complexType where there are children or attributes, xs:simpleContent where an element carries both text and attributes, and a guessed type for every value. It runs on this site's own scanner in a Web Worker, with no engine to download.

You reach for it when a partner has sent sample payloads and no schema, when you need something to feed a code generator, or when you want a written description of a format that has so far existed only as "whatever the other system sends".

What makes this one different is that it tells you what it cannot know. Inference from one document describes that document and nothing else: not which elements are optional, not the real value ranges, not which orderings the format allows. The output carries that warning in a comment, and the sections below say which parts to review.

What it actually emits

The document is scanned first, and inference is refused if it is not well-formed. Then every element is visited and a shape recorded: which attributes appeared and whether each appeared every time, which children appeared and how many of each under one parent, and what the text looked like. Elements with no children and no attributes are emitted inline with a type; elements with children get a complexType wrapping a sequence, in the order the sample used; elements with text and attributes get simpleContent with an extension.

Three details matter. Shapes are keyed on the element name alone, not the path, so a <name> under <customer> and a <name> under <product> merge into one declaration. When an element with children appears twice, the second is written as <xs:element ref="..."/>, which XSD resolves only against a global declaration, so you have to hoist it. And if an element holds both text and children, the children win and the text is dropped: mixed content needs mixed="true", which nothing here adds.

How types are guessed, and where the guess goes wrong

Types come from the characters in the sample and nothing else. Where the same name carries different-looking values the guess widens: mixed integer forms become xs:integer, an integer beside a decimal becomes xs:decimal, anything else falls back to xs:string. An attribute that looks like a different type on a second occurrence drops to xs:string.

The failure mode follows. A status field holding "1" and "2" is typed as an integer, and the schema then rejects the "N/A" that field carries once a month. A product code of "0123" becomes a number, which loses the leading zero and makes "0123" and "123" the same value. One guard runs the other way: digits too long to be a safe integer, a 20-digit account number for instance, stay a string rather than becoming a number that loses precision.

  • An empty element gives xs:string, because nothing can be inferred from nothing.
  • Digits only give xs:nonNegativeInteger, or xs:integer with a leading minus sign. Digits with a decimal point give xs:decimal; scientific notation falls through to xs:string.
  • Exactly "true" or "false" gives xs:boolean. "1", "yes" and "Y" do not.
  • YYYY-MM-DD gives xs:date, and the same followed by T and a time gives xs:dateTime. Any other date format is a string.
  • A value starting http:// or https:// gives xs:anyURI. Other schemes and relative paths do not.

What one sample cannot tell you

Occurrence is the biggest gap. An element is marked minOccurs="0" only where the sample showed it missing: it first appeared under a later parent, or a parent that had it once was seen again without it. A field that is optional in the format but present throughout your sample comes out required, and the schema will reject a legal document tomorrow. maxOccurs is coarse too: anything seen more than once under one parent becomes unbounded, so a pair that is always exactly two becomes unlimited.

Order is asserted rather than inferred. xs:sequence says the children must appear in that order, which is what the sample showed and may not be what the format requires. If order does not matter, use xs:all, which XSD 1.0 allows only at the top of a content model with each element at most once; if the children are alternatives, use xs:choice. Nothing infers ranges, enumerations, patterns, xs:key or xs:keyref, and those carry the business rules.

Namespaces are the sharpest limit. Element names are taken exactly as written, prefix included, so a document containing <dc:title> yields <xs:element name="dc:title">, and the name of an element declaration must be an NCName, which cannot contain a colon. That schema will not compile.

The review checklist

Before a generated schema goes near a build pipeline or a partner, walk this list. Start by running the output back through the XSD validator against the same sample: a schema that cannot validate the document it came from has hit one of the cases above.

  • Every minOccurs. Which fields are genuinely required, and which were merely present in your sample?
  • Every maxOccurs="unbounded". Is there a real upper bound?
  • Every type, hardest on codes, identifiers and statuses that came out as integers or booleans.
  • xs:sequence, and whether it should be xs:all or xs:choice.
  • targetNamespace and the element names, if the sample used namespaces.
  • mixed="true" on any element holding text alongside its children.
  • Any xs:element ref, which needs a global declaration to point at.
  • Merged names: an element name meaning two things in two places needs two types.
  • What nothing can infer: enumerations, patterns, ranges, keys and keyrefs.

Inferring a schema in code

Schema inference is not in any standard library except .NET's, so these tools differ more than usual. Give the same document to all of them and you get different schemas: the differences are all in the guesses.

// No dependency: DOMParser is enough to collect the shape of a document.
// This prints the inventory a schema is built from, which is the part worth
// reading before you trust any generator's output.
function inventory(xmlText) {
  const doc = new DOMParser().parseFromString(xmlText, 'application/xml');
  if (doc.querySelector('parsererror')) throw new Error('not well-formed');

  const shapes = new Map();

  const visit = (el) => {
    let shape = shapes.get(el.tagName);
    if (!shape) {
      shape = { count: 0, attrs: new Map(), children: new Map(), values: new Set() };
      shapes.set(el.tagName, shape);
    }
    shape.count++;

    for (const a of el.attributes) {
      if (a.name === 'xmlns' || a.name.startsWith('xmlns:')) continue;
      shape.attrs.set(a.name, (shape.attrs.get(a.name) ?? 0) + 1);
    }

    const kids = [...el.children];
    const seen = new Map();
    for (const c of kids) seen.set(c.tagName, (seen.get(c.tagName) ?? 0) + 1);
    for (const [name, n] of seen) {
      const m = shape.children.get(name) ?? { min: Infinity, max: 0 };
      shape.children.set(name, { min: Math.min(m.min, n), max: Math.max(m.max, n) });
    }
    if (!kids.length && el.textContent.trim()) shape.values.add(el.textContent.trim());

    kids.forEach(visit);
  };
  visit(doc.documentElement);

  for (const [name, s] of shapes) {
    // An attribute seen fewer times than its element is optional. Present on
    // every occurrence proves nothing: it may still be optional in the format.
    const optional = [...s.attrs].filter(([, n]) => n < s.count).map(([a]) => a);
    console.log(name, 'x' + s.count, 'optional attrs:', optional.join(', ') || 'none');
  }
}
# pip install defusedxml
# The standard library parser is not safe on input you did not write.
from collections import defaultdict
from defusedxml.ElementTree import parse

def inventory(path):
    root = parse(path).getroot()
    shapes = defaultdict(lambda: {'count': 0, 'attrs': defaultdict(int),
                                  'children': {}, 'values': set()})

    def visit(el):
        s = shapes[el.tag]
        s['count'] += 1
        for name in el.attrib:
            s['attrs'][name] += 1

        counts = defaultdict(int)
        for c in el:
            counts[c.tag] += 1
        for name, n in counts.items():
            lo, hi = s['children'].get(name, (n, n))
            s['children'][name] = (min(lo, n), max(hi, n))
        if len(el) == 0 and (el.text or '').strip():
            s['values'].add(el.text.strip())

        for c in el:
            visit(c)

    visit(root)

    for tag, s in shapes.items():
        optional = [a for a, n in s['attrs'].items() if n < s['count']]
        print(f"{tag} x{s['count']}  optional attributes: {optional or 'none'}")
        for name, (lo, hi) in s['children'].items():
            # lo == 0 is never inferable from a single occurrence of the parent.
            print(f"  {name}: seen {lo}..{hi} per parent")

inventory('sample.xml')
// Apache XMLBeans: org.apache.xmlbeans:xmlbeans:5.2.1
// Inst2Xsd is the closest thing Java has to a standard inference tool, and it
// takes several instance documents, which is the main thing this page cannot.
import org.apache.xmlbeans.XmlObject;
import org.apache.xmlbeans.impl.inst2xsd.Inst2Xsd;
import org.apache.xmlbeans.impl.inst2xsd.Inst2XsdOptions;
import org.apache.xmlbeans.impl.xb.xsdschema.SchemaDocument;
import java.io.File;

XmlObject[] instances = new XmlObject[] {
    XmlObject.Factory.parse(new File("sample-1.xml")),
    XmlObject.Factory.parse(new File("sample-2.xml")),   // more samples, better guesses
};

Inst2XsdOptions options = new Inst2XsdOptions();
// RUSSIAN_DOLL nests everything; SALAMI_SLICE makes every element global,
// which is easier to hand-edit afterwards.
options.setDesign(Inst2XsdOptions.DESIGN_SALAMI_SLICE);
options.setSimpleContentTypes(Inst2XsdOptions.SIMPLE_CONTENT_TYPES_SMART);
options.setUseEnumerations(Inst2XsdOptions.ENUMERATION_NEVER);

SchemaDocument[] schemas = Inst2Xsd.inst2xsd(instances, options);
for (int i = 0; i < schemas.length; i++) {
    schemas[i].save(new File("inferred-" + i + ".xsd"));
}
using System.Xml;
using System.Xml.Schema;

// XmlSchemaInference is in the framework: no package needed. It is also the
// only one of these that will refine an existing schema with a new sample.
var settings = new XmlReaderSettings
{
    DtdProcessing = DtdProcessing.Prohibit,
    XmlResolver = null,
};

var inference = new XmlSchemaInference
{
    // Relaxed: string everywhere. Restricted: guess int, date, boolean and so
    // on, with all the risk that implies for codes and identifiers.
    TypeInference = XmlSchemaInference.InferenceOption.Restricted,
    Occurrence = XmlSchemaInference.InferenceOption.Relaxed,
};

XmlSchemaSet schemas;
using (var reader = XmlReader.Create("sample-1.xml", settings))
{
    schemas = inference.InferSchema(reader);
}
using (var reader = XmlReader.Create("sample-2.xml", settings))
{
    schemas = inference.InferSchema(reader, schemas);   // widen with a second sample
}

using var output = new XmlTextWriter("inferred.xsd", null) { Formatting = Formatting.Indented };
foreach (XmlSchema schema in schemas.Schemas())
{
    schema.Write(output);
}
# Trang, from the RELAX NG authors, is the best command-line option and takes
# as many samples as you can give it.
java -jar trang.jar -I xml -O xsd sample-1.xml sample-2.xml sample-3.xml inferred.xsd

# It will emit a DTD or a RELAX NG schema from the same input:
java -jar trang.jar -I xml -O dtd sample-1.xml inferred.dtd

# Then do the step most people skip: check the schema against the documents it
# was inferred from before trusting it.
xmllint --noout --nonet --schema inferred.xsd sample-1.xml

The capability worth chasing elsewhere is multiple samples. Trang, XMLBeans and XmlSchemaInference all accept several documents and widen the result, which turns "this field was always present" into "this field is sometimes absent" without you having to guess.

Common questions

Is my sample document uploaded to generate the schema?

No. The inference runs on this site's own scanner, in JavaScript, in a Web Worker in this tab. There is no server involved and, unlike the schema validator, not even a WebAssembly engine to download.

That is worth more here than it looks: a sample is by definition a real payload, with real customer names, account numbers and prices in it.

Can I generate a schema from more than one sample document?

Not here. This page infers from the single document in the editor, which is the honest limit of a tool built to answer in one paste.

More samples do produce a better schema, because they are the only way to learn that a field is optional. Trang takes any number of input files, XMLBeans' Inst2Xsd takes an array of instances, and XmlSchemaInference refines an existing schema with a further sample. Otherwise, generate from your largest sample and relax the minOccurs values by hand.

Why did my postcode come out as an integer?

Because in the sample it was digits, and nothing in one document distinguishes a number from a code that happens to be numeric.

This is the most common thing to fix in generated output. Postcodes, product codes, phone numbers, account references and anything with a leading zero should almost always be xs:string, sometimes with an xs:pattern.

My document uses namespaces and the schema will not compile. What now?

That is expected. Element names are taken exactly as written, so <dc:title> becomes <xs:element name="dc:title">, and the name of an element declaration must be an NCName, which cannot contain a colon.

Treat the output as a structural inventory. Add a targetNamespace to the xs:schema element and remove the prefixes from the name attributes. elementFormDefault is written as "qualified": children in the same namespace as their parent.

Is the output XSD 1.0 or 1.1?

XSD 1.0, and nothing in it uses a construct 1.1 added. That is deliberate: 1.0 is what libxml2, .NET, lxml and xmllint all implement, while a 1.1 schema works in Xerces-J, Saxon-EE and the Python xmlschema package and nowhere else.

A cross-field rule such as "end must not be before start" needs xs:assert, which is 1.1 only. Add those by hand, in whichever version your consumers can process.

How do I check the generated schema is any good?

Validate the sample against it, in the XSD validator on this site. A schema that cannot validate its own source has hit one of the known cases: a namespaced element name, mixed content, or a ref pointing at a declaration that is not global.

Then try it against a document it has never seen, ideally from a different day or customer. That is where the wrong minOccurs values surface.

Related tools

Background reading