XML versus JSON

The useful comparison is not which format is better, it is which one your problem already is. XML and JSON have different data models, and almost every argument about them is really an argument about whether the extra structure XML carries is load bearing for what you are building.

This is the comparison for someone making that call: where the models diverge, what mixed content and namespaces cost, how JSON Schema stacks up against XSD in 2026, and what happens to the verbosity argument once the response is gzipped.

#Where the data models diverge

The XML Infoset carries ordered children, attributes, mixed content, namespaces, comments, processing instructions and a document type declaration. RFC 8259 gives objects, arrays, strings, numbers, booleans and null. The second set is a strict subset, so every generic XML to JSON conversion discards or invents something.

  • Order. An RFC 8259 object is an unordered collection of pairs. XML children are ordered, and in a document format that order is the content.
  • Duplicate names. RFC 8259 section 4 says names SHOULD be unique, and where they are not "the behavior of software that receives such an object is unpredictable". Repeated sibling elements are ordinary XML.
  • Numbers. XML content is text, so 007 and 1.10 survive. JSON has one numeric type, and section 6 warns interoperability holds only within IEEE 754 binary64. A 19-digit identifier does not round-trip.
  • Attributes. XML has two places to put a scalar, JSON has one. Attribute or child element is an argument you never have in JSON.
  • Comments. XML has them. JSON has no production for one, which is why JSON5 and JSONC exist and why neither is JSON.
  • Encoding. RFC 8259 section 8.1 requires UTF-8 outside a closed ecosystem. XML declares its own, which is more flexible and more to get wrong.

Michael Kay put it precisely in Schema-Aware Conversion of XML to JSON: a generic converter "is guessing what the semantics of the object model are that lie behind the lexical XML, and it is guessing wrong". The everyday version is the singleton problem. XML has no cardinality without a schema, so one item element becomes an object and two become an array, and code that maps over the result breaks on single-item responses. fast-xml-parser hands you an isArray callback; xml2js wraps everything in arrays instead.

#Mixed content is why document formats are XML

A paragraph of prose with emphasis inside it has three ordered children: text, an element, more text. XML expresses that natively. JSON objects cannot: members are unordered and a key holds one value.

What a generic converter produces
{
  "p": {
    "b": "is important",
    "#text": "Some text ."
  }
}
What the XML said
<p>Some text <b>is important</b>.</p>

The text fragments have been concatenated and their position relative to the b element is gone. This is not one library misbehaving: the author of BadgerFish lists the same defect in his own specification. Only a positional encoding such as JsonML survives it, at the cost of arrays like ["p", "Some text ", ["b", "is important"], "."].

This is why every serious document format is XML and stays XML: DocBook, DITA, JATS, TEI, OOXML, ODF, EPUB, SVG. If your payload is a record, mixed content costs you nothing. If it is a document, JSON cannot hold it without a markup convention inside strings, which is how you end up parsing HTML with regular expressions.

#Namespaces, and what having none costs

Under Namespaces in XML 1.0 an element name is a pair, a namespace URI and a local name, and the prefix is a document-local label with no significance. Two vocabularies can be mixed in one document with no coordination between their authors, which is what a SOAP envelope and an XHTML page with embedded SVG both rely on.

JSON has nothing equivalent. Keys are strings in one flat space, so merging two vocabularies means prefixing keys by hand, nesting them under a wrapper, or hoping they never collide. JSON-LD adds an @context mapping terms to IRIs, but it is an opt-in layer for linked data rather than a guarantee of the format.

The honest counterweight is that namespaces cost XML users time every week. The default namespace does not apply to unprefixed attributes, and XPath 1.0 has no notion of one at all, so an expression copied from a document matches nothing until you bind a prefix. JSON developers pay none of this because they gave up the capability.

#Schemas: XSD against JSON Schema

JSON Schema is genuinely good now, and any comparison written before about 2019 understates it. Draft 2020-12 has composition, references, conditionals and annotation vocabularies, and OpenAPI 3.1 aligned with it, so an OpenAPI description carries a real schema rather than a shape suggestion.

XSDJSON Schema 2020-12
StatusW3C Recommendation since 2001IETF Internet-Draft, never an RFC
Written inXMLJSON
DatatypesNineteen primitives including date, duration, decimalThe six JSON types; date and email live in format
Conditionalsxs:assert, XSD 1.1 only, barely implementedif/then/else and dependentSchemas, in core
CompositionType derivation, substitution groupsallOf, anyOf, oneOf, not, $ref
OrderingNative; xs:sequence is the default modelOnly inside arrays, via prefixItems
Typically runIn the transport, by the SOAP stackOpt-in, in application code

Two things reliably surprise people. In 2020-12 the format keyword is an annotation by default, not an assertion: unless the validator opts into the format-assertion vocabulary, "format": "email" reports nothing on a value that is not an email address. And XSD 1.1 assertions, the feature that would answer JSON Schema conditionals, are implemented by Xerces-J and Saxon-EE and almost nothing else, so you are really comparing 2020-12 against XSD 1.0.

The remaining structural advantage of XSD is that validation happens before your code runs. A SOAP stack rejects a non-conforming message at the boundary; a JSON Schema is a file someone has to remember to apply. That is a process difference rather than a language one, and it is why regulated formats specify XSD.

#Verbosity, and what gzip does to it

XML is more verbose, mostly because a closing tag repeats the element name, and over the wire that is the cheapest kind of redundancy there is. DEFLATE, behind gzip and ancestral to brotli, replaces a repeated byte sequence with a reference to its earlier occurrence, and a closing tag is a verbatim repeat of a string from a few bytes ago. Compressed, the two formats land far closer than raw byte counts suggest.

What compression does not remove is the parse cost. Building a DOM with namespace resolution and attribute nodes is more work than building plain objects, and JSON.parse is a native, heavily optimised path with no XML equivalent. If you are choosing on performance, that is the argument, not bytes on the wire.

#Where each is entrenched in 2026

DomainFormatWhy
Payments and financeXMLISO 20022 carries SEPA, CBPR+ and FedNow; FIXML and XBRL filings are XML
Publishing and office filesXMLJATS, DocBook, DITA, OOXML, ODF and EPUB all need mixed content
Enterprise integrationXMLSOAP and WSDL, validating in the transport
Identity and signingXMLSAML, and XML Signature can sign a subtree rather than a whole payload
Web and mobile APIsJSONNative parsing, OpenAPI 3.1 tooling, no namespace layer
Logs and eventsJSONJSON Lines streams one object per line, with no wrapper element
ConfigurationSplitJSON where machines write it, YAML or TOML where humans need comments, XML in Maven

Missing from the JSON side: anything where a human edits a long document, or a regulator specifies the message. Missing from the XML side: anything new that is browser-facing. FHIR is the clearest split, publishing both serialisations of the same resources.

  • Choose JSON for a new API consumed by browsers or mobile clients, described with JSON Schema through OpenAPI 3.1.
  • Choose XML when the payload is a document, when a counterparty publishes an XSD, or when you must sign part of a message rather than all of it.
  • Converting between them, settle the singleton-versus-array rule and the coercion rule first. Those two cause most of the bugs.

Common questions

Is JSON replacing XML?

It replaced XML for web APIs, and that turnover finished years ago. It has made no inroads into document formats, financial messaging or office files, because those need capabilities JSON does not have.

The read for 2026 is that the two stopped competing. New browser-facing work is JSON, new document and regulated-message work is XML, and the amount of XML in existence keeps growing because every DOCX, XLSX, EPUB and SVG file is one.

Can XML be converted to JSON without losing anything?

Not by a generic converter. Attributes, child order, mixed content, comments and namespaces must each be encoded into a convention, and every common convention drops something: Parker discards attributes, BadgerFish wraps every value in an object, JsonML keeps order at the cost of positional arrays.

A hand-written, schema-aware mapping is lossless in the way that matters. That is why the converter here exposes its choices as options rather than pretending one right answer exists.

Does JSON Schema do everything XSD does?

Close to it for record-shaped data, and it handles conditionals better than XSD 1.0 can. The gaps are element order, which JSON Schema expresses only inside arrays via prefixItems, and datatypes: xs:date, xs:duration and xs:decimal are types in XSD, whereas dates in JSON Schema live in format, an annotation rather than an assertion by default.

Sources

Try it