Converting XML to JSON
Feed the same document to three XML-to-JSON converters and you get three different results. That is not sloppiness in two of them. XML's data model is strictly richer than JSON's, so every converter has to throw something away or invent something, and the choices it makes are rarely written down where you can see them.
This page enumerates every one of those choices with a worked example, compares the six named conventions that have accumulated since 2006, and then states exactly what the converter on this site does and why. If you are debugging a consumer that works on Tuesday and breaks on Wednesday, the section on elements that appear once is almost certainly your bug.
#The mismatch is in the data models
The XML Information Set (Second Edition) section 2 defines eleven kinds of information item: document, element, attribute, processing instruction, unexpanded entity reference, character, comment, document type declaration, unparsed entity, notation and namespace. JSON, per RFC 8259, has objects, arrays, strings, numbers, booleans and null. Nothing else.
Five properties of an XML document have no JSON counterpart at all, and each one forces a decision:
- Order. XML children are an ordered sequence. RFC 8259 section 1 defines a JSON object as "an unordered collection of zero or more name/value pairs", so sibling order is not representable.
- Attributes. XML has two ways to attach a scalar to an element. JSON has one, so the two must either be merged (and risk colliding) or distinguished by a naming convention.
- Cardinality. XML has no way to say "this element may repeat" without a schema. JSON distinguishes a value from a one-element array structurally.
- Mixed content. Text interleaved with elements is ordinary in XML and impossible in a JSON object.
- Namespaces. A name in XML is a URI plus a local part. A JSON key is a string.
Michael Kay put the general problem well in "Schema-Aware Conversion of XML to JSON" (Balisage volume 28): a generic converter "is guessing what the semantics of the object model are that lie behind the lexical XML, and it's guessing wrong", because the object model cannot in general be reverse-engineered from the XML representation. A converter that offers no options is not simpler: it has made the same guesses, silently.
#Attributes, and elements that appear once
These two are worth taking together because they are the ones that reach production. The attribute question is the easier of the two. Consider two documents most people would call equivalent:
<user id="7"/>
<user><id>7</id></user>Merging them onto one key gives the intuitive `{"user":{"id":"7"}}` for both, which is what Parker, GData and xml2js with `mergeAttrs: true` produce. The cost appears when a document has both, as `<user id="7"><id>8</id></user>` does. One key, two values, and RFC 8259 section 4 warns that when names within an object are not unique "the behavior of software that receives such an object is unpredictable". Prefixing attributes, the `@_` of fast-xml-parser or the `@` of BadgerFish, removes the collision at the cost of a key that is not a valid identifier in most languages.
The cardinality problem is worse, because nothing fails at conversion time. XML carries no occurrence information; `minOccurs` and `maxOccurs` live in the schema, not the instance. So a converter looking only at the document has to infer cardinality by counting, and a list that happens to hold one item on this request looks exactly like a scalar.
<items><item>a</item></items>
-> { "items": { "item": "a" } }
<items><item>a</item><item>b</item></items>
-> { "items": { "item": ["a", "b"] } }
// items.item.map(f) throws on the first response:
// TypeError: items.item.map is not a function// alwaysArray: ['item']
<items><item>a</item></items>
-> { "items": { "item": ["a"] } }
<items><item>a</item><item>b</item></items>
-> { "items": { "item": ["a", "b"] } }fast-xml-parser concedes the point in its own documentation: "Whether a single tag should be parsed as an array or an object, it can't be decided by FXP." Its answer is the `isArray(tagName, jPath, isLeafNode, isAttribute)` callback, which asks you. xml2js takes the opposite route and defaults `explicitArray: true`, wrapping every child in an array whether it repeats or not. That is consistent and ugly: `result.user[0].name[0]` for a document with exactly one of each.
#Mixed content, whitespace, comments and empty elements
Mixed content is the case with no good answer. An element with text and element children has ordered children of two different kinds, and a JSON object cannot hold that order.
<root>35<nested>34</nested>46</root>
{ "root": { "nested": "34", "#text": "3546" } }Note that "35" and "46" have fused into "3546". fast-xml-parser does exactly the same thing. Goessner's original 2006 convention gave up differently and degraded mixed content to a string of raw markup, treating it as unknown. BadgerFish's author lists the gap as a known defect: it "doesn't provide a way to get the different text bits in an element whose content is a mix of text and children". Only JsonML gets this right, because it uses arrays, which are ordered.
Whitespace-only text nodes are the same problem in miniature. Pretty-printing puts a newline and an indent between every element, and those are character information items like any other. XML 1.0 (Fifth Edition) section 2.10 lets an author declare intent with `xml:space="preserve"`, but almost nobody does, so converters guess. Trimming is the near-universal default: right for data documents, wrong for poetry, code listings and anything where the layout is the content.
Comments and processing instructions have no JSON representation at all. Most converters drop both; this one will put comments under a `#comment` key if you ask, and discards processing instructions unconditionally, which matters if your document carries an `<?xml-stylesheet?>` something downstream depends on.
Empty elements have four plausible mappings and no principled way to choose: `null` (Goessner), `""` (xml2js `emptyTag`), `{}`, or `{"#text": ""}`. The one hard rule is that `<e/>` and `<e></e>` are the same infoset, so a converter that distinguishes them has a bug. This site emits the empty string, because `null` invites a null check the XML never justified.
#Namespaces, and what stripping prefixes costs
A prefix is not part of an element's name. The name is the namespace URI plus the local part, and the prefix is a document-local abbreviation an author may change without changing the meaning. JSON has neither, so every converter picks a lossy encoding: keep the prefixed name (`soap:Body`), strip it (`Body`), carry the namespace scope in a sidecar (BadgerFish's `@xmlns`, repeated on every element where it is in scope), or mangle the colon (GData turns `openSearch:startIndex` into `openSearch$startIndex`).
Keeping the prefix is honest and awkward, because `obj["soap:Body"]` is the only way to reach the key in JavaScript. Stripping it is pleasant right up to the moment two namespaces collide, and then it is silent data loss:
<record xmlns:dc="http://purl.org/dc/elements/1.1/"
xmlns:cust="urn:acme:fields">
<dc:title>Report</dc:title>
<cust:title>Q3 internal</cust:title>
</record>
// removeNamespacePrefix: false
{ "record": {
"@_xmlns:dc": "http://purl.org/dc/elements/1.1/",
"@_xmlns:cust": "urn:acme:fields",
"dc:title": "Report",
"cust:title": "Q3 internal" } }
// removeNamespacePrefix: true
{ "record": {
"@_dc": "http://purl.org/dc/elements/1.1/",
"@_cust": "urn:acme:fields",
"title": ["Report", "Q3 internal"] } }Two elements from different vocabularies have become an array of two strings, indistinguishable from a genuinely repeated element. There is a second-order effect too: `xmlns:dc` is an attribute like any other, so stripping prefixes rewrites the declaration into a key called `dc`. Strip prefixes when you control the vocabulary and know it is single-namespace. Never on a SOAP envelope.
#Type coercion is where data actually dies
XML without a schema is text. Every value is a string, and a converter that turns "123" into the number 123 is guessing at a type the document never claimed. Usually the guess is convenient and right. The failures are not rare edge cases: they are identifiers.
| XML value | Naive result | What broke |
|---|---|---|
| `<zip>01730</zip>` | `1730` | Leading zeros are not decoration. ZIP codes, UK sort codes, GTINs and SKUs are fixed-width strings. strnum, which fast-xml-parser uses, has `leadingZeros: true` by default. |
| `<version>1.20</version>` | `1.2` | The trailing zero is significant to a version comparator and to a price. "6.00" becomes 6, and "1.0" becomes 1. |
| `<id>12345678901234567890</id>` | `12345678901234567000` | RFC 8259 section 6 notes interoperability holds only within IEEE 754 binary64, that is -(2^53)+1 to (2^53)-1. Snowflake and Twitter IDs are 19 digits. The low digits are gone and no error was raised. |
Booleans have a quieter version of the same problem. `<active>true</active>` should almost certainly be a boolean. `<answer>true</answer>` in a quiz application should almost certainly be a string. The document does not distinguish them and neither can the converter.
Coercion is off by default here, and the `coerce` function is conservative even when it is on. It accepts only values matching `^-?(0|[1-9][0-9]*)(\.[0-9]+)?([eE][-+]?[0-9]+)?$`, which excludes leading zeros by construction, then applies a round-trip test: the string becomes a number only if `String(n)` is byte-identical to the input. That one check rejects "1.20", "1.0" and every integer too long for a double.
#The named conventions, compared
Six conventions have names. Using `<alice charlie="david">bob</alice>` as the common input, here is what each produces:
// Goessner / fast-xml-parser (attributeNamePrefix: "@_")
{ "alice": { "@_charlie": "david", "#text": "bob" } }
// xml2js defaults (attrkey "$", charkey "_", explicitArray true)
{ "alice": { "$": { "charlie": "david" }, "_": "bob" } }
// BadgerFish
{ "alice": { "$": "bob", "@charlie": "david" } }
// Parker (attributes discarded, root absorbed)
"bob"
// GData
{ "alice": { "$t": "bob", "charlie": "david" } }
// JsonML
["alice", { "charlie": "david" }, "bob"]| Convention | Attributes | Round trip | Verbosity | Readability |
|---|---|---|---|---|
| Goessner `@_` / `#text` (fast-xml-parser, xml2js, AWS SDK) | Kept, prefixed | Good, except mixed content, sibling order, comments and singleton arrays | Low to medium | High |
| BadgerFish | Kept as `@name`, plus full in-scope namespaces at `@xmlns` | High, except mixed content and whitespace-only text | High: every scalar is wrapped in a `$` | Low |
| Parker | Discarded entirely | Very low, one way only | Lowest | Highest |
| GData | Kept but merged with child keys, so they can collide | Medium | Low | High |
| JsonML | Kept, in a position-two object | Highest for element and text order, including mixed content; drops comments, PIs and namespace semantics | Medium | Low: everything is positional |
| `fn:xml-to-json` vocabulary | Not applicable | Exact, but only for JSON round-tripped through XML | Very high | Low |
#What the converter on this site does
The XML to JSON tool defaults to the Goessner shape with `@_` for attributes and `#text` for text, arrays only where an element genuinely repeats, whitespace trimmed, prefixes kept, comments dropped, and type coercion off. The reasoning: it is the convention engineers already recognise from fast-xml-parser, xml2js and the AWS SDK, the output is readable enough to paste into a bug report, and it is the only widely used convention that keeps attributes without wrapping every scalar in an object.
Two defaults are chosen against the crowd, and both are about not lying to you:
- Coercion is off. A converter that silently turns 01730 into 1730 is worse than useless for exactly the values people most often convert. You can turn it on, and the result panel says what it did.
- A text-only element with no attributes collapses to a plain string, so `<name>Alice</name>` becomes "Alice" rather than `{ "#text": "Alice" }`. Entity references are resolved, so `A & B` becomes "A & B" in both text and attribute values.
The tool reports every decision it made: the namespaces it found and the risk of stripping them, the fact that values were kept as strings, and a prompt to declare the repeating elements when arrays were auto-detected. Hiding those decisions is how a converter becomes a source of bugs several services downstream.
Going the other way, the JSON to XML tool has the mirror problem: XML 1.0 section 2.3 restricts element names to the `Name` production, so JSON keys like "2024", "first name" and "user@email" are illegal. It sanitises rather than escaping, keeps distinct keys distinct, and lists every rename in the output panel.
Common questions
Why did my code break when a list came back with one item?
Because XML has no cardinality without a schema, and your converter inferred it by counting. One `<item>` looks like a scalar, two look like a list, so the JSON shape changes with the data rather than with the contract.
The fix is to declare it. In fast-xml-parser use the `isArray` callback, in xml2js leave `explicitArray` at its default of true, and in the tool on this site add the element name to "always array". Then write a test that exercises a single-item response, because that is the case your fixtures almost certainly lack.
Should I use the @_ prefix or merge attributes into the object?
Prefix them unless you have checked the whole vocabulary. Merging is more readable and produces keys you can write as `obj.id` rather than `obj["@_id"]`, which is a real benefit. The risk is that an element with an attribute and a child of the same name, such as `<user id="7"><id>8</id></user>`, silently loses one of them.
If you control the schema and no element has an attribute name that collides with a child element name, merging is safe and nicer to work with. If you are consuming someone else's XML, particularly anything with an extension point, keep the prefix.
Can I convert JSON to XML and back and get the original document?
Not in general, and not with any convention listed here except JsonML, which round-trips element and text order at the cost of readability. The Goessner shape loses sibling order between differently named elements, mixed content positions, comments, processing instructions, and the distinction between an element that repeats and one that does not.
There is a further trap on the XML side: XML 1.0 section 3.3.3 requires attribute-value normalisation, so a tab or newline inside an attribute value is replaced with a space when the document is re-parsed. A JSON string containing a newline therefore cannot survive a round trip through an attribute. Put it in an element.
Which convention should I pick for a new API?
The one your consumers already have a parser for, which in practice means the Goessner or fast-xml-parser shape. It is the default in the most widely deployed libraries, it is what tutorials assume, and the output is legible to a human reading a log.
Pick BadgerFish only if you genuinely need namespace URIs preserved per element, and accept that every scalar gains a `$` wrapper. Pick JsonML only if the documents contain mixed content that matters, such as prose markup, since it is the only option that keeps order. Pick Parker only for throwaway display code where attributes are known to be absent.
Is the conversion done in my browser?
Yes. The scanner, the converter and the emitters are JavaScript running in this tab, with no server component to send anything to. You can confirm it by opening your browser's network panel and watching it stay empty while you convert.
That matters more for conversion than for validation, because the documents people convert tend to be API responses, and API responses tend to carry tokens, customer records and internal identifiers.
Sources
- W3C XML 1.0 (Fifth Edition)
- W3C XML Information Set (Second Edition), section 2: Information Items
- RFC 8259: The JavaScript Object Notation (JSON) Data Interchange Format
- W3C XPath and XQuery Functions and Operators 3.1, section 17.4: Conversion to and from JSON
- JsonML syntax and grammar
- fast-xml-parser: XML parse options