XML to JSON Converter

Convert to JSON, with every mapping decision under your control.

Input
Output
WaitingPaste a document to check it. Validation runs as you type.

Everything runs in this tab. Nothing you paste is uploaded, logged or sent anywhere. Open your network panel and check.

Paste XML above and the JSON appears as you type. The document is scanned for well-formedness first, because a converter that guesses its way through broken markup produces JSON that is quietly wrong rather than obviously wrong. If the syntax fails you get the line, the column and the fix instead of half an object.

You reach for this when a service you talk to speaks XML and everything downstream of you speaks JSON: a SOAP response you want to assert on in a test, a supplier feed you are pulling into a script, an ONIX file you need to poke at before writing the importer. Nothing is uploaded, which matters when the envelope carries a bearer token.

What is different here is that the mapping is not hidden. There is no single correct way to turn XML into JSON, every converter makes half a dozen decisions for you, and almost none of them say which. This page names each decision, shows you the toggle, and reports what the conversion cost in a notes panel beside the output.

Why there is no correct answer, only a chosen one

The XML data model is strictly richer than the JSON one. XML has ordered children, attributes, text interleaved with elements, namespaces, comments and CDATA. JSON has unordered objects, arrays, strings, numbers, booleans and null. Any function from the first to the second must discard something or invent something.

Michael Kay put it plainly in his Balisage paper on schema-aware conversion: a generic converter "is guessing what the semantics of the object model are that lie behind the lexical XML, and it's guessing wrong". These are the seven places it has to guess:

  • Attributes versus child elements. <user id="7"/> and <user><id>7</id></user> are different documents most people want as the same JSON. Merge them and <user id="7"><id>8</id></user> yields a duplicate key, which RFC 8259 calls unpredictable.
  • One occurrence or many. Nothing in <items><item>a</item></items> says whether item can repeat, so a converter guesses by counting.
  • Mixed content. <p>Some text <b>is important</b>.</p> has three ordered children, and a JSON object cannot express that order.
  • Whitespace-only text. The newline and indent between elements in pretty-printed XML are real text nodes, and keeping them is faithful but unusable.
  • Namespaces. A prefix is not the name, the URI is, and JSON has no namespace concept at all.
  • Comments and processing instructions. Neither exists in JSON.
  • Empty elements. <e/> maps plausibly to null, "", {} or an empty text key, and <e/> and <e></e> are the same document.

The convention used here, and the switches that change it

The default is the attribute-prefix convention Stefan Goessner described in 2006, which fast-xml-parser, xml2js and the AWS SDK all use variants of. Attributes take the prefix @_ so they cannot collide with a child element of the same name. Text sharing an element with attributes or children goes under #text. An element appearing once is a value; one appearing twice becomes an array. A text-only element collapses to a plain string, so <name>Alice</name> is "Alice".

The root element stays as the outermost key, comments are dropped, whitespace is trimmed, and entity references are resolved, so an attribute written "A &amp; B" arrives as "A & B". The controls above the editor set the attribute prefix, the text key, an "always an array" list of element names, and three checkboxes: strip namespace prefixes, drop attributes, coerce types. All three are off.

<order id="00042">
  <total currency="GBP">19.90</total>
  <line sku="0071">Widget</line>
  <line sku="0072">Gasket</line>
  <note/>
</order>

{
  "order": {
    "@_id": "00042",
    "total": { "@_currency": "GBP", "#text": "19.90" },
    "line": [
      { "@_sku": "0071", "#text": "Widget" },
      { "@_sku": "0072", "#text": "Gasket" }
    ],
    "note": ""
  }
}
Defaults, on a document containing all four of the awkward cases.

Why type coercion is off by default

XML without a schema is text all the way down. Turning "123" into a number is convenient right up to the point where it destroys an identifier, and identifiers are most of what moves through XML integrations.

When you do switch coercion on it is deliberately narrow. A value becomes a number only if it matches a strict JSON number grammar and then survives a round trip: the parsed result is serialised back and compared with the original, character for character. That check is what stops the failures other converters ship with. Even with coercion on, these stay strings:

  • Leading zeros. "00042" and "01730" fail the grammar outright, because a JSON number cannot start with a zero followed by more digits. Postcodes, sort codes and SKUs survive.
  • Trailing zeros after a decimal point. "19.90" parses to 19.9, which serialises back as "19.9", so the string is kept. "1.10" does not become 1.1.
  • Integers a double cannot hold. "9007199254740993" parses to a value ending 992, the round trip fails, and the string is kept. This is what corrupts nineteen-digit order identifiers elsewhere.
  • Non-canonical exponents. "1e5" parses to 100000, which is not "1e5", so it stays text.
  • Anything not exactly true, false or null. "TRUE", "yes" and "Y" stay strings.

What is lost, and the named alternatives

Three things do not survive. Document order between differently named siblings is gone, so if <line> and <discount> alternate the JSON says nothing about how. Mixed content is flattened: text fragments are concatenated and the elements between them move into their own keys, so <root>35<nested>34</nested>46</root> gives a text value of "3546". Comments are discarded.

Namespaces are kept verbatim rather than resolved, so soap:Body becomes the key "soap:Body". Stripping prefixes gives "Body", but two elements from different namespaces sharing a local name then collide on one key, and JSON offers no third option.

If this convention is not what your consumer expects, the alternatives have names. BadgerFish puts text under $, attributes under @name and carries every in-scope namespace: better round-tripping, close to unreadable. Parker drops attributes and absorbs the root, the leanest and most one-way output of the lot. JsonML writes each element as [name, attributes, children], the only common convention keeping mixed content and child order intact.

Doing this in code

The same conversion in the languages that actually consume XML. Each sample disables entity resolution, because the default in several of these stacks will fetch a URL named in a DOCTYPE, which is the XXE vulnerability. Each also names the mapping decision that library makes for you.

import { XMLParser } from 'fast-xml-parser';

const parser = new XMLParser({
  ignoreAttributes: false,       // default is true: attributes are DROPPED
  attributeNamePrefix: '@_',
  textNodeName: '#text',
  trimValues: true,
  parseTagValue: false,          // keep values as strings
  parseAttributeValue: false,
  processEntities: false,        // do not expand DOCTYPE-declared entities
  // FXP cannot know whether a tag repeats, so tell it which ones are lists.
  isArray: (name) => ['line', 'item', 'entry'].includes(name),
});

const json = parser.parse(xmlSource);

// fast-xml-parser never fetches anything over the network, so XXE is not
// reachable. It does expand entities declared in an internal DTD unless you
// set processEntities: false, so leave that off for untrusted input and cap
// the input size before parsing.
import json
import xmltodict

# disable_entities=True is the default in current xmltodict and blocks the
# expat entity-expansion attacks. Pass it explicitly so a downgrade of the
# dependency cannot silently re-enable them.
doc = xmltodict.parse(
    xml_source,
    disable_entities=True,
    attr_prefix='@_',
    cdata_key='#text',
    force_list=('line', 'item', 'entry'),   # the singleton fix
)

print(json.dumps(doc, indent=2, ensure_ascii=False))

# xmltodict returns dicts in document order, but that ordering has no meaning
# once serialised: JSON objects are unordered. Values are always strings.
# There is no coercion, which is the right default.
import com.fasterxml.jackson.databind.JsonNode;
import com.fasterxml.jackson.databind.ObjectMapper;
import com.fasterxml.jackson.dataformat.xml.XmlFactory;
import com.fasterxml.jackson.dataformat.xml.XmlMapper;
import javax.xml.stream.XMLInputFactory;

XMLInputFactory input = XMLInputFactory.newFactory();
// Neither of these is off by default. Both must be, for untrusted XML.
input.setProperty(XMLInputFactory.SUPPORT_DTD, false);
input.setProperty(XMLInputFactory.IS_SUPPORTING_EXTERNAL_ENTITIES, false);

XmlMapper xml = new XmlMapper(new XmlFactory(input));
JsonNode tree = xml.readTree(xmlSource);

String json = new ObjectMapper()
    .writerWithDefaultPrettyPrinter()
    .writeValueAsString(tree);

// Jackson's XML module merges attributes in with child elements: there is no
// prefix, so <user id="7"><id>8</id></user> loses one of the two. If your
// documents put data on attributes, bind to a class annotated with
// @JacksonXmlProperty(isAttribute = true) instead of reading a tree.
using System.Xml;
using Newtonsoft.Json;

var settings = new XmlReaderSettings
{
    DtdProcessing = DtdProcessing.Prohibit,
    XmlResolver = null,
    MaxCharactersFromEntities = 1024 * 1024,
    MaxCharactersInDocument = 20L * 1024 * 1024,
};

using var reader = XmlReader.Create(new StringReader(xmlSource), settings);
var document = new XmlDocument { XmlResolver = null };
document.Load(reader);

// omitRootObject: false keeps the root element as the outer key.
string json = JsonConvert.SerializeXmlNode(
    document, Newtonsoft.Json.Formatting.Indented, omitRootObject: false);

// Json.NET prefixes attributes with "@" and uses "#text" for text, close to
// the convention on this page. It has the singleton problem and no isArray
// hook: the only fix is a json:Array="true" attribute in the source XML,
// which you usually do not control.
<?php
libxml_use_internal_errors(true);

// LIBXML_NONET blocks network access for any DTD the document references.
// LIBXML_NOENT is deliberately NOT passed: it would substitute entities.
$xml = simplexml_load_string($source, 'SimpleXMLElement', LIBXML_NONET);

if ($xml === false) {
    foreach (libxml_get_errors() as $e) {
        fprintf(STDERR, "line %d col %d: %s\n", $e->line, $e->column, trim($e->message));
    }
    exit(1);
}

echo json_encode($xml, JSON_PRETTY_PRINT | JSON_UNESCAPED_SLASHES), "\n";

// Two things to know first. json_encode() puts attributes under an
// "@attributes" object, not a prefix. And when an element has both text and
// child elements, SimpleXML drops the text entirely: mixed content does not
// survive this route at all.
# yq v4 (Mike Farah) reads XML and writes JSON with no Python dependency.
yq -p=xml -o=json '.' document.xml

# Attributes are prefixed with + by default; match this page's convention:
yq -p=xml -o=json --xml-attribute-prefix='@_' '.' document.xml

# yq resolves nothing over the network. It also has the singleton problem and
# no per-tag array option, so a list of one comes back as a scalar. Normalise
# on the way into jq:
yq -p=xml -o=json '.' document.xml \
  | jq '.order.line |= (if type == "array" then . else [.] end)'

Every library above makes at least one mapping decision silently. Jackson merges attributes with children, SimpleXML discards mixed-content text, Json.NET and yq cannot be told which elements are lists. That is the ambiguity showing through rather than a bug in any of them. Whichever you pick, write a test that runs a single-item response and a multi-item response through the same code path.

Common questions

Is my XML sent to a server?

No. The scanner, the mapper and the JSON serialiser are all JavaScript running in this tab, and there is no backend for the conversion to reach.

Open your developer tools, switch to the Network tab, then paste and watch: this page's own assets load once and nothing follows. That matters more here than on most converters, because the XML people convert is usually an integration payload, complete with a WS-Security header or an API key.

Why does one <item> give me an object and two give me an array?

Because XML has no cardinality without a schema. Nothing in the document says whether item can repeat, so a converter guesses by counting, which makes the shape of your JSON depend on the data rather than on the contract. This is the most common way an XML integration breaks: code written against a three-item test response calls items.item.map() and works until a customer with one order arrives.

The fix is the "always an array" field above the editor. Do the same in code: fast-xml-parser has isArray, xmltodict has force_list, and xml2js defaults explicitArray to true for exactly this reason.

How do I keep XML attributes when converting to JSON?

They are kept by default, under keys prefixed with @_, so <user id="7"/> becomes {"user": {"@_id": "7"}}. The prefix exists so an attribute and a child element of the same name cannot overwrite each other.

You can change the prefix or clear it to merge attributes in among the children. Merged output reads better and silently loses a value when a name clashes, as in <user id="7"><id>8</id></user>. To drop attributes entirely there is a checkbox, which is what the Parker convention does: fine for one-way extraction, wrong for anything you convert back.

What happens to namespaces and prefixes like soap:?

The prefixed name is used verbatim, so soap:Body becomes the key "soap:Body". Nothing is resolved, because JSON cannot carry a namespace URI, and the notes panel says how many namespaces were declared.

Ticking "strip namespace prefixes" gives "Body", usually what you want when pulling one value out of a SOAP response. The risk is real: if the document has both soap:Header and wsse:Header, stripping merges them onto one key and one wins.

Should I turn on "convert numbers and booleans"?

Only if you know your data has no identifiers in it. The coercion here is stricter than most, requiring a strict number grammar and then a round-trip check, so "00042", "19.90", "1.10" and any integer too large for a double stay strings rather than being mangled.

What it cannot cover is a field holding "1" in your sample and "N/A" next Tuesday. A reasonable middle path is to leave it off and coerce the two or three fields you actually need in your own code, where the decision is written down.

What does it do with comments, CDATA and empty elements?

Comments and processing instructions are dropped. JSON has nowhere to put them, and inventing a key makes the output harder to consume for content that is by definition not data.

CDATA is treated as text: <![CDATA[a < b]]> and a &lt; b are the same content written two ways. An empty element becomes the empty string, so <note/> and <note></note> both give "note": "". Null was rejected because it reads as "no value known" rather than "present and empty".

Related tools

Background reading