CSV to XML Converter
Convert to XML, with headings turned into legal element names.
Everything runs in this tab. Nothing you paste is uploaded, logged or sent anywhere. Open your network panel and check.
Paste a CSV and each row becomes an element. The first row supplies the element names, and you choose the root name, the row name, the delimiter, and whether values are written as child elements or as attributes. The parser, the name sanitiser and the serialiser run in a Web Worker in this tab, so nothing is uploaded.
You reach for this when a spreadsheet has to feed something that only accepts XML: a bulk import into an older ERP, a payload for a SOAP endpoint, a catalogue feed, or a fixture for a test suite that reads XML. It is usually a one-way trip, which is why getting the escaping right the first time matters more than the round trip does.
Two things separate this from a split on commas. The parser implements RFC 4180, so quoted fields, embedded commas, embedded line breaks and doubled quotes all survive. And the two places where CSV and XML genuinely disagree, headings that are not legal XML names and rows with the wrong field count, are reported rather than fixed quietly behind your back.
Splitting on commas is not parsing CSV
RFC 4180 allows a field to be wrapped in double quotes, and a quoted field may contain commas, line breaks and quotes, with a literal quote written as two quotes. A converter built on a split by comma mangles all four, and mangles them silently: the row count still looks plausible, so the corruption surfaces later in whatever consumed the XML.
The parser handles those cases and is deliberate about the ones the RFC does not cover, because real exports are not always conformant. CRLF and lone CR endings are normalised to LF, a single trailing newline is dropped rather than producing a phantom empty row, and blank lines are skipped. A quote only opens a field when it is that field's first character, so a stray quote inside an unquoted value is kept as data rather than throwing the rest of the file out of alignment.
sku,description,qty,unit price
WID-9,"Widget, large",3,12.50
GRM-2,"Grommet ""heavy duty""",1,27.35
<row>
<sku>GRM-2</sku>
<description>Grommet "heavy duty"</description>
<qty>1</qty>
<unit_price>27.35</unit_price>
</row>Column headings become element names, and most of them cannot
XML 1.0 is strict about names: a name starts with a letter, an underscore or a colon, then continues with letters, digits, hyphens, full stops or colons. Spreadsheet headings almost never comply. "unit price", "2024 total" and "price (GBP)" are all illegal, and a converter has to do something about each one.
Every heading is checked and, when it fails, sanitised the same way: characters outside the legal set become underscores, a name that would start with a digit gets a leading underscore, and a name beginning with the letters xml in any case gets one too, because XML 1.0 reserves that prefix. An empty heading becomes column1, column2 and so on.
Every rename is listed in the panel as "unit price to unit_price", so you can see what your downstream XPath will have to match. That list is printed rather than counted for a reason: sanitising is not reversible, so two headings differing only in the characters being replaced can land on the same name, and XML lets sibling elements share a name, so nothing errors.
Elements or attributes
Elements are the default and are usually right. Attributes are more compact and are what some legacy importers expect; the Columns as attributes checkbox writes every value as an attribute on a self-closing row element. The root and row names you type are used exactly as entered, without sanitising, so keep them legal.
The trade is real in one direction. XML 1.0 normalises attribute values on parse, replacing a literal tab or line break with a space, so a multi-line cell would not survive a round trip. The serialiser therefore writes tabs, carriage returns and line feeds inside attributes as the character references 	, and , which are not normalised. Text content needs no such treatment.
Two other differences matter. An attribute cannot repeat on the same element, so a duplicate heading produces a document that is not well-formed, and an empty attribute is indistinguishable from an absent one.
Ragged rows are reported, not padded
A CSV where some rows carry fewer or more fields than the header is usually a symptom: an unescaped delimiter, a broken export, or two files concatenated. Padding it silently produces XML that looks fine and is wrong, so the count is reported instead: "4 rows have a different number of fields than the header."
The output is still produced. Missing values are written as empty elements so the record structure stays uniform, and extra values are kept under generated columnN names rather than dropped. Both stay visible in the output, and both mean the same thing: go and look at those rows in the source.
Why the CSV parser is hand-written
The obvious answer is a library, and it was rejected on weight. This page promises to load fast on a phone on hotel wifi, and a CSV parser plus a YAML emitter plus an XML builder adds a few hundred kilobytes to every visit for about eighty lines of behaviour. The parser is a character loop with one boolean for "inside quotes", which is the whole of RFC 4180.
The cost is dialect coverage. There is no delimiter sniffing beyond the three choices above, no quote character other than the double quote, no comment-line convention, and no type inference, because every CSV value is a string as far as XML is concerned. If your file uses a pipe delimiter or a backslash escape, normalise it before pasting.
Doing this in code
The same conversion in the languages that consume XML. The risk here is the opposite of the XML-to-CSV direction: nothing is parsed unsafely, but everything is written, and a value containing an ampersand or an angle bracket produces a broken document unless it is escaped. Each sample uses a writer that escapes for you rather than concatenating strings.
// No dependencies. The parser is RFC 4180: one loop and one flag.
function parseCsv(text, delimiter = ',') {
const src = text.replace(/\r\n?/g, '\n').replace(/\n$/, '');
const rows = [];
let row = [], field = '', quoted = false;
for (let i = 0; i < src.length; i++) {
const c = src[i];
if (quoted) {
if (c === '"' && src[i + 1] === '"') { field += '"'; i++; }
else if (c === '"') quoted = false;
else field += c;
continue;
}
if (c === '"' && field === '') quoted = true;
else if (c === delimiter) { row.push(field); field = ''; }
else if (c === '\n') { row.push(field); rows.push(row); row = []; field = ''; }
else field += c;
}
if (field !== '' || row.length) { row.push(field); rows.push(row); }
return rows;
}
// XML 1.0 names: start with a letter, underscore or colon; the 'xml'
// prefix is reserved in any case combination.
function xmlName(heading, index) {
const base = heading.trim() || 'column' + (index + 1);
if (/^[A-Za-z_][\w.\-]*$/.test(base)) return base;
let name = base.replace(/[^\w.\-]/g, '_');
if (!/^[A-Za-z_]/.test(name)) name = '_' + name;
if (/^xml/i.test(name)) name = '_' + name;
return name;
}
const esc = (s) => s.replace(/&/g, '&').replace(/</g, '<').replace(/>/g, '>');
function csvToXml(text, { root = 'rows', row = 'row', delimiter = ',' } = {}) {
const rows = parseCsv(text, delimiter);
const headers = rows[0].map(xmlName);
const out = ['<?xml version="1.0" encoding="UTF-8"?>', '<' + root + '>'];
for (const r of rows.slice(1)) {
if (r.length === 1 && r[0] === '') continue; // blank line
if (r.length !== headers.length) {
console.warn('row has ' + r.length + ' fields, header has ' + headers.length);
}
out.push(' <' + row + '>');
headers.forEach((h, i) => out.push(' <' + h + '>' + esc(r[i] ?? '') + '</' + h + '>'));
out.push(' </' + row + '>');
}
out.push('</' + root + '>');
return out.join('\n');
}# Standard library only. csv handles RFC 4180 and ElementTree escapes
# the output, so neither quoting nor & and < are your problem.
import csv
import re
import sys
import xml.etree.ElementTree as ET
NAME_OK = re.compile(r"^[A-Za-z_][\w.\-]*$")
def xml_name(heading, index):
base = heading.strip() or f"column{index + 1}"
if NAME_OK.match(base):
return base
name = re.sub(r"[^\w.\-]", "_", base)
if not re.match(r"^[A-Za-z_]", name):
name = "_" + name
if name[:3].lower() == "xml": # reserved by XML 1.0 section 2.3
name = "_" + name
return name
def csv_to_xml(path, root_name="rows", row_name="row", delimiter=","):
# newline="" lets csv handle line breaks inside quoted fields itself.
with open(path, newline="", encoding="utf-8-sig") as f:
rows = list(csv.reader(f, delimiter=delimiter))
headers = [xml_name(h, i) for i, h in enumerate(rows[0])]
root = ET.Element(root_name)
for r in rows[1:]:
if not any(field.strip() for field in r):
continue
if len(r) != len(headers):
print(f"row has {len(r)} fields, header has {len(headers)}", file=sys.stderr)
row = ET.SubElement(root, row_name)
for i, name in enumerate(headers):
ET.SubElement(row, name).text = r[i] if i < len(r) else ""
tree = ET.ElementTree(root)
ET.indent(tree, space=" ") # Python 3.9+
tree.write(sys.stdout.buffer, encoding="UTF-8", xml_declaration=True)
csv_to_xml(sys.argv[1])// Apache Commons CSV for RFC 4180 parsing, StAX for escaped output.
// <dependency>org.apache.commons:commons-csv:1.11.0</dependency>
import java.io.*;
import java.nio.charset.StandardCharsets;
import java.nio.file.*;
import java.util.*;
import javax.xml.stream.*;
import org.apache.commons.csv.*;
public class CsvToXml {
static String xmlName(String heading, int index) {
String base = heading == null ? "" : heading.trim();
if (base.isEmpty()) return "column" + (index + 1);
if (base.matches("[A-Za-z_][\\w.\\-]*")) return base;
String name = base.replaceAll("[^\\w.\\-]", "_");
if (!name.matches("^[A-Za-z_].*")) name = "_" + name;
if (name.regionMatches(true, 0, "xml", 0, 3)) name = "_" + name;
return name;
}
public static void main(String[] args) throws Exception {
CSVFormat format = CSVFormat.RFC4180.builder()
.setHeader().setSkipHeaderRecord(true).build();
try (Reader in = Files.newBufferedReader(Path.of(args[0]), StandardCharsets.UTF_8);
CSVParser parser = CSVParser.parse(in, format)) {
List<String> raw = parser.getHeaderNames();
List<String> names = new ArrayList<>();
for (int i = 0; i < raw.size(); i++) names.add(xmlName(raw.get(i), i));
XMLStreamWriter out = XMLOutputFactory.newInstance()
.createXMLStreamWriter(new OutputStreamWriter(
new FileOutputStream("out.xml"), StandardCharsets.UTF_8));
out.writeStartDocument("UTF-8", "1.0");
out.writeStartElement("rows");
for (CSVRecord record : parser) {
if (record.size() != names.size()) {
System.err.printf("row %d has %d fields, header has %d%n",
record.getRecordNumber(), record.size(), names.size());
}
out.writeStartElement("row");
for (int i = 0; i < names.size(); i++) {
out.writeStartElement(names.get(i));
// writeCharacters escapes &, < and >. Never concatenate.
out.writeCharacters(i < record.size() ? record.get(i) : "");
out.writeEndElement();
}
out.writeEndElement();
}
out.writeEndElement();
out.writeEndDocument();
out.close();
}
}
}// dotnet add package CsvHelper
using System.Globalization;
using System.Text;
using System.Text.RegularExpressions;
using System.Xml;
using CsvHelper;
using CsvHelper.Configuration;
static string XmlName(string heading, int index)
{
var b = (heading ?? "").Trim();
if (b.Length == 0) return "column" + (index + 1);
if (Regex.IsMatch(b, @"^[A-Za-z_][\w.\-]*$")) return b;
var name = Regex.Replace(b, @"[^\w.\-]", "_");
if (!Regex.IsMatch(name, @"^[A-Za-z_]")) name = "_" + name;
if (name.StartsWith("xml", StringComparison.OrdinalIgnoreCase)) name = "_" + name;
return name;
}
var config = new CsvConfiguration(CultureInfo.InvariantCulture)
{
Delimiter = ",",
HasHeaderRecord = true,
BadDataFound = ctx => Console.Error.WriteLine($"bad quoting on row {ctx.RawRecord}"),
};
using var reader = new StreamReader(args[0], Encoding.UTF8, detectEncodingFromByteOrderMarks: true);
using var csv = new CsvReader(reader, config);
csv.Read();
csv.ReadHeader();
var names = csv.HeaderRecord!.Select(XmlName).ToArray();
var settings = new XmlWriterSettings
{
Indent = true,
Encoding = new UTF8Encoding(false),
// Entitize is the default and is what keeps a newline inside an
// attribute value from being normalised away on the next parse.
NewLineHandling = NewLineHandling.Entitize,
};
using var writer = XmlWriter.Create("out.xml", settings);
writer.WriteStartDocument();
writer.WriteStartElement("rows");
while (csv.Read())
{
writer.WriteStartElement("row");
for (var i = 0; i < names.Length; i++)
{
// WriteElementString escapes the value for you.
writer.WriteElementString(names[i], csv.TryGetField<string>(i, out var v) ? v : "");
}
writer.WriteEndElement();
}
writer.WriteEndElement();
writer.WriteEndDocument();<?php
// fgetcsv implements RFC 4180 quoting, and DOMDocument escapes text nodes,
// so the two halves that usually break are both handled for you.
function xml_name(string $heading, int $index): string {
$base = trim($heading);
if ($base === '') return 'column' . ($index + 1);
if (preg_match('/^[A-Za-z_][\w.\-]*$/', $base)) return $base;
$name = preg_replace('/[^\w.\-]/', '_', $base);
if (!preg_match('/^[A-Za-z_]/', $name)) $name = '_' . $name;
if (stripos($name, 'xml') === 0) $name = '_' . $name;
return $name;
}
$handle = fopen($argv[1], 'r');
$header = fgetcsv($handle);
// Strip a UTF-8 BOM from the first heading; Excel writes one.
$header[0] = preg_replace('/^\xEF\xBB\xBF/', '', $header[0]);
$names = [];
foreach ($header as $i => $h) $names[] = xml_name($h, $i);
$doc = new DOMDocument('1.0', 'UTF-8');
$doc->formatOutput = true;
$root = $doc->createElement('rows');
$doc->appendChild($root);
$line = 1;
while (($fields = fgetcsv($handle)) !== false) {
$line++;
if ($fields === [null] || $fields === ['']) continue; // blank line
if (count($fields) !== count($names)) {
fprintf(STDERR, "line %d has %d fields, header has %d\n",
$line, count($fields), count($names));
}
$row = $doc->createElement('row');
foreach ($names as $i => $name) {
// createTextNode escapes; createElement($name, $value) does not.
$el = $doc->createElement($name);
$el->appendChild($doc->createTextNode($fields[$i] ?? ''));
$row->appendChild($el);
}
$root->appendChild($row);
}
fclose($handle);
echo $doc->saveXML();# Import-Csv parses RFC 4180 quoting, including embedded newlines.
# ConvertTo-Xml escapes the values.
Import-Csv -Path .\data.csv -Delimiter ',' |
ConvertTo-Xml -As String -NoTypeInformation |
Out-File -FilePath .\out.xml -Encoding utf8
# Note the shape it produces. Headings become Name attributes, not
# element names, so nothing needs sanitising and nothing round-trips
# into the schema you probably wanted:
#
# <Objects>
# <Object>
# <Property Name="unit price">12.50</Property>
# </Object>
# </Objects>
#
# For heading-as-element-name output, build it explicitly:
$rows = Import-Csv -Path .\data.csv
$doc = New-Object System.Xml.XmlDocument
$root = $doc.AppendChild($doc.CreateElement('rows'))
foreach ($r in $rows) {
$row = $root.AppendChild($doc.CreateElement('row'))
foreach ($p in $r.PSObject.Properties) {
$name = $p.Name -replace '[^\w.\-]', '_'
if ($name -notmatch '^[A-Za-z_]') { $name = "_$name" }
$el = $row.AppendChild($doc.CreateElement($name))
$el.InnerText = $p.Value # InnerText escapes, InnerXml does not
}
}
$doc.Save((Join-Path $PWD 'out.xml'))The rule shared by all six: let the writer escape. Building XML by concatenating strings works until the first supplier name containing an ampersand, and then produces a document that no parser will accept.
Common questions
Is my CSV uploaded anywhere?
No. The CSV parser, the name sanitiser and the XML serialiser run in a Web Worker in this tab, and there is no back end to send anything to. Open your network panel and convert something: the page assets load once, then nothing.
CSV is the format that carries the sensitive exports: a spreadsheet pasted into a converter is typically a customer list, a payroll extract or a price file under NDA. Your input is kept in this browser's localStorage so a refresh does not lose it, and Clear removes it.
What happens to a column heading like "unit price" or "2024 total"?
Both are illegal XML element names, so both are sanitised. Characters outside the legal set become underscores, giving unit_price; a name that would start with a digit gets a leading underscore, giving _2024_total. A heading starting with the letters xml in any case gets one too, because XML 1.0 reserves that prefix.
Every rename is listed in the result panel with the original beside the new name, so you can copy the real element names into whatever XPath or XSD consumes the file. To choose the names yourself, edit the header row before converting.
My values contain commas. Will the columns still line up?
Yes, as long as those values are quoted, which is what every spreadsheet and database export does automatically. The parser follows RFC 4180: a quoted field may contain commas, line breaks and quotes, and two quotes in a row inside it mean one literal quote.
A line break inside a quoted field is kept, so a multi-line address stays one value. Under Columns as attributes it is written as rather than a literal newline, because XML normalises literal breaks in attribute values into spaces and the character reference is the only form that survives.
What if some rows have more or fewer fields than the header?
You are told. The panel reports how many rows do not match the header width, because a ragged CSV nearly always means something upstream is broken: an unescaped delimiter, a truncated export, or two files concatenated.
The conversion still runs. Missing values are written as empty elements so every record keeps the same shape, and extra values are kept under generated columnN names rather than dropped. Both leave the anomaly visible in the output instead of hiding it.
Should the values be elements or attributes?
Elements unless something downstream requires attributes. Elements can repeat, hold multi-line text without encoding tricks, are easier to constrain in an XSD, and distinguish an empty value from an absent one. Attributes are more compact and are what some older importers expect.
If you do switch, a duplicate heading becomes a duplicate attribute on the same element, which is not well-formed XML. Elements have no such restriction.
My export uses semicolons, not commas. Does that work?
Yes. Set the Delimiter control to Semicolon. Excel in most European locales writes semicolon-separated files while still calling them CSV, because the comma is the decimal separator there, and a converter that assumes commas turns 12,50 into two columns.
Tab-separated files work the same way. A pipe or caret delimiter is not offered and needs a find-and-replace first. A UTF-8 byte order mark, which Excel writes at the start of the file, is trimmed off the first heading rather than becoming part of the element name.