Validatore di feed Atom

RFC 4287 controllata sul serio, ereditarietà di author inclusa.

Scaricato direttamente dall’origine. Molti server lo bloccano; se succede, incolla il contenuto.
Input
In attesaIncolla un documento per controllarlo. La convalida gira mentre scrivi.

Tutto viene eseguito in questa scheda. Nulla di ciò che incolli viene caricato, registrato o inviato da qualche parte. Apri il pannello di rete e verifica.

Incolla qui sopra un feed Atom, oppure scaricalo da una URL, e viene controllato rispetto alla RFC 4287. Ogni rilievo cita la sezione da cui proviene. I documenti di entrata Atom isolati, in cui la radice è <entry> invece di <feed>, vengono riconosciuti e controllati con le regole di entrata.

Atom è più severo di RSS, ed è proprio questo il suo scopo: quasi tutto ciò che RSS ha lasciato all’interpretazione, Atom lo ha deciso. Esattamente un id, un title e un updated sul feed e su ogni entrata. Un solo formato di data, senza varianti. Ogni costruzione di testo dichiara se è testo, HTML o XHTML.

Qui c’è una regola implementata come si deve, che la maggior parte degli strumenti salta: l’ereditarietà di author. Un’entrata non ha bisogno di un autore proprio se possiede un <atom:source> che ne porta uno, o se il feed ne ha uno. Controllare solo entry/author produce una pagina di errori falsi su un feed valido; non controllare nulla lascia passare una mancanza vera.

Perché esiste Atom, e in che cosa differisce da RSS

RSS 2.0 è stato congelato nel 2002, e con lui le sue ambiguità. La più grande era <description>: la specifica non ha mai detto se contenga testo semplice o HTML, quindi ogni lettore tirava a indovinare. L’identità dell’item era facoltativa. Le date seguivano la RFC 822, che ammette anni a due cifre e nessun fuso orario.

Atom è stato scritto all’IETF come risposta e pubblicato come RFC 4287 nel dicembre 2005. Il testo porta un type esplicito. Ogni feed e ogni entrata hanno un id obbligatorio e permanente. Le date sono RFC 3339 con fuso obbligatorio. Tutto vive in un unico namespace, così un’estensione non può collidere con il vocabolario del nucleo.

  • RSS: description può essere testo o HTML, nessuno sa quale. Atom: il type è dichiarato.
  • RSS: guid facoltativo e spesso instabile. Atom: id obbligatorio, IRI assoluta, permanente.
  • RSS: date RFC 822, fuso facoltativo. Atom: RFC 3339, fuso obbligatorio.

Che cosa devono contenere un feed e le sue entrate

Esattamente un <id>, un <title> e un <updated> sul feed, e gli stessi tre su ogni entrata. Non uno o più: esattamente uno, quindi anche un doppione è un errore. Il guasto più comune è un feed con titolo ed entrate ma senza id.

Un’entrata priva di <content> deve portare almeno un <link rel="alternate">: deve offrire al lettore o la cosa stessa o un modo per raggiungerla. Se <content> ha un attributo src, l’elemento dev’essere vuoto e l’entrata ha bisogno di un <summary>, il che vale anche per contenuti in base64.

Ogni <link> ha bisogno di un href, e nulla può portare due link rel="alternate" con lo stesso type e lo stesso hreflang, perché un lettore non saprebbe scegliere fra i due. Un rel assente vale alternate. Un feed dovrebbe inoltre portare un <link rel="self">, che la RFC 4287 formula come SHOULD: è quindi un avviso etichettato.

La regola di ereditarietà di author

La RFC 4287 richiede che per ogni entrata sia risolvibile un autore, non che l’elemento si trovi sull’entrata. La regola è una catena di ripiego: il <author> proprio dell’entrata, oppure un <author> dentro il suo <atom:source>, oppure un <author> sul <feed>. Il caso source esiste per gli aggregatori che ripubblicano entrate raccolte altrove.

Quasi nessun validatore gratuito implementa tutti e tre i passaggi. Quelli che guardano solo entry/author segnalano un errore su ogni entrata di un feed di blog a firma singola perfettamente valido, cioè la forma di feed Atom più diffusa che esista. Questo strumento percorre la catena e segnala un errore soltanto quando falliscono tutti e tre.

Un documento di entrata Atom isolato non ha un feed da cui ereditare, quindi deve portare il proprio autore. Ogni costruzione di persona ha inoltre bisogno di un <name>: un <author> con il solo indirizzo e-mail è un errore ai sensi della sezione 3.2.1.

Date e id, esattamente come li scrive la RFC

I timestamp di Atom sono RFC 3339, e la sezione 3.3 aggiunge due requisiti. Il separatore dev’essere una T maiuscola, e senza scostamento numerico il fuso dev’essere una Z maiuscola. Quindi 2026-03-14T09:30:00Z è valido e 2026-03-14t09:30:00z no, benché qualsiasi libreria di date lo analizzerebbe. Le minuscole hanno un messaggio proprio. Anche una data secca è invalida.

Un id dev’essere una IRI assoluta, quindi gli serve uno schema: https:, tag: e urn: vanno tutti bene; un percorso relativo o una stringa nuda no. Id duplicati fra le entrate sono un errore, dato che i lettori deduplicano sull’id.

Id che differiscono solo per maiuscole, per una barra finale o per una porta predefinita producono un avviso, perché gli id si confrontano carattere per carattere: /p/1 e /p/1/ sono due entrate. Lo schema tag, la RFC 4151, evita tutto questo, e tag:example.com,2026:post-4192 sopravvive a un trasloco di dominio.

Produrre e controllare Atom da codice

Le quattro regole che vale la pena automatizzare sono proprio quelle che uno schema non può esprimere: le maiuscole della RFC 3339, l’unicità degli id, l’ereditarietà di author e l’alternativa fra content e link alternate. Ogni esempio controlla queste, nella forma di analisi sicura.

// Node 18 or later.  npm i @xmldom/xmldom
import { DOMParser } from '@xmldom/xmldom';

const ATOM = 'http://www.w3.org/2005/Atom';

// RFC 3339 as RFC 4287 section 3.3 requires it: uppercase T, uppercase Z.
const RFC3339 = /^\d{4}-\d{2}-\d{2}T\d{2}:\d{2}:\d{2}(\.\d+)?(Z|[+-]\d{2}:\d{2})$/;

const xml = await (await fetch('https://example.com/atom.xml')).text();
const doc = new DOMParser().parseFromString(xml, 'text/xml');
const feed = doc.documentElement;

const kids = (el, name) => Array.from(el.getElementsByTagNameNS(ATOM, name))
  .filter((n) => n.parentNode === el);

for (const required of ['id', 'title', 'updated']) {
  if (kids(feed, required).length !== 1) {
    console.error('feed must have exactly one <' + required + '>');
  }
}

const feedHasAuthor = kids(feed, 'author').length > 0;
const seen = new Map();

for (const [i, entry] of kids(feed, 'entry').entries()) {
  const n = i + 1;

  const updated = kids(entry, 'updated')[0];
  if (updated && !RFC3339.test(updated.textContent.trim())) {
    console.error('entry ' + n + ' updated is not RFC 3339: ' + updated.textContent.trim());
  }

  const id = kids(entry, 'id')[0];
  const value = id ? id.textContent.trim() : '';
  if (!/^[A-Za-z][A-Za-z0-9+.-]*:/.test(value)) {
    console.error('entry ' + n + ' id is not an absolute IRI: ' + value);
  } else if (seen.has(value)) {
    console.error('entry ' + n + ' repeats the id of entry ' + seen.get(value));
  } else {
    seen.set(value, n);
  }

  // The inheritance chain: entry/author, else entry/source/author, else feed/author.
  const source = kids(entry, 'source')[0];
  const hasAuthor = kids(entry, 'author').length > 0
    || (source ? kids(source, 'author').length > 0 : false)
    || feedHasAuthor;
  if (!hasAuthor) console.error('entry ' + n + ' has no resolvable author');

  const hasAlternate = kids(entry, 'link')
    .some((l) => (l.getAttribute('rel') || 'alternate') === 'alternate');
  if (kids(entry, 'content').length === 0 && !hasAlternate) {
    console.error('entry ' + n + ' has neither <content> nor a <link rel="alternate">');
  }
}

// Producing a correct timestamp:
console.log(new Date().toISOString());   // 2026-03-14T09:30:00.000Z
# pip install lxml requests
import re
import sys
import requests
from datetime import datetime, timezone
from lxml import etree

ATOM = 'http://www.w3.org/2005/Atom'
NS = {'a': ATOM}

# datetime.fromisoformat is too permissive for this: it accepts a lowercase
# separator, which RFC 4287 section 3.3 forbids. Check the shape directly.
RFC3339 = re.compile(r'^\d{4}-\d{2}-\d{2}T\d{2}:\d{2}:\d{2}(\.\d+)?(Z|[+-]\d{2}:\d{2})$')
IRI = re.compile(r'^[A-Za-z][A-Za-z0-9+.\-]*:')

raw = requests.get('https://example.com/atom.xml', timeout=30).content
parser = etree.XMLParser(resolve_entities=False, no_network=True, load_dtd=False)
feed = etree.fromstring(raw, parser)

for required in ('id', 'title', 'updated'):
    if len(feed.findall('a:%s' % required, NS)) != 1:
        print('feed must have exactly one <%s>' % required)

feed_has_author = len(feed.findall('a:author', NS)) > 0
seen = {}

for n, entry in enumerate(feed.findall('a:entry', NS), start=1):
    updated = entry.findtext('a:updated', namespaces=NS) or ''
    if not RFC3339.match(updated.strip()):
        print('entry %d updated is not RFC 3339: %s' % (n, updated))

    ident = (entry.findtext('a:id', namespaces=NS) or '').strip()
    if not IRI.match(ident):
        print('entry %d id is not an absolute IRI: %s' % (n, ident))
    elif ident in seen:
        print('entry %d repeats the id of entry %d' % (n, seen[ident]))
    else:
        seen[ident] = n

    has_author = (len(entry.findall('a:author', NS)) > 0
                  or len(entry.findall('a:source/a:author', NS)) > 0
                  or feed_has_author)
    if not has_author:
        print('entry %d has no resolvable author' % n)

    alternates = [l for l in entry.findall('a:link', NS)
                  if l.get('rel', 'alternate') == 'alternate']
    if entry.find('a:content', NS) is None and not alternates:
        print('entry %d has neither <content> nor a <link rel="alternate">' % n)

# Producing a correct timestamp. Both spellings below are valid RFC 3339.
print(datetime.now(timezone.utc).isoformat(timespec='seconds'))          # ...+00:00
print(datetime.now(timezone.utc).strftime('%Y-%m-%dT%H:%M:%SZ'))         # ...Z
import javax.xml.XMLConstants;
import javax.xml.parsers.DocumentBuilderFactory;
import java.io.ByteArrayInputStream;
import java.net.URI;
import java.net.http.HttpClient;
import java.net.http.HttpRequest;
import java.net.http.HttpResponse;
import java.time.Instant;
import java.time.OffsetDateTime;
import java.time.format.DateTimeFormatter;
import java.time.format.DateTimeParseException;
import java.util.HashMap;
import java.util.Map;
import org.w3c.dom.Document;
import org.w3c.dom.Element;
import org.w3c.dom.NodeList;

final String ATOM = "http://www.w3.org/2005/Atom";

byte[] bytes = HttpClient.newHttpClient()
    .send(HttpRequest.newBuilder(URI.create("https://example.com/atom.xml")).build(),
          HttpResponse.BodyHandlers.ofByteArray())
    .body();

DocumentBuilderFactory f = DocumentBuilderFactory.newInstance();
f.setNamespaceAware(true);   // without this, every Atom lookup below finds nothing
f.setFeature(XMLConstants.FEATURE_SECURE_PROCESSING, true);
f.setFeature("http://apache.org/xml/features/disallow-doctype-decl", true);
f.setXIncludeAware(false);

Document doc = f.newDocumentBuilder().parse(new ByteArrayInputStream(bytes));
Element feed = doc.getDocumentElement();

boolean feedHasAuthor = childCount(feed, ATOM, "author") > 0;
Map<String, Integer> seen = new HashMap<>();

NodeList entries = feed.getElementsByTagNameNS(ATOM, "entry");
for (int i = 0; i < entries.getLength(); i++) {
    Element entry = (Element) entries.item(i);
    int n = i + 1;

    String updated = firstText(entry, ATOM, "updated");
    // java.time matches literals case-sensitively, but check explicitly so the
    // lowercase case gets its own message: RFC 4287 requires "T" and "Z".
    if (updated.indexOf('t') >= 0 || updated.indexOf('z') >= 0) {
        System.err.println("entry " + n + " updated uses a lowercase T or Z: " + updated);
    } else {
        try {
            OffsetDateTime.parse(updated);
        } catch (DateTimeParseException e) {
            System.err.println("entry " + n + " updated is not RFC 3339: " + updated);
        }
    }

    String id = firstText(entry, ATOM, "id");
    if (!id.matches("[A-Za-z][A-Za-z0-9+.\\-]*:.+")) {
        System.err.println("entry " + n + " id is not an absolute IRI: " + id);
    } else if (seen.containsKey(id)) {
        System.err.println("entry " + n + " repeats the id of entry " + seen.get(id));
    } else {
        seen.put(id, n);
    }

    // entry/author, else entry/source/author, else feed/author.
    NodeList sources = entry.getElementsByTagNameNS(ATOM, "source");
    boolean sourceAuthor = sources.getLength() > 0
        && childCount((Element) sources.item(0), ATOM, "author") > 0;
    if (childCount(entry, ATOM, "author") == 0 && !sourceAuthor && !feedHasAuthor) {
        System.err.println("entry " + n + " has no resolvable author");
    }
}

// Producing a correct timestamp:
System.out.println(DateTimeFormatter.ISO_INSTANT.format(Instant.now()));
using System.Globalization;
using System.Xml;
using System.Xml.Linq;

XNamespace atom = "http://www.w3.org/2005/Atom";

// RFC 4287 section 3.3: uppercase T, and uppercase Z when there is no offset.
// The literals are quoted so they are matched exactly rather than as specifiers.
string[] rfc3339 =
{
    "yyyy-MM-dd'T'HH:mm:ssK",
    "yyyy-MM-dd'T'HH:mm:ss.FFFFFFFK",
};

var settings = new XmlReaderSettings
{
    DtdProcessing = DtdProcessing.Prohibit,  // external entities are never fetched
    XmlResolver = null,
};

var bytes = await new HttpClient().GetByteArrayAsync("https://example.com/atom.xml");
using var stream = new MemoryStream(bytes);
using var reader = XmlReader.Create(stream, settings);
var feed = XDocument.Load(reader).Root!;

foreach (var required in new[] { "id", "title", "updated" })
{
    if (feed.Elements(atom + required).Count() != 1)
        Console.Error.WriteLine("feed must have exactly one <" + required + ">");
}

bool feedHasAuthor = feed.Elements(atom + "author").Any();
var seen = new Dictionary<string, int>(StringComparer.Ordinal);
int n = 0;

foreach (var entry in feed.Elements(atom + "entry"))
{
    n++;

    var updated = ((string?)entry.Element(atom + "updated") ?? string.Empty).Trim();
    if (!DateTimeOffset.TryParseExact(updated, rfc3339, CultureInfo.InvariantCulture,
                                      DateTimeStyles.None, out _))
    {
        Console.Error.WriteLine("entry " + n + " updated is not RFC 3339: " + updated);
    }

    var id = ((string?)entry.Element(atom + "id") ?? string.Empty).Trim();
    if (!Uri.IsWellFormedUriString(id, UriKind.Absolute))
        Console.Error.WriteLine("entry " + n + " id is not an absolute IRI: " + id);
    else if (seen.TryGetValue(id, out int first))
        Console.Error.WriteLine("entry " + n + " repeats the id of entry " + first);
    else
        seen[id] = n;

    // entry/author, else entry/source/author, else feed/author.
    bool hasAuthor = entry.Elements(atom + "author").Any()
        || entry.Elements(atom + "source").Elements(atom + "author").Any()
        || feedHasAuthor;
    if (!hasAuthor) Console.Error.WriteLine("entry " + n + " has no resolvable author");

    bool hasAlternate = entry.Elements(atom + "link")
        .Any(l => ((string?)l.Attribute("rel") ?? "alternate") == "alternate");
    if (entry.Element(atom + "content") is null && !hasAlternate)
        Console.Error.WriteLine("entry " + n + " has no content and no alternate link");
}

// Producing a correct timestamp: "o" on a UTC DateTime ends in Z.
Console.WriteLine(DateTime.UtcNow.ToString("o", CultureInfo.InvariantCulture));
# RFC 4287 Appendix B carries a RELAX NG Compact schema for Atom. trang
# converts it to the XML syntax xmllint understands.
trang atom.rnc atom.rng
xmllint --noout --nonet --relaxng atom.rng feed.xml

# That schema is informative and cannot express every rule in the RFC. Author
# inheritance and id uniqueness are two of the rules it cannot state, so check
# them separately. Entries carrying no author of their own:
xmlstarlet sel -N a=http://www.w3.org/2005/Atom \
  -t -v 'count(//a:entry[not(a:author) and not(a:source/a:author)])' -n feed.xml
# If that is not zero, the <feed> element itself needs an <author>.

# Duplicate entry ids. Any output at all is a bug:
xmlstarlet sel -N a=http://www.w3.org/2005/Atom \
  -t -m '//a:entry/a:id' -v . -n feed.xml | sort | uniq -d

# Timestamps with a lowercase t or z, which RFC 4287 section 3.3 forbids:
grep -n -E '<(updated|published)>[^<]*[tz]' feed.xml

Lo schema RELAX NG della RFC 4287 vale la pena di essere eseguito, ma conoscine i limiti. È informativo e non normativo, e una grammatica non può esprimere vincoli che attraversano più elementi: l’ereditarietà di author, l’unicità degli id e la regola del summary vanno tutte verificate in codice.

Domande frequenti

Qual è la differenza fra RSS e Atom?

Atom è il formato successivo e il più severo. È uno standard IETF, la RFC 4287, pubblicata nel dicembre 2005; RSS 2.0 è stato congelato nel 2002 ed è mantenuto come documento di specifica più che come standard.

Le differenze che contano riguardano l’ambiguità. In RSS nessuno può dire se una <description> contenga testo o HTML; in Atom ogni costruzione di testo dichiara il proprio tipo. RSS rende facoltativa l’identità dell’item e ammette più scritture di data; Atom pretende un id permanente e un unico formato di data con fuso.

Perché il validatore dice che la mia entrata non ha autore se il feed ce l’ha?

Non dovrebbe, e se lo fa l’elemento author probabilmente non si trova dove pensi. La regola implementata qui è la catena della RFC 4287: il <author> proprio dell’entrata, oppure uno dentro il suo <atom:source>, oppure uno sull’elemento <feed>.

La causa abituale della sorpresa è il namespace. Un <author> che non stia in http://www.w3.org/2005/Atom, per esempio un dc:creator di Dublin Core, non è un author Atom e non soddisfa la regola. L’altra è la collocazione: un <author> dentro la prima <entry> copre soltanto quell’entrata.

Perché 2026-03-14 non è una data Atom valida?

Perché la sezione 3.3 della RFC 4287 pretende una data-ora RFC 3339 completa, non una data. I timestamp Atom hanno bisogno dell’ora e di un fuso: 2026-03-14T09:30:00Z, oppure 2026-03-14T09:30:00+05:30 per uno scostamento.

Altri due dettagli sono MUST espliciti. Il separatore dev’essere una T maiuscola, e senza scostamento numerico il fuso dev’essere una Z maiuscola. Una t o una z minuscola viene analizzata senza storie da qualsiasi libreria di date e resta comunque invalida: per questo ha un messaggio di errore dedicato.

Che cosa dovrei usare come id?

Qualcosa di globalmente univoco che non avrai mai bisogno di cambiare. La specifica è categorica: un id non deve cambiare quando l’entrata viene spostata o ripubblicata, perché i lettori confrontano gli id carattere per carattere per decidere che cosa è nuovo.

Lo schema di tag URI, la RFC 4151, è stato progettato per questo. tag:example.com,2026:post-4192 è univoco, non si pretende che si risolva, e non lega l’entrata al posto in cui vive oggi. Una URL di pagina funziona ma ti impegna: aggiungere una barra finale o passare a https cambia l’identità di ogni entrata, e gli abbonati ricevono di nuovo l’archivio.

Il mio feed viene inviato da qualche parte quando lo convalido qui?

No. Il parser e tutte le regole della RFC 4287 sono JavaScript che gira in questa scheda. Non c’è alcun componente lato server né analytics con accesso all’editor. Apri il pannello di rete e convalida qualcosa: le risorse della pagina si caricano una volta, e poi più nulla.

I feed che la gente incolla in un validatore sono proprio quelli che ancora non funzionano: un sito non lanciato, un feed i cui link portano token di accesso, un feed cliente sotto accordo di riservatezza. Il servizio del W3C è lato server, quindi tutto ciò che si controlla lì viene trasmesso. Qui Scarica va dal tuo browser direttamente all’host che hai indicato.

Può controllare un feed Atom 0.3?

No, e lo dice chiaramente. Atom 0.3 usava il namespace http://purl.org/atom/ns#, e un feed che lo dichiari viene segnalato con un errore che nomina la versione, insieme al namespace 1.0 che ti serve al suo posto.

Atom 0.3 era una bozza pre-standard che la RFC 4287 ha sostituito nel 2005, e per strada sono cambiati nomi di elementi: 0.3 ha <modified> e <issued> dove 1.0 ha <updated> e <published>, e <tagline> è diventato <subtitle>. Cambiare solo il namespace produce un documento invalido in un modo nuovo, perciò il messaggio invita a rivedere anche i nomi degli elementi.

Strumenti correlati

Approfondimenti