RSS and Atom feed structure

Two formats do the same job and share almost no rules. RSS 2.0 is a short, informal specification maintained by the RSS Advisory Board, with no XML namespace and a date format inherited from 1982 email. Atom 1.0 is RFC 4287, an IETF standards-track document that states its requirements in capitalised MUSTs and namespaces everything.

A publisher cares about one thing: the feed works in the readers people use, and keeps working after an edit. This covers what each format requires, where they differ in ways that matter, and the four mistakes behind most broken feeds.

#The two formats, side by side

RSS 2.0 reached its current form in 2002 and has been frozen since; version 2.0.11 of 30 March 2009 adds clarifications rather than changes. Atom was written afterwards by an IETF working group, partly in response to how much of RSS was left to interpretation. Every serious reader handles both, so the choice is yours rather than your audience's.

RSS 2.0Atom 1.0
SpecificationRSS Advisory Board, informal proseRFC 4287, IETF standards track
NamespaceNone. Core elements are unqualified.http://www.w3.org/2005/Atom on everything
Rootrss, with version="2.0"feed, or entry for a standalone document
Required at the topchannel with title, link, descriptionexactly one each of id, title, updated
Item containeritementry
Required per itemtitle or description, at least oneid, title, updated, plus content or a rel="alternate" link
DatesRFC 822: Thu, 04 Oct 2007 23:59:45 +0000RFC 3339: 2026-08-14T09:30:00Z
Item identityguid, optional, assumed to be a URLid, mandatory, a permanent absolute IRI
Body markupAmbiguous. Escaped HTML by convention.Declared: type="text", "html" or "xhtml"
AuthorOptional, an email address in a stringStructured, with an inheritance rule
The differences that change what you have to write.

Atom is stricter, and the strictness is on the publisher's side: the rules it enforces are the ones that make an item identifiable and dateable. RSS is more forgiving, which means more of its failure modes are silent.

#What RSS 2.0 actually requires

The mandatory surface is small. The root is rss with a version attribute reading 2.0, and subordinate to it is a single channel. The channel must carry title, link and description; those three are the only required elements in the format. Everything on an item is optional except that, in the specification's words, "at least one of title or description must be present".

<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>Example Feed</title>
    <link>https://example.com/</link>
    <description>Notes on kettles and descaling.</description>
    <language>en-gb</language>
    <lastBuildDate>Thu, 14 Aug 2026 09:30:00 +0000</lastBuildDate>
    <atom:link href="https://example.com/rss.xml" rel="self"
               type="application/rss+xml"/>
    <item>
      <title>Descaling without vinegar</title>
      <link>https://example.com/p/1</link>
      <description><![CDATA[<p>Citric acid works and does not linger.</p>]]></description>
      <pubDate>Mon, 11 Aug 2026 07:00:00 +0000</pubDate>
      <guid isPermaLink="false">tag:example.com,2026:post-1</guid>
      <enclosure url="https://example.com/audio/1.mp3"
                 length="12216320" type="audio/mpeg"/>
    </item>
  </channel>
</rss>
A valid RSS 2.0 feed, including the conventions the spec does not require.

pubDate and lastBuildDate must be RFC 822 date-times, with RSS 2.0's one amendment that "the year may be expressed with two characters or four characters (four preferred)". Two-digit years are therefore legal and still break aggregators, so the validator here reports them as a warning. RFC 822 does not require the day name to agree with the date, but a feed saying Mon for a Tuesday is a hand-rolled formatter with a bug in it, so that is reported as an error.

A detail that trips up validators as often as feeds: because RSS core elements are in no namespace, an atom:link is not an RSS link even though the local name matches. Tools matching on the local name alone report a duplicate link on nearly every real feed, since almost all carry a self link. The RSS validator here matches only unprefixed elements for core RSS rules.

#guid, enclosure, and the self link

guid is optional in the specification and mandatory in practice. It is the string a reader uses to decide whether it has seen an item before. Without one, aggregators match on the link or the title, so correcting a headline republishes the post as new, at the top of every subscriber's list.

For podcasts the payload is the enclosure element, and all three of its attributes are required: url, length and type. length is the size of the file in bytes, not its duration, and a length of 0 almost always means the generator could not stat the file. The Best Practices Profile adds that an item should carry no more than one enclosure, which is not in the specification but is how clients behave.

The atom:link with rel="self" is not in the RSS specification either. It comes from the Best Practices Profile, and it exists because a feed that has been copied, proxied or cached has no other way to state its canonical address. Declare xmlns:atom on the rss element, then put the link inside channel with the feed's absolute URL.

#What RFC 4287 requires of Atom

Everything in Atom lives in the namespace http://www.w3.org/2005/Atom, fixed by RFC 4287 section 1.2. The pre-standard Atom 0.3 namespace, http://purl.org/atom/ns#, is a different vocabulary with a different element set, and a document using it is not Atom 1.0 in any sense a strict reader accepts.

Section 4.1.1 gives the feed its requirements as plain MUSTs: exactly one atom:id, one atom:title, one atom:updated. Exactly one is stricter than at least one, so two titles is as much an error as none. Section 4.1.2 imposes the same three on every entry.

The author rule is the one people get wrong, because it is conditional. Section 4.1.1 requires the feed to contain one or more atom:author elements "unless all of the atom:feed element's child atom:entry elements contain at least one atom:author element". Section 4.1.2 states the mirror image for entries, with a middle step: the fallback chain is entry author, then entry source author, then feed author, and a document fails only when all three are empty. The validator here implements that chain, which is why it passes a feed where no entry names an author but the feed does.

Section 4.1.2 also requires that an entry with no atom:content carry at least one atom:link with rel="alternate": an entry must give the reader the body or somewhere to find it. The same section requires an atom:summary when the content is remote (a content element with a src attribute, which must then be empty) or Base64-encoded, for the same reason.

atom:id, section 4.2.6, must be an absolute IRI, must be constructed so as to be unique, and must not change when the document is relocated or republished. That last clause is the point of the element. An id built from the current URL breaks the moment you move to https or change a slug, and readers compare ids literally, so every entry then arrives again as new. A tag: URI such as tag:example.com,2026:post-1 is the usual answer: stable by construction, and it never has to resolve to anything.

Two smaller rules with sharp edges. Text constructs, section 3.1.1, take a type of exactly text, html or xhtml, and when the type is xhtml "the content of the Text construct MUST be a single XHTML div element", not a fragment and not two siblings. A link, section 4.2.7, must carry an href, must be empty, and is rel="alternate" by default; nothing may have two alternate links sharing the same type and hreflang, because a reader could not choose between them.

<?xml version="1.0" encoding="UTF-8"?>
<feed xmlns="http://www.w3.org/2005/Atom">
  <id>tag:example.com,2026:feed</id>
  <title>Example Feed</title>
  <updated>2026-08-14T09:30:00Z</updated>
  <link rel="self" type="application/atom+xml"
        href="https://example.com/atom.xml"/>
  <link rel="alternate" type="text/html" href="https://example.com/"/>
  <author><name>Jane Doe</name></author>
  <entry>
    <id>tag:example.com,2026:post-1</id>
    <title>Descaling without vinegar</title>
    <updated>2026-08-11T07:00:00Z</updated>
    <published>2026-08-11T07:00:00Z</published>
    <link rel="alternate" type="text/html" href="https://example.com/p/1"/>
    <content type="html">&lt;p&gt;Citric acid works.&lt;/p&gt;</content>
  </entry>
</feed>
The minimum Atom 1.0 feed that satisfies every MUST in section 4.1.

#The four things that actually break feeds

Across both formats, most real breakage is one of four things, and the first is more common than the other three combined.

1. An unescaped ampersand in a URL, written straight into a link, a guid or an enclosure. This is not a feed error, it is a fatal XML well-formedness error, so the parser stops and the reader gets nothing at all rather than one bad item. A feed with 200 good entries and one raw ampersand shows as empty.

Fatal. Nothing after this line is read.
<link>https://example.com/p?id=1&utm_source=feed</link>
Ampersand escaped as an entity
<link>https://example.com/p?id=1&amp;utm_source=feed</link>

2. Raw HTML inside description or content. Written straight into the element, an img tag looks like a feed element, and an unclosed one makes the document malformed. Three answers are legitimate: escape the markup so a less-than becomes an entity, wrap it in a CDATA section (which cannot itself contain the sequence that ends it), or in Atom use type="html" and escape, or type="xhtml" with a single wrapping div. Never put markup in a title: titles are plain text in both formats, and readers show the tags literally.

3. The wrong date format, usually the other format's. An ISO 8601 timestamp in an RSS pubDate and an RFC 822 string in an Atom updated are both invalid, and both come from a template shared between two feed writers. A date the reader cannot parse tends to sort the item to the epoch or to the top.

RSS pubDate given an ISO 8601 value
<pubDate>2026-08-14T09:30:00Z</pubDate>
RFC 822, four-digit year, explicit zone
<pubDate>Fri, 14 Aug 2026 09:30:00 +0000</pubDate>

4. Unstable identifiers. A guid or an atom:id that changes between builds makes every item new again. The usual causes are an id derived from the title, an id carrying a cache-busting parameter or a session token, an id that is the URL and so changed when the site moved to https, and a generator emitting the same id twice. Readers compare these strings literally, so an id differing only by a trailing slash or by case is a different item. Derive it once from a database key and treat it as permanent: if the ids you have published are unstable, changing them one last time is a single duplicate flood, and leaving them is a flood on every deploy.

#Validating a feed today

The reference implementation here has always been the W3C Feed Validator, and it is in a poor state. feedvalidator.org currently serves an expired TLS certificate, so browsers interrupt with a security warning before you reach it, and the surviving mirror still offers Atom 0.3, a draft RFC 4287 superseded in December 2005. It is still excellent software: 151 error and 108 warning message types, and the closest thing this area has to a shared vocabulary.

Not everything it reports is a specification violation, and a validator that does not say which is which is not much use. Missing self links, missing guids, two-digit years, day names that disagree with their dates and duplicate enclosures are Best Practices Profile items or folklore rather than spec rules. Every finding the validators here emit names the clause it enforces, or states that it is aggregator behaviour found in no specification.

One class of check cannot run in a browser at all. Whether rel="self" matches the address the feed is served from, whether an enclosure URL is reachable and matches its declared length, whether a host resolves: each needs a request to a third-party origin, which no client-side tool can make. Those are listed as not checked rather than quietly dropped, because a rule that could not run must never be presented as one that passed.

Common questions

Should a new feed be RSS or Atom?

If nothing pushes you either way, Atom. Its requirements are stricter in the places that matter for correctness: every entry must have a permanent id, every entry must have a real timestamp with a timezone, and the markup in a body is declared rather than guessed at. Those three rules remove most of the ways a feed misbehaves after publication.

RSS is the right answer when something downstream expects it. Podcast directories are built on RSS 2.0 plus Apple's iTunes namespace, and a podcast in Atom will not be accepted anywhere useful. Publishing both is legitimate, but then both have to stay correct, and a stale second feed is worse than no second feed.

How do I put HTML in an item body without breaking the feed?

In RSS, either escape it so that every less-than becomes an entity, or wrap it in a CDATA section. Escaping is safer for generated content, because a CDATA section cannot contain the three-character sequence that terminates it, and a body that happens to include that sequence will break the document.

In Atom the ambiguity is gone: set type="html" on the content element and escape the markup, or set type="xhtml" and put well-formed XHTML inside a single wrapping div. Either way, keep markup out of the title in both formats.

Old posts keep reappearing at the top of readers. What is wrong?

The identifier is changing between builds. Readers dedupe on guid in RSS and on atom:id in Atom, comparing the strings literally, so any change at all makes the item new. Common causes are an id derived from the title, so an edit republishes it, an id that is the page URL, so a move to https or a change of slug republishes everything, or a generator producing a fresh random id each run.

If there is no guid at all, most readers fall back to the link or the title and you get the same symptom for the same reason. Fix it by deriving the identifier from something permanent, usually a database key plus the publication year in a tag: URI, and never touching it again.

Does a valid feed mean my podcast will be accepted?

No, and the gap catches people out. RSS 2.0 validity covers the channel, the items, the dates and the enclosures. The elements podcast directories actually read live in a separate namespace, http://www.itunes.com/dtds/podcast-1.0.dtd, and are outside the specification entirely.

A feed can pass every RSS 2.0 rule and still be rejected for a missing itunes:category, artwork outside the required dimensions, or an enclosure whose declared length does not match the file. Validate the RSS layer first, because an XML error stops everything else being read, then check the directory's own requirements separately.

Sources

Try it