Atom 피드 검사기
author 상속까지 포함해 RFC 4287을 제대로 확인합니다.
모든 처리는 이 탭 안에서 이루어집니다. 붙여넣은 내용은 업로드되거나 기록되거나 전송되지 않습니다. 네트워크 패널을 열어 확인하세요.
위에 Atom 피드를 붙여넣거나 URL로 가져오면 RFC 4287에 비추어 검사합니다. 모든 지적은 근거가 된 절을 함께 밝힙니다. 루트가 <feed>가 아니라 <entry>인 단독 Atom 엔트리 문서도 알아보고 엔트리 규칙으로 검사합니다.
Atom은 RSS보다 엄격하며, 바로 그 점이 존재 이유입니다. RSS가 해석에 맡긴 것 대부분을 Atom은 정해 두었습니다. 피드에도 각 엔트리에도 id, title, updated가 정확히 하나씩. 날짜 형식은 하나뿐이고 변형이 없습니다. 모든 텍스트 구성 요소가 자신이 텍스트인지 HTML인지 XHTML인지 스스로 밝힙니다.
여기에는 대부분의 도구가 건너뛰는 규칙 하나가 제대로 구현되어 있습니다. author 상속입니다. 엔트리에 author를 담은 <atom:source>가 있거나 피드에 author가 있다면, 엔트리 자체에 author가 없어도 됩니다. entry/author만 보면 멀쩡한 피드에 거짓 오류가 한 페이지 가득 뜨고, 아무것도 보지 않으면 진짜 누락을 놓칩니다.
Atom이 생겨난 이유, 그리고 RSS와 다른 점
RSS 2.0은 2002년에 동결되었고, 그 모호함도 함께 굳었습니다. 가장 큰 것이 <description>이었습니다. 명세가 그것이 평문인지 HTML인지 끝내 말하지 않아, 리더마다 추측했습니다. 항목의 동일성은 선택이었습니다. 날짜는 RFC 822를 따랐는데, 두 자리 연도와 시간대 없는 표기까지 허용합니다.
Atom은 그에 대한 답으로 IETF에서 작성되어 2005년 12월 RFC 4287로 공개되었습니다. 텍스트는 type을 명시합니다. 모든 피드와 엔트리에 필수이며 영구적인 id가 있습니다. 날짜는 시간대가 필수인 RFC 3339입니다. 모든 것이 하나의 이름공간 안에 있어, 확장이 핵심 어휘와 충돌할 수 없습니다.
- RSS: description이 텍스트인지 HTML인지 아무도 모른다. Atom: type이 선언된다.
- RSS: guid는 선택이고 자주 불안정하다. Atom: id는 필수이며 절대 IRI이고 영구적이다.
- RSS: RFC 822 날짜, 시간대는 선택. Atom: RFC 3339, 시간대 필수.
피드와 엔트리에 반드시 있어야 하는 것
피드에 <id>, <title>, <updated>가 정확히 하나씩, 각 엔트리에도 같은 셋이 하나씩. 하나 이상이 아니라 정확히 하나이므로 중복도 오류입니다. 가장 흔한 실패는 제목과 엔트리는 있는데 id가 없는 피드입니다.
<content>가 없는 엔트리는 <link rel="alternate">를 적어도 하나 가져야 합니다. 독자에게 그 대상 자체를 주든지, 거기에 이르는 길을 주든지 해야 하기 때문입니다. <content>에 src 속성이 있으면 그 요소는 비어 있어야 하고 엔트리에는 <summary>가 필요한데, base64 내용에도 똑같이 적용됩니다.
모든 <link>에는 href가 필요하고, type과 hreflang이 같은 rel="alternate" 링크를 둘 가질 수는 없습니다. 리더가 둘 중 무엇을 고를지 알 수 없기 때문입니다. rel이 없으면 alternate로 봅니다. 피드에는 <link rel="self">도 있어야 하는데, RFC 4287이 이를 SHOULD로 적었으므로 표시가 붙은 경고로 다룹니다.
author 상속 규칙
RFC 4287은 모든 엔트리에 대해 author를 해석해 낼 수 있기를 요구하지만, 그 요소가 엔트리에 있어야 한다고는 하지 않습니다. 규칙은 되돌아가는 사슬입니다. 엔트리 자신의 <author>, 그 <atom:source> 안의 <author>, 또는 <feed>의 <author>. source 경우는 다른 곳에서 모은 엔트리를 다시 배포하는 애그리게이터를 위해 있습니다.
세 단계를 모두 구현한 무료 검사기는 거의 없습니다. entry/author만 보는 것들은 멀쩡한 단독 저자 블로그 피드, 즉 가장 흔한 형태의 Atom 피드에서 모든 엔트리에 오류를 냅니다. 이 도구는 사슬을 따라가며 셋이 모두 실패할 때만 오류로 알립니다.
단독 Atom 엔트리 문서에는 상속할 피드가 없으므로 자기 author를 가져야 합니다. 사람 구성 요소에는 모두 <name>이 필요합니다. 이메일 주소만 있는 <author>는 3.2.1절에 비추어 오류입니다.
날짜와 id를, RFC가 적은 그대로
Atom의 타임스탬프는 RFC 3339이며, 3.3절이 두 가지를 더 얹습니다. 구분자는 대문자 T여야 하고, 숫자 오프셋이 없다면 시간대는 대문자 Z여야 합니다. 그래서 2026-03-14T09:30:00Z는 유효하고 2026-03-14t09:30:00z는 무효인데, 어떤 날짜 라이브러리든 후자를 파싱해 버립니다. 소문자에는 전용 메시지가 있습니다. 날짜만 적은 것도 무효입니다.
id는 절대 IRI여야 하므로 스킴이 필요합니다. https:, tag:, urn:은 모두 되지만 상대 경로나 맨 문자열은 안 됩니다. 엔트리 사이에 id가 중복되면 오류입니다. 리더가 id로 중복을 제거하기 때문입니다.
대소문자만, 끝의 슬래시만, 기본 포트만 다른 id는 경고가 됩니다. id는 한 글자씩 비교되므로 /p/1과 /p/1/은 서로 다른 두 엔트리입니다. tag 스킴(RFC 4151)은 이를 피하게 해 주며, tag:example.com,2026:post-4192는 도메인 이전에도 살아남습니다.
코드로 Atom 만들고 검사하기
자동화할 값어치가 있는 네 가지 규칙은 모두 스키마로는 표현할 수 없는 것들입니다. RFC 3339의 대소문자, id의 유일성, author 상속, 그리고 content냐 alternate 링크냐의 양자택일. 각 예제는 안전한 파싱 형태로 이것들을 검사합니다.
// Node 18 or later. npm i @xmldom/xmldom
import { DOMParser } from '@xmldom/xmldom';
const ATOM = 'http://www.w3.org/2005/Atom';
// RFC 3339 as RFC 4287 section 3.3 requires it: uppercase T, uppercase Z.
const RFC3339 = /^\d{4}-\d{2}-\d{2}T\d{2}:\d{2}:\d{2}(\.\d+)?(Z|[+-]\d{2}:\d{2})$/;
const xml = await (await fetch('https://example.com/atom.xml')).text();
const doc = new DOMParser().parseFromString(xml, 'text/xml');
const feed = doc.documentElement;
const kids = (el, name) => Array.from(el.getElementsByTagNameNS(ATOM, name))
.filter((n) => n.parentNode === el);
for (const required of ['id', 'title', 'updated']) {
if (kids(feed, required).length !== 1) {
console.error('feed must have exactly one <' + required + '>');
}
}
const feedHasAuthor = kids(feed, 'author').length > 0;
const seen = new Map();
for (const [i, entry] of kids(feed, 'entry').entries()) {
const n = i + 1;
const updated = kids(entry, 'updated')[0];
if (updated && !RFC3339.test(updated.textContent.trim())) {
console.error('entry ' + n + ' updated is not RFC 3339: ' + updated.textContent.trim());
}
const id = kids(entry, 'id')[0];
const value = id ? id.textContent.trim() : '';
if (!/^[A-Za-z][A-Za-z0-9+.-]*:/.test(value)) {
console.error('entry ' + n + ' id is not an absolute IRI: ' + value);
} else if (seen.has(value)) {
console.error('entry ' + n + ' repeats the id of entry ' + seen.get(value));
} else {
seen.set(value, n);
}
// The inheritance chain: entry/author, else entry/source/author, else feed/author.
const source = kids(entry, 'source')[0];
const hasAuthor = kids(entry, 'author').length > 0
|| (source ? kids(source, 'author').length > 0 : false)
|| feedHasAuthor;
if (!hasAuthor) console.error('entry ' + n + ' has no resolvable author');
const hasAlternate = kids(entry, 'link')
.some((l) => (l.getAttribute('rel') || 'alternate') === 'alternate');
if (kids(entry, 'content').length === 0 && !hasAlternate) {
console.error('entry ' + n + ' has neither <content> nor a <link rel="alternate">');
}
}
// Producing a correct timestamp:
console.log(new Date().toISOString()); // 2026-03-14T09:30:00.000Z# pip install lxml requests
import re
import sys
import requests
from datetime import datetime, timezone
from lxml import etree
ATOM = 'http://www.w3.org/2005/Atom'
NS = {'a': ATOM}
# datetime.fromisoformat is too permissive for this: it accepts a lowercase
# separator, which RFC 4287 section 3.3 forbids. Check the shape directly.
RFC3339 = re.compile(r'^\d{4}-\d{2}-\d{2}T\d{2}:\d{2}:\d{2}(\.\d+)?(Z|[+-]\d{2}:\d{2})$')
IRI = re.compile(r'^[A-Za-z][A-Za-z0-9+.\-]*:')
raw = requests.get('https://example.com/atom.xml', timeout=30).content
parser = etree.XMLParser(resolve_entities=False, no_network=True, load_dtd=False)
feed = etree.fromstring(raw, parser)
for required in ('id', 'title', 'updated'):
if len(feed.findall('a:%s' % required, NS)) != 1:
print('feed must have exactly one <%s>' % required)
feed_has_author = len(feed.findall('a:author', NS)) > 0
seen = {}
for n, entry in enumerate(feed.findall('a:entry', NS), start=1):
updated = entry.findtext('a:updated', namespaces=NS) or ''
if not RFC3339.match(updated.strip()):
print('entry %d updated is not RFC 3339: %s' % (n, updated))
ident = (entry.findtext('a:id', namespaces=NS) or '').strip()
if not IRI.match(ident):
print('entry %d id is not an absolute IRI: %s' % (n, ident))
elif ident in seen:
print('entry %d repeats the id of entry %d' % (n, seen[ident]))
else:
seen[ident] = n
has_author = (len(entry.findall('a:author', NS)) > 0
or len(entry.findall('a:source/a:author', NS)) > 0
or feed_has_author)
if not has_author:
print('entry %d has no resolvable author' % n)
alternates = [l for l in entry.findall('a:link', NS)
if l.get('rel', 'alternate') == 'alternate']
if entry.find('a:content', NS) is None and not alternates:
print('entry %d has neither <content> nor a <link rel="alternate">' % n)
# Producing a correct timestamp. Both spellings below are valid RFC 3339.
print(datetime.now(timezone.utc).isoformat(timespec='seconds')) # ...+00:00
print(datetime.now(timezone.utc).strftime('%Y-%m-%dT%H:%M:%SZ')) # ...Zimport javax.xml.XMLConstants;
import javax.xml.parsers.DocumentBuilderFactory;
import java.io.ByteArrayInputStream;
import java.net.URI;
import java.net.http.HttpClient;
import java.net.http.HttpRequest;
import java.net.http.HttpResponse;
import java.time.Instant;
import java.time.OffsetDateTime;
import java.time.format.DateTimeFormatter;
import java.time.format.DateTimeParseException;
import java.util.HashMap;
import java.util.Map;
import org.w3c.dom.Document;
import org.w3c.dom.Element;
import org.w3c.dom.NodeList;
final String ATOM = "http://www.w3.org/2005/Atom";
byte[] bytes = HttpClient.newHttpClient()
.send(HttpRequest.newBuilder(URI.create("https://example.com/atom.xml")).build(),
HttpResponse.BodyHandlers.ofByteArray())
.body();
DocumentBuilderFactory f = DocumentBuilderFactory.newInstance();
f.setNamespaceAware(true); // without this, every Atom lookup below finds nothing
f.setFeature(XMLConstants.FEATURE_SECURE_PROCESSING, true);
f.setFeature("http://apache.org/xml/features/disallow-doctype-decl", true);
f.setXIncludeAware(false);
Document doc = f.newDocumentBuilder().parse(new ByteArrayInputStream(bytes));
Element feed = doc.getDocumentElement();
boolean feedHasAuthor = childCount(feed, ATOM, "author") > 0;
Map<String, Integer> seen = new HashMap<>();
NodeList entries = feed.getElementsByTagNameNS(ATOM, "entry");
for (int i = 0; i < entries.getLength(); i++) {
Element entry = (Element) entries.item(i);
int n = i + 1;
String updated = firstText(entry, ATOM, "updated");
// java.time matches literals case-sensitively, but check explicitly so the
// lowercase case gets its own message: RFC 4287 requires "T" and "Z".
if (updated.indexOf('t') >= 0 || updated.indexOf('z') >= 0) {
System.err.println("entry " + n + " updated uses a lowercase T or Z: " + updated);
} else {
try {
OffsetDateTime.parse(updated);
} catch (DateTimeParseException e) {
System.err.println("entry " + n + " updated is not RFC 3339: " + updated);
}
}
String id = firstText(entry, ATOM, "id");
if (!id.matches("[A-Za-z][A-Za-z0-9+.\\-]*:.+")) {
System.err.println("entry " + n + " id is not an absolute IRI: " + id);
} else if (seen.containsKey(id)) {
System.err.println("entry " + n + " repeats the id of entry " + seen.get(id));
} else {
seen.put(id, n);
}
// entry/author, else entry/source/author, else feed/author.
NodeList sources = entry.getElementsByTagNameNS(ATOM, "source");
boolean sourceAuthor = sources.getLength() > 0
&& childCount((Element) sources.item(0), ATOM, "author") > 0;
if (childCount(entry, ATOM, "author") == 0 && !sourceAuthor && !feedHasAuthor) {
System.err.println("entry " + n + " has no resolvable author");
}
}
// Producing a correct timestamp:
System.out.println(DateTimeFormatter.ISO_INSTANT.format(Instant.now()));using System.Globalization;
using System.Xml;
using System.Xml.Linq;
XNamespace atom = "http://www.w3.org/2005/Atom";
// RFC 4287 section 3.3: uppercase T, and uppercase Z when there is no offset.
// The literals are quoted so they are matched exactly rather than as specifiers.
string[] rfc3339 =
{
"yyyy-MM-dd'T'HH:mm:ssK",
"yyyy-MM-dd'T'HH:mm:ss.FFFFFFFK",
};
var settings = new XmlReaderSettings
{
DtdProcessing = DtdProcessing.Prohibit, // external entities are never fetched
XmlResolver = null,
};
var bytes = await new HttpClient().GetByteArrayAsync("https://example.com/atom.xml");
using var stream = new MemoryStream(bytes);
using var reader = XmlReader.Create(stream, settings);
var feed = XDocument.Load(reader).Root!;
foreach (var required in new[] { "id", "title", "updated" })
{
if (feed.Elements(atom + required).Count() != 1)
Console.Error.WriteLine("feed must have exactly one <" + required + ">");
}
bool feedHasAuthor = feed.Elements(atom + "author").Any();
var seen = new Dictionary<string, int>(StringComparer.Ordinal);
int n = 0;
foreach (var entry in feed.Elements(atom + "entry"))
{
n++;
var updated = ((string?)entry.Element(atom + "updated") ?? string.Empty).Trim();
if (!DateTimeOffset.TryParseExact(updated, rfc3339, CultureInfo.InvariantCulture,
DateTimeStyles.None, out _))
{
Console.Error.WriteLine("entry " + n + " updated is not RFC 3339: " + updated);
}
var id = ((string?)entry.Element(atom + "id") ?? string.Empty).Trim();
if (!Uri.IsWellFormedUriString(id, UriKind.Absolute))
Console.Error.WriteLine("entry " + n + " id is not an absolute IRI: " + id);
else if (seen.TryGetValue(id, out int first))
Console.Error.WriteLine("entry " + n + " repeats the id of entry " + first);
else
seen[id] = n;
// entry/author, else entry/source/author, else feed/author.
bool hasAuthor = entry.Elements(atom + "author").Any()
|| entry.Elements(atom + "source").Elements(atom + "author").Any()
|| feedHasAuthor;
if (!hasAuthor) Console.Error.WriteLine("entry " + n + " has no resolvable author");
bool hasAlternate = entry.Elements(atom + "link")
.Any(l => ((string?)l.Attribute("rel") ?? "alternate") == "alternate");
if (entry.Element(atom + "content") is null && !hasAlternate)
Console.Error.WriteLine("entry " + n + " has no content and no alternate link");
}
// Producing a correct timestamp: "o" on a UTC DateTime ends in Z.
Console.WriteLine(DateTime.UtcNow.ToString("o", CultureInfo.InvariantCulture));# RFC 4287 Appendix B carries a RELAX NG Compact schema for Atom. trang
# converts it to the XML syntax xmllint understands.
trang atom.rnc atom.rng
xmllint --noout --nonet --relaxng atom.rng feed.xml
# That schema is informative and cannot express every rule in the RFC. Author
# inheritance and id uniqueness are two of the rules it cannot state, so check
# them separately. Entries carrying no author of their own:
xmlstarlet sel -N a=http://www.w3.org/2005/Atom \
-t -v 'count(//a:entry[not(a:author) and not(a:source/a:author)])' -n feed.xml
# If that is not zero, the <feed> element itself needs an <author>.
# Duplicate entry ids. Any output at all is a bug:
xmlstarlet sel -N a=http://www.w3.org/2005/Atom \
-t -m '//a:entry/a:id' -v . -n feed.xml | sort | uniq -d
# Timestamps with a lowercase t or z, which RFC 4287 section 3.3 forbids:
grep -n -E '<(updated|published)>[^<]*[tz]' feed.xmlRFC 4287의 RELAX NG 스키마는 돌려 볼 값어치가 있지만 한계를 알아 두세요. 그것은 규범이 아니라 참고이며, 문법은 여러 요소에 걸친 제약을 표현하지 못합니다. author 상속, id의 유일성, summary 규칙은 모두 코드로 확인해야 합니다.
자주 묻는 질문
RSS와 Atom의 차이는 무엇인가요?
Atom이 나중에 나왔고 더 엄격합니다. IETF 표준 RFC 4287로 2005년 12월에 공개되었습니다. RSS 2.0은 2002년에 동결되었고 표준이라기보다 명세 문서로 관리됩니다.
실제로 문제가 되는 차이는 모호함에 관한 것입니다. RSS에서는 <description>이 텍스트인지 HTML인지 누구도 단정할 수 없지만, Atom에서는 모든 텍스트 구성 요소가 자기 타입을 밝힙니다. RSS는 항목의 동일성을 선택으로 두고 여러 날짜 표기를 허용하지만, Atom은 영구적인 id와 시간대가 붙은 단일 날짜 형식을 요구합니다.
피드에 author가 있는데도 엔트리에 author가 없다고 하는 이유는 무엇인가요?
그렇게 말해서는 안 되며, 그렇게 나온다면 author 요소가 생각하는 자리에 없을 가능성이 큽니다. 여기 구현된 규칙은 RFC 4287의 사슬입니다. 엔트리 자신의 <author>, 그 <atom:source> 안의 <author>, 또는 <feed> 요소의 <author>.
놀라움의 흔한 원인은 이름공간입니다. http://www.w3.org/2005/Atom에 있지 않은 <author>, 이를테면 더블린 코어의 dc:creator는 Atom의 author가 아니어서 규칙을 만족시키지 못합니다. 다른 하나는 위치입니다. 첫 <entry> 안에 있는 <author>는 그 엔트리 하나만 덮습니다.
왜 2026-03-14는 유효한 Atom 날짜가 아닌가요?
RFC 4287 3.3절이 날짜가 아니라 완전한 RFC 3339 날짜·시각을 요구하기 때문입니다. Atom의 타임스탬프에는 시각과 시간대가 필요합니다. 2026-03-14T09:30:00Z처럼, 또는 오프셋이라면 2026-03-14T09:30:00+05:30처럼.
두 가지가 더 명시적인 MUST입니다. 구분자는 대문자 T여야 하고, 숫자 오프셋이 없을 때 시간대는 대문자 Z여야 합니다. 소문자 t나 z는 어느 날짜 라이브러리에서도 무리 없이 파싱되지만 여전히 무효이며, 그래서 전용 오류 메시지를 둡니다.
id로는 무엇을 쓰면 되나요?
다시는 바꿀 일이 없는, 전역적으로 유일한 것입니다. 명세는 단호합니다. 엔트리가 옮겨지거나 다시 배포되어도 id는 바뀌어서는 안 됩니다. 리더는 무엇이 새것인지 정하려고 id를 한 글자씩 비교하기 때문입니다.
tag URI 스킴(RFC 4151)이 바로 이를 위해 설계되었습니다. tag:example.com,2026:post-4192는 유일하고, 해석될 것을 기대하지 않으며, 엔트리를 지금 있는 자리에 묶지 않습니다. 페이지 URL도 되지만 발목을 잡습니다. 끝에 슬래시를 붙이거나 https로 옮기면 모든 엔트리의 동일성이 바뀌고, 구독자는 보관함을 다시 받습니다.
여기서 검증하면 피드가 어디론가 전송되나요?
아닙니다. 파서도 RFC 4287의 모든 규칙도 이 탭에서 도는 자바스크립트입니다. 서버 쪽 구성 요소도, 편집기에 접근하는 분석 도구도 없습니다. 네트워크 패널을 열고 무언가를 검증해 보세요. 페이지 자원이 한 번 불려 오고, 그 뒤로는 아무 일도 없습니다.
사람들이 검사기에 붙여넣는 피드는 아직 제대로 돌지 않는 것들입니다. 공개 전 사이트, 링크에 접근 토큰이 담긴 피드, 비밀 유지 계약 아래 있는 고객 피드. W3C 서비스는 서버 쪽이라 거기서 검사한 것은 전송됩니다. 여기서의 가져오기는 여러분의 브라우저에서 여러분이 지정한 호스트로 곧장 갑니다.
Atom 0.3 피드도 검사할 수 있나요?
할 수 없고, 그 점을 분명히 말해 줍니다. Atom 0.3은 이름공간 http://purl.org/atom/ns#를 썼고, 그것을 선언한 피드는 버전을 짚어 주는 오류로 보고하면서 대신 필요한 1.0 이름공간도 함께 알려 줍니다.
Atom 0.3은 표준화 이전의 초안이었고 RFC 4287이 2005년에 대체했으며, 그사이 요소 이름도 바뀌었습니다. 0.3의 <modified>와 <issued>는 1.0에서 <updated>와 <published>이고, <tagline>은 <subtitle>이 되었습니다. 이름공간만 바꾸면 새로운 방식으로 무효인 문서가 되므로, 메시지는 요소 이름도 함께 살펴보라고 일러 줍니다.