An RSS item is a structured XML record, and every field in it carries a different kind of information — a headline, a summary, the full article body, a canonical link, a timestamp, and zero or more images. Reading those fields correctly means knowing which element to trust for which job: title for display, description for previews, content:encoded for full-text, link for attribution, pubDate for ordering, and enclosure or media:content for images that actually render.
Most feed problems — duplicate posts, missing thumbnails, out-of-order timelines, broken attribution links — trace back to one field being used where another belongs. This guide walks through each element, what it reliably contains, and how to handle it in a production feed reader or content pipeline.

An RSS feed is a single XML document containing a channel and a list of items. Each is one discrete piece of content: a blog post, a news story, a podcast episode, a product update. RSS 2.0 defines a fixed set of core elements per item, and extension modules add everything the core spec leaves out.
That modular design is the reason feed parsing is harder than it looks. A title is always a title, but "the image" might live in , , , , or simply as the first tag buried inside the content payload. A production parser needs to check all of them, in a defined order.
The Atom format (standardized in RFC 4287{:target="_blank" rel="noopener"}) solves some of this with stricter rules, but a large share of the web still publishes RSS 2.0, so any reader worth building handles both.

| Field | Typical element | What it gives you | Common pitfall |
|---|---|---|---|
| Title | | Headline, often plain text | May contain HTML entities that need decoding |
| Description | | Summary or full content, depending on the publisher | Truncated snippets, encoded markup, or entire articles |
| Content | | Full article body in HTML | Frequently missing; not part of the RSS 2.0 core spec |
| Source URL | and | Canonical permalink and unique identifier | guid is not always a URL |
| Publication date | , , | When the item was published or last modified | Multiple conflicting formats in the same feed |
| Images | , , | Image URLs with optional dimensions and MIME types | Thumbnails are hotlink-protected or low resolution |
The title is the field you display and the field readers scan. Treat it as text, not markup: strip any tags that appear, decode HTML entities such as & and ’, and normalize whitespace.
Two quirks show up repeatedly. First, some publishers prefix every title with the site name, producing "Article Headline – Site Name" in every entry. Second, podcast and newsletter feeds sometimes leave the title empty and put everything in the description. A robust reader has a fallback: if the title is blank, derive one from the first line of the description, then from the URL slug.

In RSS 2.0 the spec calls description "the item synopsis," but in practice it is one of three things depending on the publisher:
Because of this ambiguity, never assume the description is short. Measure it. If a description routinely exceeds a few thousand characters, the publisher is treating it as full content and your truncation logic should kick in before rendering.
Note that description content is usually escaped:
arrives as <p>. Decoding it before display is mandatory, and sanitizing it afterward is not optional — feed content is untrusted input.

The content:encoded element comes from the RSS 1.0 Content Module (http://purl.org/rss/1.0/modules/content/), not from RSS 2.0 itself. When it exists, it holds the full post in HTML — headings, images, code blocks, embedded media.
Prefer content:encoded over description whenever both are present and your use case calls for full text. If it is absent, fall back to description, then to fetching the source URL directly.
Atom feeds handle this more cleanly with a element that declares its own type attribute (html, text, or xhtml), so you know exactly what you are decoding before you decode it.
Two elements overlap here, and conflating them causes real bugs.
is the item's permalink — the page a reader should open to read the original. It is what you attribute, what you canonicalize against, and what you send in a newsletter. Always resolve it against the feed's base URL if it is relative, and strip tracking parameters such as utm_source before storing it.
is a permanent identifier. It signals "this is the same item I sent you before." Critically, the isPermaLink attribute determines whether the GUID is a URL:
isPermaLink="true" (the default) — the GUID is probably a URL, but you should still not treat it as the clickable destinationisPermaLink="false" — the GUID is an opaque string, possibly a UUID or a database keyUse guid for deduplication. Use link for navigation. When guid is missing, which happens in loose Atom-to-RSS conversions, fall back to a hash of the link plus the title.
Dates are where feed parsing breaks most often. RSS 2.0 specifies RFC 822 date formatting, which looks like Tue, 14 May 2024 09:30:00 GMT. Atom uses RFC 3339 timestamps such as 2024-05-14T09:30:00Z. Dublin Core contributes dc:date in yet another shape. Some feeds include timezone offsets, some omit them entirely, and some send nonsense.
Practical handling:
pubDate and updated separate. The first is when the item appeared; the second is when it last changed. Sorting a timeline by updated reshuffles old posts to the top, which surprises readers.There is no single "image" element in RSS 2.0, so extraction follows a cascade. Check in roughly this order:
— the original podcast-era mechanism, still the most reliable when present — Media RSS, which adds width, height, and type — explicitly a small preview image, often 144px or smaller — podcast artwork, usually square and site-level rather than per-episode
inside content:encoded or description — the least reliable option, but sometimes the only oneTwo rules make this manageable. Prefer the largest available image, since downscaling looks better than upscaling; and honor media:thumbnail only as a last resort, because thumbnails frequently point to CDN-hosted assets that block hotlinking or expire.
Also watch for media:group, which bundles multiple renditions of the same asset. Pick one from the group rather than treating each rendition as a separate image.
When you ingest a new item, process the fields in this sequence to avoid rework:
guid, or a hash of link plus title, against your store. Skip anything already seen before doing any other work.link to an absolute URL and strip tracking parameters.content:encoded exists, use it. Otherwise use description. Otherwise queue the source URL for full-text extraction.pubDate and updated as separate columns.Feed content is syndicated content, and syndication carries obligations. The link element is your attribution channel: any republished excerpt should point back to it, and any AI-generated summary should cite it. Getting the source URL wrong — pointing readers to a cached copy, a redirect chain, or a tracking-laden variant — breaks the implicit agreement between publisher and aggregator.
The element, when present, names the original publication, and or dc:creator identifies the writer. Both should travel with the content wherever it goes. Atom's element is more structured, containing name, email, and URI as separate sub-elements, which is why Atom feeds tend to attribute more cleanly.
The same six fields power a surprising range of systems:
pubDate and keyword-matching the descriptioncontent:encoded as the input, link as the citation, pubDate as the freshness signallink as the targetguid preventing duplicate entries when feeds re-publishIn each case the failure modes are identical: missing images, duplicated items, and scrambled chronology. Fixing the field-level handling fixes all three.
- Treating description as a summary. Half the web treats it as the article. Measure before you truncate.
isPermaLink. It changes how you should interpret the identifier.What is the difference between description and content:encoded?
description is part of the RSS 2.0 core specification and holds a synopsis — though many publishers put the entire article in it. content:encoded comes from the RSS 1.0 Content Module and, when present, is intended to carry the complete post body. In practice: prefer content:encoded for full text, fall back to description for previews, and treat both as untrusted HTML that must be decoded and sanitized.
Is content:encoded part of standard RSS 2.0?
No. It belongs to the RSS 1.0 Content Module, namespaced at http://purl.org/rss/1.0/modules/content/. Most major feed readers support it, but some feeds omit it entirely, which is why your parser needs a description fallback.
Should I use guid or link for deduplication?
Use guid when it exists, since it is designed as a permanent identifier and is not guaranteed to be a URL. When guid is missing, fall back to a hash of link plus title. Never use guid as the destination readers click through to — that is what link is for.
Why do some feeds have publication dates in the future?
It is almost always a publisher misconfiguration — a scheduled post that leaked into the feed, or a CMS exporting a timezone-incorrect timestamp. Reject or clamp dates more than a few days ahead, otherwise those items pin themselves to the top of your timeline indefinitely.
Why do feed images sometimes fail to load later?
Because most feeds only hand you a URL, not the asset. Thumbnails and media:thumbnail entries often point to CDN hosts that enforce hotlink protection or rotate URLs. Downloading and re-hosting images at ingest time is the only reliable fix.
Can I trust the HTML inside a feed item?
No. Feed content is third-party input delivered over a channel nobody authenticates. Decode the entities, sanitize the markup, strip scripts, iframes, and inline event handlers, and restrict tags to a known-safe allowlist before it ever reaches a renderer.