
Reading RSS Feed Items: Title, Description, Content, Source URL, Publication Date, and Images
An RSS item is a structured XML record, and every field in it carries a different kind of information — a headline, a summary, the full article body, a canonical link, a timestamp, and zero or more images. Reading those fields correctly means knowing which element to trust for which job: title for display, description for previews, content:encoded for full-text, link for attribution, pubDate for ordering, and enclosure or media:content for images that actually render.
Most feed problems — duplicate posts, missing thumbnails, out-of-order timelines, broken attribution links — trace back to one field being used where another belongs. This guide walks through each element, what it reliably contains, and how to handle it in a production feed reader or content pipeline.

What an RSS Item Actually Is
An RSS feed is a single XML document containing a channel and a list of items. Each is one discrete piece of content: a blog post, a news story, a podcast episode, a product update. RSS 2.0 defines a fixed set of core elements per item, and extension modules add everything the core spec leaves out.
That modular design is the reason feed parsing is harder than it looks. A title is always a title, but “the image” might live in , , , , or simply as the first tag buried inside the content payload. A production parser needs to check all of them, in a defined order.
The Atom format (standardized in RFC 4287{:target=”_blank” rel=”noopener”}) solves some of this with stricter rules, but a large share of the web still publishes RSS 2.0, so any reader worth building handles both.

The Core Fields, One by One
| Field | Typical element | What it gives you | Common pitfall |
|—|—|—|—|
| Title | | Headline, often plain text | May contain HTML entities that need decoding |
| Description | | Summary or full content, depending on the publisher | Truncated snippets, encoded markup, or entire articles |
| Content | | Full article body in HTML | Frequently missing; not part of the RSS 2.0 core spec |
| Source URL | and | Canonical permalink and unique identifier | guid is not always a URL |
| Publication date | , , | When the item was published or last modified | Multiple conflicting formats in the same feed |
| Images | , , | Image URLs with optional dimensions and MIME types | Thumbnails are hotlink-protected or low resolution |
Title
The title is the field you display and the field readers scan. Treat it as text, not markup: strip any tags that appear, decode HTML entities such as & and ’, and normalize whitespace.
Two quirks show up repeatedly. First, some publishers prefix every title with the site name, producing “Article Headline – Site Name” in every entry. Second, podcast and newsletter feeds sometimes leave the title empty and put everything in the description. A robust reader has a fallback: if the title is blank, derive one from the first line of the description, then from the URL slug.

Description
In RSS 2.0 the spec calls description “the item synopsis,” but in practice it is one of three things depending on the publisher:
- A short hand-written summary, ideal for previews and cards
- A truncated excerpt with a “read more” link
- The complete article body, HTML-encoded into the XML
isPermaLink="true"(the default) — the GUID is probably a URL, but you should still not treat it as the clickable destinationisPermaLink="false"— the GUID is an opaque string, possibly a UUID or a database key- Parse against a list of known formats in priority order rather than a single strict pattern.
- Treat a missing timezone as UTC rather than local server time — otherwise your ordering shifts by hours.
- Keep
pubDateandupdatedseparate. The first is when the item appeared; the second is when it last changed. Sorting a timeline byupdatedreshuffles old posts to the top, which surprises readers. - Reject dates more than a few days in the future. Future-dated items are almost always a publisher misconfiguration and will pin themselves permanently to the top of the feed.
— the original podcast-era mechanism, still the most reliable when present— Media RSS, which addswidth,height, andtype— explicitly a small preview image, often 144px or smaller— podcast artwork, usually square and site-level rather than per-episode- First
insidecontent:encodedordescription— the least reliable option, but sometimes the only one - Deduplicate first. Compare the
guid, or a hash oflinkplustitle, against your store. Skip anything already seen before doing any other work. - Establish the canonical URL. Resolve
linkto an absolute URL and strip tracking parameters. - Determine content depth. If
content:encodedexists, use it. Otherwise usedescription. Otherwise queue the source URL for full-text extraction. - Sanitize the HTML. Decode entities, strip scripts, iframes, and inline event handlers, and restrict allowed tags to what your design supports.
- Resolve the primary image. Walk the cascade above and store the winning URL along with its dimensions if the feed supplied them.
- Normalize the timestamp. Parse to UTC, validate plausibility, and store both
pubDateandupdatedas separate columns. - Fetch and cache images locally. Hotlinked images disappear. Downloading and re-hosting them keeps your archive intact and gives you a stable URL for social cards.
- Newsletters and digests — title plus description plus image, assembled into cards on a schedule
- Internal content operations — monitoring competitor blogs, industry news, or release notes by filtering
pubDateand keyword-matching the description - AI summarization pipelines —
content:encodedas the input,linkas the citation,pubDateas the freshness signal - Social automation — title as the post text, image as the attachment,
linkas the target - Search and archive — full content indexed, with
guidpreventing duplicate entries when feeds re-publish - Using guid as a link. The spec explicitly allows non-URL GUIDs.
- Trusting a single date format. Feeds mix RFC 822, RFC 3339, and ISO 8601 variants within the same document.
- Rendering feed HTML unsanitized. Feed content is third-party input and should be treated as untrusted at every step.
- Hotlinking images. CDN protection and link rot will silently empty your archive over time.
- Ignoring
isPermaLink. It changes how you should interpret the identifier.
Because of this ambiguity, never assume the description is short. Measure it. If a description routinely exceeds a few thousand characters, the publisher is treating it as full content and your truncation logic should kick in before rendering.
Note that description content is usually escaped:
arrives as <p>. Decoding it before display is mandatory, and sanitizing it afterward is not optional — feed content is untrusted input.

Content
The content:encoded element comes from the RSS 1.0 Content Module (http://purl.org/rss/1.0/modules/content/), not from RSS 2.0 itself. When it exists, it holds the full post in HTML — headings, images, code blocks, embedded media.
Prefer content:encoded over description whenever both are present and your use case calls for full text. If it is absent, fall back to description, then to fetching the source URL directly.
Atom feeds handle this more cleanly with a element that declares its own type attribute (html, text, or xhtml), so you know exactly what you are decoding before you decode it.
Source URL and Identifier
Two elements overlap here, and conflating them causes real bugs.
is the item’s permalink — the page a reader should open to read the original. It is what you attribute, what you canonicalize against, and what you send in a newsletter. Always resolve it against the feed’s base URL if it is relative, and strip tracking parameters such as utm_source before storing it.
is a permanent identifier. It signals “this is the same item I sent you before.” Critically, the isPermaLink attribute determines whether the GUID is a URL:
Use guid for deduplication. Use link for navigation. When guid is missing, which happens in loose Atom-to-RSS conversions, fall back to a hash of the link plus the title.
Publication Date
Dates are where feed parsing breaks most often. RSS 2.0 specifies RFC 822 date formatting, which looks like Tue, 14 May 2024 09:30:00 GMT. Atom uses RFC 3339 timestamps such as 2024-05-14T09:30:00Z. Dublin Core contributes dc:date in yet another shape. Some feeds include timezone offsets, some omit them entirely, and some send nonsense.
Practical handling:
Image URLs
There is no single “image” element in RSS 2.0, so extraction follows a cascade. Check in roughly this order:
Two rules make this manageable. Prefer the largest available image, since downscaling looks better than upscaling; and honor media:thumbnail only as a last resort, because thumbnails frequently point to CDN-hosted assets that block hotlinking or expire.
Also watch for media:group, which bundles multiple renditions of the same asset. Pick one from the group rather than treating each rendition as a separate image.
A Practical Reading Order for Feed Parsers
When you ingest a new item, process the fields in this sequence to avoid rework:
Why Attribution Depends on Getting These Fields Right
Feed content is syndicated content, and syndication carries obligations. The link element is your attribution channel: any republished excerpt should point back to it, and any AI-generated summary should cite it. Getting the source URL wrong — pointing readers to a cached copy, a redirect chain, or a tracking-laden variant — breaks the implicit agreement between publisher and aggregator.
The element, when present, names the original publication, and or dc:creator identifies the writer. Both should travel with the content wherever it goes. Atom’s element is more structured, containing name, email, and URI as separate sub-elements, which is why Atom feeds tend to attribute more cleanly.
Uses Beyond Reading
The same six fields power a surprising range of systems:
In each case the failure modes are identical: missing images, duplicated items, and scrambled chronology. Fixing the field-level handling fixes all three.
Common Pitfalls Worth Memorizing
– Treating description as a summary. Half the web treats it as the article. Measure before you truncate.
Frequently Asked Questions
What is the difference between description and content:encoded?
description is part of the RSS 2.0 core specification and holds a synopsis — though many publishers put the entire article in it. content:encoded comes from the RSS 1.0 Content Module and, when present, is intended to carry the complete post body. In practice: prefer content:encoded for full text, fall back to description for previews, and treat both as untrusted HTML that must be decoded and sanitized.
Is content:encoded part of standard RSS 2.0?
No. It belongs to the RSS 1.0 Content Module, namespaced at http://purl.org/rss/1.0/modules/content/. Most major feed readers support it, but some feeds omit it entirely, which is why your parser needs a description fallback.
Should I use guid or link for deduplication?
Use guid when it exists, since it is designed as a permanent identifier and is not guaranteed to be a URL. When guid is missing, fall back to a hash of link plus title. Never use guid as the destination readers click through to — that is what link is for.
Why do some feeds have publication dates in the future?
It is almost always a publisher misconfiguration — a scheduled post that leaked into the feed, or a CMS exporting a timezone-incorrect timestamp. Reject or clamp dates more than a few days ahead, otherwise those items pin themselves to the top of your timeline indefinitely.
Why do feed images sometimes fail to load later?
Because most feeds only hand you a URL, not the asset. Thumbnails and media:thumbnail entries often point to CDN hosts that enforce hotlink protection or rotate URLs. Downloading and re-hosting images at ingest time is the only reliable fix.
Can I trust the HTML inside a feed item?
No. Feed content is third-party input delivered over a channel nobody authenticates. Decode the entities, sanitize the markup, strip scripts, iframes, and inline event handlers, and restrict tags to a known-safe allowlist before it ever reaches a renderer.
