Skip to content

fix(feed/search): parse feed once + content-based search cache - #225

Merged
chhoumann merged 2 commits into
masterfrom
chhoumann/deepsec-feed-cache-bugs
Jun 29, 2026
Merged

chhoumann merged 2 commits into
masterfrom
chhoumann/deepsec-feed-cache-bugs

Conversation

@chhoumann

Copy link
Copy Markdown
Owner

Fixes two deepsec-confirmed data-layer bugs in the feed/search path.

getEpisodes fetched the feed twice (slug: other-redundant-fetch)

FeedParser.getEpisodes(url) parsed the feed once via getFeed (to populate
channel metadata) and then parsed it again directly to read the items - two
network round-trips and two XML parses per cold call. This was actively
reachable: the URIHandler resume-link path ({{episodelink}} / share links)
constructs new FeedParser() with no cached feed, so both fetches fired.

The fix fetches and parses the document once and reuses it for both the channel
metadata and the items. Metadata extraction moves to a private
extractFeed(body, url) that getFeed and getEpisodes share, so metadata
population, the Invalid RSS feed guard, and this.feed caching are unchanged -
only the redundant second fetch is gone.

Search cache returned stale results on in-place mutation (slug: other-stale-cache)

searchEpisodes cached a Fuse index in a WeakMap keyed by the episodes array
and reused it whenever the cached size equalled the array length. Length is a
weak fingerprint: the same array reference mutated in place at the same length
(an entry swapped or edited) returned the stale index built from the old
contents.

The fix validates the cache with a content signature (a JSON-framed list of each
episode's title and streamUrl) instead of length. Any content or ordering
change rebuilds the index; an unchanged list keeps reusing it across keystrokes,
preserving the index-reuse optimization from #149. streamUrl is folded in so
the cache is keyed on episode identity rather than coupling correctness to
fuse.js's internal live-reference behavior.

Tests

  • feedParser.test.ts: the getEpisodes tests now assert a single fetch
    (toHaveBeenCalledTimes(1)); a shared feedResponse helper replaces the
    duplicated mock setup that previously implied two fetches.
  • searchEpisodes.test.ts (new): covers the fuzzy match, empty/whitespace
    short-circuits, cache reuse for an unchanged list, fresh results after an
    in-place same-length mutation, and a rebuild on a same-title identity swap. The
    cache tests use a Fuse-construction spy and were verified to fail against the
    old implementation.

Gates

lint, format:check, typecheck, build, and the unit suite pass. Note: the
pre-existing PodcastView.integration.test.ts > reopening a show reuses full in-memory cache without refetching test is flaky on master (timing-based,
fails intermittently without these changes); it mocks getEpisodes and does not
exercise the code touched here.

Scope is limited to src/parser/feedParser.ts and src/utility/searchEpisodes.ts
(plus their tests).

getEpisodes fetched and parsed the feed twice on a cold call: once via
getFeed (to populate channel metadata) and again directly to read the
items. This doubled the network round-trips and XML parses on every cold
path, e.g. the URIHandler resume-link flow that constructs a FeedParser
with no cached feed.

Fetch and parse the document a single time and reuse it for both the
channel metadata and the episode items. The metadata extraction moves to
a private extractFeed(body, url) helper that getFeed and getEpisodes both
call, so behavior (metadata population, the "Invalid RSS feed" guard,
this.feed caching) is unchanged - only the redundant second fetch is
removed.

Resolves deepsec finding other-redundant-fetch.
The Fuse search index was cached in a WeakMap keyed by the episodes array
and reused whenever the cached size matched the array length. Length is a
weak fingerprint: the same array reference mutated in place at the same
length (an entry swapped or edited) would return the stale index built
from the old contents.

Validate the cache with a content signature (a JSON-framed list of each
episode's title and streamUrl) instead of length, so any content or
ordering change rebuilds the index while an unchanged list keeps reusing
it across keystrokes (preserving the #149 optimization). Adds the first
unit tests for searchEpisodes, including a Fuse-construction spy proving
the index is reused when unchanged and rebuilt when the content changes.

Resolves deepsec finding other-stale-cache.
@cloudflare-workers-and-pages

Copy link
Copy Markdown

Deploying podnotes with  Cloudflare Pages  Cloudflare Pages

Latest commit: c57fae3
Status: ✅  Deploy successful!
Preview URL: https://de662011.podnotes.pages.dev
Branch Preview URL: https://chhoumann-deepsec-feed-cache.podnotes.pages.dev

View logs

@chhoumann
chhoumann marked this pull request as ready for review June 29, 2026 07:20
@chhoumann
chhoumann merged commit 053d51f into master Jun 29, 2026
2 checks passed
@chhoumann
chhoumann deleted the chhoumann/deepsec-feed-cache-bugs branch June 29, 2026 07:23
github-actions Bot pushed a commit that referenced this pull request Jul 9, 2026
## [2.17.3](2.17.2...2.17.3) (2026-07-09)

### Bug Fixes

* **feed/search:** parse feed once + content-based search cache ([#225](#225)) ([053d51f](053d51f)), closes [#149](#149)
* make episode identity key collision-resistant and prototype-safe ([#226](#226)) ([a5683db](a5683db))
* **opml:** correct import progress math and saved-count reporting ([#221](#221)) ([a79e529](a79e529))
* **security:** validate feed/URI URLs and cap download size ([#223](#223)) ([edef281](edef281))
* **template:** neutralize feed-controlled note injection ([#228](#228)) ([ef4ecbd](ef4ecbd))
* **timestamp:** escape live table-cell pipe after an escaped backslash ([#227](#227)) ([a34dfca](a34dfca))
* **transcription:** resolve three deepsec transcription-pipeline bugs ([#224](#224)) ([83c34e7](83c34e7))
@github-actions

github-actions Bot commented Jul 9, 2026

Copy link
Copy Markdown
Contributor

🎉 This PR is included in version 2.17.3 🎉

The release is available on GitHub release

Your semantic-release bot 📦🚀

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant