Repository navigation
Release persisted crawl page bodies on request - #501
PrzemyslawKlys wants to merge 7 commits into
Conversation
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
Codecov Report❌ Patch coverage is Additional details and impacted files@@ Coverage Diff @@
## feature/crawl-page-observer #501 +/- ##
===============================================================
+ Coverage 58.97% 59.09% +0.12%
===============================================================
Files 490 491 +1
Lines 36178 36267 +89
Branches 7301 7322 +21
===============================================================
+ Hits 21335 21433 +98
+ Misses 12413 12406 -7
+ Partials 2430 2428 -2 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
c401849 to
b4d192b
Compare
b4d192b to
0b6893d
Compare
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 0b6893d41e
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
a071c44 to
1103435
Compare
Set
HtmlCrawlOptions.RetainPageContent = falseor useInvoke-HtmlCrawl -ReleasePageContentto release HTML, text, Markdown and HTTP cache bodies after their full records have been saved.The result keeps metadata and structured data. Exports hydrate stored bodies one page at a time, and the final manifest remains self-contained. Resume, conditional refresh, copied datasets and repeated saves preserve content. Private page identities keep reordered pages and cached skipped records attached to their own stored bodies.
With
-StreamPages, emitted copies keep their content while the crawl releases its retained copies. A C# observer can callpage.CreateSnapshot()during the callback to retain the current strings. It is a shallow copy: collections and structured data remain shared. It neither reloads released content nor keeps a hidden dependency on its source content file.The option requires an output or resume path. Keep the stored dataset available while using a result with released bodies, or load it with
HtmlCrawler.LoadResultAsyncfor an editable result with retained content. Metadata, deduplication data and queued streamed copies still consume memory; loading a manifest or refresh source can materialize its bodies.Merge streamed chunk export #497, then page observers #492, before this layer. The updated stack passes 50 focused .NET 10 cases and all three streaming cases on each PowerShell version. Earlier validation covers persistence, refresh and resumed checkpoints on all three supported .NET targets, with mutation proof and independent content-lifetime review. Current hosted CI remains a merge gate.