HtmlTinkerX is the shared .NET engine for parsing, extracting, auditing, formatting, crawling, and rendering web content. PSParseHTML exposes the same engine as PowerShell cmdlets.
Dependency guardrails, including the ChartForgeX-backed screenshot image-processing path, are documented in Docs/Dependencies.md.
- HTML parsing with AngleSharp and Html Agility Pack
- object-first page reading with headings, paragraphs, tables, links, resources, and inferred repeated collections
- tables, lists, forms, metadata, JSON-LD, microdata, Open Graph, application state, tokens, image candidates, and API endpoint extraction
- static and rendered document audits for duplicate IDs, document metadata, accessible names, unsafe URL schemes, and heading order
- bounded website crawling to offline HTML, text, Markdown, JSON, JSONL, graph, and asset datasets
- Playwright sessions, interaction, screenshots, PDFs, HAR files, traces, browser recipes, cookies, storage, and SSO handoff inspection
- HTML, CSS, JavaScript, and email formatting or optimization
- .NET Framework 4.7.2, .NET 8, and .NET 10
dotnet add package HtmlTinkerXInstall-Module -Name PSParseHTML -AllowClobber -ForceThese commands install the current stable releases. Upstream dependency version labels do not change HtmlTinkerX or PSParseHTML release channels.
- π¦ NuGet Package:
HtmlTinkerX- Core .NET library - π§ PowerShell Module:
PSParseHTML- PowerShell cmdlets wrapper - π― Target Frameworks: .NET Framework 4.7.2, .NET 8.0, and .NET 10.0
- π» PowerShell Compatibility: Windows PowerShell 5.1 and PowerShell 7.4+
Start with a local HTML string so you can inspect the result immediately:
$html = @'
<!doctype html>
<html><head><title>Products</title></head><body>
<main>
<h1>Products</h1>
<p>Two items are available.</p>
<table><tr><th>Name</th><th>Price</th></tr>
<tr><td>Desk</td><td>120</td></tr>
<tr><td>Chair</td><td>40</td></tr>
</table>
</main>
</body></html>
'@
$page = Get-HtmlPage -Content $html
$page.Headings
$page.Paragraphs
$page.TablesFor a live page, use Get-HtmlPage -Url $url. The result also exposes links,
forms, resources, inferred repeated collections, readable text and Markdown.
The object workflow guide
shows how to inspect tables, choose collections by their fields and reuse
selectors when you need a stable extraction recipe.
Skip analyses you do not need:
$page = Get-HtmlPage -Content $html -NoReadableText -NoMarkdown -NoWebData -NoCollections
$page.Headings
$page.TablesFor client-rendered pages, read a browser snapshot:
$snapshot = Invoke-HtmlRendering -Url $url -Snapshot
$page = Get-HtmlPage -RenderedSnapshot $snapshotusing System;
using HtmlTinkerX;
HtmlPageDocument page = HtmlPageReader.Read(
"<html><head><title>Products</title></head><body>"
+ "<main><h1>Products</h1><p>Two items are available.</p></main></body></html>");
Console.WriteLine(page.Headings[0].Text);| Workflow | Guide |
|---|---|
| Read semantic objects and inferred collections | Object workflows |
| Crawl pages, mirror assets, choose content and resume a dataset | Crawling and exports |
| Parse and audit documents or automate a browser in PowerShell | PowerShell workflows |
| Call the .NET engine and inspect its public APIs | .NET examples and API overview |
| Parse local files, format resources and capture browser output | More examples |
| Check page errors, network requests, media and browser installation | Browser testing |
| Configure request interception, profiles and saved state | Browser sessions |
| Set HTTP response limits and understand encoding or cancellation | HTTP response controls |
The generated command reference contains the complete PowerShell parameter and pipeline documentation. See MIGRATION.md for API changes when upgrading.
URL parsing, HTTP form submission and related shared readers enforce a response size limit and honor cancellation. See HTTP response controls for the defaults, explicit overrides and decoding order.
Use the command reference for individual commands and the PowerShell workflow guide for combinations of commands.
- HTML/CSS/JavaScript Processing
- Browser Automation & Interaction
- Browser Extraction Mode
- Browserless Extraction Mode
- Screenshots & Media
- Network & Debugging
- Cookies & State Management
- Content & Resources
- JavaScript AST parsing
- CSS and HTML workflow audits
- React Server Component payload extraction
- Modern page parsing helpers
The .NET guide covers the parsing, extraction, crawl and browser APIs with examples.
- Core Classes
- HtmlParser
- HtmlFormatter
- HtmlOptimizer
- HtmlBrowser (Browser Automation)
- HtmlUtilities
- PreMailerClient
- Extension Methods
- HtmlParserExtensions
See workflow examples, PowerShell scripts and the .NET example project.
- PowerShell Examples
- Table Extraction
- Resource Optimization
- Browser Automation
- C# Examples
- Document Processing
- Resource Optimization
- Browser Automation
The browser testing guide covers page checks, console and network logs, performance metrics, browser setup and troubleshooting.
- PowerShell Browser Testing
- Basic Testing
- Testing Local HTML Files
- Testing for Console Errors
- CSS Resource Testing
- Performance Testing
- Advanced Testing with Proxy
- Batch Testing Multiple URLs
- Integration with Pester Tests
- Testing Local HTML Reports
- Monitoring and Alerting
- C# Browser Testing
- Basic Testing
- Testing Local HTML Files
- Network Request Analysis
- Console Error Detection
- Performance Analysis
- Testing Local HTML Files
- Integration Testing Examples
- Batch Testing Multiple Pages
- Test Result Properties
- HtmlBrowserTestResult
- HtmlNetworkEntryDetailed
- HtmlConsoleEntryDetailed
- HtmlPerformanceMetrics
- Playwright Auto-Setup
- How Auto-Download Works
- Linux: Avoiding sudo prompts
- Cleaning Playwright Cache
- C# Cache Cleaning
- Integration with Test Frameworks
- xUnit Example
- Pester Example
The browser session guide covers context settings, request interception and saved state.
The PowerShell commands are thin surfaces over HtmlTinkerX. The generated command reference covers every cmdlet.
| Task | PowerShell | C# owner |
|---|---|---|
| Read a page as objects without selectors | Get-HtmlPage |
HtmlPageReader.Read |
| Parse a document | ConvertFrom-Html |
HtmlParser.ParseWithAngleSharp, HtmlParser.ParseWithHtmlAgilityPack |
| Extract tables | ConvertFrom-HtmlTable |
HtmlParser.ParseTablesWithAngleSharpDetailed, HtmlParser.ParseTablesWithHtmlAgilityPackDetailed |
| Extract lists | ConvertFrom-HtmlList |
HtmlParser.ParseListsWithAngleSharpDetailed, HtmlParser.ParseListsWithHtmlAgilityPackDetailed |
| Extract forms | ConvertFrom-HtmlForm |
HtmlParser.ParseFormsWithAngleSharp |
| Extract metadata | ConvertFrom-HtmlMeta |
HtmlParser.ParseMetaTags |
| Extract Open Graph data | ConvertFrom-HtmlOpenGraph |
HtmlParser.ParseOpenGraph |
| Extract microdata | ConvertFrom-HtmlMicrodata |
HtmlParser.ParseMicrodataItems |
| Normalize mixed page data | Select-HtmlData |
HtmlParsingToolbox.SelectData |
| Build a page workbench and audit | Invoke-HtmlPageWorkbench |
HtmlPageWorkbench.AnalyzeAsync, HtmlDocumentAudit.Analyze |
| Find interaction surfaces | Find-HtmlInteractionSurface |
HtmlParsingToolbox.FindInteractionSurfaceAsync |
| Discover API endpoints | Find-HtmlApiEndpoint |
HtmlApiEndpointInventory.Build |
| Compare static and rendered HTML | Compare-HtmlStaticRendered |
HtmlParsingToolbox.CompareStaticRendered |
| Crawl and export a dataset | Invoke-HtmlCrawl |
HtmlCrawler.CrawlAsync |
| Open or navigate a browser | Start-HtmlBrowserSession, Invoke-HtmlBrowserNavigation |
HtmlBrowser.OpenSessionAsync, HtmlBrowser.NavigateAsync |
| Click or fill an element | Invoke-HtmlBrowserClick, Set-HtmlBrowserInput |
HtmlBrowser.ClickSelectorAsync, HtmlBrowser.FillInputAsync |
| Capture a screenshot or PDF | Save-HtmlBrowserScreenshot, Save-HtmlBrowserPdf |
HtmlBrowser.CaptureScreenshotAsync, HtmlBrowser.SavePagePdfAsync |
| Export HAR or trace data | Export-HtmlBrowserHar, Start-HtmlBrowserTracing, Stop-HtmlBrowserTracing |
HtmlBrowser.ExportHarAsync, HtmlBrowser.StartTracingAsync, HtmlBrowser.StopTracingAsync |
| Test a rendered page | Test-HtmlBrowser |
HtmlBrowserTester |
| Inline email CSS | Optimize-Email |
PreMailerClient.MoveCssInline, PreMailerClient.MoveCssInlineAsync |
| Format or minify resources | Format-Html, Format-Css, Format-JavaScript, Optimize-* |
HtmlFormatter, HtmlOptimizer |
HtmlTinkerX utilizes several high-quality open-source libraries:
- AngleSharp - MIT License - Modern HTML5 parser
- AngleSharp.Css - MIT License - CSS parsing and styling
- AngleSharp.Js - MIT License - JavaScript engine integration
- AngleSharp.Diffing - MIT License - HTML document comparison
- Html Agility Pack - MIT License - Alternative HTML parser
- NUglify - BSD 2-Clause License - HTML/CSS/JS minification
- Jsbeautifier - MIT License - JavaScript formatting
- PreMailer.Net - Apache 2.0 License - Email CSS inlining
- Microsoft.Playwright - Apache 2.0 License - Browser automation
- ChartForgeX - MIT License - Dependency-free screenshot image composition
Screenshot image post-processing is routed through ChartForgeX. See Docs/Dependencies.md before changing that dependency path.
- System.Net.Http - MIT License - HTTP client
- System.IO.Compression - MIT License - Compression support
- System.Threading.Channels - MIT License - Async communication
All dependencies are distributed under permissive licenses. Refer to each project's repository for complete license information.
- π Examples: Check the
Examplesfolder for comprehensive usage samples - π Issues: Report bugs and request features on GitHub Issues
- π¬ Discord: Join our Discord community for support and discussions
- π Blog: Read detailed tutorials on evotec.xyz
Update-Module -Name PSParseHTMLdotnet add package HtmlTinkerXHtmlTinkerX and PSParseHTML are available under the MIT License. Third-party components remain under their own license terms.