Skip to content

About

HtmlTinkerX is a powerful async C# library for HTML, CSS, and JS processing, parsing, formatting, and optimization. It provides web content processing capabilities including browser automation, document parsing with multiple engines, resource optimization, and more. PSParseHTML is the PowerShell module exposing HtmlTinkerX to PowerShell.

Topics

Resources

Stars

139 stars

Watchers

3 watching

Forks

Latest commit

Β 

History

1,609 Commits

Folders and files

Repository files navigation

HtmlTinkerX & PSParseHTML - HTML processing for .NET and PowerShell

HtmlTinkerX is the shared .NET engine for parsing, extracting, auditing, formatting, crawling, and rendering web content. PSParseHTML exposes the same engine as PowerShell cmdlets.

πŸ“¦ NuGet Package

nuget downloads nuget version

πŸ’» PowerShell Module

powershell gallery version powershell gallery platforms powershell gallery downloads

πŸ› οΈ Project Information

.NET Tests PowerShell Tests top language codecov

Dependency guardrails, including the ChartForgeX-backed screenshot image-processing path, are documented in Docs/Dependencies.md.

πŸ‘¨β€πŸ’» Author & Social

Twitter Follow Blog LinkedIn Discord

What it covers

  • HTML parsing with AngleSharp and Html Agility Pack
  • object-first page reading with headings, paragraphs, tables, links, resources, and inferred repeated collections
  • tables, lists, forms, metadata, JSON-LD, microdata, Open Graph, application state, tokens, image candidates, and API endpoint extraction
  • static and rendered document audits for duplicate IDs, document metadata, accessible names, unsafe URL schemes, and heading order
  • bounded website crawling to offline HTML, text, Markdown, JSON, JSONL, graph, and asset datasets
  • Playwright sessions, interaction, screenshots, PDFs, HAR files, traces, browser recipes, cookies, storage, and SSO handoff inspection
  • HTML, CSS, JavaScript, and email formatting or optimization
  • .NET Framework 4.7.2, .NET 8, and .NET 10

πŸ“¦ Installation & Packages

πŸ“¦ NuGet Package (C#/.NET)

dotnet add package HtmlTinkerX

πŸ”§ PowerShell Module

Install-Module -Name PSParseHTML -AllowClobber -Force

These commands install the current stable releases. Upstream dependency version labels do not change HtmlTinkerX or PSParseHTML release channels.

πŸ“‹ Package Information

  • πŸ“¦ NuGet Package: HtmlTinkerX - Core .NET library
  • πŸ”§ PowerShell Module: PSParseHTML - PowerShell cmdlets wrapper
  • 🎯 Target Frameworks: .NET Framework 4.7.2, .NET 8.0, and .NET 10.0
  • πŸ’» PowerShell Compatibility: Windows PowerShell 5.1 and PowerShell 7.4+

πŸš€ Quick Start

Read a page as objects

Start with a local HTML string so you can inspect the result immediately:

$html = @'
<!doctype html>
<html><head><title>Products</title></head><body>
<main>
  <h1>Products</h1>
  <p>Two items are available.</p>
  <table><tr><th>Name</th><th>Price</th></tr>
    <tr><td>Desk</td><td>120</td></tr>
    <tr><td>Chair</td><td>40</td></tr>
  </table>
</main>
</body></html>
'@
$page = Get-HtmlPage -Content $html
$page.Headings
$page.Paragraphs
$page.Tables

For a live page, use Get-HtmlPage -Url $url. The result also exposes links, forms, resources, inferred repeated collections, readable text and Markdown. The object workflow guide shows how to inspect tables, choose collections by their fields and reuse selectors when you need a stable extraction recipe.

Skip analyses you do not need:

$page = Get-HtmlPage -Content $html -NoReadableText -NoMarkdown -NoWebData -NoCollections
$page.Headings
$page.Tables

For client-rendered pages, read a browser snapshot:

$snapshot = Invoke-HtmlRendering -Url $url -Snapshot
$page = Get-HtmlPage -RenderedSnapshot $snapshot

Read a page in C#

using System;
using HtmlTinkerX;

HtmlPageDocument page = HtmlPageReader.Read(
    "<html><head><title>Products</title></head><body>"
    + "<main><h1>Products</h1><p>Two items are available.</p></main></body></html>");
Console.WriteLine(page.Headings[0].Text);

Workflow guides

Workflow Guide
Read semantic objects and inferred collections Object workflows
Crawl pages, mirror assets, choose content and resume a dataset Crawling and exports
Parse and audit documents or automate a browser in PowerShell PowerShell workflows
Call the .NET engine and inspect its public APIs .NET examples and API overview
Parse local files, format resources and capture browser output More examples
Check page errors, network requests, media and browser installation Browser testing
Configure request interception, profiles and saved state Browser sessions
Set HTTP response limits and understand encoding or cancellation HTTP response controls

The generated command reference contains the complete PowerShell parameter and pipeline documentation. See MIGRATION.md for API changes when upgrading.

HTTP response controls

URL parsing, HTTP form submission and related shared readers enforce a response size limit and honor cancellation. See HTTP response controls for the defaults, explicit overrides and decoding order.

πŸ”§ PowerShell Cmdlets

Use the command reference for individual commands and the PowerShell workflow guide for combinations of commands.

🎯 C# API Reference

The .NET guide covers the parsing, extraction, crawl and browser APIs with examples.

πŸ“š Examples

See workflow examples, PowerShell scripts and the .NET example project.

πŸ§ͺ Browser Testing & Network Monitoring

The browser testing guide covers page checks, console and network logs, performance metrics, browser setup and troubleshooting.

πŸ”§ Advanced Features

The browser session guide covers context settings, request interception and saved state.

PowerShell and C# entry points

The PowerShell commands are thin surfaces over HtmlTinkerX. The generated command reference covers every cmdlet.

Task PowerShell C# owner
Read a page as objects without selectors Get-HtmlPage HtmlPageReader.Read
Parse a document ConvertFrom-Html HtmlParser.ParseWithAngleSharp, HtmlParser.ParseWithHtmlAgilityPack
Extract tables ConvertFrom-HtmlTable HtmlParser.ParseTablesWithAngleSharpDetailed, HtmlParser.ParseTablesWithHtmlAgilityPackDetailed
Extract lists ConvertFrom-HtmlList HtmlParser.ParseListsWithAngleSharpDetailed, HtmlParser.ParseListsWithHtmlAgilityPackDetailed
Extract forms ConvertFrom-HtmlForm HtmlParser.ParseFormsWithAngleSharp
Extract metadata ConvertFrom-HtmlMeta HtmlParser.ParseMetaTags
Extract Open Graph data ConvertFrom-HtmlOpenGraph HtmlParser.ParseOpenGraph
Extract microdata ConvertFrom-HtmlMicrodata HtmlParser.ParseMicrodataItems
Normalize mixed page data Select-HtmlData HtmlParsingToolbox.SelectData
Build a page workbench and audit Invoke-HtmlPageWorkbench HtmlPageWorkbench.AnalyzeAsync, HtmlDocumentAudit.Analyze
Find interaction surfaces Find-HtmlInteractionSurface HtmlParsingToolbox.FindInteractionSurfaceAsync
Discover API endpoints Find-HtmlApiEndpoint HtmlApiEndpointInventory.Build
Compare static and rendered HTML Compare-HtmlStaticRendered HtmlParsingToolbox.CompareStaticRendered
Crawl and export a dataset Invoke-HtmlCrawl HtmlCrawler.CrawlAsync
Open or navigate a browser Start-HtmlBrowserSession, Invoke-HtmlBrowserNavigation HtmlBrowser.OpenSessionAsync, HtmlBrowser.NavigateAsync
Click or fill an element Invoke-HtmlBrowserClick, Set-HtmlBrowserInput HtmlBrowser.ClickSelectorAsync, HtmlBrowser.FillInputAsync
Capture a screenshot or PDF Save-HtmlBrowserScreenshot, Save-HtmlBrowserPdf HtmlBrowser.CaptureScreenshotAsync, HtmlBrowser.SavePagePdfAsync
Export HAR or trace data Export-HtmlBrowserHar, Start-HtmlBrowserTracing, Stop-HtmlBrowserTracing HtmlBrowser.ExportHarAsync, HtmlBrowser.StartTracingAsync, HtmlBrowser.StopTracingAsync
Test a rendered page Test-HtmlBrowser HtmlBrowserTester
Inline email CSS Optimize-Email PreMailerClient.MoveCssInline, PreMailerClient.MoveCssInlineAsync
Format or minify resources Format-Html, Format-Css, Format-JavaScript, Optimize-* HtmlFormatter, HtmlOptimizer

πŸ—οΈ Third-Party Dependencies

HtmlTinkerX utilizes several high-quality open-source libraries:

πŸ“¦ HTML & DOM Processing

🎨 Resource Optimization

  • NUglify - BSD 2-Clause License - HTML/CSS/JS minification
  • Jsbeautifier - MIT License - JavaScript formatting
  • PreMailer.Net - Apache 2.0 License - Email CSS inlining

🌐 Browser Automation

Screenshot image post-processing is routed through ChartForgeX. See Docs/Dependencies.md before changing that dependency path.

πŸ”§ System Libraries

All dependencies are distributed under permissive licenses. Refer to each project's repository for complete license information.

πŸ“– Documentation & Support

  • πŸ“š Examples: Check the Examples folder for comprehensive usage samples
  • πŸ› Issues: Report bugs and request features on GitHub Issues
  • πŸ’¬ Discord: Join our Discord community for support and discussions
  • πŸ“ Blog: Read detailed tutorials on evotec.xyz

πŸ”„ Updates & Versioning

PowerShell Module Updates

Update-Module -Name PSParseHTML

NuGet Package Updates

dotnet add package HtmlTinkerX

⚠️ Important: Always test updates in a development environment before deploying to production. Breaking changes may occur between versions.

πŸ”§ Troubleshooting

πŸ“„ License

HtmlTinkerX and PSParseHTML are available under the MIT License. Third-party components remain under their own license terms.

About

HtmlTinkerX is a powerful async C# library for HTML, CSS, and JS processing, parsing, formatting, and optimization. It provides web content processing capabilities including browser automation, document parsing with multiple engines, resource optimization, and more. PSParseHTML is the PowerShell module exposing HtmlTinkerX to PowerShell.

Topics

Resources

Stars

139 stars

Watchers

3 watching

Forks

Releases

Sponsor this project

Used by

Contributors

Languages