01 · DOM-Based HTML Extraction
Uses native browser DOMParser to build inert trees safely without executing scripts.
Accurately converts malformed markup, entities, and complex tables.
02 · Edge-Case Resilient Transforms
Transforms handle unicode Base64 (UTF-8 safe), URL-safe Base64 without padding, camelCase
acronym boundaries (parseHTTPResponse), and JSON escaping.
03 · Transparent Output Logs
Reports every discarded tag, stripped container, and table degradation so you know exactly what was modified or retained.
Operational Limits & Assumptions
- · Client-Rendered SPAs: Parses static source HTML; for JavaScript-rendered SPAs, paste the outer HTML directly from DevTools inspector.
- · Merged Table Cells: Tables with colspans/rowspans convert into annotated description lists to avoid broken Markdown pipe formatting.
- · ASCII Slug Decomposition: Unicode diacritics are normalized; non-Latin alphabets without ASCII equivalents are stripped.
- · Abbreviation Sentence Splits: Regex-based sentence splitters may divide text on periods in uncommon abbreviations.