HTML DOM Structure JSON Parser
Paste the following prompt into your AI chat to install this skill:
Please follow https://skillhub.cn/install/skillhub.md to install @user_223dc0b0/ksmy004.
About this skill
Problem
HTML pages often bury useful data inside nested tags, utility classes, and repetitive markup. Pulling values directly from the raw DOM forces engineers to handle table, li, div.item, form controls, and context-dependent text fields separately, which adds ongoing parsing and maintenance work.
How It Works
The parse skill accepts an HTML string or a local HTML file path and uses an LLM with parsing tools to produce structured output:
- Tables: detect table and extract headers plus row data.
- Lists: identify repeated containers such as div.item or li, then extract key fields.
- Forms: read inputs, dropdowns, and current values.
- Semantic mapping: convert context phrases such as 订单号:123 into meaningful key-value pairs like order_id mapped to 123.
- Cleaning: remove redundant tags and whitespace to produce JSON suitable for storage or analysis.
Boundaries
It works best on pages with recognizable structure and semantically nameable fields. For very complex markup, pre-clean with BeautifulSoup before semantic extraction. The generated JSON still needs validation against downstream schema and field-naming conventions.
Use Cases
- Extract names, prices, and stock from repeated product list containers into unified JSON.
- Pull order table headers and row data from an admin page for downstream reporting scripts.
- Read current input and dropdown values from a settings form and map them to config JSON.
- Clean a detail page text and convert order numbers and merchant fields into key-value pairs.
Best For
- Backend engineers maintaining scrapers who need stable JSON from tables and lists.
- Data analysts cleaning data who need to extract current form configuration values.
- Product engineers building dashboards who need semantic fields like order IDs as keys.
- Test engineers verifying page structures who need consistent item-level field extraction.
Related Skills
Fetches Baidu Hot Search Top 10 titles using web_fetch first, validates same-day data, and falls back to browser automation when stale.
Generates an evening A-share policy and trading opportunity daily report by collecting same-day index, policy, and capital data, then applying a fixed template to highlight beneficiary sectors, drivers, and price directions.
Maps natural-language TikTok requests to KeyAPI REST scenarios, validates endpoints against docs, and executes data queries and analysis.
Extract city-specified AI jobs from BOSS Zhipin, save CSV/table data, mark new postings, and summarize salary trends, application advice, and HTML reports.