AI Agent Hub
Back to skills
HTML DOM Structure JSON Parser icon

HTML DOM Structure JSON Parser

Data Analysis Updated 2026.08.30

Paste the following prompt into your AI chat to install this skill:

Please follow https://skillhub.cn/install/skillhub.md to install @user_223dc0b0/ksmy004.

About this skill

Problem

HTML pages often bury useful data inside nested tags, utility classes, and repetitive markup. Pulling values directly from the raw DOM forces engineers to handle table, li, div.item, form controls, and context-dependent text fields separately, which adds ongoing parsing and maintenance work.

How It Works

The parse skill accepts an HTML string or a local HTML file path and uses an LLM with parsing tools to produce structured output:
- Tables: detect table and extract headers plus row data.
- Lists: identify repeated containers such as div.item or li, then extract key fields.
- Forms: read inputs, dropdowns, and current values.
- Semantic mapping: convert context phrases such as 订单号:123 into meaningful key-value pairs like order_id mapped to 123.
- Cleaning: remove redundant tags and whitespace to produce JSON suitable for storage or analysis.

Boundaries

It works best on pages with recognizable structure and semantically nameable fields. For very complex markup, pre-clean with BeautifulSoup before semantic extraction. The generated JSON still needs validation against downstream schema and field-naming conventions.

Use Cases

  • Extract names, prices, and stock from repeated product list containers into unified JSON.
  • Pull order table headers and row data from an admin page for downstream reporting scripts.
  • Read current input and dropdown values from a settings form and map them to config JSON.
  • Clean a detail page text and convert order numbers and merchant fields into key-value pairs.

Best For

  • Backend engineers maintaining scrapers who need stable JSON from tables and lists.
  • Data analysts cleaning data who need to extract current form configuration values.
  • Product engineers building dashboards who need semantic fields like order IDs as keys.
  • Test engineers verifying page structures who need consistent item-level field extraction.