Introduction

In DeepSeek Harness (DSH), read_url is suitable for answering “What does this page say?”, but after parsing the page into text or Markdown, it may lose the association between metadata (such as rankings, prices, and likes) and specific items.

dsh-fetch-data aims to answer “What exactly is this data?”. It intercepts the page’s real data APIs (XHR/fetch JSON), extracts structured fields, and preserves precise field ownership.

Plugin Overview

dsh-fetch-data is a structured data extraction plugin for DeepSeek Harness. It uses a browser engine (Playwright) to capture network requests during page loading, extracts data from JSON APIs, and returns it to DSH for use.

  • Plugin Name: dsh-fetch-data
  • Maintainer: 2672243194
  • License: MIT

Core Features

This plugin mainly provides the following capabilities:

  1. API Interception: Captures the page’s XHR and fetch requests and obtains the real JSON responses.
  2. Extraction Modes:
    • Structural Mode: Returns the structure of the auto-selected endpoint and a list of all JSON APIs (path, size, whether it contains arrays), making it easy for the model to choose a specific API.
    • Field Mode: Extracts the values of specified field paths (such as data.list[].title).
  3. Automatic Selection: Automatically selects the largest endpoint containing an array from the captured JSON as the default data source.
  4. Enhanced Capture: Supports automatic scrolling to trigger lazy-loaded data and automatic unwrapping of JSONP responses.

Installation and Enablement

Before use, install Playwright as the browser engine and then add the plugin through DSH commands.

# 1. 进入 DSH 配置目录
cd <DSH profile dir>

# 2. 安装 Playwright 并安装 Chromium 浏览器
npm i playwright && npx playwright install chromium

# 3. 添加插件
dsh plugin --profile web add github:2672243194/dsh-fetch-data

Usage Examples

The plugin provides a fetch_data function that supports the following parameters:

fetch_data(url, api?, fields?, maxItems?)

1. Structural Mode (Default)

If the fields parameter is not provided, it returns an overview of the API structure and a list of all available JSON APIs. This helps the model determine the specific data API.

页面 25 个 JSON 接口
选中: /x/web-interface/ranking/v2 (137763B)
结构: { code: number, message: string, ttl: number, data: { note: string, list: [100] { aid: number, ... } } }

接口清单(可传 api=<路径> 指定其中一个再提取字段):
  /x/web-interface/nav · 249B
  /x/vip/ads/materials · 1140B · 含数组
  ...

2. Field Mode

Provide the fields parameter to extract concrete field values. [] is used to represent array paths.

data.list[].title (前5条):
  用MC还原《神的随波逐流》 【B萌应援】
  WasteTheFallen丨首曝PV&实机演示:凝视深渊,人性渐泯
  ...

data.list[].stat.view (前5条):
  2428730
  10326863
  ...

3. Fixed Endpoint Mode

Use the api parameter to force a specific endpoint and ignore the automatic selection logic.

fetch_data(url, "/x/web-interface/ranking/v2")

Applicable Scenarios and Limitations

  • Applicable Scenarios: Suitable for scraping precise numerical values and list data, such as rankings, product prices, article view counts, and so on.
  • Login Limitations: Cannot access APIs behind a login wall (consistent with the behavior of read_url).
  • Static Sites: Purely static sites return a “no JSON API captured” error.
  • Dependency Requirements: Requires Node.js >= 20.

The plugin is designed with zero runtime dependencies (except for Node.js built-in modules), but it requires Playwright as the browser engine to perform network interception.

Conclusion

dsh-fetch-data mitigates the lack of data precision in text extractors by intercepting real APIs. It is suitable for scenarios that require exact values and structured data.