Introduction¶
In the development and debugging of DeepSeek Harness (DSH), assessing the value of data assets is often a challenge. File size or download counts alone do not directly reflect data value. The signal that truly reflects the value of data assets is how they are used in real-world scenarios.
DataTally is a DSH plugin. Its core role is to aggregate public usage signals of data assets (such as downloads, citations, and Stars) into a single, verifiable profile. It does not calculate value or rankings; it only records and presents facts.
Plugin Overview¶
- Plugin Name:
daamaao/datatally - Maintainers: DAAMAAO
- Category: Memory
- License: MIT
DataTally consolidates usage data from distributed public sources and provides AI agents with tools to query and compare how data assets are used.
Core Features¶
The plugin provides three main tools for read-side data analysis:
search_assets: Search data asset profiles by keyword or domain.get_asset_profile: Retrieve the complete usage profile of a single asset, including metrics by source, timeline, and citation information.compare_assets: Compare two assets at the same signal level.
Installation and Activation¶
Before installing, ensure DeepSeek Harness is installed (dsh is in PATH) and Node.js version is 22.19 or higher.
Run the following command to install the plugin:
dsh plugin --profile web add datatally
After installation, the plugin can be loaded and used through the DSH Web context.
Typical Usage¶
When using it, you can instruct the agent in natural language to call the tools above. For example:
“Find datasets for sentiment analysis and compare the two most used ones.”
The agent will execute search_assets, find relevant datasets, and then call compare_assets to compare them.
Output Example (records only, no subjective evaluation):
Agent (calling search_assets("sentiment")):
- imdb (stanfordnlp/imdb) | domain: nlp | deep: downloads 195,669 | shallow: likes 725 | fetched_at: 2026-09-08T14:35:50Z
- glue (nyu-mll/glue) | domain: nlp | deep: downloads 826,859 ...
Agent (calling compare_assets("nyu-mll/glue", "stanfordnlp/imdb")):
- downloads [deep]: 791,429 vs 197,595 | same_source: true | same_access: true
- forks [deep]: 178 vs 0 | same_source: false | same_access: true
- likes [shallow]: 741 vs 722 | same_source: true | same_access: true
note: cross-source metrics are not directly comparable
The agent’s response only states the recorded data (citation counts, sources, fetch times) and does not conclude which one is better.
Use Cases and Limitations¶
Use Cases: When you need to verify the actual popularity, citation status, or usage frequency of data assets in public channels.
Notes:
* The CLI provided by the plugin is read-only and does not include a data refresh pipeline.
* The plugin does not calculate asset value, weighted signals, or rankings.
* The plugin does not recommend which one is better, nor does it sell data or fabricate numbers.
* Missing sources are explicitly marked; missing licenses are displayed as null.
Summary¶
DataTally focuses on recording the sources and facts of data, not value judgments. It provides a transparent data asset usage signal query layer for the DSH ecosystem, suitable for developers who need to objectively assess data asset activity.