When DeepSeek Harness or other agents handle tasks involving a large number of input files, an Agent often needs to repeatedly perform the same checks: What is the file format? How many columns does the table have? What is the nested structure of the JSON/YAML? If these checks are rerun in every session, it not only wastes time and context window space but also makes it easy for temporary debug statements to be left in the final code.
The file-brief plugin aims to solve this problem. It is an Agent-agnostic skill (compatible with OpenAI Codex, Claude Code, and DeepSeek Harness) that can turn repetitive file checks into reusable, task-level structured documentation.
Plugin Positioning¶
file-brief is a task-level file catalog tool. It scans files under the task root directory and generates Markdown documentation and a SQLite retrieval index. The generated documentation includes only file structure metadata (such as field names, types, size, modification time, and statistical counts), and does not store raw data or cell examples. This protects privacy and saves storage space. The index is stored in the .file-catalog folder under the task directory.
Core Features¶
The plugin focuses on providing “structural knowledge” rather than the data itself:
- Privacy-preserving indexing: It records only file structure metadata (file name, format, size, SHA-256, field names, key names, types, dimensions, missing values, etc.), and does not save raw data, text paragraphs, code snippets, or cell values.
- Broad format support: It supports CSV, TSV, XLSX, Parquet, SQLite, JSON, YAML, ZIP, XML, HTML, Jupyter Notebook, Stata, R data files (RDS/RDA), various code files (Python, R, JS, etc.), as well as documents and images (PDF, DOCX).
- Cross-platform compatibility: As a DeepSeek Harness plugin, it can be installed and used together with the official CLI.
Installation and Activation¶
Install it through the official plugin command of DeepSeek Harness:
dsh plugin --profile web add file-brief
After installation, start a new Agent session so that the skill list is reloaded. The plugin itself is loaded via SKILL.md, and no additional configuration is required.
Usage¶
After installation, you need to run Python scripts in the task directory to generate and manage the index. <skill-dir> refers to the path of the installed skill.
1. Initial Cataloging¶
Scan all files under the task root directory and generate a structural index.
python "<skill-dir>/scripts/file_catalog.py" catalog --task-root "/work/my-task"
2. Query File Status¶
Check whether a specific file exists and whether it needs to be refreshed. A return value of fresh means it does not need to be reprocessed, while stale or missing means you need to run catalog to refresh it.
python "<skill-dir>/scripts/file_catalog.py" lookup --task-root "/work/my-task" "data/observations.csv"
3. Refresh Changes¶
Only restructure newly added or modified files.
python "<skill-dir>/scripts/file_catalog.py" catalog --task-root "/work/my-task" "data/observations.csv"
4. Search and Overview¶
Search for keywords in the index, or view the overall index statistics for the task directory.
python "<skill-dir>/scripts/file_catalog.py" search --task-root "/work/my-task" "species"
python "<skill-dir>/scripts/file_catalog.py" info --task-root "/work/my-task"
Environment Requirements and Notes¶
- Python version: Python 3.9 or higher is required.
- Dependencies: Core functionality depends on the standard library, but to support the full range of formats (such as Excel, Parquet, PDF, etc.), additional libraries must be installed:
python -m pip install pandas openpyxl pyarrow PyYAML pypdf Pillow tomli
- R data support: Parsing RDS/RData formats requires
Rscriptand the R packagelite(orjsonlite) to be installed on the system. - Permissions and security: The plugin runs with the permissions of the current DSH process. Check the source code and license before installation.
Applicable Scenarios¶
This plugin is suitable for any task that requires frequent reading and inspection of local files, especially engineering scenarios such as data analysis, code generation, and file conversion. By maintaining the .file-catalog directory, it can prevent the Agent from repeatedly reading the same files, allowing more tokens to be used for core business logic.
Summary¶
file-brief provides a standardized way to manage task file structure. By generating lightweight structural indexes, it allows Agents to quickly understand file overviews and reprocess files only when they change. This is very helpful for improving Agent execution efficiency and code quality.
- GitHub: https://github.com/Zhiyi-Zhao/file-brief
- DSH Market: https://www.skillhub.cn/plugins/Zhiyi-Zhao/file-brief