Introduction

In the agent development workflow, ensuring that model outputs meet expectations is a critical step. Traditional testing relies on manual verification, which is inefficient and prone to omissions. DeepSeek Harness (DSH), as a plugin-based ecosystem, pushes evaluation capabilities into reusable Skills. satan9394/dsh-llm-eval is a tool designed for this purpose. It standardizes the LLM evaluation process and helps developers quickly identify issues during model iteration.

What Is This

This is a DSH skill plugin belonging to the model reasoning category. It is maintained by the user satan9394. It is mainly used for automated evaluation of large language model (LLM) outputs, covering multiple dimensions ranging from faithfulness and relevance to hallucination detection. It is an important auxiliary tool for model quality assurance.

Core Features

This plugin implements the following evaluation dimensions:
* Faithfulness: Detects whether model output is faithful to the given instructions or context.
* Relevance: Evaluates the degree of relevance between the response content and the user query.
* Test Sets: Supports batch evaluation workflows based on specific test sets.
* Hallucination Detection: Identifies fabricated information in the model-generated content that is not present in the factual basis.
* Regression Guard: After model iteration or fine-tuning, quickly detects whether performance has degraded.

Installation & Enablement

According to the directory record, this plugin does not provide a preset installation command. DSH plugins are usually introduced through a directory link or a GitHub repository URL.
1. Visit the plugin directory page: https://www.skillhub.cn/plugins/satan9394/dsh-llm-eval
2. Or visit the source repository: https://github.com/satan9394/dsh-llm-eval
3. Add the corresponding plugin path or repository URL in the DSH configuration.

Typical Usage

Use it as a DSH Skill. In the evaluation workflow, this plugin loads the specified test set (if the configuration supports it) and scores or classifies the model’s responses. Refer to the documentation in the plugin directory or the source code for specific command invocation details.

Applicable Scenarios & Precautions

  • Applicable Scenarios: Agent projects that require automated evaluation pipelines; development scenarios that need regression testing to prevent model performance degradation; business logic with strict output quality requirements.
  • Precautions: DSH plugins usually run with the permissions of the current DSH process. When dealing with sensitive data, pay attention to permission isolation. Because the project license information is empty, be sure to check the source code and license terms before use.

Summary

satan9394/dsh-llm-eval provides DSH users with a toolkit for evaluating LLM outputs, covering core metrics such as faithfulness and hallucination detection. For developers who need stable model performance, this is a practical solution for regression testing and quality checking. For detailed documentation and source code, please refer to GitHub.