DSH Desktop 的本地智能体开发中,语音交互是一个常见需求。目前大多数方案需要调用云端 API,存在延迟和隐私顾虑。dshtools-sensevoice-input 插件提供了一种完全本地化的解决方案。它利用 FunASR 和 SenseVoiceSmall 模型,在用户点击输入框旁的麦克风按钮后,直接将识别出的文本写入草稿区。
核心功能¶
- 集成麦克风与草稿写入:在 DSH Desktop 的 composer 输入框左下角集成了麦克风按钮。点击后录音,识别文本会自动通过
inputActions.setDraft写入输入框草稿。 - 语言与情感标签解析:识别结果包含语言和情感标签(例如
zh · NEUTRAL),方便后续处理。 - 纯本地推理:基于 FunASR 和 SenseVoiceSmall 模型进行推理,不依赖任何云服务,音频数据不出本机。
- 模型懒加载与自愈:模型采用懒加载机制,首次加载耗时 10–30 秒,后续单次识别仅需几秒。Python 进程崩溃后具备自愈能力。
- 主题自适应:UI 界面会自动适配当前的主题。
安装与启用¶
1. 准备推理环境¶
插件需要本地 Python 环境。首次使用需运行安装脚本,该过程会下载模型,占用约 5GB 空间。
.\scripts\setup.ps1
默认安装路径为 ~\.dsh-sensevoice,模型缓存于 ~\.cache\modelscope。也可通过环境变量 DSHSENSEVOICE_PYTHON 指定 Python 解释器。
2. 安装插件¶
安装后需重启 DSH Desktop。
dsh plugin --profile desktop add dshtools-sensevoice-input
若需要从 GitHub 源码安装:
dsh plugin --profile desktop add github:ilovedyou6666-hub/dshtools-sensevoice-input
3. 卸载¶
dsh plugin --profile desktop remove dshtools-sensevoice-input
典型用法¶
- 点击输入框左下角的 🎤 按钮开始录音(最长 60 秒)。
- 等待界面提示「♪ SenseVoice 正在识别…」,识别完成的文本会自动填入输入框。
- 悬浮按钮可显示模型就绪状态。若识别失败,会显示具体原因,点击即可重试。
注意事项¶
- 离线批处理:目前不支持流式输出,为离线批处理模式,单段录音时长限制为 60 秒。
- 首次识别延迟:首次使用时,CPU 加载模型需要 10–30 秒,这是正常现象。
- 权限范围:插件仅服务于 127.0.0.1 的桌面进程,由 DSH Desktop 自带访问控制,无需额外配置。
- 版本兼容性:若出现 HTTP 405 或 403 错误,通常是因为时序问题,需更新到最新版并重启。
- 依赖检查:若提示
python_not_found,需先执行scripts\setup.ps1;若提示model_load_failed,需重跑该脚本下载模型。
结语¶
该插件为 DSH Desktop 开发者提供了一个轻量级的本地语音输入方案。通过 inputActions.setDraft 的集成,语音交互可以无缝融入现有的工作流中。
Wait that final is Chinese. Need produce English. Oops final must be translation. Need replace. Need no more than final? We can now final English. Ensure no Chinese. Let’s craft final.
Need maybe preserve Markdown exactly. Use:
In local agent development with DSH Desktop, voice interaction is a common requirement. Most current solutions require calling cloud APIs, which introduces latency and privacy concerns. The dshtools-sensevoice-input plugin provides a fully localized solution. It uses the FunASR and SenseVoiceSmall models to write recognized text directly into the draft area after the user clicks the microphone button next to the input box.
Core Features¶
- Integrated microphone and draft writing: A microphone button is integrated into the lower-left corner of the DSH Desktop composer input box. After clicking it to start recording, the recognized text is automatically written into the input box draft through
inputActions.setDraft. - Language and emotion tag parsing: Recognition results include language and emotion tags (for example,
zh · NEUTRAL), which simplifies downstream processing. - Pure local inference: Inference is based on the FunASR and SenseVoiceSmall models, does not depend on any cloud services, and audio data never leaves the local machine.
- Model lazy loading and self-healing: Models are lazily loaded. The first load takes 10–30 seconds, while subsequent single recognitions only take a few seconds. It can self-heal after a Python process crash.
- Theme adaptation: The UI automatically adapts to the current theme.
Installation and Enablement¶
1. Set Up the Inference Environment¶
The plugin requires a local Python environment. On first use, run the setup script. This process downloads the model and occupies approximately 5 GB of space.
.\scripts\setup.ps1
The default installation path is ~\.dsh-sensevoice, and models are cached in ~\.cache\modelscope. You can also specify the Python interpreter via the environment variable DSHSENSEVOICE_PYTHON.
2. Install the Plugin¶
After installation, restart DSH Desktop.
dsh plugin --profile desktop add dshtools-sensevoice-input
If you need to install from GitHub source:
dsh plugin --profile desktop add github:ilovedyou6666-hub/dshtools-sensevoice-input
3. Uninstall¶
dsh plugin --profile desktop remove dshtools-sensevoice-input
Typical Usage¶
- Click the 🎤 button in the lower-left corner of the input box to start recording (up to 60 seconds).
- Wait for the interface to show “♪ SenseVoice is recognizing…”; the recognized text will automatically fill the input box.
- Hovering over the button shows the model readiness status. If recognition fails, the specific reason is shown; click it to retry.
Notes¶
- Offline batch processing: Streaming output is not currently supported. It runs in offline batch mode, with a maximum recording length of 60 seconds per segment.
- First-recognition latency: On first use, loading the model on CPU takes 10–30 seconds, which is normal.
- Permission scope: The plugin only serves the desktop process on 127.0.0.1. DSH Desktop provides built-in access control, so no additional configuration is required.
- Version compatibility: If HTTP 405 or 403 errors occur, they are usually caused by timing issues; update to the latest version and restart.
- Dependency check: If
python_not_foundis reported, runscripts\setup.ps1first; ifmodel_load_failedis reported, rerun the script to download the model.
Conclusion¶
This plugin provides DSH Desktop developers with a lightweight local voice input solution. Through integration with inputActions.setDraft, voice interaction can be seamlessly incorporated into existing workflows.
- Catalog page: dshtools-sensevoice-input
- GitHub repository: ilovedyou6666-hub/dshtools-sensevoice-input
Need check if any code spans should stay same. scripts\setup.ps1 unchanged. Good.
Final only this.
In local agent development with DSH Desktop, voice interaction is a common requirement. Most current solutions require calling cloud APIs, which introduces latency and privacy concerns. The dshtools-sensevoice-input plugin provides a fully localized solution. It uses the FunASR and SenseVoiceSmall models to write recognized text directly into the draft area after the user clicks the microphone button next to the input box.
Core Features¶
- Integrated microphone and draft writing: A microphone button is integrated into the lower-left corner of the DSH Desktop composer input box. After clicking it to start recording, the recognized text is automatically written into the input box draft through
inputActions.setDraft. - Language and emotion tag parsing: Recognition results include language and emotion tags (for example,
zh · NEUTRAL), which simplifies downstream processing. - Pure local inference: Inference is based on the FunASR and SenseVoiceSmall models, does not depend on any cloud services, and audio data never leaves the local machine.
- Model lazy loading and self-healing: Models are lazily loaded. The first load takes 10–30 seconds, while subsequent single recognitions only take a few seconds. It can self-heal after a Python process crash.
- Theme adaptation: The UI automatically adapts to the current theme.
Installation and Enablement¶
1. Set Up the Inference Environment¶
The plugin requires a local Python environment. On first use, run the setup script. This process downloads the model and occupies approximately 5 GB of space.
.\scripts\setup.ps1
The default installation path is ~\.dsh-sensevoice, and models are cached in ~\.cache\modelscope. You can also specify the Python interpreter via the environment variable DSHSENSEVOICE_PYTHON.
2. Install the Plugin¶
After installation, restart DSH Desktop.
dsh plugin --profile desktop add dshtools-sensevoice-input
If you need to install from GitHub source:
dsh plugin --profile desktop add github:ilovedyou6666-hub/dshtools-sensevoice-input
3. Uninstall¶
dsh plugin --profile desktop remove dshtools-sensevoice-input
Typical Usage¶
- Click the 🎤 button in the lower-left corner of the input box to start recording (up to 60 seconds).
- Wait for the interface to show “♪ SenseVoice is recognizing…”; the recognized text will automatically fill the input box.
- Hovering over the button shows the model readiness status. If recognition fails, the specific reason is shown; click it to retry.
Notes¶
- Offline batch processing: Streaming output is not currently supported. It runs in offline batch mode, with a maximum recording length of 60 seconds per segment.
- First-recognition latency: On first use, loading the model on CPU takes 10–30 seconds, which is normal.
- Permission scope: The plugin only serves the desktop process on 127.0.0.1. DSH Desktop provides built-in access control, so no additional configuration is required.
- Version compatibility: If HTTP 405 or 403 errors occur, they are usually caused by timing issues; update to the latest version and restart.
- Dependency check: If
python_not_foundis reported, runscripts\setup.ps1first; ifmodel_load_failedis reported, rerun the script to download the model.
Conclusion¶
This plugin provides DSH Desktop developers with a lightweight local voice input solution. Through integration with inputActions.setDraft, voice interaction can be seamlessly incorporated into existing workflows.
- Catalog page: dshtools-sensevoice-input
- GitHub repository: ilovedyou6666-hub/dshtools-sensevoice-input