Baidu Yijian Vision Agent
Paste the following prompt into your AI chat to install this skill:
Please install @user_eb2b0784/wtf by following https://skillhub.cn/install/skillhub.md.
About this skill
What Problem It Addresses
Surveillance images and live scenes often need decisions such as person falls, vehicle crossings, occupied zones, or OCR. Raw frames alone are not enough; detection scope, counting rules, and reviewable output must be structured. This skill connects Baidu Yijian vision detection capabilities into a local workflow, turning image or video frames into inspectable structured results.
How It Works
- Skill selection and registration: reads the platform skill list and matches user intent, such as person detection, vehicle recognition, OCR, pose estimation, or tracking.
- Detection invocation: uses the
YIJIAN_API_KEYenvironment variable to call a registered skill with an image or video frame; scripts write JSON to stdout and errors to stderr. - Area constraints: defines an ROI electronic fence to limit detection scope, or a tripwire to count crossing events.
- Visualization: draws bounding boxes, generates grid references, and previews ROI/tripwire placement for coordinate confirmation.
Limits and Caveats
It requires access to the Baidu Yijian platform and a valid YIJIAN_API_KEY. It is intended for scenarios that already rely on vision detection services, not for deploying models, storing business data, or replacing a full monitoring backend; video workloads usually require frame extraction before per-frame processing.
Use Cases
- Define an ROI in a corridor camera image and count only people entering a target room.
- Extract frames from a 30-second video, then detect and track people frame by frame.
- Draw a tripwire on a surveillance image and count crossings by people or vehicles.
- Read the Yijian skill list, select a fall-detection skill from user intent, and invoke it.
Best For
- Algorithm engineer integrating video feeds who needs local JSON calls to Yijian detection APIs.
- Safety analyst monitoring live scenes who needs quick ROI or tripwire setup and event previews.
- Inspection platform engineer who needs OCR, person, or vehicle skills for images and video frames.
- Application integrator who needs intent-based skill selection and standard JSON output.
Related Skills
Provides Claw with character-library selection, switching, saving, and global SOUL.md style sync for role-based conversation.
An AIONE Agentic AI Infrastructure SDK wrapper for building production AI agents with memory, skills, workflows, and hooks.
A Python/TypeScript SDK wrapper for the DeepSeek-Reasonix native AI coding agent, with prefix-cache support.
A browser automation tool for analysts, operators, and developers that locates elements, fills forms, extracts structured content, and supports no-code scheduling and export.