AI Agent Hub
Back to skills
Baidu Yijian Vision Agent icon

Baidu Yijian Vision Agent

AI Agent Updated 2026.08.29

Paste the following prompt into your AI chat to install this skill:

Please install @user_eb2b0784/wtf by following https://skillhub.cn/install/skillhub.md.

About this skill

What Problem It Addresses

Surveillance images and live scenes often need decisions such as person falls, vehicle crossings, occupied zones, or OCR. Raw frames alone are not enough; detection scope, counting rules, and reviewable output must be structured. This skill connects Baidu Yijian vision detection capabilities into a local workflow, turning image or video frames into inspectable structured results.

How It Works

  • Skill selection and registration: reads the platform skill list and matches user intent, such as person detection, vehicle recognition, OCR, pose estimation, or tracking.
  • Detection invocation: uses the YIJIAN_API_KEY environment variable to call a registered skill with an image or video frame; scripts write JSON to stdout and errors to stderr.
  • Area constraints: defines an ROI electronic fence to limit detection scope, or a tripwire to count crossing events.
  • Visualization: draws bounding boxes, generates grid references, and previews ROI/tripwire placement for coordinate confirmation.

Limits and Caveats

It requires access to the Baidu Yijian platform and a valid YIJIAN_API_KEY. It is intended for scenarios that already rely on vision detection services, not for deploying models, storing business data, or replacing a full monitoring backend; video workloads usually require frame extraction before per-frame processing.

Use Cases

  • Define an ROI in a corridor camera image and count only people entering a target room.
  • Extract frames from a 30-second video, then detect and track people frame by frame.
  • Draw a tripwire on a surveillance image and count crossings by people or vehicles.
  • Read the Yijian skill list, select a fall-detection skill from user intent, and invoke it.

Best For

  • Algorithm engineer integrating video feeds who needs local JSON calls to Yijian detection APIs.
  • Safety analyst monitoring live scenes who needs quick ROI or tripwire setup and event previews.
  • Inspection platform engineer who needs OCR, person, or vehicle skills for images and video frames.
  • Application integrator who needs intent-based skill selection and standard JSON output.