AI Agent Hub
Back to skills
Python Crawler Development Guide icon

Python Crawler Development Guide

Data Analysis Updated 2026.08.30

Paste the following prompt into your AI chat to install this skill:

Please install @user_5548e881/python-crawler-guide into the current AI assistant according to https://skillhub.cn/install/skillhub.md.

About this skill

Problem

Non-technical users often ask for a crawler, but the hard part is not the first script: it is turning goals into a durable project with clear fields, cadence, compliance, logging, storage, and acceptance criteria. This skill frames that request as a process instead of a one-off code dump, while still avoiding overengineering for simple one-page reads.

How it works

It starts with a workspace check for files like requirements.txt, crawler.py, config.yaml, and db.py, then routes to either a new-build path or an audit path. In development mode, it checks Python version, existing framework choices, database type, and hardcoded credentials. If the user says “you decide,” it states a reasonable assumption before continuing. It selects a stack by scenario: requests + BeautifulSoup for static pages, Playwright for dynamic pages, and Scrapy only for large, long-running jobs. The workflow enforces a contract: timeouts, retries, per-item exception handling, rate limiting, robots.txt, config separation, checkpointing, streaming data, rotating logs, and uniform error responses. Storage defaults are lightweight: Markdown for documents, Excel for structured rows, and database support only when explicitly requested, with connection pooling, parameterized queries, and upsert deduplication.

Boundaries

It is not intended for one-time static URL reads, general Python debugging, local HTML file parsing, or projects where the user has already chosen a dedicated tool. File safety rules are strict: do not delete user files, confirm before overwriting same-name files, keep changes inside the current project, and avoid unrelated refactors. When a blocker appears—JS rendering, blocking, or scheduled runs—it prefers free local options first, such as Playwright, cron, or systemd, and treats paid services only as a last resort.

Use Cases

  • When you need to scrape public product prices and names weekly into local Excel files.
  • When you want to audit an existing Python crawler for hardcoded secrets, unlimited logs, or uncontrolled requests.
  • When you need to add a crawler UI with start/stop controls and live log viewing.
  • When a crawler hits JS rendering or blocking and should first try local Playwright with rate limits.

Best For

  • Non-technical operators who describe fields and cadence in plain language without choosing frameworks.
  • Data analysts who need public web data saved to Excel or databases and refreshed regularly.
  • Project maintainers who want to review credentials, log rotation, error handling, and compliance risks.
  • Independent developers who need standardized pooling, checkpointing, and delivery checklists for new crawlers.