AI Agent Hub
Back to skills
TencentOS Kernel Crash Analysis Expert icon

TencentOS Kernel Crash Analysis Expert

IT Ops & Security Updated 2026.08.29

Paste the following prompt into your AI chat to install this skill:

Please install @user_7ac3e421/tencentos-crash-expert according to https://skillhub.cn/install/skillhub.md.

About this skill

Problem Solved

Analyzing Linux kernel crashes (vmcore) typically relies on manually executing crash commands, which presents two main issues: global states and local call stacks are easily confused, leading to "anchoring bias"; and the lack of a structured validation chain when searching for fix commits and evaluating mitigations often results in unverified, ineffective recommendations. This skill breaks down the analysis into 7 strictly controlled phases, enforcing "evidence-driven" investigation and "causal validation".

Core Capabilities & Key Steps

Powered by the AiCrasher MCP tool for automated crash debugging, the core workflow includes:
* Global Scanning First: In Soft/Hard Lockup scenarios, it mandates a global CPU state check (bt -a) first, preventing the misidentification of the Panic CPU's local snapshot as the root cause.
* Root Cause Routing: After analysis, it forces a "Kernel Defect" or "Non-Kernel Defect" classification. If it is a kernel defect, it proceeds to search for upstream fix commits and distro patch status; if it is a business configuration or user-space issue, it skips patch searching to avoid invalid retrieval.
* Parameter & Unit Cross-Validation: Built-in strict error-prevention mechanisms, such as verifying the RSS unit in ps output (KB, not pages), and requiring a three-step validation (source code, disassembly, and struct offsets) to confirm the actual runtime value of sysctl parameters.
* Standardized Reporting: Finally, it uses the crash_report_generator.py script to render the conclusions and crash command logs (@cmd[] references) into a Chinese HTML report, automatically releasing the MCP session resources.

Boundaries & Caveats

  • Environment Dependency: Strongly depends on the registration and session startup of AiCrasher MCP Server; a session restart is required after initial registration for hot-loading.
  • Source Code & Disassembly: For sysctl parameter troubleshooting or patch validation, a matching kernel-debuginfo or kernel source git repository is required. The sym command alone cannot locate per-netns struct fields.
  • Long-Context Forgetting Prevention: The skill enforces a "phase-switching rule", requiring the agent to re-read the corresponding phase*.md step file before entering each new phase to prevent early instructions from being diluted in long conversations.

Use Cases

  • A production server records a vmcore after a soft lockup, and you need to test whether the panic CPU stack is only local.
  • You receive an OOM crash dump and need to decide whether a kernel leak or a bad workload setting made processes unkillable.
  • With a kernel source repo available, you need to find the upstream fix commit and identify which distro tag contains it.
  • After analyzing a NULL pointer dereference, you need to produce validated sysctl mitigation advice and an HTML report.

Best For

  • Ops engineers responsible for postmortems on production Linux server crashes who need vmcore evidence in a deliverable root-cause report.
  • Infrastructure engineers handling kernel crash tickets who need to separate kernel defects from workload or user-space triggers.
  • Distro kernel patch maintainers who need to verify fix commits, release tags, and disassembly evidence in a local git repo.
  • SREs doing security and availability troubleshooting who need vmcore-based validation before recommending mitigations.